Communication signal classification and identification method based on multi-sequence feature fusion

Through the multi-sequence feature fusion method, combined with the time-domain and frequency-domain feature extraction modules, feature fusion is solved by using the gating attention mechanism, and the problems of insufficient feature extraction and poor noise robustness in the traditional modulation recognition method are significantly improved, and the modulation recognition accuracy is significantly improved.

CN120223485APending Publication Date: 2025-06-27JILIN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510497179.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional modulation recognition methods rely on a single time series and cannot fully characterize the signal, resulting in insufficient feature extraction, low recognition rate, and poor noise robustness.

Method used

Using the multi-sequence feature fusion method, I/Q sequence, A/P sequence and PSD sequence are input to the multi-channel feature fusion network model, and feature fusion is performed through the time domain and frequency domain feature extraction modules, combined with the gating attention mechanism.

Benefits of technology

By characterizing signals from multiple angles, the cross-domain feature fusion mechanism enhances the complementarity of signal features, significantly improves the modulation recognition accuracy at low signal-to-noise ratio, and solves the problems of insufficient feature extraction and poor noise robustness in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223485A_ABST
    Figure CN120223485A_ABST
Patent Text Reader

Abstract

The invention discloses a communication signal classification and identification method based on multi-sequence feature fusion, and belongs to the technical field of communication, and the method takes an I / Q sequence, an A / P sequence and a PSD sequence of a signal as the input of a multi-channel feature fusion network model: the I / Q sequence retains the complete time domain in-phase / orthogonal component information of the signal, and provides basic waveform features; the A / P sequence highlights the instantaneous amplitude fluctuation and phase jump characteristics of the signal through polar coordinate transformation; the PSD sequence reveals the frequency domain energy distribution characteristics of the signal, and effectively captures a modulation mode with remarkable frequency domain characteristics. Three-dimensional complementation of a time domain, a transform domain and a frequency domain is formed, and the representation limitation of a single domain feature in a complex channel environment is overcome. And constructing a multi-channel feature fusion network model by using complementary gain information in different data to classify and identify a communication signal modulation mode. And deep processing is performed on different types of input data by using different network structures, so that the overall model recognition performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and specifically, relates to a communication signal classification and recognition method based on multi-sequence feature fusion. Background Art

[0002] Automatic Modulation Classification (AMC) technology is an important means to obtain the modulation type and its parameters of received signals, aiming to effectively obtain the modulation type of signals in a non-cooperative scenario, and lay a foundation for subsequent signal demodulation and further processing. Whether in the military or civilian fields, the automatic modulation recognition technology has very important research significance. In the modern military field, the competition for the control of the electromagnetic spectrum in the informationized battlefield is becoming increasingly fierce. To achieve this goal, the reconnaissance and analysis of enemy communication signals have become crucial tasks. As an important foundation for electromagnetic spectrum operations, AMC technology can help military personnel intercept and identify enemy communication signals to obtain military intelligence, and effectively interfere with and counter unknown enemy signals, providing strong support for the decision-making and implementation of military operations. In the civilian field, with the rapid growth of mobile communication devices, the spectrum resources have become increasingly crowded, and the problem of mutual interference between devices has become more prominent. To ensure communication security and the reasonable utilization of spectrum resources, real-time monitoring and analysis of communication signals through AMC technology can accurately identify the modulation type of signals, thereby effectively distinguishing different users, preventing illegal users from occupying precious spectrum resources, ensuring the normal operation and stability of the communication system, and providing a solid technical guarantee for the healthy development of modern communication society.

[0003] Currently, in the field of modulation recognition, there are mainly modulation recognition methods based on likelihood theory, modulation recognition methods based on feature extraction, and automatic modulation recognition methods based on deep learning. Among them, the modulation recognition method based on likelihood theory has a high computational complexity, and it is necessary to calculate the likelihood function for each possible modulation type, involving complex integral operations and probability density function estimation, and the dependence on prior knowledge, which may be difficult to obtain in practical applications; the modulation recognition method based on feature extraction depends on extracting specific features from signals, and the feature selection and extraction process requires professional knowledge. Different features have different sensitivities to different modulation types, and the generalization of the recognition method is poor.

[0004] Most traditional deep learning-based debugging recognition methods use a single sequence of information as the input of the network and improve the recognition accuracy by improving the network model. This method has the following disadvantages: 1. Using a single time-form signal representation ignores other potential important feature information and the complementary information between different modal features, resulting in confusion among modulation categories with similar single features; 2. With the rapid development of wireless communication technology, the modulation schemes of signals will become more complex and diverse, and the classification accuracy of single-sequence single-channel network models in low signal-to-noise ratio communication scenarios is poor. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem that the single time series in traditional modulation recognition fails to comprehensively represent the signal; the single-sequence single-channel network model has insufficient signal feature extraction, low modulation recognition rate, and poor noise robustness, and a communication signal classification and recognition method based on multi-sequence feature fusion is proposed.

[0006] The specific technical solution adopted by the present invention to achieve the above purpose is: a communication signal classification and recognition method based on multi-sequence feature fusion, including obtaining a communication signal to be classified and recognized; inputting the communication signal to be classified and recognized into a pre-trained multi-channel feature fusion network model to obtain a recognition result;

[0007] The multi-channel feature fusion network model includes a signal conversion module, a feature extraction module, and a feature fusion module, wherein the feature extraction module includes a frequency-domain feature extraction module and a time-domain feature extraction module arranged in parallel;

[0008] The signal conversion module is used to convert the communication signal to be classified and recognized into I / Q sequences, A / P sequences, and PSD sequences;

[0009] The time-domain feature extraction module takes time-domain features as input and outputs feature G(X i ), and the time-domain features include I / Q sequences and A / P sequences; the time-domain feature extraction module includes a first time convolutional layer, a second time convolutional layer, a first Dropout layer, a first Bi-LSTM layer, a multi-head attention layer, a second Bi-LSTM layer, and a first Dropout layer arranged in sequence, and the dilation factor of the second time convolutional layer is greater than that of the first time convolutional layer;

[0010] The frequency-domain feature extraction module takes the PSD sequence as input, and the frequency-domain feature extraction module is configured as:

[0011] First, use one-dimensional convolution to perform feature mapping on the PSD sequence to obtain coarse-grained feature F cg ;

[0012] Then, use dilated convolution on the coarse-grained feature F cgExtract fine-grained feature F fg ;

[0013] Then, input F cg and F fg into the parallel channel feature refinement network and spatial feature refinement network; in the channel feature refinement network, add F cg and F fg element-wise on the channels to obtain a fused feature, use the global average pooling layer to compress the fused feature into an aggregated channel, and activate the aggregated channel through the Sigmoid function to obtain a scaled channel weight vector and multiply W channel with F cg and F fg respectively, and then stack the two multiplied results on the channels to obtain the channel-refined feature In the spatial feature refinement network, concatenate F cg and F fg on the channels to obtain a concatenated feature, and use a one-dimensional convolution for channel compression; then use the Sigmoid function to generate an adaptive spatial weight vector Multiply W space with the concatenated feature to generate a spatially refined feature

[0014] Finally, add the channel-refined feature F cfr element-wise to the spatially refined feature F sfr to generate a fused refined feature F(IMF j ),

[0015] The feature fusion module is configured to perform feature fusion using the gated attention mechanism.

[0016] Furthermore, the generation process of the coarse-grained feature F cg is expressed as:

[0017] F cg = ReLU(Conv(PSD))

[0018] The process of extracting the fine-grained feature F fg is expressed as:

[0019]

[0020] where ReLU represents the activation function of the neural network; Conv represents convolution; and represent the weights of the dilated convolutions with dilation rates of 1, 2, and 5 respectively; represents element-wise addition.

[0021] Furthermore, in the channel feature refinement network, the process of channel feature refinement is expressed as:

[0022]

[0023] Among them, GAP represents the global average pooling layer, and Sigmiod represents the normalized Logistic function.

[0024] Furthermore, in the spatial feature refinement network, the process of spatial feature refinement is expressed as:

[0025] W space = sigmoid(Conv(Concat(F cg , F fg )))

[0026] F sfr = Concat(F cg , F fg )W channel

[0027] Among them, Conv represents convolution; Concat(F cg , F fg ) represents the merging of F cg and F fg .

[0028] Furthermore, the feature fusion module is specifically configured as:

[0029] The features G(X i ) output by the time-domain feature extraction module and the features F(IMF j ) output by the frequency-domain feature extraction module are pairwise fused on the channel to obtain the fused features:

[0030]

[0031] According to the importance of different information, the gating attention is used to allocate reasonable weights to the feature information from different sources; when the dynamic attention mechanism is activated, GAP will compress the input feature sequence into an aggregated channel where is a one-dimensional vector, represents the feature sequence after average pooling; two fully connected layers and the ReLU activation function are used to perform a non-linear transformation on to obtain the dependencies between channels of the input feature sequence, resulting in representing the feature sequence after non-linear activation; the process is expressed as:

[0032]

[0033] Among them, w fc1 and w fc2 respectively represent the weight matrices of two fully connected layers;

[0034] Through the Sigmoid function and activate respectively to obtain two gating weights λ g and λ f , λ g and λ f both range from 0 to 1 and satisfy λ g +λ f = 1; Apply the obtained gating weights to the input G(X i ) and F(IMF j ), Z′ j represents the gated fusion feature, and its process is expressed as:

[0035]

[0036] Through the above design scheme, the present invention can bring the following beneficial effects: The present invention characterizes the signal from multiple perspectives of time domain, transform domain, and frequency domain. The cross-domain feature fusion mechanism enhances the complementary gain information in different data, and the features of the signal are fully mined, solving the problem of insufficient signal feature extraction existing in traditional single sequence as input. Design corresponding feature extraction modules using the signal features of different sequences, and make full use of the differences and complementarities between different sequences to construct a multi-channel feature fusion network model, solving the problem of poor noise robustness in modulation recognition and significantly improving the modulation recognition accuracy under low signal-to-noise ratio. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to understand the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0038] Figure 1 is a flowchart of a communication signal classification and recognition method based on multi-sequence feature fusion;

[0039] Figure 2 is a structural diagram of a multi-channel feature fusion network model;

[0040] Figure 3 is a network diagram of a frequency domain feature extraction module;

[0041] Figure 4 is a network diagram of a time domain feature extraction module;

[0042] Figure 5 is a network diagram of a feature fusion module;

[0043] Figure 6 It is a graph showing the modulation recognition accuracy results of 11 modulation signals in RML2016.10a under different models. Detailed implementation manners

[0044] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention and the accompanying drawings. Obviously, the present invention is not limited by the following embodiments, and the specific implementation manners can be determined according to the technical solutions of the present invention and the actual situation. To avoid confusing the essence of the present invention, well-known methods, processes, and procedures are not described in detail. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0045] The communication signal classification and recognition method based on multi-sequence feature fusion proposed by the present invention takes the I / Q sequence, A / P sequence, and PSD sequence of the signal as the input of the multi-channel feature fusion network model: the I / Q sequence retains the complete time-domain in-phase / quadrature component information of the signal and provides basic waveform features; the A / P sequence highlights the instantaneous amplitude fluctuation and phase jump features of the signal through polar coordinate transformation and has strong representational ability for phase-sensitive modulations such as PSK and QAM; the PSD power spectral density reveals the frequency-domain energy distribution characteristics of the signal and effectively captures modulation modes with significant frequency-domain features. The three form a three-dimensional complementarity of time domain - transform domain - frequency domain, overcoming the representational limitations of single-domain features in complex channel environments. The multi-channel feature fusion network model is constructed using the complementary gain information in different data to classify and recognize the modulation modes of communication signals. Different network structures are used to deeply process different types of input data, and the overall model recognition performance is improved by reasonably organizing the network structure.

[0046] As Figure 1 shown, the specific implementation scheme of the present invention is as follows:

[0047] 1. Data preprocessing

[0048] Preprocess the discrete complex signals in the dataset and convert them into time-domain features and frequency-domain features, namely the in-phase / quadrature component I / Q sequence, amplitude / phase vector A / P sequence, and PSD power spectral density sequence.

[0049] 1.1 The dataset used in this invention is RML2016.10a, which is an open-source dataset widely used in the fields of wireless communication and machine learning. The RML2016.10a dataset contains 11 different modulation signals, with the SNR range from -20dB to 18dB at an interval of 2dB, including 11 modulation signals commonly used in modern communication such as 8PSK, AM-DSB, AM-SSB, BPSK, CPFSK, GFSK, 4PAM, 16QAM, 64QAM, QPSK, and WBFM. Among them, the three analog modulation types are AM-DSB, AM-SSB, and WBFM, and the eight digital modulation signals are 8PSK, BPSK, CPFSK, GFSK, 4PAM, 16QAM, 64QAM, and QPSK. This dataset contains a total of 220,000 signal samples. There are 1000 samples for each modulation method at each SNR. Among them, the maximum sampling rate offset is 50Hz, the maximum carrier frequency offset is 500Hz, and the sampling rate is 200Hz. This invention divides the RML2016.10a dataset according to the allocation ratio of 60% training set, 20% validation set, and 20% test set, and randomly shuffles the training set data.

[0050] 1.2 Convert the discrete complex signals in the dataset into in-phase / quadrature component I / Q sequences. The discrete complex signals are represented in complex form:

[0051] Z[n] = I[n] + jQ[n], n = 1, 2..., N;

[0052] where I[n] represents the in-phase component of the nth sample (in-phase component: the component in phase with the carrier); Q[n] represents the quadrature component of the nth sample (quadrature component: the component orthogonal to the carrier); N is the sampling length of the signal; j represents the imaginary part;

[0053] The I / Q sequence is represented as:

[0054]

[0055] 1.3 Convert the in-phase / quadrature component I / Q sequence into an amplitude / phase vector A / P sequence:

[0056]

[0057] where, Z A [n] and Z P [n] are the instantaneous amplitude and phase of the nth sample of Z[n] respectively;

[0058]

[0059] 1.4 Convert the in-phase / quadrature component I / Q sequence into a power spectral density PSD sequence:

[0060] Combine the in-phase component I[n] and the quadrature component Q[n] into a complex baseband signal: Z[n] = I[n] + jQ[n], n = 1, 2, ..., N; first calculate the mean of the complex signal and then remove the mean to eliminate the DC offset: Z′[n] = Z[n] - μ; Apply a Hamming window function to the mean-removed signal to suppress spectral leakage, Z win [n] = Z′[n] · w[n], n = 1, 2, ..., N, and the Hamming window function is: Calculate the power spectral density PSD, with the calculation formula:

[0061]

[0062]

[0063] where s(f) is the power spectral density function; F S is the sampling frequency (unit: Hz); Z win [n] is the complex signal after mean removal and windowing; U is the energy of the window function (for the Hamming window U ≈ 0.54N); Arrange the calculation result s(f) along the frequency axis to obtain a complete PSD sequence:

[0064]

[0065] The I / Q data completely preserves the time-domain information of the signal in complex form without any information loss. Its complex nature naturally avoids the phase wrapping problem. The I / Q sequence is the original baseband signal before demodulation by a communication receiver and is naturally suitable for preprocessing operations such as time-domain filtering and spectral analysis, providing a basic data foundation for joint time-frequency domain feature extraction.

[0066] The A / P signal separates the amplitude and phase into independent channels through feature decoupling, enabling the model to focus on the sensitive features of different modulation types (such as the amplitude change of AM modulation and the phase jump of PSK) respectively, reducing interference between features.

[0067] The power spectral density (PSD) of a signal represents the relationship between the power energy of the signal and frequency and is an important feature for distinguishing different modulation signals. Therefore, the power spectral density of the signal is used as one of the inputs to the neural network to provide more comprehensive signal features.

[0068] 2. Model construction

[0069] Figure 2The network architecture diagram of the multi-channel feature fusion network model in the present invention is shown. It is mainly divided into a signal conversion module: converting the original discrete complex signal into I / Q sequences, A / P sequences, and PSD power spectral density sequences; a feature extraction module: in terms of time-domain feature extraction, introducing a temporal convolutional network based on the multi-head attention mechanism; in terms of frequency-domain feature extraction, adopting a network architecture based on multi-scale dilated convolution and an efficient channel attention mechanism to extract features from the PSD sequence; a feature fusion module: feature fusion based on the gated attention mechanism is used to enhance the information gain and importance of the representative features of the signal.

[0070] 2.1 Frequency-domain Feature Extraction Module

[0071] To fully extract the regular features of the PSD sequence signal, a feature extraction module based on dilated convolution is constructed, and its structure is as Figure 3 shown. First, use one-dimensional convolution to perform feature mapping on the PSD sequence to obtain coarse-grained features F cg . Then, adopt dilated convolution to extract fine-grained features F cg from the coarse-grained features F fg to capture the detailed feature information distribution, and its operation process is expressed as:

[0072] F cg = ReLU(Conv(PSD));

[0073]

[0074] Among them, ReLU represents the activation function of the neural network; and respectively represent the weights of dilated convolutions with dilation rates of 1, 2, and 5, represents element-wise addition. To prevent the loss of local correlation of information, the dilation rates of the convolution are set to rate = 1, rate = 2, and rate = 5 respectively in a zigzag structure. Subsequently, use channel feature refinement and spatial feature refinement to further extract the refined features of the sub-modal signals.

[0075] In the channel feature refinement network, add F cg and F fg element-wise on the channel to obtain the fused features. Use the global average pooling layer to compress the fused features into aggregated channels, and activate the aggregated channels through the Sigmoid function to obtain the scaled channel weight vector C represents the number of channels of the feature; respectively multiply with F cg , F fg to refine the coarse-grained features F cg and the fine-grained features F fgImportant information. Stack these important feature vectors on the channels to obtain channel-refined features C represents the number of channels of the feature, and T represents the time dimension of the feature; the process of channel feature refinement is expressed as:

[0076]

[0077] Among them, GAP represents the global average pooling layer, and Sigmiod represents the normalized Logistic function. In the spatial feature refinement network, F cg and F fg are concatenated on the channels, and one-dimensional convolution is used for channel compression. Subsequently, an adaptive spatial weight vector is generated using the Sigmoid function Multiply W space by the concatenated features respectively to obtain refined spatial features The process of spatial feature refinement is expressed as:

[0078] W space = sigmoid(Conv(Concat(F cg ,F fg )));

[0079] F sfr = Concat(F cg ,F fg )W channel ;

[0080] Finally, add the channel-refined feature F cfr element-wise to the spatial-refined feature F sfr to generate the fused refined feature F(IMF j ), and its process is expressed as:

[0081]

[0082] 2.2 Time-domain Feature Extraction Module

[0083] For the I / Q sequence and A / P sequence, in order to accurately learn the long-term dependence relationship of signal features and fully extract the time-frequency features of the signal, a network architecture based on time convolution and recurrent neural network is constructed, as Figure 4 shown. The time-domain feature extraction module includes a first time convolution layer, a second time convolution layer, a first Dropout layer, a first Bi-LSTM layer, a multi-head attention layer, a second Bi-LSTM layer, and a second Dropout layer arranged in sequence.

[0084] The signals are input in I / Q and A / P formats. The first temporal convolutional layer of the temporal convolutional network captures local temporal patterns with a small dilation factor (such as dilation factor d = 1, 2). The second temporal convolutional layer increases the dilation factor to expand the receptive field and extract stride-dependent relationships, and constructs a multi-scale feature pyramid by setting different dilation factors, reducing the sequence length while preserving temporal causality. After the convolutional output, the first Dropout layer (Dropout means discarding units) randomly masks 25% of the neurons to enhance the feature robustness and prevent the model from overfitting. The first Bi-LSTM layer (Bi-LSTM means bidirectional long short-term memory network) integrates historical and future context information to capture the dynamic evolution law of temporal data, such as the continuous change of signal phase. The multi-head attention mechanism better captures and processes the relationships and global information between different parts of the time series, adaptively enhancing the feature contribution degree of key time slices while suppressing noise interference. On this basis, the concatenation and linear transformation of the multi-head outputs achieve spatial feature fusion, enabling the model to capture both phase continuity and resolve amplitude mutations. Then, based on the attention-weighted features, the second Bi-LSTM layer further refines the temporal context representation to filter out redundant information. Finally, the second Dropout layer is applied before the classifier for secondary regularization to improve the decision-making stability.

[0085] 2.3 Feature Fusion Module

[0086] To effectively integrate the time-domain features and frequency-domain features and enhance the information gain and importance of the representative features of the signals, a feature weighted fusion strategy is designed, and the structure is as Figure 5 shown. First, the features G(X i ) output by the time-domain feature extraction module and the features F(IMF j ) output by the frequency-domain feature extraction module are fused pairwise in the channels to obtain the fused features:

[0087]

[0088] According to the importance of different information, gated attention is used to assign reasonable weights to the feature information from different sources; when the dynamic attention mechanism is activated, GAP will compress the input feature sequence into an aggregated channel where is a one-dimensional vector, represents the feature sequence after average pooling; two fully connected layers (FC1 and FC2) and the ReLU activation function are used to perform non-linear transformation on to obtain the dependencies between channels of the input feature sequence, and represents the feature sequence after non-linear activation; the process is expressed as:

[0089]

[0090] Among them, w fc1 and w fc2 respectively represent the weight matrices of two fully connected layers;

[0091] Through the Sigmoid function and activate respectively to obtain two gating weights λ g and λ f , λ g and λ f both range from 0 to 1 and satisfy λ g +λ f = 1; Apply the obtained gating weights to the input G(X i ) and F(IMF j ), Z′ j represents the gated fusion feature, and its process is expressed as:

[0092]

[0093] Among them

[0094] That is:

[0095] Use the training set and the validation set to train and optimize the constructed multi-channel feature fusion network model.

[0096] 3. Experimental analysis

[0097] Test the final model with the test set and obtain the recognition result; The software used in the experiment is pycharm, the deep learning framework is Keras and Tesnorflow2.0, and the server GPU is 24GB NVIDIA GTX 3090TI. The simulation experiment training parameters are set as shown in the table.

[0098]

[0099]

[0100] Such as Figure 6Shows the modulation recognition accuracies of 11 modulation signals in RML2016.10a under different models. As the signal-to-noise ratio increases, the noise power in the signal relatively decreases, and the time-domain / frequency-domain features of the effective signal become clearer. For example, the fluctuation amplitudes of the instantaneous amplitude (A) and phase (P) decrease, and the time-series features of the signal are more easily captured by the model, so the accuracies of each model increase rapidly. A comparative experiment was conducted between the method of the present invention and the existing residual neural network ResNet, temporal convolutional network TCN, densely connected neural network DenseNet, convolutional-long short-term memory neural network CNN-LSTM, and k-nearest neighbor model KNN in the prior art. The modulation classification accuracies of different network models are as Figure 6 shown. As can be Figure 6 seen, the classification level of the present invention is significantly higher than that of the other five models at each signal-to-noise ratio. In particular, at low signal-to-noise ratios from -6 dB to 0 dB, since the present invention fully considers the complementarity between different modalities, the features of the signal are fully mined, and the cross-domain feature fusion mechanism enhances the complementary gain information in different data, overcoming the representational limitations of single-domain features in complex communication environments. Therefore, its classification level leads that of other models. In a high signal-to-noise ratio environment above 0 dB, the recognition accuracy of the present invention is the highest, and its classification accuracy remains at a level of 82% and above.

[0101] The advantages of the present invention are as follows:

[0102] 1. The present invention uses the A / P sequence, I / Q sequence, and power spectral density sequence PSD of the signal as the input of the neural network. The I / Q sequence retains the complete time-domain in-phase / quadrature component information of the signal, providing basic waveform features; the A / P sequence highlights the instantaneous amplitude fluctuation and phase jump features of the signal through polar coordinate transformation, and has strong representational ability for phase-sensitive modulations such as PSK and QAM; the PSD power spectral density reveals the frequency-domain energy distribution characteristics of the signal and effectively captures modulation modes with significant frequency-domain features. The three form a three-dimensional complementarity of time-domain - transform-domain - frequency-domain, the features of the signal are fully mined, and the cross-domain feature fusion mechanism enhances the complementary gain information in different data, overcoming the representational limitations of single-domain features in complex communication environments.

[0103] 2. Construct parallel network branches to process different modality data respectively: In terms of time-domain feature extraction, a temporal convolutional network based on the multi-head attention mechanism is introduced. The temporal convolutional network can accurately extract the long-term dependencies in the signal time-series features, and the multi-head attention mechanism allows the model to focus on different parts of the sequence from multiple perspectives, and can capture the correlations between different frequency components of the modulation signal;

[0104] In terms of frequency-domain feature extraction, a network architecture based on multi-scale dilated convolution and efficient channel attention mechanism is adopted to extract features from the PSD sequence, making full use of the coarse-grained and fine-grained features of the signal. The multi-scale dilated convolution can capture local details and global trends, while the efficient channel attention mechanism module further optimizes the combination of these multi-scale features, enabling the model to more comprehensively understand the complex patterns in the PSD sequence.

[0105] Characterize the signal from multiple perspectives of time domain, transform domain, and frequency domain. The cross-domain feature fusion mechanism enhances the complementary gain information in different data, and the features of the signal are fully exploited, solving the problems such as the lack of frequency-domain features and weak phase decoupling ability existing in the traditional I / Q sequence as a single input. Use different network structures to deeply process different types of input data, solve the problems of poor noise robustness in modulation recognition, and significantly improve the modulation recognition accuracy under low signal-to-noise ratio.

Claims

1. A communication signal classification and recognition method based on multi-sequence feature fusion, characterized in that: The method includes obtaining a communication signal to be classified and identified; inputting the communication signal to be classified and identified into a pre-trained multi-channel feature fusion network model to obtain an identification result; The multi-channel feature fusion network model includes a signal conversion module, a feature extraction module and a feature fusion module. The feature extraction module includes a frequency domain feature extraction module and a time domain feature extraction module arranged in parallel; The signal conversion module is used to convert the communication signal to be classified and identified into an I / Q sequence, an A / P sequence and a PSD sequence; The time domain feature extraction module takes the time domain feature as input and outputs the feature G(X i ), the time domain features include an I / Q sequence and an A / P sequence; the time domain feature extraction module includes a first time convolution layer, a second time convolution layer, a first Dropout layer, a first Bi-LSTM layer, a multi-head attention layer, a second Bi-LSTM layer and a first Dropout layer, which are arranged in sequence, and the expansion factor of the second time convolution layer is greater than the expansion factor of the first time convolution layer; The frequency domain feature extraction module takes the PSD sequence as input, and the frequency domain feature extraction module is configured as follows: First, one-dimensional convolution is used to perform feature mapping on the PSD sequence to obtain the coarse-grained feature F cg ; Next, dilated convolution is used to extract the coarse-grained features F cg Extract fine-grained features F fg ; Then, F cg and F fg Input into the parallel channel feature refinement network and spatial feature refinement network; in the channel feature refinement network, F cg and F fg Add each element on the channel to get the fused features, use the global average pooling layer to compress the fused features into aggregated channels, and activate the aggregated channels through the Sigmoid function to get the scaled channel weight vector And W channel Respectively with F cg 、F fg Multiply them, and then superimpose the two multiplied results on the channel to obtain the channel refinement feature In the spatial feature refinement network, F cg and F fg A concatenation operation is performed on the channel to obtain the concatenation feature, and a one-dimensional convolution is used for channel compression; then the Sigmoid function is used to generate an adaptive spatial weight vector W space Multiply with the concatenated features to generate spatially refined features Finally, the channel refinement feature F cfr Add the spatial refinement feature F element by element sfr Generate fused refined features F(IMF j ), The feature fusion module is configured to perform feature fusion using a gated attention mechanism.

2. The communication signal classification and recognition method based on multi-sequence feature fusion according to claim 1 is characterized in that: Coarse-grained feature F cg The generation process of is expressed as: F cg =ReLU(Conv(PSD)) Extract fine-grained features F fg The process is expressed as: Among them, ReLU represents the activation function of the neural network; Conv represents convolution; and Represent the weights of the dilated convolutions with dilation rates of 1, 2, and 5, respectively; Represents addition of elements.

3. The communication signal classification and recognition method based on multi-sequence feature fusion according to claim 1 is characterized in that: In the channel feature refinement network, the process of channel feature refinement is expressed as: Among them, GAP represents the global average pooling layer, and Sigmiod represents the normalized Logistic function.

4. The communication signal classification and recognition method based on multi-sequence feature fusion according to claim 1 is characterized in that: In the spatial feature refinement network, the process of spatial feature refinement is expressed as: W space =sigmoid(Conv(Concat(F cg ,F fg ))) F sfr =Concat(F cg ,F fg )W channel Among them, Conv means convolution; Concat(F cg ,F fg ) indicates the merge F cg and F fg .

5. The communication signal classification and recognition method based on multi-sequence feature fusion according to claim 1 is characterized in that: The specific configuration of the feature fusion module is: The feature G(X) output by the time domain feature extraction module i ) and the feature F(IMF j ) are fused in pairs on the channel to obtain the fusion feature: According to the importance of different information, gated attention is used to assign reasonable weights to feature information from different sources; When the dynamic attention mechanism is activated, GAP will input the feature sequence Compression into aggregate channels in is a one-dimensional vector, Represents the feature sequence after average pooling; two fully connected layers and ReLU activation function are used to Perform nonlinear transformation to obtain the dependency of input feature sequence between channels, and obtain Represents the feature sequence after nonlinear activation; The process is expressed as: Among them, w fc1 and w fc2 Represent the weight matrices of the two fully connected layers respectively; Through the Sigmoid function and Activate and obtain two gating weights λ g and λ f ,λ g and λ f The range is between 0 and 1, and satisfies λ g +λ f =1; the obtained gating weights are applied to the input G(X i ) and F(IMF j ), Z′ j represents the gated fusion feature, and its process is expressed as:

Citation Information

Cited By

  • Physically guided deep learning signal modulation identification method and system

    CN122339908A