Complex electromagnetic signal demodulation method based on multi-task learning
By using a hybrid architecture of ResNet-18 and Transformer and a dynamic weight allocation mechanism, a multi-task learning model is constructed, which overcomes the limitations of existing technologies in demodulating complex electromagnetic signals. It achieves accuracy and robustness in modulation type identification and symbol sequence, improves the model's adaptability and robustness in complex environments, and solves the problems of feature fragmentation and task conflict in multi-task learning.
Patent Information
- Application Number
- CN202511417969.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing deep learning methods have limitations when dealing with signal overlap, noise interference, and multipath propagation in complex electromagnetic environments. They also tend to focus on single tasks and lack joint optimization for multiple tasks, making it difficult for models to meet the needs of multiple tasks simultaneously. Furthermore, they increase computational overhead when processing variable-length sequence data and may lead to position encoding shifts.
A hybrid architecture of ResNet-18 and Transformer is adopted, combined with a dynamic weight allocation mechanism, to construct a multi-task learning model. Through cross-branch feature interaction gating units and dynamic weight allocation mechanism, the model achieves collaborative optimization of modulation type classification and code sequence generation, and suppresses the negative transfer effect caused by task conflict.
It improves the accuracy of modulation type identification and symbol sequence generation, enhances the model's adaptability and robustness in complex environments, and solves the problems of feature fragmentation and task conflict in multi-task learning through collaborative optimization.
Smart Images

Figure CN120896824A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electromagnetic signal demodulation, in particular to a complex electromagnetic signal demodulation method based on multi-task learning. BACKGROUND
[0002] Electromagnetic signals are information carriers that propagate through electromagnetic waves and are widely used in communication, radar, navigation, remote sensing, and other fields. Modulation and demodulation of electromagnetic signals are key technologies for information transmission and reception. Traditional demodulation techniques are mainly based on the frequency domain or time domain characteristics of signals, and are implemented through hardware devices such as filters and demodulators. However, with the increasing complexity of the electromagnetic environment, traditional demodulation techniques have obvious limitations in dealing with noise interference, multipath propagation, and signal overlap. In recent years, artificial intelligence and machine learning techniques have been widely applied in the field of electromagnetic signal demodulation. Deep learning-based demodulation methods, such as convolutional neural networks and recurrent neural networks, can effectively extract signal features and improve the accuracy and robustness of demodulation. However, existing deep learning methods still have limitations in dealing with signal overlap, noise interference, and multipath propagation in complex electromagnetic environments, and are mostly focused on single tasks, lacking joint optimization of multiple tasks, which makes it difficult for the model to meet the needs of multiple tasks simultaneously, and the conflict and dependency between tasks are not fully utilized. In addition, existing methods usually use padding or truncation when dealing with variable-length sequence data, which not only increases computational overhead but also may cause position encoding offset, affecting the performance of the model.
[0003] Therefore, a complex electromagnetic signal demodulation method based on multi-task learning is provided to solve the above problems by introducing the conflict and dependency between multiple tasks to improve the accuracy of prediction. SUMMARY
[0004] To solve the above problems, the present application provides a complex electromagnetic signal demodulation method based on multi-task learning, which realizes the collaborative optimization of modulation type classification and code sequence generation by constructing a multi-task learning framework with ResNet-18 and Transformer as the backbone, effectively suppresses the negative transfer effect caused by task conflict through a dynamic weight distribution mechanism, and improves the accuracy of modulation type recognition and the accuracy of code sequence generation through collaborative optimization.
[0005] To achieve the above purpose, the present application provides a complex electromagnetic signal demodulation method based on multi-task learning, comprising the following steps: S1: Obtain signal data sets under different modulation modes, and divide the signal data sets under different modulation modes into training sets and validation sets; S2: adopting a hybrid architecture of ResNet-18 and Transformer as a backbone network, combining a dynamic weight distribution mechanism to construct a multi-task learning model; S3: training the multi-task learning model using the training set and the validation set to obtain a trained multi-task learning model; S4: inputting the complex electromagnetic signal to be tested into the trained multi-task learning model to obtain a modulation type recognition result and a symbol sequence prediction result.
[0006] Preferably, S1 further comprises pre-processing the signal data in the data set, the pre-processing comprising: removing noise and clutter in the signal data by using a band-pass filtering method, and visualizing the denoising result by constellation diagram.
[0007] Preferably, S2 specifically comprises: The backbone network adopts a hybrid architecture of ResNet-18 and Transformer, the ResNet-18 branch is used to extract local time sequence features and modulation type discriminative patterns, and the Transformer branch captures long-term spectral correlation and modulation parameter evolution law through a multi-head self-attention mechanism; The dynamic weight distribution mechanism is used to adaptively adjust the loss weights of the modulation type recognition task and the symbol sequence prediction task; A cross-branch feature interaction gate unit is set to dynamically adjust the fusion ratio of ResNet-18 deep features and Transformer sequence features.
[0008] Preferably, the dynamic weight distribution mechanism in S2 specifically comprises a dynamic EMA weight balancing mechanism, a sequence alignment and bit error rate integration, and a composite loss quantification system; The dynamic EMA weight balancing mechanism is used to dynamically adjust the loss weight of each task through an exponential moving average (EMA) algorithm; The sequence alignment and bit error rate integration are used to utilize a dynamic sequence alignment algorithm to realize length matching of the predicted sequence and the target sequence through permute operation and zero padding, and embed the bit error rate as a soft constraint into the loss function; The composite loss quantification system is used to weight the modulation type classification loss, the sequence generation probability distribution loss, and the bit error rate loss scaled by a temperature coefficient to form a total loss function.
[0009] Preferably, the exponential moving average (EMA) algorithm is expressed as: ; Wherein, is a decay factor, is the loss weight of the t-th time step, is the loss weight of the t-1-th time step, is the original loss value.
[0010] Preferably, the bit error rate is embedded into the loss function as a soft constraint, denoted as: ; wherein, is an indicator function, is the number of valid bits, is the target sequence, is the predicted sequence.
[0011] Preferably, the total loss function adopts a hierarchical multi-task joint modeling framework to decouple the modulation type classification and the sequence generation probability distribution into two parallel sub-tasks; the main path adopts a cross-entropy loss to quantify the modulation type classification error, and the loss weight is adaptively adjusted through a dynamic EMA mechanism; the auxiliary path optimizes the sequence generation probability distribution through a cross-entropy loss and monitors the performance of the model through the bit error rate calculation.
[0012] Preferably, the total loss function is denoted as: ; wherein, is the modulation type classification loss weight, is the sequence generation probability distribution loss weight, is the bit error rate loss weight scaled by a temperature coefficient, is the modulation type classification loss, is the sequence generation probability distribution loss, is the bit error rate loss scaled by a temperature coefficient.
[0013] Preferably, the model training adopts a dynamic batch construction strategy, an optimizer setting, and a multi-dimensional stability enhancement mechanism; The dynamic batch construction strategy is used to ensure consistent sequence length within each batch through width clustering and dynamic batch construction; The optimizer setting is used to adopt the AdamW optimizer algorithm combined with the cosine annealing learning rate scheduler to balance the momentum update and the adaptive learning rate correction.
[0014] Preferably, the multi-dimensional stability enhancement mechanism includes an adversarial data augmentation strategy, a gradient clipping technique, and a model checkpoint saving strategy; The adversarial data augmentation strategy adds Gaussian white noise and transient interference pulses at the input end to make the model learn noise invariance features; The gradient clipping technique is used to limit the risk of gradient explosion, with a maximum norm constraint of 1.0, and the joint gradient of the cross-branch feature interaction gating unit is normalized through the clip_grad_norm function; Early stopping mechanism and model checkpoint saving strategy are used to terminate training when the weighted total loss on the validation set does not decrease for 15 consecutive cycles and to preserve the optimal parameter state.
[0015] Therefore, the application adopts the above-mentioned complex electromagnetic signal demodulation method based on multi-task learning. In the electromagnetic signal demodulation process, the modulation mode of the signal is first judged, and then the original code sequence is predicted. It can be seen that modulation type identification is the basis of code sequence prediction, and the accuracy of modulation type identification directly affects the final result of code sequence prediction. Based on this, the multi-task learning model designed by the application improves the accuracy of modulation type identification through the cooperative optimization of two tasks, and at the same time improves the accuracy of code sequence generation. Through the ResNet-18 network, the local time sequence features and modulation type discriminative patterns of the signal are extracted, and the long-range spectral correlation and modulation parameter evolution law are captured by using the Transformer architecture, forming a dual-task branch of modulation classification and code sequence generation. To solve the common feature fragmentation problem in multi-task learning, a cross-branch feature interaction gating unit is innovatively designed to dynamically adjust the fusion ratio of ResNet-18 deep features and Transformer sequence features, realizing the implicit cooperative optimization of modulation feature learning and sequence generation tasks. Further, the dynamic weight distribution mechanism effectively suppresses the negative transfer effect caused by task conflict, thereby further improving the adaptability and robustness of the model in complex environments.
[0016] The technical solutions of the application will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 A flowchart of a complex electromagnetic signal demodulation method based on multi-task learning in the application; Figure 2 A model framework diagram in an embodiment of the application; Figure 3 An I signal time domain graph in an embodiment of the application; Figure 4 A Q signal time domain graph in an embodiment of the application; Figure 5 A filtered I signal time domain graph in an embodiment of the application; Figure 6 A filtered Q signal time domain graph in an embodiment of the application; Figure 7 A standard QPSK constellation graph in an embodiment of the application; Figure 8 A constellation graph of filtered I / Q signals in an embodiment of the application; Figure 9The loss function of the embodiment of the present application changes with the number of training rounds. Among them, (a) is the loss change line chart of the training set, and (b) is the loss change line chart of the verification set. Figure 10 The confusion matrix diagram of the classification result in the embodiment of the present application. DETAILED DESCRIPTION
[0018] The detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but only represents selected embodiments of the application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0019] EMBODIMENT A complex electromagnetic signal demodulation method based on multi-task learning, as shown in Figures 1-2 includes the following steps: S1: Obtain signal data sets under different modulation modes, and divide the signal data sets under different modulation modes into training sets and verification sets; The data set comes from the third "Electromagnetic Big Data Challenge", and the data set has 10 known modulation signal sequence samples, and the ID mapping relationship is as follows: 1: BPSK, 2: QPSK, 3: 8PSK, 4: MSK, 5: 8QAM, 6: 16-QAM, 7: 32-QAM, 8: 8-APSK, 9: 16-APSK, 10: 32-APSK, the signal generation process covers different signal lengths, a wide range of pulse widths, different symbol widths, and simulates different signal-to-noise ratio SNR conditions and multipath, fading and other interference environments, and the answer of the sequence sample is the modulation type, the symbol width and the modulation code sequence.
[0020] The sampling rate of the training data setting signal is 20MHz, and it contains 10 folders, each folder has 1.8W different signal lengths, symbol widths, SNR and multipath sample data, the data format is csv file, the first two columns are simulation IQ waveform, the third column is simulation code sequence, the fourth column is simulation modulation type, and the fifth column is simulation symbol width. The data sample format is shown in Table 1.
[0021] Table 1 Data sample format ;
[0022] The dataset simulates a real wireless environment, and the data classification difficulties are as follows: (1) there are many modulation methods, including various common digital modulation methods; (2) the symbol width and symbol sequence length are not fixed, resulting in a large difference in I / Q signal sequence length; (3) the electromagnetic data constellation is not significantly different, and to increase the difficulty of classification, multipath, fading and other interference environments are added, and noise and other actual influencing factors are added. The above three factors bring certain difficulty to modulation recognition, and are also the problems that the field tries to solve.
[0023] Due to the limitation of computing power, only the data sets corresponding to the following five modulation methods in the data set are taken for training in this embodiment: 1: BPSK, 2: QPSK, 3: 8PSK, 4: MSK, and 5: 8QAM. Meanwhile, the I / Q signal sequence and the symbol sequence are divided into a certain length, reducing the length of the sequence. The training set and the validation set are divided according to the ratio of 8:2.
[0024] S1 also includes preprocessing the signal data in the data set, which includes: removing noise and clutter in the signal data using a band-pass filtering method, and visualizing the denoising result through constellation diagram.
[0025] First, use the traditional filtering method to preprocess the data, remove noise, clutter and other influencing factors in the original I / Q signal data, and visualize the denoising result through constellation diagram.
[0026] Take a data modulated by QPSK as an example. QPSK (Quadrature Phase Shift Keying) is a digital modulation method that transmits data by changing the phase of the carrier signal. QPSK carries two bits of data on each symbol, so its spectrum utilization efficiency is higher, and it is a common modulation technique, especially widely used in wireless communication. Its carrier phase has four possible values, corresponding to the four bit combinations 00, 01, 10 and 11.
[0027] As can be seen from the I / Q signal time domain diagram as shown in Figures 3-4 The I / Q signal waveform fluctuates very obviously, indicating that the original signal sequence contains a lot of noise.
[0028] First, use the band-pass filtering method to simply process the data. A band-pass filter is a filter that only allows middle frequencies to pass through. It only retains signals with frequencies within a certain middle range, and filters out high and low frequency signals above and below that range, so that the passing signal only contains middle frequency components. The frequency response characteristic of such a filter is relatively flat in the middle frequency band, and rapidly decays in the high and low frequency bands. The lower cutoff frequency of the band-pass filter is set to 10MHz, the upper cutoff frequency is set to 50MHz, and the filter order is set to 100, to obtain Figures 5-6The filtered I / Q signal time domain diagram is shown. As can be seen from the figure, the band-pass filter well filters out some noise in the original signal sequence, making the signal sequence more orderly, but some noise in the sequence is not filtered out.
[0029] The constellation diagram drawn with the filtered I / Q signal is shown in Figure 8 As shown, the signal sequence contains a lot of noise. The constellation diagram of the standard QPSK modulated signal is shown in Figure 7 As shown, all signal points converge around four points; and in the constellation diagram drawn according to the data set, the signal points are scattered and there is no clear convergence center, indicating that a large amount of noise data is contained therein. It can be seen therefrom that the ordinary filtering method is difficult to directly process the noise in the data set, and a deep learning method needs to be used to classify the data set.
[0030] S2: A hybrid architecture of ResNet-18 and Transformer is used as the backbone network, a dynamic weight distribution mechanism is combined, and a multi-task learning model is constructed; In the electromagnetic signal demodulation process, the modulation mode of the signal needs to be judged first, and then the original code sequence is predicted. Therefore, modulation type identification is the basis of code sequence prediction, and the accuracy of modulation type identification directly affects the final result of code sequence prediction. Based on this, the embodiment designs a multi-task learning model, which improves the accuracy of modulation type identification through the cooperative optimization of two tasks, and improves the accuracy of code sequence generation.
[0031] Specifically, the backbone network adopts a hybrid architecture of ResNet-18 and Transformer, the ResNet-18 branch is used to extract local time sequence features and modulation type discriminative patterns, and the Transformer branch captures long-term spectral correlation and modulation parameter evolution law through a multi-head self-attention mechanism; ResNet-18 is composed of 4 residual block groups, each containing different numbers of residual blocks, each containing multiple convolutional layers, batch normalization layers, and activation functions. Taking the common two-layer residual block as an example, the input x first passes through a convolutional layer, then undergoes batch normalization and ReLU activation, followed by a second convolutional layer and batch normalization, obtaining the residual F(x). Finally, the residual F(x) is added to the input x, and then undergoes ReLU activation to obtain the output of the residual block. When the dimensions of the input and the residual are inconsistent, an additional convolutional layer (usually a 1x1 convolution) is needed to adjust the dimensions of the input x so that the addition operation can be performed. ResNet-18, as a deep neural network model, has strong feature extraction and classification capabilities. In practical applications, the network structure can be optimized by adjusting the number of residual blocks, the parameters of convolutional layers, etc., to improve the performance of the model. In addition, other techniques such as data augmentation, regularization, etc. can be combined to further improve the generalization ability of the model, which is very suitable for feature extraction of I / Q signal sequences.
[0032] Transformer is a sequential data modeling architecture with long-distance feature capture capability, feature extraction capability, and efficient parallel computing capability. The Transformer architecture introduces an attention mechanism, allowing the model to consider all positions in the input sequence simultaneously, rather than processing them sequentially. The self-attention mechanism allows the model to weight each position based on the relationship between different positions in the input sequence, capturing global context information and being widely used in sequence-to-sequence tasks. The self-attention mechanism generates three vector representations of query (Query), key (Key), and value (Value) for each element in the input sequence, constructs a similarity measurement matrix across positions, and then dynamically integrates context information through weighted aggregation.
[0033] In mathematical form, given the input sequence where is the sequence length, is the feature dimension, the self-attention mechanism first generates the query matrix , the key matrix and the value matrix , where , , are learnable parameter matrices, and are the dimensions of the query and key, and the value, respectively. Then, the scaled dot-product attention score of the query and the key is calculated: ; where The softmax function normalizes the scores into a probability distribution, implementing a soft selection of attention weights. This mechanism explicitly models the similarity of all pairs of positions in the sequence through matrix operations, and then generates the context-aware representation of each position through weighted summation.
[0034] To enhance the model's representation ability of multi-granularity information, the Transformer introduces a multi-head attention mechanism, i.e., the input features are divided into independent subspaces, and after performing the above attention calculation in each subspace in parallel, the results are spliced and fused through linear transformation: (3.3); where , is the output projection matrix. The multi-head mechanism pays attention to different feature dimensions of the input sequence through different subspaces, significantly improving the model's representation ability and robustness. The advantage of self-attention mechanism lies in its parallelization and global dependency modeling capability. Self-attention allows simultaneous processing of full sequence information in a single forward propagation, and its computational complexity grows quadratically with the sequence length, but can be optimized to linear complexity through sparse attention patterns. In addition, the self-attention mechanism has position awareness, which injects sequence position information into the input vector by adding position encoding, allowing the model to distinguish the context relationships of different positions and capture the mutual relationships between different time sampling data in the I / Q signal sequence.
[0035] The dynamic weight distribution mechanism is used to adaptively adjust the loss weights of the modulation type recognition task and the symbol sequence prediction task; the dynamic weight distribution mechanism specifically includes a dynamic EMA weight balancing mechanism, sequence alignment and bit error rate integration, and a composite loss quantification system; The dynamic EMA weight balancing mechanism is used to dynamically adjust the loss weight of each task through the exponential moving average (EMA) algorithm; To address the negative transfer problem in multi-task learning, a dynamic weight distribution system based on exponential moving average (EMA) is designed. Three learnable normalization factors mod_loss_scale, seq1_loss_scale and seq2_loss_scale fuse historical loss statistics through the EMA algorithm: ; where, is the decay factor, which can be 0.1, is the loss weight at the t-th time step, is the loss weight at the t-1-th time step, is the original loss value. This mechanism enables the loss weight to adapt to the data distribution changes of different tasks and the characteristics of the training stage. Compared with static weight setting, the dynamic adjustment of EMA improves the sensitivity of task weight adjustment of the model in the dynamic electromagnetic interference scene, effectively suppressing the gradient cancellation phenomenon caused by task conflict.
[0036] Sequence alignment and bit error rate integration is used to utilize the dynamic sequence alignment algorithm to match the lengths of the predicted sequence and the target sequence through permute operation and zero padding, and to embed the bit error rate as a soft constraint into the loss function. For the variable-length sequence generation task, a dynamic sequence alignment algorithm is designed to match the lengths of the predicted sequence and the target sequence through permute operation and zero padding, while minimizing the interpolation error while maintaining temporal continuity. Further innovatively, the bit error rate (BER) is embedded as a soft constraint into the loss function: ; wherein, is the indicator function, is the number of valid bits (filtering padding values), is the target sequence, is the predicted sequence. This design enables the model to optimize the sequence generation probability distribution while directly aligning the bit-level performance indicators, solving the scale difference problem between traditional cross-entropy loss and actual bit error rate.
[0037] Composite loss quantification system is used to weight the modulation type classification loss, sequence generation probability distribution loss, and temperature coefficient scaled bit error rate loss to form the total loss function.
[0038] The total loss function is represented as: ; wherein, is the modulation type classification loss weight, is the sequence generation probability distribution loss weight, is the temperature coefficient scaled bit error rate loss weight, is the modulation type classification loss, is the sequence generation probability distribution loss, is the temperature coefficient scaled bit error rate loss.
[0039] The loss function adopts a hierarchical multi-task joint modeling framework to decouple the modulation type classification (discrete label prediction) and the simulation code sequence reconstruction (continuous sequence generation) into two parallel sub-tasks. The main path adopts a cross-entropy loss (CrossEntropyLoss) to quantify the modulation classification error, and its loss weight is adaptively adjusted through a dynamic EMA mechanism. The auxiliary path innovatively introduces a double constraint mechanism: on the one hand, it directly optimizes the sequence generation probability distribution through a cross-entropy loss, and on the other hand, it realizes bit-level performance monitoring through a bit error rate (BER) calculation. This triple loss architecture breaks through the limitations of traditional single-objective optimization and forms an implicit collaboration between modulation feature learning and sequence generation tasks.
[0040] A cross-branch feature interaction gating unit is set up to dynamically adjust the fusion ratio of ResNet-18 deep features and Transformer sequence features.
[0041] To address the feature fragmentation problem in multi-task learning, a cross-branch feature interaction gating unit is designed to dynamically adjust the fusion ratio of ResNet-18 deep features and Transformer sequence features, achieving implicit collaborative optimization of modulation classification tasks and code sequence generation tasks.
[0042] This embodiment constructs a dual-flow multi-task network based on cross-modal feature interaction and dynamic task coordination. Based on the shared I / Q orthogonal signal time-frequency feature space, a hierarchical multi-task collaborative mechanism is adopted. The input signal is processed by a dual-channel preprocessing module to generate time-domain waveform features and frequency-domain features, and the cross-modal feature fusion network is used to realize the complementary information of time and frequency domains.
[0043] When predicting the length of the symbol sequence, the symbol width and the sampling rate are considered as known constants, and the product of the two is used as the length of the token to divide the I / Q signal sequence, forming a [batch_size, seq_len, token_len, 2] dimensional dataset, which converts the long sequence prediction task into a task of predicting one symbol with one token. Then, the information of the modulation type is added to learn the different mapping methods corresponding to different modulation types. The transformer is used to learn the relationship between tokens to capture the mutual relationship between symbols and reduce the influence of noise, thereby improving the accuracy of symbol sequence prediction and realizing dynamic prediction of the length of the symbol sequence.
[0044] Multi-task learning (MTL) is a method that improves the generalization performance and robustness of a system on a single task by constructing a shared parameterized model architecture and implementing knowledge transfer and representation learning during the joint training process of multiple related tasks. The theoretical basis of this method can be traced back to the joint risk minimization principle in statistical learning theory, that is, by jointly optimizing the loss functions of multiple tasks, the inherent correlation between tasks is used to alleviate the risk of overfitting, and the potential shared feature representation in the data is mined. At the mathematical framework level, multi-task learning can be expressed as a multi-objective optimization problem, and the objective function is usually a weighted combination of the loss functions of each task, that is: ; where represents shared parameters, is a private parameter of task t, and represent task weights and regularization coefficients, respectively, and R is a model complexity constraint term. This parameterization design enables the model to capture cross-task common features through shared layers, while adapting to the differentiated needs of each task through task-specific layers.
[0045] The core of multi-task learning is parameter sharing strategy, loss function design and model architecture innovation. Parameter sharing strategy covers two paradigms of hard sharing and soft sharing. Hard sharing extracts shared features through the first few layers of a unified neural network, followed by task-specific output layers. Soft sharing achieves implicit feature interaction through parameterized latent variable models such as multi-task Gaussian processes or variational autoencoders, which has the advantage of avoiding negative transfer caused by task conflict. Loss function design focuses on dynamic adjustment mechanism of task weights, including uncertainty weighting and gradient normalization. The former optimizes the weights by minimizing the weighted variance of the uncertainty of each task, and the latter avoids the problem of gradient suppression of secondary tasks by gradient amplitude equalization. Model architecture innovation is reflected in heterogeneous network design and dynamic network routing, which enhances the efficiency of collaborative learning by enhancing the information interaction path between tasks.
[0046] The multi-task learning framework effectively promotes knowledge transfer between modulation type recognition and code sequence generation tasks by sharing time-frequency features of I / Q signals. The local time-frequency features extracted by the ResNet-18 branch provide context information related to modulation type for the Transformer branch, while the long-range dependency modeling capability of the Transformer benefits the analysis of temporal correlation of the modulation classification task. This bidirectional feature enhancement mechanism enables the model to exhibit stronger robustness when facing complex electromagnetic interference.
[0047] S3: Train the multi-task learning model using the training set and the validation set to obtain a trained multi-task learning model; The model training adopts a dynamic batch construction strategy, an optimizer setting, and a multi-dimensional stability enhancement mechanism. In view of the variable length characteristic of the I / Q signal sequence, a dynamic batch construction strategy is implemented in the data loading stage. In the preprocessing stage, the original signal is clustered by width, and samples with the same signal width (i.e., the same token division length) are allocated to the same batch to ensure the consistency of the sequence length in each batch, thereby avoiding the additional calculation overhead and position encoding offset problem caused by dynamic padding. This strategy can improve the GPU memory utilization while maintaining the global modeling effectiveness of the self-attention mechanism of the Transformer branch.
[0048] The model training adopts AdamW optimizer as the core optimization algorithm, and the parameter setting follows the characteristics of the mixed task. The initial learning rate is set to 3x10 −4 , the weight decay coefficient is configured to 1x10 −4 , and the beta parameter is set to (0.9, 0.999) to balance the momentum update and adaptive learning rate correction. To cope with the gradient scale difference of multi-task learning, a cosine annealing learning rate scheduler is introduced, and the cycle length is set to 50 epochs. The learning rate is automatically reduced to 10% of the initial value in the later training stage, which effectively alleviates the oscillation phenomenon of the modulation classification and sequence generation tasks in the parameter space update.
[0049] Multi-dimensional stability enhancement mechanism: First, for noise robustness training in complex electromagnetic environments, an adversarial data augmentation strategy is implemented to add Gaussian white noise and transient interference pulses at the input end to force the model to learn noise invariant features. Second, the gradient clipping technique is used to limit the risk of gradient explosion, and the maximum norm constraint is set to 1.0. The joint gradient of the cross-modal feature interaction gating unit is normalized by the clip_grad_norm function. Finally, an early stopping mechanism and model checkpoint saving strategy are designed. When the weighted total loss on the validation set does not decrease for 15 consecutive epochs, the training is terminated, and the optimal parameter state is preserved.
[0050] S4: input the complex electromagnetic signal to be tested into the trained multi-task learning model to obtain the modulation type recognition result and the symbol sequence prediction result.
[0051] Embodiment 1 In this embodiment, classification accuracy, bit error rate, and cosine similarity are used as evaluation indicators.
[0052] Classification Accuracy is a commonly used evaluation metric in machine learning and statistical classification tasks, used to measure the degree of consistency between the model's prediction results and the true labels. Its core definition is the ratio of correctly predicted samples to the total number of samples. Assuming there is a dataset containing N samples, the model's prediction results for each sample are , and the true labels are , then the accuracy is represented as: ; where is an indicator function that takes a value of 1 when the condition is met, otherwise 0.
[0053] Bit Error Rate (BER) is a core indicator in the field of communication systems and digital transmission, used to measure the reliability of data transmission, representing the proportion of error bits detected at the receiving end to the total number of transmitted bits. The bit error rate is defined as: ; For example, if 1000 bits of data are transmitted, and 100 bits are in error, then (i.e. 1 bit in error per billion bits).
[0054] Cosine similarity measures the directional similarity of two vectors by calculating the cosine of the angle between them in space, represented as: ; where is the dot product of the vectors, and are the lengths of the vectors. The value range is [-1, 1], the closer to 1 indicates the more consistent direction, the closer to 0 indicates orthogonal, and the closer to -1 indicates opposite direction.
[0055] In the task of modulation signal recognition and symbol sequence prediction, this embodiment compares the performance of three deep learning architectures in modulation type classification and symbol sequence prediction, and the results are shown in Table 2: Table 2 ;
[0056] The experimental results show that the Resnet18-Transformer hybrid model combining convolutional neural network and self-attention mechanism achieves a classification accuracy of 56.51% in the modulation type classification task, which is 2.71 percentage points higher than the Resnet18 benchmark model with pure convolutional architecture, and also shows better symbol sequence prediction capability, with a bit error rate 3.53% lower than the Transformer benchmark model (a relative decrease of 6.87%), and a cosine similarity index of 0.072 (a relative increase of 11.57%). It is worth noting that although the pure Transformer architecture performs well in the sequence similarity measurement of symbol prediction (cosine similarity of 0.6225), its classification accuracy decreases by 24.55 percentage points compared with the hybrid model, reflecting the limitations of single attention mechanism in local feature extraction. The traditional Resnet18 model performs second in the classification task, with a bit error rate 21.83% higher than the hybrid model, indicating that the pure convolutional architecture has shortcomings in time series dependency modeling. The above comparisons verify the advantages of the hybrid architecture in the joint task, which effectively captures local spectral features through convolutional layers and establishes long-time dependency relationships through self-attention mechanisms, achieving a synergistic improvement in classification accuracy and sequence prediction performance.
[0057] Embodiment 2 In the ablation experiment of the multi-task joint optimization framework, this embodiment analyzes the synergistic mechanism in the modulation type classification and symbol sequence prediction tasks through a systematic module removal strategy. The experiment design includes three key configurations: simultaneous execution of dual tasks, only focusing on the classification task, and only focusing on the sequence prediction task. The experimental results are shown in Table 3.
[0058] Table 3 Experimental results ;
[0059] The model shows significant advantages in the multi-task cooperative training mode, with a modulation type classification accuracy of 56.51%, which is 2.71 percentage points higher than only performing the classification task; the symbol sequence prediction bit error rate is as low as 0.4960, and the cosine similarity is 0.6943, which is 3.53% and 11.57% lower than and higher than the only sequence prediction task optimization, respectively.
[0060] This experiment verifies the superiority of the multi-task learning framework from the task coupling perspective. It achieves feature reuse through parameter sharing mechanism, maintains performance advantage in classification task compared with single task model, and achieves great similarity improvement in prediction task. It is worth noting that the single task framework does not reach the joint optimization level of the multi-task model in the dedicated task, indicating that the synergistic effect of multi-task not only does not cause negative transfer between tasks, but also enhances the generalization representation ability of the model through feature interaction.
[0061] Example 3 The preprocessed signal data is input into the built multi-task model, and a line chart of the loss function value changing with the training round number is drawn as shown in Figure 9 The loss function value shows a clear trend of change with the increase of the training round number. In the early stage of training, the loss function value decreases rapidly, and the model is rapidly learning the basic features and patterns in the data, and the prediction error is significantly reduced. The rapid decline in this stage reflects the preliminary adaptation of the model to the training data and the formation of feature extraction capability.
[0062] As the training goes deeper, the decline rate of the loss function value gradually slows down and starts to fluctuate. This fluctuation is due to the model starting to learn more complex features, which may be affected by noise, uneven data distribution and other factors. In the framework of multi-task learning, the model needs to optimize both modulation classification and symbol sequence generation tasks, increasing the complexity of training. The feature fragmentation and task conflict between different tasks may cause the loss function value to fluctuate. In addition, the adjustment of the dynamic EMA weight allocation mechanism during training may also affect the loss function value, so that the model dynamically allocates resources between different tasks to seek the overall optimum.
[0063] In the later stage of training, the loss function value tends to be stable and gradually converges to a smaller value. This indicates that the model has learned most of the useful information in the data and can achieve good performance on the validation set or test set. The achievement of the convergence state means that the model has reached a relatively stable state in the training process and can better generalize to new data.
[0064] The change of the loss function value is closely related to the model architecture and the dynamic weight allocation mechanism. The joint demodulation model based on multi-task learning optimizes both modulation classification and symbol sequence generation tasks through the design of the compound loss function. This design makes the model need to balance the learning progress between different tasks during training, thereby increasing the difficulty of training. However, the introduction of the dynamic EMA weight allocation mechanism enables the model to adaptively adjust the loss weights of different tasks, ensuring that the model always focuses on the most optimized task during training. This mechanism helps the model better balance the feature fragmentation and task conflict problems in multi-task learning, thereby improving the overall performance.
[0065] The confusion matrix of the classification results is drawn as shown in Figure 10It can be seen from the figure that the algorithm of the present research has high recognition accuracy for BPSK and MSK, but is easy to recognize other types as BPSK, indicating that the algorithm cannot well distinguish other modulation types from BPSK. PSK in BPSK represents the use of phase shift keying, which is a form of phase modulation, used to express a series of discrete states, BPSK corresponds to 2 states, QPSK corresponds to 4 states, and 8PSK corresponds to 8 states, so BPSK, QPSK and 8PSK are closely related, which leads to the model being unable to distinguish well. Therefore, the model needs to further study the differentiation of different PSK modulation methods to improve the accuracy.
[0066] Therefore, the complex electromagnetic signal demodulation method based on multi-task learning adopts the above-mentioned ResNet-18 network to extract the local time sequence features and modulation type discriminative patterns of the signal, uses the Transformer architecture to capture the long-range spectral correlation and modulation parameter evolution law, and forms a dual-task branch of modulation classification and code sequence generation. In solving the common feature fragmentation problem in multi-task learning, the cross-branch feature interaction gate unit is designed innovatively to dynamically adjust the fusion ratio of ResNet deep features and Transformer sequence features, realizing the implicit collaborative optimization of modulation feature learning and sequence generation tasks. In the dynamic weight distribution mechanism, an exponential moving average (EMA)-based dynamic weight balancing system is proposed to realize the adaptive adjustment of task weights by fusing historical loss statistics. Compared with static weight setting, the EMA mechanism improves the sensitivity of task weight adjustment of the model in the dynamic electromagnetic interference scene, effectively inhibiting the gradient cancellation phenomenon caused by task conflict. For the variable-length code sequence generation task, a dynamic sequence alignment algorithm and bit error rate integration strategy are designed to directly associate the traditional cross-entropy loss and bit-level performance indicators, so that the model can optimize the sequence generation probability distribution while realizing the explicit constraint of the bit error rate. The multi-task learning model is designed to improve the accuracy of modulation type recognition through the collaborative optimization of dual tasks, and to improve the accuracy of code sequence generation.
[0067] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application rather than limit them, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can still be modified or replaced by equivalents, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.
Claims
1. A method for demodulating complex electromagnetic signals based on multi-task learning, characterized in that: Includes the following steps: S1: Obtain signal datasets under different modulation schemes, and divide the signal datasets under different modulation schemes into training sets and validation sets; S2: Employs a hybrid architecture of ResNet-18 and Transformer as the backbone network, combined with a dynamic weight allocation mechanism, to construct a multi-task learning model; specifically including: The backbone network adopts a hybrid architecture of ResNet-18 and Transformer. The ResNet-18 branch is used to extract local temporal features and modulation type discriminative patterns, while the Transformer branch captures long-term spectral correlation and modulation parameter evolution through a multi-head self-attention mechanism. A dynamic weight allocation mechanism is used to adaptively adjust the loss weights for modulation type identification and symbol sequence prediction tasks. Set up a cross-branch feature interaction gating unit to dynamically adjust the fusion ratio of ResNet-18 deep features and Transformer sequence features; The dynamic weight allocation mechanism specifically includes a dynamic EMA weight balancing mechanism, sequence alignment and bit error rate integration, and a composite loss quantization system; The dynamic EMA weight balancing mechanism is used to dynamically adjust the loss weight of each task through the exponential moving average EMA algorithm. Sequence alignment and bit error rate integration is used to leverage dynamic sequence alignment algorithms, achieving length matching between the predicted and target sequences through permute operations and zero padding, and embedding the bit error rate as a soft constraint into the loss function; The composite loss quantization system is used to weight the modulation type classification loss, sequence generation probability distribution loss, and bit error rate loss scaled by temperature coefficient to form the total loss function; S3: Train the multi-task learning model using the training set and validation set to obtain a trained multi-task learning model; S4: Input the complex electromagnetic signal to be tested into the trained multi-task learning model to obtain the modulation type identification result and the symbol sequence prediction result.
2. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 1, characterized in that: S1 also includes preprocessing the signal data in the dataset. The preprocessing includes: removing noise and clutter from the signal data using a mid-pass filter method, and visualizing the denoising results using a constellation diagram.
3. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 1, characterized in that: The Exponential Moving Average (EMA) algorithm is expressed as follows: ; in, As the attenuation factor, Let be the loss weight at time step t. The loss weight at time step t-1, This is the original loss value.
4. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 3, characterized in that: The bit error rate is embedded as a soft constraint into the loss function and expressed as: ; in, For indicator functions, For the number of effective bits, For the target sequence, This is the predicted sequence.
5. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 4, characterized in that: The overall loss function adopts a hierarchical multi-task joint modeling framework, decoupling modulation type classification and sequence generation probability distribution into two parallel sub-tasks; the main path uses cross-entropy loss to quantify modulation type classification error and adaptively adjusts the loss weight through a dynamic EMA mechanism; the auxiliary path optimizes the sequence generation probability distribution through cross-entropy loss and monitors the model's performance through bit error rate calculation.
6. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 5, characterized in that, The total loss function is expressed as: ; in, Weights for modulation type classification loss, Loss weights are used to generate a probability distribution for the sequence. The bit error rate loss weights are scaled by a temperature coefficient. For modulation type classification loss, The loss is used to generate a probability distribution for the sequence. This is the bit error rate loss after temperature coefficient scaling.
7. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 6, characterized in that: The model training employs a dynamic batch construction strategy, optimizer settings, and a multi-dimensional stability enhancement mechanism; The dynamic batch construction strategy is used to ensure that the sequence length is consistent within each batch by combining width clustering and dynamic batch construction. The optimizer settings are used to employ the AdamW optimizer algorithm, combined with a cosine annealing learning rate scheduler, to balance momentum updates and adaptive learning rate corrections.
8. The method for demodulating complex electromagnetic signals based on multi-task learning as described in claim 7, characterized in that: Multi-dimensional stability enhancement mechanisms include adversarial data augmentation strategies, gradient pruning techniques, and model checkpoint preservation strategies; Adversarial data augmentation strategies add Gaussian white noise and transient interference pulses to the input, enabling the model to learn noise-invariant characteristics. Gradient clipping is used to limit the risk of gradient explosion. The maximum norm constraint is set to 1.0, and the joint gradient of the cross-branch feature interaction gate unit is normalized by the clip_grad_norm function. The early stopping mechanism and model checkpoint saving strategy are used to terminate training and preserve the optimal parameter state when the weighted total loss on the validation set has not decreased for 15 consecutive epochs.
Citation Information
Patent Citations
Interference signal modulation recognition method for communication carrier monitoring system
CN112347871A
Signal modulation identification method for coupling depth feature and physical feature and related device
CN117544461A
Distributed system navigation interference intelligent identification method, system and device under non-Gaussian noise and medium
CN118033681A
Radio modulation signal identification method and system based on hybrid neural network
CN120050146A
Method for improving evaluation accuracy of various indexes of non-neoplastic diseases of stomach in histopathological image based on multi-task learning model
CN120411019A
Cited By
Electromagnetic task processing method, related device, equipment and medium
CN121051711A