Communication signal modulation identification method based on multi-modal characteristic orthogonalization adaptive fusion

By extracting multimodal features and building an adaptive fusion strategy, the problems of single modal feature limitations and modal feature redundancy in the prior art are solved, and the high accuracy and robustness of communication signal modulation recognition are achieved.

CN120075004AActive Publication Date: 2025-05-30SICHUAN JIUZHOU SOFTWARE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510140073.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The existing communication signal modulation identification methods have the limitations of single modular features, redundancy and interference between modal features, the lack of efficient feature fusion strategies, and the single model optimization target, resulting in limited classification performance and insufficient generalization capabilities.

Method used

By extracting multimodal features of IQ time domain signals and amplitude/phase information, and constructing an adaptive fusion strategy of orthogonal constraints, an adaptive feature fusion module with dynamic weight allocation is adopted, and feature extraction and classification is combined with a deep learning model.

Benefits of technology

It effectively improves the accuracy, robustness and generalization capabilities of modulation recognition, enhances the adaptability of the model, and can provide reliable modulation recognition solutions in complex communication environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075004A_ABST
    Figure CN120075004A_ABST
Patent Text Reader

Abstract

The invention discloses a communication signal modulation identification method based on multi-modal feature orthogonalization adaptive fusion, which comprises the following steps: acquiring communication signal data of various modulation types, sampling rates and bandwidths, and preprocessing to generate a training set; constructing a communication signal modulation recognition model comprising a feature extractor for different modal features, an adaptive feature fusion module and a classifier; inputting the training set into a communication signal modulation recognition model, and completing model training by utilizing back propagation iteration updating; and using the trained communication signal modulation identification model for the modulation type identification of the target communication signal data. According to the method, signal features are comprehensively mined, and the integrity and classification performance of feature expression are improved. Meanwhile, redundant interference between modal features is reduced through an orthogonalization loss function, the generalization ability of the model is enhanced, high adaptability is shown in a complex communication environment, and a reliable modulation identification scheme can be provided for scenes with low signal-to-noise ratio, multipath interference and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of modulation recognition, and particularly to a communication signal modulation recognition method based on orthogonal adaptive fusion of multi-modal features. Background Art

[0002] With the rapid development of wireless communication technology and the continuous expansion of application scenarios, the types of communication signals are increasing day by day, and the modulation methods of signals have become more complex. In modern communication systems, signal modulation recognition, as an important link in signal processing, is widely used in radio monitoring, spectrum management, cognitive radio, electronic countermeasure and other fields. Its purpose is to determine the modulation method of a signal by analyzing and classifying the received signal.

[0003] Existing communication signal modulation recognition methods are mainly divided into two categories: likelihood function-based methods and feature extraction-based methods. Likelihood function-based methods classify by statistically analyzing the characteristic parameters of signals and constructing a likelihood function. They are theoretically optimal, but their performance highly depends on the accuracy of signal parameters, and the computational complexity is relatively high, making them unsuitable for real-time scenarios or large-scale data processing. Feature extraction-based methods: extract specific statistical features (such as amplitude, phase, power spectral density, etc.) from signals and use traditional classification algorithms (such as support vector machines, decision trees, etc.) for classification. Although this method has high flexibility, its classification accuracy highly depends on the quality of manually designed features, and it has poor robustness to low signal-to-noise ratio environments and complex channel conditions.

[0004] In recent years, deep learning technology has been widely used in the field of signal processing. Using deep learning models, especially convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), can automatically extract multi-level features from communication signals, significantly improving the performance of modulation recognition. However, the existing deep learning modulation recognition methods still have the following main problems:

[0005] 1. Limitations of single-modal features: Many studies only use IQ time-domain signals as input, ignoring important information in other modalities (such as amplitude, phase, etc.) of communication signals, resulting in insufficient integrity of feature representation.

[0006] 2. Redundancy and interference between modal features: In the application of multi-modal features, there may be problems of redundancy or overly strong correlation between different modal features, resulting in limited classification performance of the model.

[0007] 3. Lack of efficient feature fusion strategies: Existing methods usually adopt simple concatenation or weighted methods with fixed weights when fusing multi-modal features, and fail to dynamically adjust weights according to the importance of modal features, restricting the effectiveness of the fused features.

[0008] 4. Single model optimization objective: Many models are trained only based on classification loss, ignoring the constraints on the distribution of modal features, resulting in insufficient generalization ability of the models. Summary of the Invention

[0009] In view of the deficiencies in the above-mentioned existing technologies, the present invention aims to provide a communication signal modulation recognition method based on orthogonalization and adaptive fusion of multi-modal features. By extracting multi-modal features of IQ time-domain signals and amplitude / phase information from communication signals and constructing an adaptive fusion strategy with orthogonalization constraints, the present invention can effectively improve the accuracy, robustness, and generalization ability of modulation recognition, providing an efficient and reliable solution for modulation recognition in complex communication environments.

[0010] Advantages

[0011] By introducing multi-modal feature orthogonalization and an adaptive fusion strategy, the present invention effectively improves the accuracy and robustness of communication signal modulation recognition. Using feature extractors designed for IQ time-domain signals and amplitude / phase information, combined with an adaptive feature fusion module with dynamic weight allocation, the present invention comprehensively excavates signal features, improves the integrity of feature expression and classification performance. At the same time, by reducing the redundant interference between modal features through an orthogonalization loss function, the generalization ability of the model is enhanced, and it shows strong adaptability in complex communication environments, and can provide a reliable modulation recognition solution for scenarios such as low signal-to-noise ratio and multipath interference. Description of the Drawings

[0012] Figure 1 Schematic diagram of the flow of a communication signal modulation recognition method based on orthogonalization and adaptive fusion of multi-modal features provided in a preferred embodiment of the present invention;

[0013] Figure 2 Schematic diagram of the structure of the first feature extractor provided in another preferred embodiment of the present invention;

[0014] Figure 3 Schematic diagram of the structure of the second feature extractor provided in another preferred embodiment of the present invention. Detailed Embodiment

[0015] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings. In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0016] As Figure 1 shown, this embodiment discloses a communication signal modulation recognition method based on orthogonalized adaptive fusion of multi-modal features, including the steps:

[0017] S1. Obtain communication signal data of multiple modulation types, sampling rates, and bandwidths, including but not limited to digital modulation types 4ASK, 8ASK, BPSK, QPSK, OQPSK, PI4QPSK, 8QAM, 16QAM, GMSK and analog modulation types FM, AM-SSB-WC, AM-SSB-SC, AM-DSB-WC, AM-DSB-SC obtained through a receiver, and generate a training set after preprocessing.

[0018] Those skilled in the art can understand that the preprocessing refers to performing a series of processes on the obtained original communication signal data to generate high-quality training samples and input data to meet the requirements of subsequent feature extraction and classification tasks. The goal of preprocessing is to clean the noise in the signal, extract key information, and standardize the data format, thereby improving the training and classification performance of the model. In some preferred embodiments, a preferred preprocessing method is provided, specifically including:

[0019] 1. Filtering: Remove the noise and interference in the communication signal through a filter (such as a low-pass filter), retain the main components of the signal, enhance the clarity and quality of the signal, and reduce the impact of noise on model training.

[0020] 2. Segmentation: Divide the filtered communication signal into signal segments of equal length, and each signal segment is used as a sample for subsequent feature extraction and batch processing to ensure the unity of the sample size.

[0021] 3. Amplitude and phase calculation: Calculate the amplitude and phase for each signal segment, extract the important feature information of the signal, and provide support for multi-modal feature input. In some preferred embodiments, the calculation methods of amplitude a(t) and phase φ(t) are given, specifically including:

[0022]

[0023] where y I (t), y Q (t) represent the I channel and Q channel of the IQ time-domain signal sample.

[0024] 4. Normalization: Normalize the signal samples using the zero-mean normalization formula to ensure that the numerical ranges of all samples are the same, and avoid instability in the training process due to excessive differences in data amplitudes.

[0025] 5. Sample Division: According to the modulation type and channel parameters of the communication signal, the normalized samples are divided into multiple sub-training sets and combined into a complete training set to ensure the diversity of the training set and improve the generalization ability of the model. It should be understood that the training set includes a training sample set, a validation sample set, and a test sample set divided according to different sample ratios. The specific division ratio can be determined by those skilled in the art according to actual needs, and the present invention does not make further limitations.

[0026] Furthermore, the method for obtaining multi-modal communication signal samples includes: filtering the communication signal data and then segmenting it into signal segments of equal length. Each signal segment is used as a sample, and then the amplitude and / or phase of each signal segment are calculated respectively as another sample of each signal segment. Through the above process, after preprocessing each segment of the original signal, samples of two modalities are obtained, namely:

[0027] Time-domain signal samples, which retain the original waveform characteristics of the communication signal and can extract its time-domain features through convolution operations; amplitude / phase information samples, which provide the geometric characteristics (intensity and angle) of the signal and can capture its dynamic changes through a time-series model (such as LSTM).

[0028] Such multi-modal features provide more comprehensive input information for the model, which helps to improve the accuracy and robustness of signal modulation recognition.

[0029] S2. Construct a communication signal modulation recognition model including a feature extractor, an adaptive feature fusion module, and a classifier for different modal features.

[0030] Among them, the role of the feature extractor is to extract high-quality features from different modalities (such as IQ time-domain signals, amplitude / phase information) and provide input for subsequent fusion and classification. Since the structural characteristics and properties of each modal feature are different, a dedicated feature extractor is required.

[0031] In some preferred embodiments, the feature extractor includes a first feature extractor for IQ time-domain signals and a second feature extractor for amplitude and / or phase information.

[0032] The first feature extractor is used to extract high-quality features from the original IQ time-domain signal and generate a feature vector of a fixed size as the input for subsequent feature fusion and classification. As Figure 2 shown, its specific structure includes:

[0033] An input layer set to a size of (2, 1024, 1); setting the input layer size in this way is to process the time-domain signals of two channels (I and Q components), taking advantage of the two-dimensional characteristics of IQ signals (I path and Q path). This structural design enables the network to directly extract information from the original IQ signals without the need for other more complex preprocessing operations.

[0034] A number of two-dimensional convolutional layers connected in sequence and extended by patching; the two-dimensional convolutional layers are used to extract local features of the signal and extend the boundaries in combination with the patching method to retain edge information.

[0035] A residual connection layer for fusing shallow outputs and deep features; the residual connection layer adds the shallow output and deep features through weighted summation to retain the complementary information of shallow and deep features. This structure can alleviate the problem of gradient disappearance and improve the feature expression ability at the same time.

[0036] A global average pooling layer for integrating global information; the global average pooling layer is used to map high-dimensional features to a fixed dimension, reduce the number of parameters, and improve the generalization ability of the model.

[0037] It should be understood that the specific setting parameters of the above structure (such as the size and number of convolutional kernels, and the design of activation functions) and connection relationships can be flexibly set by those skilled in the art according to needs, and in order to achieve the complete function, it should also include other necessary structural layers, such as a batch normalization layer for stabilizing the training process, a flattening layer for outputting feature vectors, etc.

[0038] The second feature extractor is used to extract high-quality temporal features from the amplitude and phase features calculated from each segment of communication signal, generate a feature vector of a fixed size, and provide input for subsequent feature fusion and classification. As Figure 3 shown, its specific structure includes:

[0039] An input layer with an input size of (1024, 2); it represents 1024 time steps of each segment of signal, and each time step contains 2 values (amplitude and phase). This layer can accept amplitude and phase information simultaneously, avoiding feature loss that may occur when processing them separately.

[0040] A two-layer long short-term memory network for extracting temporal features of amplitude and / or phase signals. Specifically, the first-layer long short-term memory network is used to capture the underlying time series features of amplitude / phase information, and the second-layer long short-term memory network is used to further extract high-level dynamic features. The stacking of the two-layer long short-term memory network can effectively learn the complex correlations and long-term dependencies of amplitude / phase features in the time series, increasing the non-linear expression ability of the model and being suitable for feature extraction of complex signals.

[0041] The fully-connected output layer takes the output of the last time step of the double-layer long short-term memory network as input, and is used to reduce the dimensionality of the extracted dynamic features into a feature vector of a fixed size, facilitating subsequent modality fusion.

[0042] The adaptive feature fusion module is used to integrate multi-modal features extracted from different feature extractors (the first feature extractor and the second feature extractor) into a unified feature representation, providing an efficient input for the classifier. Those skilled in the art can know that the importance of different modality features (such as IQ signal features, amplitude / phase features) for the task may vary with data and scenarios, and there may be a certain redundancy between modality features. In complex scenarios such as low signal-to-noise ratio and multi-interference, the performance of different modalities may vary. To solve these problems, in some preferred embodiments, it is considered to construct a projection module for multi-modal features, specifically including an IQ time-domain signal feature projection module and an amplitude and / or phase information feature projection module that are identical in structure and arranged in parallel. Its structure includes:

[0043] An input layer;

[0044] A double-layer fully-connected layer. The first fully-connected layer is used to initially extract fused features, and the second fully-connected layer is used to output the weights of each modality feature. Through layer-by-layer operations, the double-layer fully-connected layer performs non-linear transformations on the input modality features, gradually capturing complex feature distributions. In order to uniformly map all modality features to a target feature space of a fixed size (such as 64 dimensions), ensuring that different modality features have consistent dimensions and distributions before fusion, the input features are processed by the second fully-connected layer for dimensionality reduction, and a feature vector of a fixed size (such as 64 dimensions) is output, thus completing the mapping.

[0045] Furthermore, the adaptive feature fusion module further includes a fusion module, which is used to calculate weights after concatenating the outputs of each feature projection module, multiply the features of each modality by the corresponding weights and then concatenate them, and output the adaptive fusion features. The specific structure includes:

[0046] An input layer with the concatenated multi-modal features as input;

[0047] A second double-layer fully-connected layer for initially extracting fused features and outputting the weights of each modality feature;

[0048] A Softmax layer that normalizes the output of the second double-layer fully-connected layer to generate the weight distribution of the two modality features. The weight magnitude reflects the importance of each modality feature.

[0049] The fusion module weights different modality features, multiplies the modality features by the corresponding weights and then concatenates them to obtain the adaptive fusion features.

[0050] The classifier is responsible for classifying the fused features processed by the feature extractor and the adaptive feature fusion module, so as to output the specific modulation type of the communication signal (such as BPSK, QPSK, 16QAM, etc.). Its specific structure can be designed by those skilled in the art according to the prior art and actual needs, and the present invention does not make further limitations.

[0051] S3. Input the training set into the communication signal modulation recognition model, and use backpropagation to iteratively update and complete the model training.

[0052] Backpropagation (abbreviated as BP) is an optimization algorithm for training neural networks. It calculates the gradients of the loss function with respect to the model parameters and uses these gradients to update the parameters in the network (such as weights and biases), so that the model can gradually approach the optimal solution. Traditional backpropagation usually only uses a single loss function (such as cross-entropy loss or mean squared error) to guide the optimization of model parameters. In some preferred embodiments of the present invention, a joint optimization strategy is considered, combining the classification loss and the orthogonality loss together to jointly optimize the model parameters:

[0053] Specifically, a conventional classification loss is considered to optimize the overall model. Specifically, the cross-entropy loss is used to construct the loss function:

[0054] where n is the number of modulation categories, y i is the true label, is the label predicted by the network.

[0055] Aiming at the possible redundancy problem between multi-modal features (IQ signal features and amplitude / phase features), an orthogonality constraint is introduced to minimize the correlation between modal features. Specifically, the orthogonal loss function L OG is used to perform backpropagation iterative optimization on the feature extractor:

[0056] L OG =λ 1 |f IQ ·f 幅 / 相 |+λ 2 (f IQ ·f 幅 / 相 ) 2 ; where f IQ and f 幅 / 相 are the output feature vectors of the first feature extractor and the second feature extractor respectively, and λ 1 , λ 2 are preset weights for controlling different regularization losses respectively.

[0057] In the backpropagation optimization algorithm, the learning rate (η) is an important hyperparameter that controls the step size of parameter updates during model training. It directly affects the training speed and effect:

[0058] If the learning rate is too large, the model may oscillate or fail to converge during optimization, resulting in unstable training; if the learning rate is too small, the model converges slowly and may even get stuck in a local optimum, unable to achieve the best performance.

[0059] During different stages of model training, the optimization requirements are different. In the initial stage, the parameter values are randomly initialized, and the model requires a relatively large learning rate to quickly approach the vicinity of the optimal solution. In the middle and late stages, the model gradually approaches the optimal solution. At this time, a smaller learning rate is needed to more finely adjust the parameters and prevent oscillations during the training process. Therefore, the traditional static learning rate (fixed value) cannot adapt to the dynamic needs of the model throughout the training process. In some preferred embodiments, a dynamic learning rate adjustment strategy is considered, specifically including:

[0060] When the decrease in the loss function value does not exceed a preset change threshold within the first preset period, reduce the subsequent learning rate to 1 / 3 of the current learning rate. It should be noted that during the training process, after a certain number of training rounds (referred to as the "first preset period", such as 5 - 10 periods), monitor whether the loss function value continues to decrease. If the loss function value does not decrease significantly within this preset period, it indicates that the current learning rate may be too large, resulting in the inability to more finely adjust the model parameters. When the trigger condition is met, dynamically adjust the subsequent learning rate to 1 / 3 of the current learning rate and continue training the model with the new learning rate, so that the parameters can be adjusted with a smaller step size, thereby more finely optimizing the model performance.

[0061] The setting of the first preset period can be designed by those skilled in the art according to actual needs. If the period is too short, the change trend of the loss function may be affected by accidental factors and may not be sufficient to reflect the true state of the model; if the period is too long, it will delay the timing of learning rate adjustment and affect the training efficiency.

[0062] The preset change threshold can be selected as an absolute change or a relative change. For example, the value of the loss function does not decrease by more than a certain threshold (such as ∈ = 0.001) within the period, or the change amplitude of the loss function is less than a certain percentage (such as 1%).

[0063] S4. Use the trained communication signal modulation recognition model to identify the modulation type of the target communication signal data.

[0064] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A communication signal modulation recognition method based on orthogonalization and adaptive fusion of multimodal features, characterized in that: Includes steps: Acquire communication signal data of various modulation types, sampling rates and bandwidths, and generate training sets after preprocessing; Construct a communication signal modulation recognition model including a feature extractor for different modal features, an adaptive feature fusion module and a classifier; Inputting the training set into the communication signal modulation recognition model, and completing the model training by back propagation iterative update; The trained communication signal modulation recognition model is used to identify the modulation type of the target communication signal data.

2. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as claimed in claim 1, characterized in that: The preprocessing method includes: filtering the communication signal data and dividing it into signal segments of equal length, taking each signal segment as a sample, and then calculating the amplitude and / or phase of each signal segment respectively as another sample of each signal segment.

3. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as claimed in claim 1, characterized in that: The feature extractor comprises a first feature extractor for an IQ time domain signal, comprising: Set the input layer to a size of (2,1024,1); Several two-dimensional convolutional layers connected in sequence and expanded in a patching manner; Residual connection layer for fusing shallow layer output with deep layer features; Global average pooling layer for integrating global information.

4. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as claimed in claim 1, characterized in that: The feature extractor includes a second feature extractor for amplitude and / or phase information, and the second feature extractor includes a double-layer long short-term memory network for extracting timing features of amplitude and / or phase signals.

5. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as claimed in claim 1, characterized in that: The adaptive feature fusion module includes: The feature projection module of each modality is used to map the high-dimensional features of each modality to the target feature space of the same dimension through a fully connected network; The fusion module is used to calculate the weights after splicing the outputs of each feature projection module, multiply the features of each modality by the corresponding weights and then splice them, and the output is the adaptive fusion feature.

6. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as claimed in claim 5, characterized in that: The feature projection module includes: an IQ time domain signal feature projection module and an amplitude and / or phase information feature projection module, which are provided with an input layer and a double-layer fully connected layer with the same input and output size and connected in sequence, and the double-layer fully connected layer is used to gradually extract the nonlinear characteristics of the modal features and map them to the target feature space.

7. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as described in claims 3 and 4 is characterized in that: The method for completing model training by back-propagation iterative updating includes: Use the orthogonal loss function L OG Perform back-propagation iterative optimization on the feature extractor: L OG =λ1|f IQ ·f 幅 / 相 |+λ2(f IQ ·f 幅 / 相 ) 2 ; Among them, f IQ and f 幅 / 相 are the output feature vectors of the first feature extractor and the second feature extractor respectively, and λ1 and λ2 are the preset weights for controlling different regularization losses respectively.

8. The communication signal modulation recognition method based on multi-modal feature orthogonalization adaptive fusion as claimed in claim 1, characterized in that: The method for completing model training using back-propagation iterative updating includes: when the decrease in the loss function value within a first preset period does not exceed a preset change threshold, reducing the subsequent learning rate to 1 / 3 of the current learning rate.

Citation Information

Patent Citations

  • Multi-mode ultrasonic image classification method and breast cancer diagnosis device

    CN110930367A

  • Automatic identification method and system based on double-channel convolution long-short-term neural network

    CN112241724A

  • Modulation identification method based on joint multi-modal information and domain adversarial neural network

    CN115392326A

  • Automatic modulation identification method based on multi-modal fusion and multi-task optimization

    CN118643460A

  • Communication signal modulation identification method based on improved ViT model and computer device

    CN119363533A