A sparse cross-modal communication radiation source identification method

By using the MHACNN network and a sparse cross-modal fusion method with joint sparse representations, the problem of underutilization of modal interaction characteristics in existing technologies is solved, the recognition accuracy of communication radiation sources is improved, and stronger feature representation capabilities are achieved.

CN118861527BActive Publication Date: 2025-12-05NAT TIME SERVICE CENT CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410996742.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-12-05
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

Existing single-modal and multi-modal fusion methods fail to fully utilize the interaction characteristics of the two modalities in communication radiation source identification, resulting in limited improvement in identification accuracy. In particular, it is difficult to extract useful modulation domain features in non-cooperative scenarios, and the modal independence acquired by multiple sensors is relatively strong, making it difficult to obtain low-redundancy feature representations.

Method used

The MHACNN network is used to extract features from the original time series and time-frequency map of the radiation source. Joint sparse representation is used to obtain modal features with stronger representation capabilities, and cross-modal fusion is performed. Redundant and noise information is eliminated by sparse cross-modal fusion method, and the feature learning process of one modality is used to guide the feature learning process of another modality.

Benefits of technology

It improves the accuracy of identifying communication radiation sources, and by utilizing the richness and modal interactivity of time-frequency information, it obtains stronger feature representation capabilities and enhances the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861527B_ABST
    Figure CN118861527B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on sparse cross-modal communication radiation source identification method, by signal pre-processing module, feature extraction module, cross-modal module and identification module composition, the signal pre-processing module inside includes the collection of radiation source signal, to radiation source signal is sliced, time series and time-frequency diagram two kinds of modal are obtained after slicing, time series and time-frequency diagram are respectively normalized, label radiation source class label and construct modal dataset, construct modal dataset including training set, verification set, test set.This method is by using MHACNN network to the original time series and time-frequency diagram two kinds of input modal of radiation source are carried out feature extraction, then utilize joint sparse representation to obtain the modal feature of stronger representation ability and carry out cross-modal fusion to improve the identification accuracy for radiation source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication signal identification technology, specifically relating to a method for identifying radiation sources based on sparse cross-modal communication. Background Technology

[0002] With the rapid development of deep learning, existing single-modal feature fusion methods (original signal time series or time-frequency graph of the signal) have achieved good recognition results on various types of communication radiation sources (e.g., ADSB, Iridium, LoRa, etc.). However, due to the lack of diversity, using only a single modality for feature fusion greatly limits further improvement in recognition accuracy. Meanwhile, existing multimodal fusion methods are mostly used to perform target detection and segmentation tasks from images obtained from multiple sensors (e.g., LiDAR, Camera, etc.), while the fusion problem of two highly correlated modalities (original time series and time-frequency graph) from a single sensor is rarely addressed. Therefore, we need a feature fusion algorithm to fully utilize the advantages of these two modalities, improve the representational power of features, and thus improve the recognition accuracy of communication radiation sources.

[0003] Existing multimodal fusion methods mainly focus on feature-level fusion and decision-level fusion. Feature-level fusion aggregates features from different modalities into a single feature set for classification and recognition, while decision-level fusion combines the decisions made by the respective classifiers of different modalities into a comprehensive decision.

[0004] For identifying communication radiation sources in non-cooperative scenarios, demodulating the communication signal to baseband is often impractical, making it difficult to extract useful modulation domain features. A more practical approach is to use feature fusion based on the original signal time series and various transform domain features (time-frequency plots, etc.). These two modalities are fused using a feature extractor based on a deep neural network. Depending on where the fusion occurs within the network, it can be categorized as "early fusion," "middle fusion," and "late fusion."

[0005] (Late Fusion) and other fusion methods fall short in that they fail to fully utilize the interaction characteristics of the two modalities, using one modality to guide the feature learning process of the other. Decision-based fusion algorithms also suffer from the need to dynamically adjust the contributions of the two modalities to classification based on actual data characteristics to determine a suitable weighting coefficient, which is difficult to determine in practical applications. Furthermore, existing fusion methods are mostly designed for heterogeneous modalities acquired from multiple sensors, where the modalities are relatively independent. However, time series and time-frequency maps used for non-cooperative communication signals exhibit strong correlation between the two modalities, requiring further acquisition of low-redundancy feature representations for radiation source classification. Summary of the Invention

[0006] The purpose of this invention is to solve the above-mentioned problems. This application proposes a radiation source identification method based on sparse cross-modal communication. By using the MHACNN network, features are extracted from the two input modalities of the radiation source: the original time series and the time-frequency graph. Then, joint sparse representation is used to obtain modal features with stronger representation capabilities and cross-modal fusion is performed to improve the accuracy of radiation source identification.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a radiation source identification method based on sparse cross-modal communication, comprising a signal preprocessing module, a feature extraction module, a cross-modal module, and an identification module. The signal preprocessing module includes acquiring radiation source signals, slicing the radiation source signals, obtaining two modalities—time series and time-frequency graphs—after slicing, normalizing the time series and time-frequency graphs respectively, labeling radiation source category labels, and constructing a modal dataset, which includes a training set, a validation set, and a test set.

[0008] The feature extraction module includes MHACNN and joint sparse representation. MHACNN is used to extract features from two modalities, while the joint sparse representation algorithm extracts low-redundancy differential features.

[0009] Furthermore: The MHACNN consists of multiple residual blocks, Transformer units, and fully connected layers. Each residual block comprises a residual unit and max pooling. The residual unit is operated through two convolutions and one skip connection. Assuming the input is Xs, the output of the first residual unit is:

[0010]

[0011] In the formula, c1 is the number of convolutional filters, σ represents the ReLU activation function, and W and b are the weights and bias parameters learned by each layer of the network.

[0012] The output of the first residual block after max pooling is:

[0013]

[0014] Assuming there are L residual blocks in total, the output of the l-th residual block is:

[0015]

[0016] Furthermore: the Transformer encoder consists of an MLP and multi-head attention. Assuming the output after position encoding is Op, the input to the Transformer encoder is O = Ol + Op, and the output of the Transformer encoder is Ot.

[0017] Furthermore: the joint sparse representation yields the optimal sparse matrix under the two modal time series and time-frequency plots, and simultaneously performs sparse cross-modal fusion of the two modalities. The sparse cross-modal fusion method is as follows:

[0018] First, the optimal sparse matrix is ​​obtained through joint sparse representation, resulting in sparse representations of the two modal time series and the time-frequency plot. The joint sparse representation is then transformed into:

[0019]

[0020] In the formula l u ' is the cost function, Let be the optimal sparse matrix to be found, d be the number of dictionary atoms, S be the number of modes, and here S = 2, α s Let x be the s-th column of A, which corresponds to the sparse representation of the s-th mode. s D represents the output features of the two input modalities after passing through the fourth convolutional unit B4. s Let be the dictionary for the s-th modality. The l2 norm is given by λ, where λ is the regularity coefficient.

[0021] Dictionary D s The solution obtained through optimization is as follows:

[0022]

[0023] The definition of convex set C in the formula is ns is the feature dimension of the s-th mode, and E(·) is the expectation operation;

[0024] The optimization problem in Equation (5) is solved using the classic projective stochastic gradient descent algorithm, which includes the following sequence update process:

[0025]

[0026] In the formula Let δ be the optimal dictionary obtained in the t-th iteration. t Π is the gradient step size. C For the orthogonal projection onto the convex set C, It is obtained by extracting from a randomly arranged feature set aggregated from features of different modalities;

[0027] Obtain the optimal dictionary D s Then, the optimal sparse matrix A is obtained by solving the ADMM algorithm, and the corresponding columns of A correspond to the corresponding sparse modes.

[0028] Furthermore, the cross-modal module performs dimensional splicing of the time series and time-frequency graph modes.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0030] The proposed sparse cross-modal communication radiation source identification method uses an MHACNN network to extract features from two input modalities: the original time series and the time-frequency graph of the radiation source. Then, it utilizes joint sparse representation to obtain more powerful modal features and performs cross-modal fusion to improve the accuracy of radiation source identification. This method leverages the rich time-frequency information contained in the communication radiation source signal, and its feature extraction using both the time and time-frequency domains exhibits better robustness than single-domain feature extraction methods. Furthermore, it utilizes joint sparse representation to obtain more powerful, low-redundancy modal features. Finally, it leverages the interaction of information from the two modalities, using one modality to guide the feature learning process of the other, and performs cross-modal fusion to further enhance feature representation capabilities. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only for more clearly illustrating the technical solutions in the embodiments of the present invention or the prior art. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 A schematic diagram of the modules that make up the present invention;

[0033] Figure 2 This is a schematic diagram of the residual block network structure of the present invention;

[0034] Figure 3 This is a schematic diagram of the Transformer encoder structure of the present invention;

[0035] Figure 4 This is a schematic diagram illustrating the principle of cross-modal feature fusion of coefficients in this invention.

[0036] Figure 5 The diagram shows the pseudocode algorithm for finding the optimal dictionary in this invention. Detailed Implementation

[0037] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be further described below with reference to specific embodiments. However, the embodiments are only for illustration and are not intended to limit the present invention.

[0038] like Figure 1 The method for identifying radiation sources based on sparse cross-modal communication shown consists of a signal preprocessing module, a feature extraction module, a cross-modal module, and an identification module.

[0039] The signal preprocessing module is used to preprocess the acquired original communication radiation source intermediate frequency signal. It includes acquiring the radiation source signal, slicing the radiation source signal, obtaining two modes after slicing: time series and time-frequency graph. The time series and time-frequency graph are normalized respectively, the radiation source category labels are labeled, and the modal dataset is constructed. The constructed modal dataset includes a training set, a validation set, and a test set.

[0040] The feature extraction module first uses a dual-channel Transformer network MHACNN (MultiheadAttention Based Convolutional Neural Network) to extract features from two modalities. Then, in order to reduce the redundancy between features caused by correlation, a joint sparse representation algorithm is used to extract low-redundancy differential features for better classification and recognition.

[0041] The MHACNN consists of multiple residual blocks, Transformer units, and fully connected layers. The network structure is shown in Table 1. B1 to B6 represent six residual blocks with the same structure. The structure of each residual block is as follows: Figure 2 As shown; B7 is a Transformer encoder unit with 4 attention heads and 16 hidden MLP units; B8 is a global average pooling unit with no hyperparameter settings, followed by a Dropout layer with a dropout ratio of 0.15, representing that 15% of the neurons in the Dropout layer have their outputs set to zero; B9 is a fully connected layer with SELU as its activation function, followed by a Dropout layer with a dropout ratio of 0.15; the last fully connected layer B10 is used to map to the output class, with softmax as its activation function.

[0042] Table 1. MHACNN Network Structure

[0043]

[0044]

[0045] The network structure of residual blocks is as follows Figure 2 As shown, one residual block consists of a residual unit and max pooling. The residual unit is the current input and the output of the input after passing through two one-dimensional convolutional layers. The two convolutional layers use ReLU and linear activation functions, respectively.

[0046] The residual unit is mainly completed through two operations: two convolutions and one skip connection. Assuming the input is Xs, the output of the first residual unit is:

[0047]

[0048] In the formula, c1 represents the number of convolutional filters, σ represents the ReLU activation function, and W and b are the weights and bias parameters learned by each layer of the network.

[0049] The output of the first residual block after max pooling is:

[0050]

[0051] Assuming there are L residual blocks in total, the output of the l-th residual block is:

[0052]

[0053] The structural diagram of the Transformer encoder is as follows: Figure 3 As shown, it mainly consists of an MLP and multi-head attention. The MLP is a simple fully connected feedforward network, while multi-head attention is a self-attention mechanism that establishes a mapping between the query and a set of key-value pairs to the output. Assuming the output after position encoding is Op, the input to the Transformer encoder is O = Ol + Op, and the output of the Transformer encoder is Ot.

[0054] Joint sparse representation provides an effective means for multimodal information fusion. It has been proven to be more effective than other fusion methods in multimodal information fusion, thus a sparse cross-modal fusion method is proposed. Its basic principle is to first obtain the optimal sparse matrix for two modalities (time series and time-frequency graph) through joint sparse representation, thereby eliminating redundant and noisy information. Simultaneously, the features of the two modalities in the sparse representation are fused across modalities. Each modal network can learn not only its own modal features but also the features of the other modality. Finally, the network parameters are iteratively optimized using the information from both modalities to obtain stronger feature representation capabilities.

[0055] The principle diagram of the sparse cross-modal fusion method is as follows: Figure 4 As shown in the figure, input modality 1 is a time series, and input modality 2 is a time-frequency plot. In the figure, "+" represents the concatenation of the dimensions of the two modalities, and "add" indicates the addition of features of the same dimension. Represents sparse modal features The output result after passing through the first modal convolutional unit B5 Cross-modal features, representing sparse modal features The output result after passing through the first modal convolution unit B5.

[0056] The sparse cross-modal fusion method first needs to obtain the optimal sparse matrix through joint sparse representation, thereby obtaining the sparse representations of the two modes. The joint sparse representation can be transformed into the following optimization problem:

[0057]

[0058] In the formula l u ' is the cost function, Let be the optimal sparse matrix to be found, d be the number of dictionary atoms, S be the number of modes, and here S = 2, α s Let x be the s-th column of A, which corresponds to the sparse representation of the s-th mode. s D represents the output features of the two input modalities after passing through the fourth convolutional unit B4. s Let be the dictionary for the s-th modality. The l2 norm is given by λ, where λ is the regularity coefficient.

[0059] Dictionary D s The following optimization problem can be solved:

[0060]

[0061] The definition of convex set C in the formula is ns represents the feature dimension of the s-th mode, and E(·) represents the expectation operation.

[0062] The optimization problem in formula (5) above is solved using the classic projective stochastic gradient descent algorithm, which includes the following sequence update process:

[0063]

[0064] In the formula Let δ be the optimal dictionary obtained in the t-th iteration. t Π is the gradient step size. C For the orthogonal projection onto the convex set C, It is obtained by extracting from a feature set aggregated from different modal features in a random arrangement.

[0065] Find the optimal dictionary D s pseudocode such as Figure 5 As shown in the figure, the alternating direction method of multipliers (ADMM) is used to solve the sparse coding, and then the gradient descent algorithm is used to update the dictionary D. After a certain number of iterations, the optimal dictionary D is returned. s .

[0066] Obtain the optimal dictionary D s Then, the optimal sparse matrix A can be solved using the ADMM algorithm based on variable splitting proposed by Afonso et al.

[0067] The corresponding columns of A correspond to the corresponding sparse modes, that is Two sparse modes are fused across modes via the fifth convolutional unit B5, yielding an output that is relevant to its own mode. and We also obtained the feed result for another mode. and Therefore, the cross-modal fusion results And the fusion results of mode 1 Fusion results of mode 2

[0068] The cross-modal output is used to predict the classification category, while the output from the fusion of modalities 1 and 2 is used as a regularization method for the loss function. This yields the corresponding outputs of the three fusions.

[0069]

[0070] In the formula y es y es1 and y es2 This is the output after passing through the second fully connected layer of the B10 unit and softmax activation. Its dimension is N*C, where N is the number of sample slices. Each row of elements corresponds to the confidence probability of the sample belonging to C classes.

[0071] The network's loss function, Loss, is:

[0072]

[0073] The first term in the formula is the cross-entropy loss, and the last two terms are regularization terms, which help prevent network overfitting and optimize network performance.

[0074] According to y es The maximum a posteriori probability criterion yields the prediction results for this sample slice. Assume the radiation source category labels y∈{1,2,...,C}, with a total of C classes. Given a sample x, w k Let be the weight of the k-th class, then the class of this sample slice is:

[0075]

[0076] This method uses the MHACNN network to extract features from two input modalities of radiation sources: the original time series and the time-frequency graph. Then, it uses joint sparse representation to obtain modal features with stronger representation capabilities and performs cross-modal fusion to improve the accuracy of radiation source identification.

[0077] All content not described in detail in this invention is prior art.

[0078] The above description is merely a preferred embodiment of the present invention and is not limited to the description in the specification and embodiments. Therefore, all equivalent changes or modifications made to the structure, features, and principles described in the claims of this invention should be included within the scope of this patent application.

Claims

1. A sparse cross-modality communication radiation source identification method based on, comprising a signal preprocessing module, a feature extraction module, a cross-modality module and an identification module, characterized in that: The signal preprocessing module includes the collection of radiation source signals, the slicing of the radiation source signals, the obtaining of time series and time-frequency graphs after slicing, the normalization of the time series and the time-frequency graphs, the labeling of radiation source class labels, and the construction of modal data sets, wherein the modal data sets include training sets, verification sets, and test sets; The feature extraction module is internally provided with MHACNN and joint sparse representation, wherein the MHACNN is used to extract the features of the two modalities, and the joint sparse representation algorithm is used to extract low-redundancy differentiated features; The joint sparse representation obtains the optimal sparse matrix under the time series and the time-frequency graph of the two modalities, and simultaneously performs sparse cross-modal fusion on the two modalities, and the sparse cross-modal fusion method is as follows: First, the optimal sparse matrix is obtained through the joint sparse representation, and the sparse representation of the time series and the time-frequency graph of the two modalities is obtained, and the joint sparse representation is converted into: (1) wherein is a cost function, is the optimal sparse matrix to be solved, d is the number of dictionary atoms, S is the number of modalities, here S = 2, is the A column of s , which corresponds to the sparse representation of the s th modality, is the output feature of the two input modalities after the fourth convolutional unit B4, is the dictionary of the s th modality, is the l 2-norm, and λ is the regularization coefficient; dictionary by solving the optimization problem as: (2) wherein the convex set C is defined as , n s is the feature dimension of the s th modality, is the desired operation; The optimization problem in formula (2) is solved by a classical projection stochastic gradient descent algorithm, which includes the following sequence update process: (3) In the formula is the first t The optimal dictionary obtained by step iteration, is the gradient step size, is the orthogonal projection on the convex set C , is obtained by extracting in a randomly arranged feature set aggregated by different modal features; Obtaining optimal dictionary Then, the optimal sparse matrix A is obtained by solving the ADMM algorithm, and the corresponding column of A corresponds to the corresponding sparse mode, .

2. The method of claim 1, wherein: The MHACNN is composed of multiple residual blocks, Transformer units and fully connected layers, wherein one residual block is composed of one residual unit and max pooling, the residual unit is completed by twice convolution and once skip connection operation, assuming that the input is X s The output of the first residual unit is: (4) wherein c 1 is the number of convolution filters, σ represents the ReLU activation function, W and b are the learned weight and bias parameters for each layer of the network; max-pooling maxpool The output of the first residual block is: (5) Assuming there are L residual blocks, the output of the l th residual block is: (6)。 3. The method of claim 2, wherein: The Transformer encoder is composed of an MLP and multi-head attention, assuming the output after position encoding is O p The input into the Transformer encoder is O = O l + O p The output of the Transformer encoder is O t .

4. The method of claim 1, wherein: The cross-modal module performs dimension concatenation on the time series and the time-frequency graph of the two modalities.

Citation Information

Patent Citations

  • Explanatable multi-source remote sensing image joint classification method based on sparse representation model

    CN115719431A

  • Radar radiation source signal identification method based on improved Transform

    CN117421654A