Radar active jamming identification method based on CNN-ViT hybrid network
By constructing a CNN-ViT hybrid network, combining adaptive convolutional kernels and a depth transposed attention module, the problems of high computational complexity and low accuracy in radar interference signal recognition are solved, achieving efficient and accurate radar active interference recognition.
Patent Information
- Application Number
- CN202510567750.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-01
AI Technical Summary
Existing radar jamming signal identification methods have high computational complexity, poor identification accuracy, and are difficult to meet real-time requirements. Traditional convolutional neural networks lack global information interaction, while ViT models have high computational costs and are not suitable for resource-constrained devices.
A radar active interference identification method based on CNN-ViT hybrid network is adopted. By constructing an active interference signal dataset, performing time-frequency transformation and multi-scale feature extraction and fusion, and using a hybrid network constructed with adaptive convolution kernel and depth transposed attention module for identification.
It achieves improved recognition accuracy while reducing computational complexity, effectively captures multi-level features, and combines local details with global context, making it suitable for real-time recognition on resource-constrained devices.
Smart Images

Figure CN120408351A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar, and particularly relates to a radar active interference recognition method based on a CNN-ViT hybrid network. Background Art
[0002] In today's complex and ever-changing radar application environment, the continuous development of radar jamming technology poses a huge threat to the effectiveness of radar systems. To ensure the effective operation of radar systems in this context, the requirements for radar anti-jamming technology are constantly increasing. Therefore, accurately and quickly identifying radar jamming signals has become a key prerequisite for ensuring the effectiveness of anti-jamming measures. Traditional methods for detecting and identifying radar jamming signals include subjective experience judgments based on expert knowledge and analysis through feature parameter extraction and pattern recognition models. These methods all have some limitations. First, judgments relying on experts and subjective experience are easily affected by individual subjective factors, resulting in poor reliability of the recognition results and low engineering practicability. Second, the method of feature parameter extraction requires manual operation and involves cumbersome feature engineering.
[0003] Compared with traditional methods, the radar jamming signal recognition method based on deep learning reduces a large amount of complex feature engineering by automatically learning features from data. Currently, for the active interference recognition method based on deep learning, a convolutional neural network is generally used to automatically extract features for recognition. However, the two-dimensional time-frequency diagram of interference signals is different from specific physical pictures. Generally, after converting the interference signal into a time-frequency grayscale diagram, the pixel size is only a few kilobytes, which is three orders of magnitude different from the size of physical pictures in megabytes. Directly using a general deep convolutional neural network for recognition has unsatisfactory results, and there are problems such as being prone to overfitting, poor recognition effect, and difficulty in meeting the real-time requirements of interference recognition.
[0004] The lightweight network based on CNN (Convolutional Neural Networks) only has a local receptive field and lacks the interaction of global information, and cannot model global information. Although introducing the self-attention mechanism can make it possible to explicitly model this global interaction, the self-attention in ViT (Vision Transformer) brings expensive computational costs, which is not conducive to deployment and real-time recognition on resource-constrained mobile devices.
[0005] Therefore, how to provide a radar active interference recognition method with low computational complexity, high recognition accuracy, and high efficiency has become an important issue. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a method for identifying radar active jamming based on a CNN-ViT hybrid network.
[0007] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0008] In a first aspect, the present invention provides a method for identifying radar active jamming based on a CNN-ViT hybrid network, the radar active jamming identification method including:
[0009] Construct an active jamming signal data set; the active jamming signal data set includes a plurality of active jamming signal samples and true classification labels corresponding to each active jamming signal sample;
[0010] Perform time-frequency transformation on each active jamming signal sample in the active jamming signal data set to obtain an interference signal time-frequency map sample corresponding to the active jamming signal sample;
[0011] Use a pre-constructed CNN-ViT hybrid network to perform multi-scale feature extraction and fusion and interference identification operations on each interference signal time-frequency map sample to obtain a predicted sample classification result; the CNN-ViT hybrid network is constructed based on an adaptive convolution kernel and a depth transposed attention module;
[0012] Train the CNN-ViT hybrid network according to the difference between the true classification label corresponding to each active jamming signal sample and the predicted sample classification result, so as to identify radar active jamming according to the trained CNN-ViT hybrid network.
[0013] Optionally, the CNN-ViT hybrid network includes:
[0014] An initial convolution module for extracting features from each active jamming signal sample in the active jamming signal data set to obtain initial sample image features of each active jamming signal sample;
[0015] A multi-layer multi-scale fusion processing module for performing multi-scale fusion on the initial sample image features of each active jamming signal sample to obtain multi-scale fusion features, and introducing a self-attention mechanism to perform weighted fusion on the multi-scale features to obtain sample optimized fusion features of each active jamming signal sample;
[0016] A fully connected module for classifying each active jamming signal sample based on the sample optimized fusion features of each active jamming signal sample to obtain a predicted sample classification result.
[0017] Optionally, the multi-layer multi-scale fusion processing module includes a first adaptive convolution module, a first depth transposed attention module, a second adaptive convolution module, a second depth transposed attention module, a third adaptive convolution module, and a third depth transposed attention module;
[0018] The first adaptive convolution module is used to focus on the local features in the initial sample image features of each active interference signal sample to extract the small-scale features of each active interference signal sample;
[0019] The first depth transposed attention module is used to process the subsets obtained by splitting the small-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed small-scale features to obtain small-scale multi-layer information representations;
[0020] The second adaptive convolution module is used to balance the small-scale multi-layer information representation features and regional context of each active interference signal sample to extract the medium-scale features of each active interference signal sample;
[0021] The second depth transposed attention module is used to process the subsets obtained by splitting the medium-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed medium-scale features to obtain medium-scale multi-layer information representations;
[0022] The third adaptive convolution module is used to integrate the medium-scale multi-layer information representations and global semantic information of each active interference signal sample to extract the large-scale features of each active interference signal sample;
[0023] The third depth transposed attention module is used to process the subsets obtained by splitting the large-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed large-scale features to obtain large-scale multi-layer information representations as sample optimized fusion features.
[0024] Optionally, the fully connected module includes a global average pooling layer, a first fully connected layer, and a second fully connected layer connected in sequence.
[0025] Optionally, performing time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample, including:
[0026] Using the CWD transformation method, perform time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample.
[0027] Optionally, the CWD transformation method is implemented using the following formula:
[0028] J CWD J(t,f) = ∫∫A(η,τ)Φ(η,τ)exp(j2π(ηt - τf))dηdτ;
[0029]
[0030] Wherein, J CWD J(t,f) represents the time-frequency diagram sample of the interference signal; A(η,τ) represents the ambiguity function; exp(·) represents the exponential function of e; j represents the imaginary unit; t represents time; f represents the frequency of the active interference signal sample; s(·) represents the time-domain form of the active interference signal sample; s * (·) represents the conjugate of s(·); Φ(η,τ) represents the kernel function; η is used to determine the frequency resolution of the kernel function; τ represents the time shift term.
[0031] In a second aspect, the present invention provides a radar active interference recognition device based on a CNN-ViT hybrid network, the radar active interference recognition device includes:
[0032] A construction module for constructing an active interference signal data set; the active interference signal data set includes a plurality of active interference signal samples and the true classification labels corresponding to each active interference signal sample;
[0033] A time-frequency transformation module for performing time-frequency transformation on each active interference signal sample in the active interference signal data set to obtain the interference signal time-frequency diagram sample corresponding to the active interference signal sample;
[0034] An identification module for performing multi-scale feature extraction and fusion and interference identification operations on each interference signal time-frequency diagram sample by using a pre-constructed CNN-ViT hybrid network to obtain a predicted sample classification result; the CNN-ViT hybrid network is constructed based on an adaptive convolution kernel and a depth transposed attention module;
[0035] A training module for training the CNN-ViT hybrid network according to the difference between the true classification label corresponding to each active interference signal sample and the predicted sample classification result, so as to perform radar active interference identification according to the trained CNN-ViT hybrid network.
[0036] In a third aspect, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0037] The memory is used to store a computer program;
[0038] A processor, when executing a computer program stored in a memory, implements the method steps of any of the above-mentioned radar active interference recognition methods based on the CNN-ViT hybrid network.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it implements the method steps of any of the above-mentioned radar active interference recognition methods based on the CNN-ViT hybrid network.
[0040] In the radar active interference recognition method based on the CNN-ViT hybrid network provided by the present invention, the CNN-ViT hybrid network is constructed based on an adaptive convolution kernel and a depth transposed attention module. The adaptive convolution kernel can significantly reduce the computational complexity while effectively capturing multi-level features. At the same time, the depth transposed attention module expands the receptive field range and realizes multi-scale feature encoding. This dual-module collaborative architecture not only retains the advantages of the ViT model in processing global features but also avoids the problem of high computational complexity of self-attention in traditional ViT. Finally, on the premise of ensuring the unified modeling ability of local details and global context, it realizes the efficient utilization of computing resources. The following will further elaborate on the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic flowchart of a radar active interference recognition method based on the CNN-ViT hybrid network provided by an embodiment of the present invention;
[0042] Figure 2 is a schematic structural diagram of the CNN-ViT hybrid network according to an embodiment of the present invention;
[0043] Figure 3 is a schematic structural diagram of the adaptive convolution module provided by an embodiment of the present invention;
[0044] Figure 4 is a schematic structural diagram of the depth transposed attention module provided by an embodiment of the present invention;
[0045] Figure 5 is a schematic structural diagram of a radar active interference recognition device based on the CNN-ViT hybrid network provided by an embodiment of the present invention;
[0046] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following further describes the present invention in detail with specific embodiments, but the embodiments of the present invention are not limited thereto.
[0048] To solve the problems of high computational complexity, poor recognition accuracy, and low recognition efficiency existing in the existing radar active interference recognition methods, an embodiment of the present invention provides a radar active interference recognition method based on a CNN-ViT hybrid network. Refer to Figure 1 , Figure 1 which is a schematic flowchart of a radar active interference recognition method based on a CNN-ViT hybrid network provided by an embodiment of the present invention, and specifically includes the following steps:
[0049] Step S101, construct an active interference signal dataset; the active interference signal dataset includes multiple active interference signal samples and the corresponding true classification labels of each active interference signal sample.
[0050] In an embodiment of the present invention, an active interference signal refers to an electromagnetic interference signal deliberately radiated by the enemy to the radar actively through a specific signal source.
[0051] In an embodiment of the present invention, an active interference signal dataset can be constructed according to 8 types of active interference signals. Among them, the 8 interference types of active interference signals include noise amplitude modulation interference, noise frequency modulation interference, noise convolution interference, noise product interference, slice reconstruction interference, intermittent sampling and forwarding interference, comb-like spectrum interference, and spectrum dispersion interference. Random noise is added to each active interference signal, and the SNR (signal-to-noise ratio) is used to control the ratio of signal power to noise power.
[0052] Specifically, in the active interference signal dataset, the range of SNR is set to (-10dB, 10dB). At each SNR, 100 active interference signal samples are generated for each type of interference according to random parameters, and their true classification labels are noise amplitude modulation interference, noise frequency modulation interference, noise convolution interference, noise product interference, slice reconstruction interference, intermittent sampling and forwarding interference, comb-like spectrum interference, or spectrum dispersion interference.
[0053] Step S102, perform time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain the interference signal time-frequency map sample corresponding to the active interference signal sample.
[0054] In an embodiment of the present invention, by performing time-frequency transformation on the active interference signal sample, the interference signal time-frequency map sample can be obtained.
[0055] In one implementation, performing time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain the interference signal time-frequency map sample corresponding to the active interference signal sample includes:
[0056] Adopt the CWD transformation method to perform time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain the interference signal time-frequency map sample corresponding to the active interference signal sample.
[0057] Through Choi-Williams distribution (CWD) time-frequency analysis, the time-varying characteristics of each active interference signal sample in the active interference signal dataset can be clearly revealed, and the time-frequency diagram samples of the interference signals are obtained. Among them, the CWD transformation method is a time-frequency analysis method, which is a characterization method obtained by a linear transformation of the joint distribution of the time domain and frequency domain of the signal.
[0058] In one implementation, the CWD transformation method is implemented using the following formula:
[0059] J CWD (t,f) = ∫∫A(η,τ)Φ(η,τ)exp(j2π(ητ - τf))dηdτ;
[0060]
[0061] where J CWD (t,f) represents the CWD transformation result, that is, the time-frequency diagram sample of the interference signal; A(η,τ) represents the ambiguity function, which characterizes the autocorrelation characteristics of the active interference signal sample; exp(·) represents the exponential function of e; j represents the imaginary unit; t represents time; f represents the frequency of the active interference signal sample; s(·) represents the time-domain form of the active interference signal sample; s * (·) represents the conjugate of s(·); Φ(η,τ) represents the kernel function, which can suppress cross-term interference and keep the self-term energy of the signal, and its definition is:
[0062]
[0063] where η is used to determine the frequency resolution of the kernel function; τ represents the time shift term; σ is a scaling factor that controls the trade-off between frequency resolution and cross-term suppression.
[0064] Specifically, σ = 1 can be selected for CWD time-frequency transformation. So finally, a dataset of 16,800 interference signal time-frequency diagrams covering 8 types of active interference and 21 SNR levels is generated, and the training set, validation set, and test set are divided according to the ratio of 70 - 15 - 15.
[0065] Step S103, use the pre-constructed CNN-ViT hybrid network to perform multi-scale feature extraction and fusion and interference recognition operations on each interference signal time-frequency diagram sample to obtain the predicted sample classification result; the CNN-ViT hybrid network is constructed based on an adaptive convolutional kernel and a depth transposed attention module.
[0066] In the embodiment of the present invention, in the CNN-ViT hybrid network, CNN is good at local features, while ViT can capture long-range dependencies. The CNN-ViT hybrid network includes:
[0067] An initial convolutional module for extracting features from each active interference signal sample in the active interference signal dataset to obtain the initial sample image features of each active interference signal sample.
[0068] See Figure 2 , Figure 2 is a schematic diagram of the structure of the CNN-ViT hybrid network according to an embodiment of the present invention. The initial convolutional module is a conventional convolutional kernel Conv with a size of 7×7 for capturing rough image features. The convolution of the convolutional kernel W with the image coordinate I(i,j) of the input interference signal time-frequency map sample I can be calculated as:
[0069]
[0070] where x(I,W) (i,j) is the image feature value at the (i,j) coordinate of the input interference signal time-frequency map sample I, i represents the abscissa of the interference signal time-frequency map sample, and j represents the ordinate of the interference signal time-frequency map sample; (m,n,k) is the coordinate index of the convolutional kernel; b is the bias.
[0071] In an embodiment of the present invention, a batch normalization (Bn) layer and an exponential linear unit (eLU) layer can also be added after the initial convolutional module.
[0072] Assume that the mean of each mini-batch of interference signal time-frequency map samples and input channels is μ b and the variance is The Bn layer first normalizes the input interference signal time-frequency map sample to:
[0073]
[0074] where, represents the first normalized eigenvalue at the (i,j) coordinate of the interference signal time-frequency map sample I; ∈ represents a constant scalar.
[0075] Calculate the output value of the Bn layer:
[0076]
[0077] where, y (i,j) represents the second normalized eigenvalue at the (i,j) coordinate of the interference signal time-frequency map sample I; λ and β are two learnable parameters for scaling and translating the normalized data respectively.
[0078] Input y (i,j) into the eLU layer, where the eLU activation function of the eLU layer returns the same value for positive inputs and the exponential operation result for negative inputs, and its expression is as follows:
[0079]
[0080] Among them, eLU(y (i,j) ) represents the third eigenvalue at the (i, j) coordinate of the time-frequency diagram sample I of the interference signal.
[0081] According to the third eigenvalue at each coordinate in the time-frequency diagram sample of the interference signal, the initial sample image features of each active interference signal sample can be obtained.
[0082] In the embodiment of the present invention, the multi-layer multi-scale fusion processing module is used to perform multi-scale fusion on the initial sample image features of each active interference signal sample to obtain multi-scale fusion features, and introduce a self-attention mechanism to perform weighted fusion on the multi-scale features to obtain the sample optimized fusion features of each active interference signal sample.
[0083] In one implementation, the multi-layer multi-scale fusion processing module includes a first adaptive convolution module, a first depth transposed attention module, a second adaptive convolution module, a second depth transposed attention module, a third adaptive convolution module, and a third depth transposed attention module.
[0084] The first adaptive convolution module is used to focus on the local features in the initial sample image features of each active interference signal sample to extract the small-scale features of each active interference signal sample;
[0085] The first depth transposed attention module is used to process the subsets obtained by splitting the small-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed small-scale features to obtain small-scale multi-layer information representations;
[0086] The second adaptive convolution module is used to balance the small-scale multi-layer information representation features and regional context of each active interference signal sample to extract the middle-scale features of each active interference signal sample;
[0087] The second depth transposed attention module is used to process the subsets obtained by splitting the middle-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed middle-scale features to obtain middle-scale multi-layer information representations;
[0088] The third adaptive convolution module is used to integrate the middle-scale multi-layer information representation and global semantic information of each active interference signal sample to extract the large-scale features of each active interference signal sample;
[0089] The third depth transposed attention module is used to process subsets obtained by splitting large-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed large-scale features, obtaining large-scale multi-layer information representations as sample optimized fusion features.
[0090] See Figure 2 and Figure 3 , Figure 3 is a schematic structural diagram of the adaptive convolution module provided by an embodiment of the present invention. Three N×N adaptive convolution modules are built. This block consists of depthwise separable convolutions with adaptive kernel sizes, and the kernel sizes change dynamically. The sizes of the first adaptive convolution module, the second adaptive convolution module, and the third adaptive convolution module are N = 3, 5, and 7 in sequence. After convolution, a batch normalization (Bn) layer and an exponential linear unit (eLU) layer are added. Finally, a residual connection is added to enable information to propagate jumpwise in the network hierarchy, where the Concat operation is a splicing operation, and Norm represents the norm. The calculation method of the adaptive convolution module can be expressed as follows:
[0091] x q+1 = x q + eLU(Pw(Bn(Dw(x q ))))
[0092] where, x q represents the input feature of the adaptive convolution module; Dw represents depth convolution; Bn is the batch normalization layer; Pw represents pointwise convolution; eLU represents the exponential linear unit; x q+1 represents the feature output by this adaptive convolution module.
[0093] In the embodiments of the present invention, the first adaptive convolution module, the second adaptive convolution module, and the third adaptive convolution module all apply to the above-mentioned calculation method of the adaptive convolution module.
[0094] In the embodiments of the present invention, through the first adaptive convolution module, the second adaptive convolution module, and the third adaptive convolution module, different-level features can be captured in the network while reducing the computational cost. Smaller kernels are used in the early stage to capture low-level features, while larger kernels are used in the later stage to capture high-level features, reducing the use of large kernels to reduce the network computational amount.
[0095] See Figure 2 and Figure 4 , Figure 4It is a schematic structural diagram of the depth transposed attention module provided by an embodiment of the present invention. Three depth transposed attention modules are built and placed sequentially after the adaptive convolution module. That is, the order is the first adaptive convolution module, the first depth transposed attention module, the second adaptive convolution module, the second depth transposed attention module, the third adaptive convolution module, and the third depth transposed attention module. The depth transposed attention module is mainly composed of two components. The first component is used for feature fusion at the multi-scale level, and then the fused information is used for global information modeling in the second component.
[0096] In the first component, first, the features of multiple active interference signal samples input are evenly split into s subsets, that is, each subset is represented by O p , and each subset has the same number of channels and spatial size, p ∈ {1, 2,..., s}. Among them, corresponding to different depth transposed attention modules, the input can be small-scale features, medium-scale features, or large-scale features of multiple active interference signal samples.
[0097] Except for the first subset, a 3×3 depth convolution is added after each subset branch. The input of each depth convolution is composed of the output of the previous branch convolution and the subset of this branch, and the output formula is as follows:
[0098]
[0099] y p represents the output feature of the current depth convolution, and y p-1 represents the output feature of the previous branch convolution.
[0100] Finally, the output features of all branches are merged in the channel dimension, so the number of channels of the finally input and output features remains unchanged.
[0101] The second component receives the output of the first component. Assuming that the output of the first component is a feature of H×W×C, then the projections of the query Q, the key K, and the value V are calculated through three linear layers, and Q = W Q Y, K = W K Y, V = W V Y, with a dimension of HW×C, where W Q , W K , and W V are the projection weights of Q, K, and V respectively. Then, a dot product is applied between Q T and K along the channel dimension instead of the spatial dimension to generate a C×C softmax-scaled attention score matrix. The computational complexity of such an operation is C 2 (HW). Compared with the dot product operation in the spatial dimension, the computational complexity is reduced to a linear complexity with respect to (HW), greatly reducing the computational cost of the network, whereFigure 4 Reshape in it is used to change the shape of the input tensor while keeping the total number of elements unchanged. Finally, the score matrix is multiplied by V to obtain the final attention map. The specific calculation formula is as follows:
[0102]
[0103] Among them, Y is the input feature of the second component, is the output feature of the second component, and T represents the transpose operation of the matrix. Then, two 1×1 pointwise convolutional layers, Bn layers, and eLU layers are added to generate non-linear features.
[0104] Specifically, the Q encoding represents the information requirements at the current position and is used to query the potential connections in different features; the K encoding represents the information identity at the current position and is used to match the similarity measure of the query; the V encoding represents the information content at the current position and is used to aggregate the actually transmitted features under the attention weights.
[0105] In the embodiment of the present invention, through the first depth transposed attention module, the second depth transposed attention module, and the third depth transposed attention module, in each depth transposed attention module, the input tensor is split into multiple channel groups, and depth convolution and self-attention across the channel dimension are used to implicitly increase the receptive field and encode multi-scale features.
[0106] In one implementation, referring to Figure 2 , positional encoding can also be introduced after the first adaptive convolution module. The positional encoding is used to provide the model with the order information of each position in the sequence because the self-attention mechanism itself does not have the ability to perceive positions.
[0107] In the embodiment of the present invention, the fully connected module is used to classify each active interference signal sample based on the sample-optimized fusion features of each active interference signal sample to obtain the predicted sample classification result.
[0108] In one implementation, the fully connected module includes a global average pooling layer, a first fully connected layer, and a second fully connected layer.
[0109] A global average pooling layer is set at the beginning of the fully connected layer. A dropout layer with a dropout rate of 0.5 is added to the first fully connected layer to alleviate overfitting. The number of hidden neurons set in the second fully connected layer is the same as the number of interference categories in the given dataset. Finally, the predicted sample classification result is output through the softmax function.
[0110] Step S104, train the CNN-ViT hybrid network according to the difference between the true classification labels corresponding to each active interference signal sample and the predicted sample classification result, so as to identify radar active interference according to the trained CNN-ViT hybrid network.
[0111] In the embodiments of the present invention, the training parameters are set as follows:
[0112] The network is trained using the momentum gradient descent method. The momentum parameter μ is set to 0.9, the L2 regularization factor λ is set to 0.0001, the maximum number of iterations is 1000, the initial learning rate α is 0.01 (decreasing by a factor of 0.1 every 45 iterations), and the mini-batch size is 256.
[0113] Randomly take 256 active interference signal samples from the active interference signal dataset constructed in step S101, and input them into the CNN-ViT hybrid network to obtain the cross-entropy loss function value
[0114]
[0115] Among them, y is the true classification label corresponding to the active interference signal sample, that is, the true value, which is a one-hot vector, that is, except for the index position corresponding to the correct category being 1, all other positions are 0; is the predicted sample classification result, that is, the model prediction value, which is a vector output by the final softmax function of the fully connected layer, and the sum of its elements is 1; t represents the t-th category value.
[0116] Let the parameters of each layer of the network be represented by θ. For the parameters θ of each layer of the network, calculate the gradient magnitude grad of the loss function with respect to θ:
[0117]
[0118] Update the network parameters θ according to the gradient grad, the momentum parameter, and the learning rate:
[0119]
[0120] Among them, g t represents the current gradient; v t represents the current momentum; v t-1 represents the momentum of the previous iteration.
[0121] During the training process, taking the accuracy of the validation set as the standard, save the model parameters with the highest accuracy after each iteration. After reaching the maximum number of training iterations, terminate the training to obtain the trained network model.
[0122] Input the test set in the active interference signal dataset generated in step S101 into the trained network model to identify 8 types of active interferences, and obtain the accuracy of the model in identifying the interferences at different SNRs.
[0123] In an embodiment of the present invention, the CNN-ViT hybrid network is constructed based on an adaptive convolutional kernel and a depth transposed attention module. The adaptive convolutional kernel can effectively capture multi-level features while significantly reducing the computational complexity. At the same time, the depth transposed attention module expands the receptive field range and realizes multi-scale feature encoding. This dual-module collaborative architecture not only retains the advantage of the ViT model in processing global features but also avoids the problem of high computational complexity of self-attention in traditional ViT. Finally, on the premise of ensuring the unified modeling ability of local details and global context, the efficient utilization of computing resources is achieved.
[0124] Based on the same inventive concept, an embodiment of the present invention also provides a radar active interference recognition device based on the CNN-ViT hybrid network. Refer to Figure 5 , Figure 5 which is a schematic structural diagram of a radar active interference recognition device based on the CNN-ViT hybrid network provided by an embodiment of the present invention. The radar active interference recognition device includes:
[0125] A construction module 501 for constructing an active interference signal dataset. The active interference signal dataset includes a plurality of active interference signal samples and the corresponding true classification labels of each active interference signal sample;
[0126] A time-frequency transformation module 502 for performing time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample;
[0127] An identification module 503 for performing multi-scale feature extraction and fusion and interference identification operations on each interference signal time-frequency map sample by using a pre-constructed CNN-ViT hybrid network to obtain a predicted sample classification result. The CNN-ViT hybrid network is constructed based on an adaptive convolutional kernel and a depth transposed attention module;
[0128] A training module 504 for training the CNN-ViT hybrid network according to the difference between the true classification label corresponding to each active interference signal sample and the predicted sample classification result, so as to identify radar active interference according to the trained CNN-ViT hybrid network.
[0129] In an embodiment of the present invention, the CNN-ViT hybrid network is constructed based on an adaptive convolutional kernel and a depth transposed attention module. The adaptive convolutional kernel can effectively capture multi-level features while significantly reducing the computational complexity. At the same time, the depth transposed attention module expands the receptive field range and realizes multi-scale feature encoding. This dual-module collaborative architecture not only retains the advantage of the ViT model in processing global features but also avoids the problem of high computational complexity of self-attention in traditional ViT. Finally, on the premise of ensuring the unified modeling ability of local details and global context, the efficient utilization of computing resources is achieved.
[0130] Optionally, the CNN-ViT hybrid network includes:
[0131] An initial convolutional module for extracting features from each active interference signal sample in the active interference signal dataset to obtain the initial sample image features of each active interference signal sample;
[0132] A multi-layer multi-scale fusion processing module for performing multi-scale fusion on the initial sample image features of each active interference signal sample to obtain multi-scale fusion features, and introducing a self-attention mechanism to perform weighted fusion on the multi-scale features to obtain the sample optimized fusion features of each active interference signal sample;
[0133] A fully connected module for classifying each active interference signal sample based on the sample optimized fusion features of each active interference signal sample to obtain the predicted sample classification result.
[0134] Optionally, the multi-layer multi-scale fusion processing module includes a first adaptive convolutional module, a first depth transposed attention module, a second adaptive convolutional module, a second depth transposed attention module, a third adaptive convolutional module, and a third depth transposed attention module;
[0135] The first adaptive convolutional module is used to focus on the local features in the initial sample image features of each active interference signal sample to extract the small-scale features of each active interference signal sample;
[0136] The first depth transposed attention module is used to process the subsets obtained by splitting the small-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed small-scale features to obtain small-scale multi-layer information representations;
[0137] The second adaptive convolutional module is used to balance the small-scale multi-layer information representation features and regional context of each active interference signal sample to extract the medium-scale features of each active interference signal sample;
[0138] The second depth transposed attention module is used to process subsets obtained by splitting the mesoscale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed mesoscale features to obtain a mesoscale multi-layer information representation;
[0139] The third adaptive convolution module is used to integrate the mesoscale multi-layer information representation and global semantic information of each active interference signal sample to extract the large-scale features of each active interference signal sample;
[0140] The third depth transposed attention module is used to process subsets obtained by splitting the large-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed large-scale features to obtain a large-scale multi-layer information representation as the sample optimized fusion feature.
[0141] Optionally, the fully connected module includes a global average pooling layer, a first fully connected layer, and a second fully connected layer connected in sequence.
[0142] Optionally, the time-frequency transformation module is specifically used for:
[0143] Using the CWD transformation method, perform time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample.
[0144] Optionally, the CWD transformation method is implemented using the following formula:
[0145] J CWD (t,f) = ∫∫A(η,τ)Φ(η,τ)exp(j2π(ητ - τf))dηdτ;
[0146]
[0147] where J CWD (t,f) represents the interference signal time-frequency map sample; A(η,τ) represents the ambiguity function; exp(·) represents the exponential function of e; j represents the imaginary unit; t represents time; f represents the frequency of the active interference signal sample; s(·) represents the time-domain form of the active interference signal sample; s * (·) represents the conjugate of s(·); Φ(η,τ) represents the kernel function; η is used to determine the frequency resolution of the kernel function; τ represents the time shift term.
[0148] An embodiment of the present invention also provides an electronic device, as Figure 6 shown, Figure 6The following is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, which includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 complete mutual communication through the communication bus 604.
[0149] The memory 603 is used to store computer programs.
[0150] When the processor 601 is used to execute the program stored on the memory 603, it realizes the method steps of any one of the above radar active interference recognition methods based on the CNN-ViT hybrid network.
[0151] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0152] The communication interface is used for communication between the above electronic device and other devices.
[0153] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0154] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0155] The present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method steps of any of the above-mentioned radar active interference recognition methods based on the CNN-ViT hybrid network are implemented.
[0156] Optionally, the computer-readable storage medium may be a non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory.
[0157] Optionally, the above computer-readable storage medium may also be at least one storage device located away from the aforementioned processor.
[0158] In another embodiment of the present invention, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the method steps of any of the above-mentioned radar active interference recognition methods based on the CNN-ViT hybrid network.
[0159] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention.
[0160] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0161] Although the present invention has been described in connection with various embodiments, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the accompanying drawings and the disclosure during the implementation of the claimed invention. In the description of the present invention, the term "comprising" does not exclude other components or steps, the indefinite article "a" or "an" does not exclude a plurality, and the meaning of "plurality" is two or more, unless otherwise specifically defined. In addition, certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0162] The method provided by the embodiments of the present invention can be applied to an electronic device. Specifically, the electronic device can be: a desktop computer, a portable computer, a smart mobile terminal, a server, etc. There is no limitation here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.
[0163] For the device / electronic device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0164] It should be noted that the device, electronic device, and storage medium of the embodiments of the present invention are respectively the device, electronic device, and storage medium that apply the above-mentioned method for identifying radar active interference based on a CNN-ViT hybrid network. Then all embodiments of the above-mentioned method for identifying radar active interference based on a CNN-ViT hybrid network are applicable to the device, electronic device, and storage medium, and can achieve the same or similar beneficial effects.
[0165] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A radar active interference recognition method based on a CNN-ViT hybrid network, characterized in that The radar active interference recognition method includes: Construct an active interference signal data set; the active interference signal data set includes multiple active interference signal samples and the corresponding true classification labels of each active interference signal sample; Perform time-frequency transformation on each active interference signal sample in the active interference signal data set to obtain the interference signal time-frequency map sample corresponding to the active interference signal sample; Use the pre-constructed CNN-ViT hybrid network to perform multi-scale feature extraction fusion and interference recognition operations on each interference signal time-frequency map sample to obtain the predicted sample classification result; the CNN-ViT hybrid network is constructed based on an adaptive convolution kernel and a depth transposed attention module; Train the CNN-ViT hybrid network according to the difference between the true classification label corresponding to each active interference signal sample and the predicted sample classification result, so as to identify radar active interference according to the trained CNN-ViT hybrid network.
2. The radar active interference recognition method according to claim 1, wherein The CNN-ViT hybrid network includes: An initial convolution module for extracting features from each active interference signal sample in the active interference signal data set to obtain the initial sample image features of each active interference signal sample; A multi-layer multi-scale fusion processing module for performing multi-scale fusion on the initial sample image features of each active interference signal sample to obtain multi-scale fusion features, and introducing a self-attention mechanism to perform weighted fusion on the multi-scale features to obtain the sample optimized fusion features of each active interference signal sample; A fully connected module for classifying each active interference signal sample based on the sample optimized fusion features of each active interference signal sample to obtain the predicted sample classification result.
3. The radar active interference recognition method according to claim 2, wherein The multi-layer multi-scale fusion processing module includes a first adaptive convolution module, a first depth transposed attention module, a second adaptive convolution module, a second depth transposed attention module, a third adaptive convolution module and a third depth transposed attention module; The first adaptive convolution module is used to focus on the local features in the initial sample image features of each active interference signal sample to extract the small-scale features of each active interference signal sample; The first depth transposed attention module is used to process the subsets obtained by splitting the small-scale features of multiple active interference signal samples with multiple channel groups, and introduce a self-attention mechanism to fuse the processed small-scale features to obtain the small-scale multi-layer information representation; The second adaptive convolution module is used to take into account the small-scale multi-layer information representation features and regional context of each active interference signal sample to extract the medium-scale features of each active interference signal sample; The second depth transposed attention module is used to process the subsets obtained by splitting the medium-scale features of multiple active interference signal samples with multiple channel groups, and introduce a self-attention mechanism to fuse the processed medium-scale features to obtain the medium-scale multi-layer information representation; The third adaptive convolution module is used to integrate the medium-scale multi-layer information representation and global semantic information of each active interference signal sample to extract the large-scale features of each active interference signal sample; The third depth transposed attention module is used to process subsets obtained by splitting the large-scale features of multiple active interference signal samples using multiple channel groups, and introduce a self-attention mechanism to fuse the processed large-scale features, obtaining large-scale multi-layer information representations as sample optimized fusion features.
4. The radar active interference recognition method according to claim 2, characterized in that The fully connected module includes a global average pooling layer, a first fully connected layer, and a second fully connected layer connected in sequence.
5. The radar active interference recognition method according to claim 1, wherein, Performing time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample, including: Using the CWD transformation method to perform time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample.
6. The radar active interference recognition method according to claim 5, characterized in that The CWD transformation method is implemented using the following formula: Among them, J CWD (t,f) represents the time-frequency diagram sample of the interference signal; A(η,τ) represents the ambiguity function; exp(·) represents the exponential function of e; j represents the imaginary unit; t represents time; f represents the frequency of the active interference signal sample; s(·) represents the time-domain form of the active interference signal sample; s * (·) represents the conjugate of s(·); Φ(η,τ) represents the kernel function; η is used to determine the frequency resolution of the kernel function; τ represents the time shift term.
7. A radar active interference recognition device based on a CNN-ViT hybrid network, characterized in that, The radar active interference recognition device includes: A construction module for constructing an active interference signal dataset; the active interference signal dataset includes multiple active interference signal samples and corresponding true classification labels for each active interference signal sample; A time-frequency transformation module for performing time-frequency transformation on each active interference signal sample in the active interference signal dataset to obtain an interference signal time-frequency map sample corresponding to the active interference signal sample; An identification module for performing multi-scale feature extraction and fusion and interference identification operations on each interference signal time-frequency map sample using a pre-constructed CNN-ViT hybrid network to obtain a predicted sample classification result; the CNN-ViT hybrid network is constructed based on an adaptive convolution kernel and a depth transposed attention module; A training module for training the CNN-ViT hybrid network according to the difference between the true classification label corresponding to each active interference signal sample and the predicted sample classification result, so as to identify radar active interference according to the trained CNN-ViT hybrid network.
8. The radar active interference recognition device according to claim 7, wherein The CNN-ViT hybrid network includes: An initial convolution module for extracting features from each active interference signal sample in the active interference signal dataset to obtain initial sample image features of each active interference signal sample; A multi-layer multi-scale fusion processing module for performing multi-scale fusion on the initial sample image features of each active interference signal sample to obtain multi-scale fusion features, and introducing a self-attention mechanism to perform weighted fusion on the multi-scale features to obtain sample optimized fusion features of each active interference signal sample; A fully connected module for classifying each active interference signal sample based on the sample optimized fusion features of each active interference signal sample to obtain a predicted sample classification result.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the radar active interference recognition method according to any one of claims 1-6 when executing the computer programs stored on the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the radar active interference recognition method according to any one of claims 1-6.