A radar working mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer

By improving the combination of wavelet transform and visual Transformer model, the problems of energy diffusion and insufficient recognition accuracy in radar working mode recognition are solved, and high-resolution multi-scale representation and accurate recognition are achieved.

CN120993333BActive Publication Date: 2026-01-23成都富元辰科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511529066.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-23
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing radar operating mode recognition methods suffer from energy diffusion when processing complex radar signals, making it difficult to accurately extract key modulation features. Furthermore, traditional methods lack sufficient accuracy and robustness in low signal-to-noise ratio and complex electromagnetic environments.

Method used

An improved wavelet transform is used to extract time-frequency map features from three different scales: between pulse groups, within pulse groups, and within pulses. A three-channel time-frequency representation image is constructed and trained using a visual Transformer model. Feature fusion is then performed by combining block embedding, block merging, and BiFormer modules.

Benefits of technology

It achieves high-resolution, multi-scale characterization of complex radar signals, improves recognition accuracy and robustness, reduces computational complexity, and can accurately identify radar operating modes under strong frequency variations and complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120993333B_ABST
    Figure CN120993333B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of radar signal processing, and discloses a radar working mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer, which comprises the following steps: performing improved wavelet transformation on the acquired radar signal, extracting time-frequency graph features of the signal from three different scales; constructing a data set of a radar working mode recognition model; constructing a block embedding module, a block merging module and a BiFormer module in the visual Transformer model; dividing the data set into a training set, a verification set and a test set according to a 3:1:4 ratio to obtain a trained radar working mode recognition model; the model outputs corresponding radar working mode recognition results according to input time-frequency graph features; and through the adoption of the WTMSST technology and iterative optimization of group time delay estimation, the energy diffusion problem existing in the analysis of strong frequency change radar signals by traditional time-frequency methods is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of radar signal processing, and more particularly to a radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer. BACKGROUND

[0002] With the rapid development of phased array radar technology, its flexible beam pointing, waveform modulation and multi-task processing capability greatly improve the adaptability and complexity of the radar system. Radar operating mode recognition is a key technology, which aims to accurately identify different operating modes of the radar through in-depth analysis of the radar signal. This technology not only requires extracting effective time-domain and frequency-domain features from the radar signal, but also needs to combine the pattern recognition algorithm for in-depth learning and training to cope with the increasingly complex environment and high dynamic changes of the radar operating state.

[0003] The existing radar operating mode recognition method mainly faces the problem that the traditional time-frequency analysis method (such as STFT, CWT) has a serious energy diffusion phenomenon when processing complex radar pulse signals with strong frequency changes (such as linear frequency modulation, nonlinear frequency modulation, phase-coded transient), which leads to ambiguity of time-frequency ridge line and difficulty in accurately extracting key modulation features. Most methods only extract features from a single time scale (such as within a pulse), which is difficult to fully capture the coordinated evolution characteristics of radar signals in three key scales: between pulse groups (mode switching), within pulse groups (pulse sequence regularity), and within pulses (fine modulation). Traditional machine learning or shallow neural network models are sensitive to changes in time-frequency features under low signal-to-noise ratio and complex electromagnetic environment interference, and have limited generalization ability, making it difficult to meet the requirements of recognition accuracy and robustness.

[0004] The wavelet-based time reassignment synchronous compression transform significantly improves the time-frequency energy concentration by fusing time reassignment and multi-synchronous compression strategies, and introducing group time delay estimation and iterative optimization mechanism, providing a new idea for complex radar signal preprocessing. However, applying it to multi-scale joint representation and driving high-performance recognition models is still a problem to be solved. SUMMARY

[0005] In order to overcome the above-mentioned defects of the prior art, the present application provides a radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer to solve the problems existing in the background art.

[0006] The present application provides the following technical solutions: a radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer, comprising the following steps:

[0007] Step one, the acquired radar signal is improved wavelet transform, the time-frequency feature of the signal is extracted from the three different scales of pulse group, pulse group and pulse;

[0008] Step two, the time-frequency graph of each scale is respectively grayed and size compressed, and a three-channel time-frequency representation image is constructed, which reflects the change rule of the signal in different time and frequency domain, and constitutes the data set of the radar working mode recognition model;

[0009] Step three, a block embedding module, a block merging module and a BiFormer module in the visual Transformer model are constructed;

[0010] Step four, the data set is divided into training set, validation set and test set according to the ratio of 3:1:4, and the training data set is input into the visual Transformer model for training to obtain the trained radar working mode recognition model;

[0011] Step five, the test set is input into the trained radar working mode recognition model, and the model outputs the corresponding radar working mode recognition result according to the input time-frequency graph feature.

[0012] Preferably, in step one, the acquired radar signal is improved wavelet transform, and the time-frequency feature of the signal is extracted from the three different scales of pulse group, pulse group and pulse, and the specific process is as follows:

[0013] The radar signal is represented by the multi-component form of the frequency domain eigenmode function modeling, that is:

[0014] ; wherein, represents the frequency domain of the radar signal; represents the signal amplitude, represents the component index; represents the total number of signal components; represents the imaginary unit; represents the phase; represents the angular frequency;

[0015] The atom is modeled by a scaling factor and a translation factor , that is, ; wherein, represents the mother wavelet function; The atom is the basic analysis unit of wavelet transform; the of the signal is:

[0016] ; wherein, denotes the complex conjugate; denotes the complex conjugate of ; denotes the space of square integrable functions; if the wavelet function is analytic, it is defined as being modulated by a real window function ; wherein denotes the center frequency; then there is

[0017] ;

[0018] When the regular expression takes into account an additional phase shift , the function definition of the improved wavelet transform is:

[0019] ; wherein denotes a local frequency parameter related to the position parameter b;

[0020] Let and consider the Parseval theorem, which leads to the frequency domain form of :

[0021] ; wherein denotes the frequency domain integration variable; denotes the complex conjugate of

[0022] is calculated by the following formula:

[0023] ; wherein there is a conversion relationship between and ; that is: ; thus there is

[0024] ; wherein denotes the frequency domain integration variable.

[0025] Preferably, the scale map of is described as:

[0026] ; wherein denotes the probability distribution function around the point ; thus, the centroid of the lower signal of is defined as:

[0027] ;

[0028] ;

[0029] where, denotes the group delay estimate; denotes the real part operation; denotes the windowed ; ; denotes the windowed ; .

[0030] The time-reassigned synchrosqueezing transform is preferably defined by the operator , i.e.:

[0031] ; where, denotes the synchrosqueezing transform result; denotes the Dirichlet function; denotes the set of parameters for which the wavelet coefficients are non-zero; ; the WTSST transforms the time-scale coefficients from the point to the new point ; and the WTSST preserves the ability to recover the original signal, i.e.:

[0032] ; for ; the performance of the WTSST in the case of strong frequency- varying group delay is analyzed; there is a signal model whose phase is locally expanded by the second-order Taylor formula, i.e. , ; the resulting is obtained according to the following formula: ;

[0033] According to the Parseval theorem, this is rewritten as:

[0034] ; where, denotes the signal frequency-domain derivative; denotes the scale dependent frequency shift parameter; the local group delay candidates of are obtained, i.e.:

[0035] ; where, , ; where, denotes the first-order derivative of the phase function; denotes the second-order derivative of the phase function; denotes the Taylor expansion coefficient.

[0036] Preferably, a fixed-point iteration strategy is introduced to compensate and​ WTMSST technique is introduced and expressed as:

[0037] ;

[0038] ;

[0039] ;

[0040] wherein, i.e. equivalent to , is the iteration number, such that ;

[0041] The relationship between and is analyzed, and is substituted into and combined with the Fubini theorem to obtain:

[0042] ; wherein, denotes the time integral variable.

[0043] Preferably, the step two is gray-scale processing and size compression of the time-frequency map of each scale respectively, to construct a three-channel time-frequency representation image, and the time-frequency map constitutes the data set of radar mode recognition, and the process is as follows:

[0044] Gray-scale processing, each time-frequency distribution is normalized and mapped to the gray value range :

[0045] ; wherein, denotes the maximum value of the current distribution, denotes the minimum value of the current distribution, denotes the pixel value after gray-scale processing; denotes the time coordinate; denotes the frequency coordinate; denotes the rounding function;

[0046] Size compression adjusts each gray-scale image to a fixed size of 224x224, denoted as:

[0047] ; three-channel time-frequency representation construction constructs the gray-scale images under three scales as the RGB channels of a picture respectively, to generate a new three-channel time-frequency representation image:

[0048] ;

[0049] Finally, a time-frequency representation with a size of 224x224x3 is obtained; each time-frequency graph is labeled according to its corresponding radar operating mode to generate a labeled sample, and the processed time-frequency graphs are classified according to the radar operating mode to construct a time-frequency graph dataset containing multiple operating modes.

[0050] Preferably, the visual Transformer model based on the step three is obtained by constructing a block embedding, block merging and a BiFormer module. The input time-frequency graph is divided into sub-blocks according to the block size PxP, and each sub-block is mapped to a D-dimensional embedding vector through linear projection:

[0051] ; wherein, is a learnable projection matrix, represents a sub-block vectorization operation; represents a bias vector; represents a sub-block index; represents the embedding vector; represents the total number of sub-blocks; represents the input sub-block;

[0052] For the feature map of the layer, it is merged into super blocks according to the 2x2 field, and each super block generates a high-dimensional feature through channel splicing and linear transformation:

[0053] ; wherein, the output feature dimension of the layer is , represents a linear transformation operation; represents a channel splicing operation; represents a row index range; represents a column index range; represents a column index variable; represents the 2x2 field feature block of the 1st layer;

[0054] The overall structure of the BiFormer module can be formally represented as:

[0055] ; wherein, represents the output feature map after pooling; represents the input feature map; represents a convolution operation; represents a deep convolution operation; ​denotes adaptive global pooling; denotes two-stage routing attention mechanism; denotes layer normalization;

[0056] The formula of adaptive global pooling is represented as:

[0057] .

[0058] Preferably, in the fourth step, the data set is divided into training set, validation set and test set in the proportion of 3:1:4, and the training data set is input into the visual Transformer model for training to obtain the trained model.

[0059] Preferably, in the fifth step, the test set is input into the trained working mode recognition model, and the model will output the corresponding radar working mode recognition result according to the input time-frequency feature.

[0060] Technical effects and advantages of the present application:

[0061] By adopting the WTMSST technology, the present application is advantageous for iterative optimization of group time delay estimation, effectively solves the energy dispersion problem existing in the analysis of strong frequency change radar signals by the traditional time-frequency method, and obtains a focused accurate and clear contour time-frequency diagram, thereby laying a high-quality underlying feature foundation for recognition.

[0062] The present application fuses and encodes the time-frequency information of three key scales (inter-pulse group (macroscopic mode), intra-pulse group (mesoscopic correlation) and intra-pulse (microscopic modulation)) into a color image (RGB channel) through a three-channel time-frequency representation framework. This representation naturally contains the multi-level space-time evolution characteristics of the radar signal, and provides a more comprehensive and rich basis for recognition.

[0063] Through the recognition model based on visual Transformer constructed by the present application, the calculation complexity is greatly reduced through efficient block processing and hierarchical feature fusion, and the two-stage routing attention mechanism can dynamically focus on the most discriminative time-frequency area, effectively eliminating the influence of noise and interference. BRIEF DESCRIPTION OF DRAWINGS

[0064] Fig. 1 The present application is a radar working mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer.

[0065] Fig. 2 The present application is a multi-scale three-channel time-frequency representation framework based on WTMSST.

[0066] Fig. 3 The present application is a visual Transformer recognition model framework.

[0067] Fig. 4 Figure 2 is a comparison curve of recognition accuracy of the present application and other methods under different signal-to-noise ratios. DETAILED DESCRIPTION

[0068] The technical solutions in the present application will be described clearly and completely in combination with the drawings in the present application. In addition, the forms of each structure described in the following embodiments are only examples, and the radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer involved in the present application is not limited to each structure described in the following embodiments. All other embodiments obtained by those skilled in the art without creative labor belong to the scope of protection of the present application.

[0069] As shown in Figs. 1 to 4 The present application provides a radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer, comprising the following steps:

[0070] Step one, performing improved wavelet transform on the acquired radar signal to extract time-frequency graph features of the signal from three different scales of inter-pulse group, intra-pulse group and intra-pulse;

[0071] Step two, respectively performing gray scale processing and size compression on the time-frequency graph of each scale to construct a three-channel time-frequency representation image, wherein the time-frequency graph reflects the change rule of the signal in different time and frequency domains to constitute a data set of the radar operating mode recognition model;

[0072] Step three, constructing a block embedding module, a block merging module and a BiFormer module based on the visual Transformer model;

[0073] Step four, dividing the data set into a training set, a validation set and a test set according to a ratio of 3:1:4, and inputting the training data set into the visual Transformer model for training to obtain the trained radar operating mode recognition model;

[0074] Step five, inputting the test set into the trained radar operating mode recognition model, and the model outputs the corresponding radar operating mode recognition result according to the input time-frequency graph features.

[0075] In the present embodiment, it needs to be specifically pointed out that in step one, the acquired radar signal is subjected to improved wavelet transform to extract time-frequency graph features of the signal from three different scales of inter-pulse group, intra-pulse group and intra-pulse, and the specific process is as follows:

[0076] The radar signal is represented in the form of multiple components modeled by the frequency domain eigenmode function, that is:

[0077] ; wherein, denotes the frequency domain of the radar signal; denotes the signal amplitude, denotes the component index; denotes the total number of signal components; denotes the imaginary unit; denotes the phase; denotes the angular frequency;

[0078] Conventional linear time-frequency analysis methods relate the signal to a dictionary of waveforms, wherein, denotes a waveform function in the dictionary of waveforms; denotes a set of multi-index parameters; denotes the definition domain of the parameter set; the dictionary of waveforms consists of a series of waveform functions with certain time-frequency localization capability, i.e. ; wherein, denotes the time-frequency analysis result; denotes the time parameter; the inner product result shows that the position and size of the Heisenberg box depend on the time-frequency center and span, when the time-frequency index varies in , the Heisenberg box covers the entire time-frequency plane, denotes a set of real numbers;

[0079] An atom is modeled by a scaling factor and a translation factor , i.e. ; wherein, denotes the mother wavelet function; The atom is the basic analysis unit of the wavelet transform; the signal has :

[0080] ; wherein, denotes the wavelet transform coefficient; * is the conjugate complex representation; denotes the conjugate complex of ; denotes the square integrable function space; assuming that the wavelet function is analytic means that it can be defined as a real window function modulated by , i.e. ; wherein, denotes the center frequency; then,

[0081] ;

[0082] When the regular ​Expression takes into account additional phase shift The function definition of the improved wavelet transform is given as:

[0083] ; wherein, denotes the frequency modulation parameter related to the position parameter b;

[0084] Let , and considering the Parseval theorem, the frequency domain form of is obtained as:

[0085] ; wherein, denotes the frequency domain integral variable; denotes the conjugate complex of ;

[0086] It is calculated by the following formula:

[0087] ; wherein, there is a conversion relationship between , that is ; that is: ; thus obtaining:

[0088] ; wherein, denotes the frequency domain integral variable;

[0089] Considering the serious energy diffusion phenomenon in , it is necessary to further explore the energy distribution law. According to the Plancherel theorem, the scale diagram of can be described as:

[0090] ; wherein, denotes the probability distribution function around the point ; thus, the centroid of the lower signal of

[0091] ;

[0092] ;

[0093] wherein, denotes the group delay estimation value; denotes the real part operation; denotes with window ; denotes with window ;

[0094] To improve the energy concentration and keep the reconstruction ability, a post-processing strategy along the time direction is considered, thus, the time reassignment synchronous synchrosqueezing transform based on can be defined by the operator , i.e.,

[0095] ; where, represents the synchronous synchrosqueezing transform result; represents the Dirichlet function; represents the parameter set of non-zero wavelet coefficients; ; the essence of WTSST is a cumulative process, by converting the time-scale coefficients from the point to the new point ; and WTSST retains the ability to recover the original signal, i.e.,

[0096] ; for , if the selected window satisfies the time-band limit condition, the time-frequency distribution of the non-zero coefficients of WTSST is concentrated in the narrow band near the trajectory, and WTSST is indeed a reliable tool when dealing with weakly frequency-modulated signals; however, radar pulse signals usually exhibit more complex modulation laws, which may make WTSST no longer have resolution, so the performance of WTSST in the case of strong frequency-modulated groups is analyzed; consider a signal model whose phase can be locally expanded by the second-order Taylor formula, i.e. , ; the resulting can be obtained according to the following formula: ;

[0097] According to the Parseval theorem, it can be rewritten as:

[0098] ; where, represents the signal frequency domain result; represents the scale related frequency offset parameter; the local group delay candidate of can be obtained, i.e.

[0099] ; where, , ; where, represents the first derivative of the phase function; represents the second derivative of the phase function; represents the Taylor expansion coefficient;

[0100] In order to conveniently observe the error of the group delay estimator, the Gaussian window is selected to specify the equation, i.e. ; re-introducing the fixed-point iteration strategy to compensate and error, thus introducing the WTMSST technique and expressing it as:

[0101] ;

[0102] ;

[0103] ;

[0104] where, i.e., equivalent to , is the iteration number, such that ;

[0105] To obtain the regularity in multiple iterations, it is necessary to analyze the relationship between and , substitute into and combine the Fubini theorem to obtain:

[0106] ; where, denotes the time integral variable;

[0107] The above formula shows that the iteration operation of is equivalent to compressing the WT coefficients into a newly generated group delay candidate If the iteration step continues, the group delay candidate can always be updated, and is further introduced to replace the original group delay candidate in WTSST, and the expression of WTMSST is obtained as:

[0108] ; can be described as:

[0109] ;

[0110] When the iteration number is large enough, the error between the new group delay candidate and may be close to 0, i.e., ; means that the new operator is more suitable for processing strong frequency change signals than In addition, WTMSST can accurately locate the energy ridge, i.e.:

[0111] ;

[0112] In addition, WTMSST not only improves the resolution of WTSST, but also retains its reconstruction ability, i.e.​

[0113] .

[0114] In this embodiment, it needs to be specifically pointed out that in step two, the gray scale processing and size compression are respectively performed on each scale time-frequency graph to construct a three-channel time-frequency representation image. These time-frequency graphs constitute a data set for radar working mode recognition, and the process is as follows:

[0115] Gray scale processing, that is, normalizing each time-frequency distribution to a gray value range :

[0116] ; wherein, represents the maximum value of the current distribution, represents the minimum value of the current distribution, represents the pixel value after gray scale processing; represents the time coordinate; represents the frequency coordinate; represents the rounding function;

[0117] Size compression, that is, adjusting each gray scale image to a fixed size of 224x224, denoted as:

[0118] ; three-channel time-frequency representation construction, that is, taking the gray scale images under three scales as the RGB channels of a picture respectively to generate a new three-channel time-frequency representation image:

[0119] ;

[0120] Finally, a time-frequency representation with a size of 224x224x3 is obtained; each time-frequency graph is labeled according to its corresponding radar working mode to generate a sample with a label. The processed time-frequency graphs are classified according to the radar working mode to construct a time-frequency graph data set containing multiple working modes.

[0121] In this embodiment, it needs to be specifically pointed out that in step three, a visual Transformer model is obtained by constructing patch embedding, block merging and BiFormer module, and the process is as follows:

[0122] In order to discretize the continuous time-frequency graph into a sequence feature suitable for Transformer processing while retaining the local time-frequency structure information, a patch embedding (PE) strategy is introduced here. For the input time-frequency graph , it is divided into sub-blocks according to the patch size PxP, and each sub-block is mapped to a D-dimensional embedding vector through linear projection:

[0123] , ; wherein, is a learnable projection matrix, denotes a sub-block vectorization operation; denotes a bias vector; denotes a sub-block index; denotes the th embedding vector; denotes the total number of sub-blocks; denotes the th input block;

[0124] The effect of the block-wise embedding is to capture the energy distribution characteristics of adjacent regions in the time-frequency graph through the block operation, and to simplify the pixel-level attention of quadratic complexity into block-level attention of linear complexity, and the operation supports a multi-scale block strategy to adapt to different radar time-frequency representation resolution requirements.

[0125] In order to continuously perform spatial downsampling and channel expansion on the feature map in the deep network, a layer-by-layer block merging using a pyramid structure is used, cross-scale feature fusion is achieved through spatial compression, and for the feature map of the th layer, , it is merged into super blocks according to a 2x2 field, and each super block generates a high-dimensional feature through channel splicing and linear transformation:

[0126] ; wherein, the output feature dimension of the th layer is , denotes a linear transformation operation; denotes a channel splicing operation; denotes a row index range; denotes a column index range; denotes a column index variable; denotes a 2x2 field feature block of the 1st layer;

[0127] The significance of the block merging operation is to gradually expand the time-frequency range covered by the features through hierarchical downsampling, to realize the deep network to capture the pattern shift between pulses, the shallow network to analyze the modulation details within the pulse, and to reduce the sequence length with increasing depth, balancing the attention calculation overhead.

[0128] The overall structure of the BiFormer module can be formally represented as:

[0129] ; wherein, denotes a pooled output feature map; denotes an input feature map; denotes a convolution operation; denotes a deep convolution operation; AGP denotes adaptive global pooling; BiRA denotes bi-level routing attention mechanism; LN denotes layer normalization;

[0130] Adaptive global pooling is a pooling method that converts input feature maps of different sizes into fixed-size outputs. Unlike traditional global pooling, AGP can automatically adjust the parameters of the pooling operation according to the size of the input feature map, so it can work effectively on inputs of different sizes. In this way, BiFormer can adapt to different input sizes without losing important spatial information. The formula of adaptive global pooling can be expressed as:

[0131] In the BiFormer module, AGP is used to reduce computational complexity while preserving the global information of the input feature map, and it plays a crucial role in extracting multi-scale features.

[0132] Depthwise separable convolution is an efficient convolution calculation method in convolutional neural networks. Traditional convolution calculates between each input channel and output channel, while depthwise separable convolution divides the convolution operation into two steps: first, perform channel-wise convolution (depthwise convolution), then perform point-wise convolution (1x1 convolution) between output channels. This method significantly reduces the amount of calculation, making the model more efficient.

[0133] Depthwise separable convolution first performs depthwise convolution operation, then point-wise convolution operation. The formula is expressed as:

[0134] ; where, denotes point-wise convolution operation, through this decomposition way, the amount of calculation of convolution operation is reduced, the model calculation is more efficient. The depthwise separable convolution in BiFormer is used to extract features, effectively reducing the computational complexity and maintaining the richness of feature representation.

[0135] Residual connection helps to alleviate the problem of gradient vanishing in deep networks by directly adding the input to the output. Specifically, in the BiFormer module, residual connection allows the features of the previous layer to be directly passed to the subsequent layer, thereby speeding up the training process and improving the performance of the model. It can improve the expression ability of the network and ensure that information is not lost between multiple layers.

[0136] Layer normalization is a commonly used normalization method, mainly used in deep learning to improve the stability of training. Unlike batch normalization, layer normalization operates on each sample individually, rather than in batches, making it suitable for small batch and dynamic input scenarios. Layer normalization can effectively alleviate the problem of training instability caused by changes in activation value distribution. In the BiFormer module, layer normalization is used to ensure the stability of the input feature mean and variance of each layer, thereby improving the training efficiency of the model.

[0137] Multi-layer perceptron (MLP) is a type of feedforward neural network that contains multiple hidden layers and is commonly used to learn non-linear mapping relationships. In BiFormer, MLP is used to perform higher-level non-linear combination on the extracted features. Through the combination of multiple neuron layers, the model can capture more complex data patterns.

[0138] The output of the multi-layer perceptron can be represented by the following formula:

[0139] ; wherein, represents the output after processing by the multi-layer perceptron; and represents the weight matrix; represents the input feature; and represents the bias term, represents the activation function, through the non-linear transformation of multiple layers, the model can capture complex patterns.

[0140] In this embodiment, it needs to be specifically pointed out that in step four, the data set is divided according to the training set, validation set and test set ratio of 3:1:4, and the training data set is input into the visual Transformer model for training to obtain the trained model.

[0141] In step five, the test set is input into the trained working mode recognition model, and the model will output the corresponding radar working mode recognition result according to the input time-frequency feature.

[0142] The present application improves the acquired radar signal by performing improved wavelet transform to obtain three-scale time-frequency graph data, and performs gray scale processing and channel splicing on the three-scale time-frequency graph data to form a multi-channel time-frequency representation image. These images constitute the data set of radar working mode recognition. Then input the training set in the data set for training into the visual Transformer model for training to obtain the trained recognition model. Finally, the test set in the data set for testing is input into the recognition model as the input, and the visual Transformer model is used to identify the current radar working mode and output the result.

[0143] Finally: the above only for the preferred embodiments of the present application, and not for limiting the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application, should be included in the scope of protection of the present application.

[0144] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer, characterized in that: Includes the following steps: Step 1: Perform improved wavelet transform on the acquired radar signal to extract the time-frequency characteristics of the signal at three different scales: between pulse groups, within pulse groups, and within pulses. Step 2: Perform grayscale processing and size compression on the time-frequency maps at each scale to construct a three-channel time-frequency representation image. The time-frequency maps reflect the variation of the signal in different time and frequency domains, forming the dataset for the radar working mode recognition model. Step 3: Construct the block embedding module, block merging module, and BiFormer module based on the visual Transformer model; Step 4: Divide the dataset into training set, validation set and test set in a ratio of 3:1:4, and input the training dataset into the visual Transformer model for training to obtain the trained radar working mode recognition model. Step 5: Input the test set into the trained radar operating mode recognition model. The model outputs the corresponding radar operating mode recognition result based on the input time-frequency map features. In step three, a visual Transformer-based model is obtained by constructing a block embedding, block merging, and BiFormer module for the input time-frequency graph. Divide it into blocks according to the block size P×P. Each of the following sub-blocks is mapped to a D-dimensional embedding vector via linear projection: , ;in, For learnable projection matrices, Indicates sub-block vectorization operation; Represents the bias vector; Indicates a sub-block index; Indicates the first Embedded vectors; Indicates the total number of sub-blocks; Indicates the first One input sub-block; For the Feature map of layer Merge into 2×2 domains Each superblock generates high-dimensional features through channel concatenation and linear transformation: Among them, the first The output feature dimension of the layer is , Represents a linear transformation operation; Indicates a channel splicing operation; Indicates the range of row indexes; Indicates the range of column indexes; Indicates column index variable; This represents a 2×2 neighborhood feature block in the first layer; The overall structure of the BiFormer module can be formally represented as follows: ;in, This represents the output feature map after pooling; Indicates the input feature map; Indicates the convolution operation; Indicates a depthwise convolution operation; Indicates adaptive global pooling; This indicates a two-level routing attention mechanism; Representation layer normalization; The formula for adaptive global pooling is expressed as: 。 2. The radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer according to claim 1, characterized in that: In step one, the acquired radar signal undergoes an improved wavelet transform to extract time-frequency features from three different scales: inter-pulse group, intra-pulse, and intra-pulse. The specific process is as follows: The radar signal is represented in a multi-component form by modeling the frequency domain eigenmode functions, namely: ;in, Represents the frequency domain of radar signals; Indicates signal amplitude. Indicates component index; Indicates the total number of signal components; Represents the imaginary unit; Indicates phase; Indicates angular frequency; Atoms are proportional to factors Translation factor Modeling, i.e. ;in, Represents the mother wavelet function; Atoms are the basic analytical units of wavelet transform; signals of for: ;in, Represents wavelet transform coefficients; * denotes conjugate complex number representation; express The conjugate of complex numbers; Let represent the space of square-integrable functions; if the wavelet function is analytic, then it is defined as [a_n] = [a_n] * ... Modulated real window function ,Right now ;in, Representing the center frequency; then we have: ; When regular expression The expression takes into account additional phase shift At that time, the function definition of the improved wavelet transform is: ;in, This represents the local frequency parameter related to the position parameter b; set up And considering Passevar's theorem, we obtain Frequency domain form: ;in, Represents the frequency domain integral variable; express The conjugate of complex numbers; It is calculated using the following formula: ;in, and There is a conversion relationship between them, that is ;Right now: Therefore, we get: ;in, This represents the frequency domain integral variable.

3. The radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer according to claim 2, characterized in that: The scale map Described as: ;in, Indicates the area around the point The probability distribution function; The centroid of the lower signal is defined as: ; ; in, This represents the group delay estimate; This indicates the operation of taking the real part; Indicates that there is a window of ; Indicates that there is a window of .

4. The radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer according to claim 3, characterized in that: based on Time redistribution synchronous compression transform operator To define, that is: ;in, This indicates the result of synchronous compression transformation; Represents the Diclave function; The set of parameters representing non-zero wavelet coefficients; WTSST changes the time scale coefficient from point... Convert to new point Furthermore, WTSST retains the ability to recover the original signal, that is: ;for ;Analyze the performance of WTSST under strong frequency variation group delay conditions;There exists a signal model whose phase is locally expanded using the second-order Taylor formula, i.e. , ; income Obtained using the following formula: ; According to Passevar's theorem, it can be rewritten as: ;in, Represents the frequency domain derivative of the signal; Representing scale The relevant frequency offset parameters; obtain The local group delay candidate is: ;in, , ;in, This represents the first derivative of the phase function; The second derivative of the phase function is represented. This represents the Taylor expansion coefficients.

5. The radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer according to claim 4, characterized in that: Introducing a fixed-point iterative strategy to compensate and To address the error between them, the WTMSST technique is introduced and expressed as: ; ; ; in, That is equivalent to , Let be the number of iterations, such that ; analyze and The relationship between them will Substitution And by combining this with Fubini's theorem, we get: ;in, This represents the time integral variable.

6. The radar operating mode recognition method based on high-resolution multi-scale time-frequency representation and visual Transformer according to claim 5, characterized in that: In step two, the time-frequency images at each scale are converted to grayscale and compressed to construct a three-channel time-frequency representation image. The time-frequency images constitute the dataset for radar operating mode recognition. The process is as follows: Grayscale processing normalizes each time-frequency distribution and maps it to a grayscale value range. : ;in, This represents the maximum value of the current distribution. This represents the minimum value of the current distribution. This represents the pixel value after grayscale processing; Represents the time coordinate; Represents frequency coordinates; This represents the rounding function; Size compression resizes each grayscale image to a fixed size of 224×224, denoted as: The three-channel time-frequency representation construction uses grayscale images at three different scales as the RGB channels of a single image to generate a new three-channel time-frequency representation image. ; The final time-frequency representation is 224×224×3. Each time-frequency map is labeled according to its corresponding radar operating mode, generating a labeled sample. The processed time-frequency maps are classified according to radar operating modes to construct a time-frequency map dataset containing multiple operating modes.

Citation Information

Patent Citations

  • Radar working mode identification method based on multi-scale visual Transform

    CN117746163A

  • Radar intra-pulse modulation type identification method and system based on lightweight deep learning

    CN119001623A