Target recognition method based on adaptive polynomial kernel function and multimodal fusion
By employing an adaptive polynomial kernel function and multimodal fusion method, combined with improved polynomial frequency modulated wavelet transform, Wigner-Ville distribution, and short-time Fourier transform, the time-frequency resolution contradiction and cross-term interference problems of nonlinear IF signals are resolved, achieving high-precision target recognition and improved signal-to-noise ratio.
Patent Information
- Application Number
- CN202511447277.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing time-frequency analysis methods suffer from time-frequency resolution contradictions, cross-term interference, and adaptability issues when processing nonlinear IF signals, making it difficult to achieve high-precision analysis and estimation.
An adaptive polynomial kernel function and multimodal fusion approach is adopted, which combines an improved polynomial frequency modulated wavelet transform (IPCT) with an adaptive kernel function, Wigner-Ville distribution (WVD) and short-time Fourier transform (STFT), and uses a convolutional neural network and gradient boosting regression tree model for signal processing.
It improves the processing accuracy of nonlinear IF signals, enhances adaptability in complex environments, improves the signal-to-noise ratio, and increases the accuracy of target recognition.
Smart Images

Figure CN120910813B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target recognition technology, specifically relating to a target recognition method based on adaptive polynomial kernel function and multimodal fusion. Background Technology
[0002] With the rapid development and progress of target recognition technology, the demand for complex signal processing in modern recognition systems is growing exponentially. Signals often exhibit nonlinear and time-varying characteristics, and accurate assessment of instantaneous frequency has become the key to analyzing the essence of signals and ensuring the quality of subsequent feature extraction and recognition decisions.
[0003] Currently, the evaluation methods for instantaneous frequency in target recognition are showing a diversified development trend, mainly including short-time Fourier transform (STFT), continuous wavelet transform (CWT), Wigner-Ville distribution (WVD), and linear frequency modulated wavelet transform (CT), etc. They play their respective roles in different scenarios, but they also have certain limitations.
[0004] Current mainstream time-frequency analysis methods each have limitations when processing nonlinear IF signals in target recognition: Short-Time Fourier Transform (STFT) is limited by the Heisenberg-Gabor inequality, making it difficult to balance time and frequency resolution; Continuous Wavelet Transform (CWT), while offering better time-frequency localization capabilities, still cannot accurately capture rapidly changing frequency components; Wigner-Ville Distribution (WVD) is susceptible to noise and cross-term interference; Linear Frequency Modulated Wavelet Transform (CT) is only applicable to linear IF signals. While Polynomial Frequency Modulated Wavelet Transform (PCT) provides a framework for nonlinear IF signal analysis, its traditional form uses a fixed polynomial kernel function, making it difficult to dynamically match the time-varying characteristics of the signal, and still suffers from insufficient accuracy when processing strongly nonlinear, multi-component target recognition signals.
[0005] In summary, while existing time-frequency distribution-based methods each have their advantages in processing nonlinear IF signals in the field of target recognition, they all have problems to varying degrees, making it difficult to achieve high-precision analysis and estimation. Further improvements and optimizations are still needed based on existing technologies. Summary of the Invention
[0006] To address the time-frequency resolution contradiction, cross-term interference, adaptability issues, and accuracy degradation problems in nonlinear instantaneous frequency signal analysis, this invention proposes a target recognition method based on adaptive polynomial kernel function and multimodal fusion. This invention uses an improved polynomial frequency modulated wavelet transform (IPCT) as its core, designs an adaptive kernel function to enhance nonlinear fitting capabilities, and fuses IPCT, WVD, and STFT in a multimodal manner to maximize the role of IPCT. By utilizing multimodal analysis techniques and combining the advantages of short-time Fourier transform and wavelet transform, accurate target recognition is achieved.
[0007] This invention proposes a target recognition method based on adaptive multinomial kernel function and multimodal fusion, specifically including the following steps:
[0008] The original signal of the target to be identified is obtained, and the original signal is preprocessed based on the wavelet threshold denoising algorithm to generate the signal to be analyzed.
[0009] Multimodal time-frequency analysis is performed on the signal to be analyzed, and the output is a first time-frequency feature map based on an adaptive polynomial kernel function, a second time-frequency feature map based on Wigner-Ville distribution (WVD), and a third time-frequency feature map based on short-time Fourier transform (STFT).
[0010] The first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map are dynamically weighted and fused to generate a fused time-frequency feature tensor, which is then processed by a convolutional neural network to generate a time-series feature vector.
[0011] The temporal feature vectors are input into a convolutional neural network and a gradient boosting regression tree (GBRT) model, respectively, and then the preliminary identification results and feature verification results based on the target category are output.
[0012] Target identification is performed based on the preliminary identification results of the target category and the characteristic verification results.
[0013] Furthermore, the steps for preprocessing the original signal based on the wavelet threshold denoising algorithm to generate the signal to be analyzed specifically include:
[0014] The original signal is decomposed into multiple wavelet components based on its kurtosis. If the kurtosis is greater than 3, the original signal contains impulse noise.
[0015] The power spectral density is obtained by performing a Fourier transform on the original signal, and the power spectral flatness is calculated based on the power spectral density.
[0016] If the power spectrum flatness is approximately equal to 1, then the original signal contains white noise;
[0017] If the power spectral flatness is less than 1, then the original signal contains colored noise;
[0018] The corresponding noise thresholds are set for white noise, colored noise and impulse noise respectively, and wavelet coefficients are processed according to the corresponding noise thresholds to generate a signal with high signal-to-noise ratio to be analyzed.
[0019] Furthermore, the steps of setting corresponding noise thresholds for white noise, colored noise, and impulse noise, and generating a high signal-to-noise ratio signal after wavelet coefficient processing based on the corresponding noise thresholds, specifically include:
[0020] Set corresponding noise thresholds for white noise, colored noise, and impulse noise respectively, as follows:
[0021] White noise using a threshold ;
[0022] Color noise uses dynamic threshold ;
[0023] Impulse noise with kurtosis adaptive hard threshold ;
[0024] Wavelet coefficients smaller than the noise threshold are set to zero or attenuated, while wavelet coefficients larger than the noise threshold are retained.
[0025] The retained wavelet coefficients are then reconstructed using an inverse transform to obtain the signal to be analyzed with a high signal-to-noise ratio.
[0026] in This represents the estimated standard deviation of noise. Indicates the length of the original signal. This represents the estimated noise power spectral density. This is the kurtosis coefficient.
[0027] Furthermore, the step of performing multimodal time-frequency analysis on the signal to be analyzed and outputting a first time-frequency feature map based on an adaptive polynomial kernel function specifically includes:
[0028] After framing and extracting local features from the signal to be analyzed, a local energy distribution is generated. and rate of change of frequency ;
[0029] Construct an adaptive polynomial kernel function containing higher-order terms , represented as:
[0030] ;
[0031] Indicates the first The polynomial coefficients of the frame. Denotes the order of a polynomial. Represents time axis variables, This represents the displacement integral variable in the wavelet transform. Energy weighting factor , for Energy distribution at time, For the first The center moment of the frame, express From frame 1 to frame 2 The maximum value of the local energy distribution within the frame;
[0032] The polynomial coefficients were analyzed using a particle swarm optimization algorithm. Optimization yields the first The optimal polynomial coefficients of the frame, and the optimal adaptive polynomial kernel function;
[0033] Based on the optimal adaptive polynomial kernel function The time-frequency distribution adapted to the nonlinear instantaneous frequency characteristics of the signal to be analyzed is obtained, and the first time-frequency feature map is output.
[0034] Furthermore, the steps of performing multimodal time-frequency analysis on the signal to be analyzed and outputting a second time-frequency feature map based on the Wigner-Ville distribution (WVD) and a third time-frequency feature map based on the short-time Fourier transform (STFT) specifically include:
[0035] The signal to be analyzed is smoothed by Gaussian window using Wigner-Ville distribution (WVD) to suppress cross terms, and after noise suppression, a second time-frequency feature map of the energy concentration region is output.
[0036] The short-time Fourier transform (STFT) is used to adaptively adjust the window length of the signal to be analyzed. If it is a slowly changing signal, a long window is used to improve the frequency resolution, and if it is a fast changing signal, a short window is used to ensure the time resolution. Finally, a third time-frequency feature map is generated as a reference.
[0037] Furthermore, the step of dynamically weighting and fusing the first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map to generate a fused time-frequency feature tensor specifically includes:
[0038] The first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map are set to the same time axis and frequency axis, and then adjusted to the same resolution by interpolation.
[0039] The energies of the first, second, and third time-frequency feature maps after resolution adjustment are normalized to eliminate differences in different energy scales.
[0040] A dynamic weighted fusion method is adopted, and the weights are determined based on the entropy values of the three time-frequency feature maps in different frequency ranges;
[0041] The energy values of the three time-frequency feature maps at each time-frequency point are weighted and summed according to a determined weight to generate a fused time-frequency feature tensor.
[0042] Furthermore, the specific steps for generating time-series feature vectors by processing the fused time-frequency feature tensor through a convolutional neural network include:
[0043] A convolutional neural network (CNN) is used to extract features from the fused time-frequency feature tensor. Local time-frequency patterns are captured by using convolutional kernels of different sizes through convolutional layers. After dimensionality reduction processing by pooling layers, local feature vectors are output.
[0044] The local feature vector is expanded into a temporal feature sequence along the time axis, input into a recurrent neural network (RNN) for temporal modeling, and the temporal feature vector is output through the gating mechanism of the RNN.
[0045] Furthermore, the steps of inputting the temporal feature vectors into the convolutional neural network and the gradient boosting regression tree (GBRT) model respectively for processing and outputting preliminary identification results and feature verification results based on the target category specifically include:
[0046] The temporal feature vector is split into two paths and input into the convolutional neural network and the boosting regression tree GBRT model, respectively.
[0047] The first temporal feature vector passes through the fully connected layer of the convolutional neural network (CNN) and the softmax function to output the target class probability. The class with the highest probability value is taken as the preliminary recognition result, and the highest probability value is taken as the probability confidence level.
[0048] The second-order time-series feature vector is processed by the Gradient Boosting Regression Tree (GBRT) model to output the intermediate frequency estimate. The target feature verification parameters carrying the target information are extracted from the intermediate frequency estimate. The target feature verification parameters are matched with the source feature verification parameters in the pre-built mapping library to generate feature verification results.
[0049] The target characteristic verification parameters include the frequency change rate based on actual values. Mid-frequency range and time-frequency ridge morphology parameters .
[0050] Furthermore, the step of matching the target feature verification parameters with the source feature verification parameters in a pre-built mapping library to generate feature verification results specifically includes:
[0051] A "target category - theoretical mid-frequency characteristic" mapping library is pre-constructed. Each target category in the mapping library is matched with theoretical values of the mean and threshold range of mid-frequency characteristics under different maneuvering states and electromagnetic environments.
[0052] Based on the rate of change of frequency Mid-frequency range and time-frequency ridge morphology parameters The actual value is matched with the theoretical value in the mapping library to generate a value based on the frequency change rate error. Mid-frequency range overlap Ridge morphology The results of the characteristic verification;
[0053] in, , This represents the mean of the theoretical frequency change rate of the initially identified categories. The smaller the value, the higher the match. , This represents the theoretical value for the intermediate frequency range. The larger the value, the higher the match. , Indicates the number of ridge sampling points. , Representing the actual and theoretical ridge lines respectively Point shape value, The smaller the value, the higher the match.
[0054] Furthermore, the specific steps for target identification based on the preliminary identification results and feature verification results of the target category include:
[0055] Obtain the probability confidence level of the preliminary identification results and the frequency change rate error of the feature verification results. Mid-frequency range overlap Ridge morphology ;
[0056] like And the probability confidence level of the preliminary identification results If the target recognition result is accurate, the final target category and recognition confidence level will be output.
[0057] like Or, the probability confidence level If so, the recognition result is deemed invalid;
[0058] The recognition confidence level is expressed as the probability confidence level minus the frequency change rate error. Mid-frequency range overlap and ridge line morphology The weighted average calculation result.
[0059] The present invention has the following technical effects:
[0060] (1) This invention proposes an improved adaptive polynomial frequency modulated wavelet transform (IPCT). By setting an adaptive polynomial kernel function and using it in the processing of nonlinear IF signals for target recognition, local segments are obtained by performing frame processing on the signal to be analyzed, and then the local energy distribution and frequency change rate are calculated. Then, the polynomial coefficients are dynamically adjusted using the particle swarm optimization algorithm to achieve accurate fitting of the nonlinear IF signal.
[0061] (2) This invention designs differentiated thresholds for different types of noise (fixed threshold for white noise, frequency-dependent threshold for colored noise, and adaptive threshold for impulse noise kurtosis), and combines neighborhood smoothing to eliminate residual distortion of impulse noise. Compared with traditional fixed threshold denoising, the signal-to-noise ratio is improved by 3-5dB in mixed noise scenarios, providing a higher quality signal foundation for subsequent time-frequency analysis.
[0062] (3) By integrating the advantages of multimodal time-frequency features, this invention makes up for the limitations of a single method, effectively improves the estimation accuracy of nonlinear instantaneous frequency signals and intermediate frequencies, and enhances adaptability in complex environments. Attached Figure Description
[0063] Figure 1 This is a flowchart of the target recognition method based on adaptive polynomial kernel function and multimodal fusion proposed in this invention;
[0064] Figure 2 This is a schematic diagram of the three time-frequency feature maps output after multimodal time-frequency analysis in an embodiment of the present invention;
[0065] Figure 3 This is a schematic diagram of the fused time-frequency feature tensor output after fusing three time-frequency feature maps in an embodiment of the present invention. Detailed Implementation
[0066] The present application will now be described in further detail with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the specific embodiments.
[0067] like Figure 1 As shown, this embodiment provides a target recognition method based on adaptive polynomial kernel function and multimodal fusion, aiming to improve the processing accuracy of nonlinear IF signals. The target recognition method specifically includes the following steps:
[0068] The original signal of the target to be identified is obtained, and the original signal is preprocessed based on the wavelet threshold denoising algorithm to generate the signal to be analyzed.
[0069] Since noise and clutter (often a mixture of white noise, colored noise, and impulse noise) are inevitably present in the transmission and reception of signals during target recognition, an adaptive noise classification and dynamic threshold optimization wavelet threshold denoising algorithm is used to preprocess the original signal. This algorithm uses wavelet transform to decompose the signal into different frequency sub-bands. Based on the differences in the characteristics of target echoes and noise and clutter in the wavelet domain, an appropriate threshold is set to process the wavelet coefficients, effectively suppressing environmental noise and providing a high-quality signal foundation for subsequent time-frequency analysis.
[0070] In this embodiment, the original signal is decomposed into multiple wavelet components based on the kurtosis of the original signal. The kurtosis K is a statistic that describes the steepness of the probability distribution of the signal and is used to distinguish impulse noise.
[0071] ;
[0072] In the above formula, Represents the mathematical expectation. The sampled values of the original signal are used. This formula can quantify the "sharpening" of the signal distribution. It should be noted that if K > 3, the noise is identified as impulse noise. It is Gaussian noise.
[0073] In this embodiment, white noise and colored noise are distinguished based on the power spectrum flatness S, an index describing the uniformity of the signal power spectrum.
[0074] ;
[0075] In the above formula, The signal power spectral density is calculated by performing a Fourier transform on the original signal, reflecting the signal at different frequencies. Power distribution at the location; It is a frequency axis variable, and the integration interval is limited to... ; The signal sampling frequency is a fundamental parameter for signal processing, determined by the parameters of the data acquisition equipment. The range of values is The closer the value is to 1, the more uniform the signal power spectrum is across all frequency ranges. When this occurs, the original signal contains white noise. If so, the original signal contains colored noise.
[0076] Through the above process, noise is ultimately classified into three categories: "white noise dominant", "colored noise dominant", and "impulse noise dominant".
[0077] In this embodiment, the threshold is adjusted according to different noise types. The specific process is as follows:
[0078] White noise using a threshold ;
[0079] Color noise uses dynamic threshold ;
[0080] Impulse noise with kurtosis adaptive hard threshold ;
[0081] Wavelet coefficients smaller than the noise threshold are set to zero or attenuated, while wavelet coefficients larger than the noise threshold are retained.
[0082] in, This represents the estimated standard deviation of noise. Indicates the length of the original signal. This represents the estimated noise power spectral density. This is the kurtosis coefficient.
[0083] For white noise, The larger the value, the longer the length of L. The larger the value, the stronger the noise suppression; for colored noise, , It is obtained by separating the signal power spectrum from the target echo power spectrum; for impulse noise, With peak To avoid residual false spikes after impulse noise removal, the wavelet coefficients after thresholding are smoothed using a 3×3 neighborhood process, increasing the linear increment. Through this process, the wavelet coefficients after thresholding for white noise, colored noise, and impulse noise are reconstructed using inverse transform to obtain the signal to be analyzed with a high signal-to-noise ratio.
[0084] It should be noted that the noise type identification process in this embodiment is designed to more accurately preserve signal features and suppress mixed noise. Different types of noise have fundamentally different characteristics in the wavelet domain, making it impossible to achieve the same effect through uniform denoising. Noise in target recognition scenarios is not of a single type, and the energy distribution and coefficient characteristics of different noises in the wavelet domain are completely different. If a uniform threshold is applied directly without classification, it will lead to signal distortion and noise residue. This embodiment designs differentiated threshold strategies for different noise characteristics. For example, a smaller threshold is used for low-frequency colored noise sub-bands (to avoid distortion of low-frequency signal features), and a larger threshold is used for high-frequency colored noise sub-bands (to enhance noise suppression). Impulse noise is first screened using a hard threshold, and then residual artifacts are eliminated using neighborhood smoothing to improve the fidelity of the denoised signal. If white noise remains, it will cause false peaks in the local energy distribution, leading to misjudgment of the signal energy core; if colored noise remains, it will interfere with the calculation of the second derivative of the frequency change rate, resulting in the incorrect selection of the order of the kernel function polynomial (n=2 / 3-4 / 5); if impulse noise remains, it will cause "pseudo ridges" in the time-frequency distribution, affecting the feature alignment of subsequent multimodal fusion (with WVD, STFT).
[0085] In the "Nonlinear IF Signal Processing for Target Recognition" scenario of this embodiment, "noise classification" is combined with "wavelet threshold denoising". A dynamic threshold is designed for mixed noise, which solves the problem that traditional fixed thresholds cannot adapt to mixed noise.
[0086] Multimodal time-frequency analysis is performed on the signal to be analyzed, and the output is a first time-frequency feature map based on an adaptive polynomial kernel function, a second time-frequency feature map based on Wigner-Ville distribution (WVD), and a third time-frequency feature map based on short-time Fourier transform (STFT).
[0087] In this embodiment, the process of performing multimodal time-frequency analysis on the signal to be analyzed mainly includes three parts: processing using polynomial frequency-modulated wavelet transform with adaptive polynomial functions, signal processing based on Wigner-Ville distribution (WVD), and signal processing based on short-time Fourier transform (STFT). Figure 2 The results of the multimodal time-frequency analysis are shown below.
[0088] (1) Processing is performed using polynomial frequency modulated wavelet transform of adaptive polynomial function.
[0089] Traditional polynomial linear frequency modulated wavelet transform (PCT) relies on a fixed-form polynomial kernel function to perform time-frequency analysis. While this fixed kernel function design ensures the determinism of the computation process, it cannot dynamically adjust the kernel function structure to match the time-varying characteristics of highly nonlinear instantaneous frequency (IF) signals widely present in real-world scenarios (such as signals with multi-component coupling, frequency abrupt changes, or high-order nonlinear variations). This can easily lead to problems such as time-frequency ambiguity and energy dispersion. Therefore, an adaptive polynomial kernel function is proposed. It can be understood that the adaptive polynomial kernel function in this embodiment is actually an improvement on polynomial linear frequency modulated wavelet transform (PCT). To clearly describe the content of this embodiment, the improved polynomial linear frequency modulated wavelet transform is defined as IPCT (Improved PCT, IPCT).
[0090] To capture the time-varying characteristics of a signal, the signal to be analyzed... Frame segmentation is performed, using a sliding window function to divide the long signal into several short-time local segments. By segmenting the data into frames, the time-frequency characteristics of each local segment can be made relatively stable over a short period of time, providing a "local benchmark" for subsequent adaptive polynomial kernel function design. For each local segment... Calculate the reflection of the first Local energy distribution of the intensity spatial distribution of the frame signal and rate of change of frequency , represents the following:
[0091] ;
[0092] ;
[0093] in, This is the intra-frame time offset. For the first Frame signal in Amplitude at time, The peak position corresponds to the energy core of the signal in that frame. Subsequent kernel functions need to enhance adaptation in this region by calculating the rate of frequency change to quantize the first... The degree of nonlinearity of the instantaneous frequency of the frame.
[0094] The first digit extracted by short-time Fourier transform (STFT) Intra-frame instantaneous frequency; second derivative It directly reflects the curvature of the frequency curve; the larger its absolute value, the stronger the frequency nonlinearity of the signal frame. Time (weakly nonlinear), order of kernel function polynomial ;when Time (nonlinear), ;when Time (strong nonlinearity), ,in, The order of the polynomial that determines the kernel function The peak region is used as a weighting factor for kernel function coefficient optimization.
[0095] The kernel function is the core of time-frequency analysis using polynomial linear frequency modulated wavelet transform, and its structure needs to be consistent with that extracted in step one. Precise matching means that "the degree of nonlinearity determines the order, and the energy distribution determines the weight." To adapt to the time-frequency coupling characteristics of nonlinear IF signals, this embodiment constructs an adaptive polynomial kernel function containing higher-order terms, expressed as:
[0096] ;
[0097] in, For the first The polynomial coefficients to be optimized in the frame kernel function. Let be the order of the polynomial. Energy weighting factor . This is the global analysis point, a time-axis variable corresponding to the entire signal's time series. In the analysis of the... Frame time fixed as , For the first The center moment of the frame is a nonvariable. For the displacement integral variable in wavelet transform, express From frame 1 to frame 2 The maximum value of the local energy distribution in the frame.
[0098] In this embodiment, the coefficients of the kernel function The key to its adaptability lies in the particle swarm optimization (PSO) algorithm used to solve for the optimal coefficients, aiming at the energy concentration of the time-frequency distribution, so that the kernel function accurately matches the signal features of the m-th frame. The specific process is as follows:
[0099] Initialize the particle swarm, with each particle corresponding to a set of coefficients. , dimension . Number of particles Inertial weight Learning factor .
[0100] For each particle Its location Corresponding to a set of coefficients Substituting into the adaptive polynomial kernel function formula, we obtain the first... Particle-specific kernel function for frames Based on this kernel function, an analytical wavelet is constructed:
[0101] ;
[0102] in It is the first Frame number The kernel function for each particle; The frequency carrier term is used to map the "time-frequency structure" of the kernel function to a specific frequency. ; For the first Frame number The analytical wavelet of an individual particle, whose time-frequency focusing characteristics are entirely determined by the kernel function. Decide.
[0103] Finally, the calculation of the first... Frames :
[0104] ;
[0105] in, Denotes the complex conjugate of the wavelet function. For frequency axis variables ( , (where the sampling frequency is used); the integral is accelerated by Fast Fourier Transform (FFT) (converting the time-domain integral into a frequency-domain product, reducing complexity).
[0106] according to Calculate particles fitness (i.e., energy concentration):
[0107] ;
[0108] The numerator is the "fourth power sum of time-frequency distribution" (to enhance the contribution of the energy concentration region), and the denominator is the "square of the square sum of time-frequency distribution" (normalization process to avoid the influence of absolute energy value). The range of values is , The closer the value is to 1, the better the fit of the kernel function for that particle.
[0109] Perform iterative particle updates and initialize random generation. The initial position of each particle and speed Record the individual optimal position of each particle. (This particle is the largest in history) corresponding and global optimal position (The largest in the history of all particles) corresponding Then update the particle velocity. With position (Constraints during iteration) (To avoid speed divergence)
[0110] ;
[0111] ;
[0112] in, This represents the current iteration number. Is A random number that is uniformly distributed within an interval; The inertia weight is used to balance the global and local search capabilities of the algorithm, and its calculation formula is as follows: , For maximum inertia weight, For minimum inertia weight, This represents the maximum number of iterations. and It is the learning factor, usually a constant greater than zero, used to adjust the degree to which a particle learns towards its individual optimal position and the global optimal position.
[0113] When the number of iterations or ( Represents the globally optimal energy concentration, globally optimal When the value tends to stabilize, stop iterating; at this point... Corresponding coefficients That is, the first Find the optimal polynomial coefficients of the frame and obtain the optimal kernel function. Based on the optimal adaptive polynomial kernel function The time-frequency distribution adapted to the nonlinear instantaneous frequency characteristics of the signal to be analyzed is obtained, and the first time-frequency feature map is output.
[0114] (2) The signal to be analyzed is smoothed by Gaussian window using Wigner-Ville distribution (WVD) to suppress cross terms. After noise suppression, a second time-frequency feature map is output to enhance the energy concentration region. The signal to be analyzed is processed by adaptive adjustment of window length using short-time Fourier transform (STFT). If it is a slow-changing signal, a long window is used to improve frequency resolution. If it is a fast-changing signal, a short window is used to ensure time resolution. Finally, a third time-frequency feature map is generated as a reference.
[0115] It should be noted that Gaussian window smoothing effectively suppresses cross-term interference, but may slightly reduce time-frequency resolution; long window improves frequency resolution and presents narrowband spectrum lines; short window improves time resolution and can capture transient changes.
[0116] The first, second, and third time-frequency feature maps are dynamically weighted and fused to generate a fused time-frequency feature tensor, which is then processed by a convolutional neural network to generate a time-series feature vector.
[0117] In this embodiment, the fusion of the above three types of feature maps specifically includes:
[0118] The first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map are set to the same time axis and frequency axis, and then adjusted to the same resolution by interpolation.
[0119] The energies of the first, second, and third time-frequency feature maps after resolution adjustment are normalized to eliminate differences in different energy scales.
[0120] A dynamic weighted fusion method is adopted, and the weights are determined based on the entropy values of the three time-frequency feature maps in different frequency ranges. The smaller the entropy value, the more concentrated the features are in that range, and the higher the weight is assigned accordingly.
[0121] like Figure 3 As shown, the energy values of the three time-frequency feature maps at each time-frequency point are weighted and summed according to a determined weight to generate a fused time-frequency feature tensor.
[0122] A convolutional neural network (CNN) is used to extract features from the fused time-frequency feature tensor. Convolutional layers use kernels of different sizes to capture local time-frequency patterns. Pooling layers are used to reduce dimensionality and enhance key features, outputting local feature vectors. Then, the local feature vectors output by the CNN are expanded into a time-series feature sequence along the time axis and input into a recurrent neural network (RNN) for time-series modeling. The gating mechanism of the RNN is used to capture the global time-varying patterns of the signal and output time-series feature vectors.
[0123] The time-series feature vectors are input into the convolutional neural network and the gradient boosting regression tree (GBRT) model respectively for processing, and then the preliminary identification results and feature verification results based on the target category are output. Target identification is performed based on the preliminary identification results and feature verification results of the target category.
[0124] In this embodiment, the temporal feature vector is split into two paths and input into the convolutional neural network and the boosting regression tree GBRT model, respectively.
[0125] The first temporal feature vector passes through the fully connected layer of the convolutional neural network (CNN) and the softmax function to output the target class probability. The class with the highest probability value is taken as the preliminary recognition result, and the highest probability value is taken as the probability confidence level.
[0126] The second-order time-series feature vector is processed by the Gradient Boosting Regression Tree (GBRT) model to output the intermediate frequency estimate. The target feature verification parameters carrying the target information are extracted from the intermediate frequency estimate. The target feature verification parameters are matched with the source feature verification parameters in the pre-built mapping library to generate feature verification results.
[0127] The target characteristic verification parameters include the frequency change rate based on actual values. Mid-frequency range and time-frequency ridge morphology parameters .
[0128] In this embodiment, during target recognition, intermediate frequency (IF) refers to the frequency range of the signal after processing that falls between high frequency and baseband, and carries key target information. During time-frequency analysis, the frequency range of the signal is used to determine the intermediate frequency. Signal-to-noise ratio Dynamically adjust the scale factor of IPCT based on features. The dataset was constructed using this as the target variable. The Gradient Boosting Regression Tree (GBRT) algorithm was employed, with the tree depth set accordingly. 5. Used to control the complexity of each regression tree, assuming the number of nodes in the decision tree is... In the case of an ideal binary tree, the relationship between the number of nodes and the depth satisfies The learning rate is set to 0.1, which controls the step size of the newly generated regression tree in each iteration for updating the model. The model update formula is as follows: ,in For the first Predicted value after the next iteration This is the predicted value from the previous iteration. For learning rate, For the first The predicted values of each regression tree; setting the number of trees to 100 determines the total number of regression trees generated during model training, and the final prediction result of the model is:
[0129] ;
[0130] in For the number of trees, For the first The prediction function for each tree was used. Finally, a 5-fold cross-validation was employed, comparing the mean squared error (MSE) across the five validation iterations. The model with the smallest average MSE is selected as the optimal model, and the optimal construction mapping function for that model is obtained:
[0131] ;
[0132] Energy concentration reflects the "frequency coverage range" of the signal. The maneuverability of different targets leads to differences in the frequency range of the nonlinear IF signal. Reflects the "degree of noise interference" of the signal. High indicates weak noise. Low indicates high noise; output variable It is a core adjustment parameter of IPCT, which directly controls the "time-frequency resolution adaptation capability" of IPCT wavelet analysis. As the frequency range increases, the frequency coverage of the wavelet expands. When the wavelet is reduced, its temporal resolution is improved.
[0133] The input signal is subjected to a Fast Fourier Transform (FFT) to convert the time-domain signal into a frequency-domain representation, and then the mapping function is applied... The current signal dynamic scaling factor is calculated. The dynamic scaling factor can be expressed as:
[0134]
[0135] in These are the coefficients obtained by fitting using the least squares method.
[0136] Next, the calculated dynamic scaling factor Substitution polynomial transformation In the intermediate frequency estimation algorithm, the formula is as follows:
[0137] ;
[0138] in, It is the output variable after IPCT transformation, representing the instantaneous frequency. For input signal, These are polynomial basis functions; their expressions are:
[0139] ;
[0140] in, It is the first The optimal kernel function complex conjugate of the frame, For frequency carrier terms, Let be the instantaneous frequency output by IPCT, and t be the global analysis time. For displacement integral variables; The Gaussian window function is expressed as:
[0141] ;
[0142] The Gaussian window width parameter, Let be a polynomial parameter vector, represented as:
[0143] ;
[0144] Indicates the kernel function order. Let k represent the optimal coefficients, where k = 0, ..., n.
[0145] Introducing dynamic scaling factor Then, the basis functions become:
[0146] ;
[0147] Perform peak search to obtain mid-frequency estimate Since the obtained mid-frequency estimate carries the frequency range of key target information, it is used to perform target "characteristic verification".
[0148] In this embodiment, the output target category probability is represented as the probability value assigned to different targets (such as "aircraft", "ship", "ground armored vehicle", "clutter interference") (the sum of all category probabilities is 1), and the category with the highest probability is taken as the "preliminary identification result". For example, if the probability of the "aircraft" category is 0.88, "ship" is 0.10, and "clutter interference" is 0.02, then the target is initially determined to be "aircraft", and the "probability confidence level" of this preliminary category (i.e., the highest probability value, which is 0.88 here) is recorded.
[0149] In this embodiment, the second-path time-series feature vector is processed by the Gradient Boosting Regression Tree (GBRT) model to output an intermediate frequency estimate, based on... Peak search of time-frequency energy distribution; extraction of three types of target characteristic verification parameters from the mid-frequency estimate; matching of these target characteristic verification parameters with source feature verification parameters in a pre-built mapping library to generate characteristic verification results. In this process, the target characteristic verification parameters include the frequency change rate based on the actual value. Mid-frequency range and time-frequency ridge morphology parameters .
[0150] Among them, the rate of change of frequency Represented as:
[0151] ;
[0152] Mid-frequency range: It is determined by the frequency range corresponding to the energy peak in the time-frequency distribution, reflecting the frequency coverage characteristics of the target echo.
[0153] Time-frequency ridge morphology parameters :right The ridge lines are fitted to extract morphological features such as ridge smoothness and number of abrupt changes, which are used to distinguish the target from clutter. (Clutter ridge lines usually have irregular abrupt changes, while target ridge lines show continuous dynamic changes).
[0154] Based on the rate of change of frequency Mid-frequency range and time-frequency ridge morphology parameters The actual value is matched with the theoretical value in the mapping library to generate a value based on the frequency change rate error. Mid-frequency range overlap Ridge morphology The results of the characteristic verification;
[0155] in:
[0156] ;
[0157] ;
[0158] ;
[0159] This represents the mean of the theoretical frequency change rate of the initially identified categories. The smaller the value, the higher the match. This represents the theoretical value for the intermediate frequency range. The larger the value, the higher the match. Indicates the number of ridge sampling points. , Representing the actual and theoretical ridge lines respectively Point shape value, The smaller the value, the higher the match.
[0160] The specific steps for target identification based on the preliminary identification results and feature verification results of the target category include:
[0161] Obtain the probability confidence level of the preliminary identification results and the frequency change rate error of the feature verification results. Mid-frequency range overlap Ridge morphology ;
[0162] like And the probability confidence level of the preliminary identification results If the target recognition result is accurate, the final target category and recognition confidence level will be output.
[0163] like Or, the probability confidence level If so, the recognition result is deemed invalid;
[0164] The recognition confidence level is expressed as the probability confidence level minus the frequency change rate error. Mid-frequency range overlap and ridge line morphology The weighted average calculation result.
[0165] It should be noted that during the construction of the mapping library, the "target category - theoretical intermediate frequency characteristics" mapping library was generated by training a large number of sample signals labeled with real target categories. For each type of target, the mean value and threshold range of its intermediate frequency characteristics under different maneuvering states and electromagnetic environments were statistically analyzed to form a theoretical standard. Through multiple optimizations, a mapping library that meets the threshold range requirements was formed.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target recognition method based on adaptive polynomial kernel function and multimodal fusion, characterized in that, Includes the following steps: The original signal of the target to be identified is obtained, and the original signal is preprocessed based on the wavelet threshold denoising algorithm to generate the signal to be analyzed. Multimodal time-frequency analysis is performed on the signal to be analyzed, and the output is a first time-frequency feature map based on an adaptive polynomial kernel function, a second time-frequency feature map based on Wigner-Ville distribution (WVD), and a third time-frequency feature map based on short-time Fourier transform (STFT). The first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map are dynamically weighted and fused to generate a fused time-frequency feature tensor, which is then processed by a convolutional neural network to generate a time-series feature vector. The temporal feature vectors are input into a convolutional neural network and a gradient boosting regression tree (GBRT) model, respectively, and then the preliminary identification results and feature verification results based on the target category are output. Target identification is performed based on the preliminary identification results of the target category and the characteristic verification results; The specific steps of performing multimodal time-frequency analysis on the signal to be analyzed and outputting a first time-frequency feature map based on an adaptive polynomial kernel function include: After framing and extracting local features from the signal to be analyzed, a local energy distribution is generated. and rate of change of frequency ; Construct an adaptive polynomial kernel function containing higher-order terms , is represented as: ; Indicates the first The polynomial coefficients of the frame. Denotes the order of a polynomial. Represents time axis variables, This represents the displacement integral variable in the wavelet transform. Energy weighting factor , for Energy distribution at time, For the first The center moment of the frame, express From frame 1 to frame 2 The maximum value of the local energy distribution within the frame; The polynomial coefficients were analyzed based on the particle swarm optimization algorithm. Optimization yields the first The optimal polynomial coefficients of the frame, and the optimal adaptive polynomial kernel function; Based on the optimal adaptive polynomial kernel function, a time-frequency distribution adapted to the nonlinear instantaneous frequency characteristics of the signal to be analyzed is obtained, and a first time-frequency feature map is output.
2. The target recognition method according to claim 1, characterized in that, The steps for preprocessing the original signal to generate the signal to be analyzed based on the wavelet threshold denoising algorithm specifically include: The original signal is decomposed into multiple wavelet components based on its kurtosis. If the kurtosis is greater than 3, the original signal contains impulse noise. The power spectral density is obtained by performing a Fourier transform on the original signal, and the power spectral flatness is calculated based on the power spectral density. If the power spectrum flatness is approximately equal to 1, then the original signal contains white noise; If the power spectral flatness is less than 1, then the original signal contains colored noise; The corresponding noise thresholds are set for white noise, colored noise and impulse noise respectively, and wavelet coefficients are processed according to the corresponding noise thresholds to generate a signal with high signal-to-noise ratio to be analyzed.
3. The target recognition method according to claim 2, characterized in that, The steps of setting corresponding noise thresholds for white noise, colored noise, and impulse noise, and then generating a high signal-to-noise ratio signal after wavelet coefficient processing based on these noise thresholds, specifically include: Set corresponding noise thresholds for white noise, colored noise, and impulse noise respectively, as follows: White noise using a threshold ; Color noise uses dynamic threshold ; Impulse noise with kurtosis adaptive hard threshold ; Wavelet coefficients smaller than the noise threshold are set to zero or attenuated, while wavelet coefficients larger than the noise threshold are retained. The retained wavelet coefficients are then reconstructed using an inverse transform to obtain the signal to be analyzed with a high signal-to-noise ratio. in This represents the estimated standard deviation of noise. Indicates the length of the original signal. This represents the estimated noise power spectral density. This is the kurtosis coefficient.
4. The target recognition method according to claim 1, characterized in that, The steps of performing multimodal time-frequency analysis on the signal to be analyzed, and outputting a second time-frequency feature map based on Wigner-Ville distribution (WVD) and a third time-frequency feature map based on short-time Fourier transform (STFT), specifically include: The signal to be analyzed is smoothed by Gaussian window using Wigner-Ville distribution (WVD) to suppress cross terms, and after noise suppression, a second time-frequency feature map of the energy concentration region is output. The short-time Fourier transform (STFT) is used to adaptively adjust the window length of the signal to be analyzed. If it is a slowly changing signal, a long window is used to improve the frequency resolution, and if it is a fast changing signal, a short window is used to ensure the time resolution. Finally, a third time-frequency feature map is generated as a reference.
5. The target recognition method according to claim 4, characterized in that, The steps of dynamically weighting and fusing the first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map to generate a fused time-frequency feature tensor specifically include: The first time-frequency feature map, the second time-frequency feature map, and the third time-frequency feature map are set to the same time axis and frequency axis, and then adjusted to the same resolution by interpolation. The energies of the first, second, and third time-frequency feature maps after resolution adjustment are normalized to eliminate differences in different energy scales. A dynamic weighted fusion method is adopted, and the weights are determined based on the entropy values of the three time-frequency feature maps in different frequency ranges; The energy values of the three time-frequency feature maps at each time-frequency point are weighted and summed according to a determined weight to generate a fused time-frequency feature tensor.
6. The target recognition method according to claim 5, characterized in that, The specific steps for generating time-series feature vectors by processing the fused time-frequency feature tensor through a convolutional neural network include: A convolutional neural network (CNN) is used to extract features from the fused time-frequency feature tensor. Local time-frequency patterns are captured by using convolutional kernels of different sizes through convolutional layers. After dimensionality reduction processing by pooling layers, local feature vectors are output. The local feature vector is expanded into a temporal feature sequence along the time axis, input into a recurrent neural network (RNN) for temporal modeling, and the temporal feature vector is output through the gating mechanism of the RNN.
7. The target recognition method according to claim 6, characterized in that, The steps of inputting the temporal feature vectors into a convolutional neural network and a gradient boosting regression tree (GBRT) model for processing, and then outputting preliminary identification results and feature verification results based on the target category, specifically include: The temporal feature vector is split into two paths and input into the convolutional neural network and the boosting regression tree GBRT model, respectively. The first temporal feature vector is passed through the fully connected layer of the convolutional neural network (CNN) and the softmax function to output the target class probability. The class with the highest probability value is taken as the preliminary recognition result, and the highest probability value is taken as the probability confidence level. The second-order time-series feature vector is processed by the Gradient Boosting Regression Tree (GBRT) model to output the intermediate frequency estimate. The target feature verification parameters carrying the target information are extracted from the intermediate frequency estimate. The target feature verification parameters are matched with the source feature verification parameters in the pre-built mapping library to generate feature verification results. The target characteristic verification parameters include the frequency change rate based on actual values. Mid-frequency range and time-frequency ridge morphology parameters .
8. The target recognition method according to claim 7, characterized in that, The steps of matching the target feature verification parameters with the source feature verification parameters in a pre-built mapping library to generate feature verification results specifically include: A "target category - theoretical mid-frequency characteristic" mapping library is pre-constructed. Each target category in the mapping library is matched with theoretical values of the mean and threshold range of mid-frequency characteristics under different maneuvering states and electromagnetic environments. Based on the rate of change of frequency Mid-frequency range and time-frequency ridge morphology parameters The actual value is matched with the theoretical value in the mapping library to generate a value based on the frequency change rate error. Mid-frequency range overlap Ridge morphology The results of the characteristic verification; in, , This represents the mean of the theoretical frequency change rate of the initially identified categories. The smaller the value, the higher the match. , This represents the theoretical value for the intermediate frequency range. The larger the value, the higher the match. , Indicates the number of ridge sampling points. , Representing the actual and theoretical ridge lines respectively Point shape value, The smaller the value, the higher the match.
9. The target recognition method according to claim 8, characterized in that, The specific steps for target identification based on the preliminary identification results and feature verification results of the target category include: Obtain the probability confidence level of the preliminary identification results and the frequency change rate error of the feature verification results. Mid-frequency range overlap Ridge morphology ; like And the probability confidence level of the preliminary identification results If the target recognition result is accurate, the final target category and recognition confidence level will be output. like Or, the probability confidence level If so, the recognition result is deemed invalid; The recognition confidence level is expressed as the probability confidence level minus the frequency change rate error. Mid-frequency range overlap and ridge line morphology The weighted average calculation result.
Citation Information
Patent Citations
Multi-modal regression analysis based hydroelectric generating set's cavitation erosion signal feature extraction method
CN106407944A
Systems and methods for seizure detection and closed-loop neurostimulation
US20250195894A1