A low, slow and small target fast detection method and system adaptive to wind noise

By constructing a high signal-to-noise ratio dataset for multi-level decomposition and feature fusion of acoustic signals, and combining it with a YOLO network optimization model, the problem of identifying low, slow, and small targets under dynamic wind noise was solved, achieving high-precision target detection and recognition.

CN121831684BActive Publication Date: 2026-06-23CHENGDU IND VOCATIONAL TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU IND VOCATIONAL TECHN COLLEGE
Filing Date
2026-03-13
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively distinguish and identify low, slow, and small targets in dynamic wind-noise environments. Radar signal detection and identification have low reliability and are susceptible to interference, while acoustic detection technologies are highly sensitive but have low reliability.

Method used

We employ an adaptive wind noise-based method for the rapid detection of small, slow targets. By constructing a high signal-to-noise ratio dataset, we perform multi-level decomposition and time-frequency domain transformation of acoustic signals to extract target waveform and spectrogram features. We then optimize feature fusion through three-dimensional parallel PSO dynamic weight search and combine it with the YOLO network for model optimization and deployment.

Benefits of technology

It improves the detection accuracy and anti-interference ability of low, slow and small targets in dynamic wind noise environment, and realizes high-precision target recognition and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121831684B_ABST
    Figure CN121831684B_ABST
Patent Text Reader

Abstract

The application provides a low, slow and small target fast detection method and system with adaptive wind noise, and belongs to the technical field of target detection. The method first collects unmanned aerial vehicle acoustic data and calls bird acoustic data set, constructs high signal-to-noise ratio data set, training set, test set and verification set; the acoustic signals of the high signal-to-noise ratio data set are subjected to synchronous compression wavelet transform, the high and low frequency information is separated, the waveform features and the spectrogram features are extracted respectively, the feature fusion is carried out, the feature fusion weight is optimized in the fusion process, and the adaptive wind noise fusion features are obtained; the YOLO network is called and its structure is optimized based on the fusion features, and the feature extraction and target detection network model is constructed; after training, testing and verification by the training set, the test set and the verification set, the model is deployed to the mobile terminal or the ground station for low, slow and small target detection. Through the dual-mode feature fusion and network optimization, the detection accuracy and timeliness of the low, slow and small target under dynamic wind noise are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target detection technology, and in particular to a method and system for rapid detection of small, slow targets with adaptive wind noise. Background Technology

[0002] The rapid growth of drones has propelled their application in various fields such as smart city construction, smart oilfield exploration, and smart railway inspection. However, this has also brought new security risks, such as unauthorized intrusion and malicious attacks by drones. Therefore, it is necessary to monitor drones, determine their flight status and target intentions, and provide crucial information for subsequent handling.

[0003] In the airspace, balloons, birds, and drones are collectively referred to as low, slow, and small targets. These targets typically coexist in the monitored airspace, requiring effective detection and identification. Currently, low, slow, and small targets are mainly detected and identified through spatial scanning of radar electromagnetic signals and target echo analysis. However, due to the similarity of radar echoes, this radar signal-based detection and identification method often struggles to distinguish these targets, resulting in low reliability. Furthermore, low, slow, and small target detection and identification technology based on radar signals and echo analysis is an active detection technique. Uncooperative intruding targets can easily detect radar waves and conduct electromagnetic interference, thus distorting the radar echo and preventing the detection and analysis of intruding targets.

[0004] In addition, there are currently passive acoustic detection technologies that can detect and identify low, slow, and small targets. However, they are extremely sensitive to wind noise, can only perform short-range detection and identification, have low reliability, and are mostly used as a supplement to radar detection and identification technologies. Summary of the Invention

[0005] In view of this, the present invention provides a method and system for rapid detection of low, slow and small targets under adaptive wind noise, so as to improve the detection accuracy of low, slow and small targets under dynamic wind noise.

[0006] The technical solution adopted in this invention is:

[0007] This invention provides a fast detection method for low-speed, small targets with adaptive wind noise, comprising:

[0008] Acoustic data from drones and bird acoustic datasets were collected to construct a high signal-to-noise ratio dataset, training set, test set, and validation set.

[0009] Synchronous compressed wavelet transform is performed on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data is then subjected to time-frequency domain transformation and energy suppression processing to obtain the target waveform features.

[0010] High-frequency information is extracted from multi-layer decomposed data and superimposed. Then, time-frequency transformation is performed on the superimposed high-frequency information to obtain the target spectrogram features.

[0011] Feature fusion is performed on the target waveform features and target spectrogram features, and the feature fusion weights are optimized by three-dimensional parallel PSO dynamic weight search to obtain adaptive wind noise fused features;

[0012] By calling the YOLO network and optimizing the network structure based on fused features, a feature extraction network model and an object detection network model are obtained.

[0013] The feature extraction network model and the object detection network model are trained using the training set and the test set, and the model parameters are saved. At the same time, the effectiveness of the model is verified using the validation set.

[0014] After the model's effectiveness is verified, the feature extraction network model and the target detection network model are deployed to mobile devices or ground stations to detect low, slow, and small targets and output the target detection results.

[0015] Based on the aforementioned adaptive wind noise-based fast detection method for low-speed, small targets, this invention further provides an adaptive wind noise-based fast detection system for low-speed, small targets, comprising:

[0016] The dataset construction module is used to collect drone acoustic data and retrieve bird acoustic datasets to build a high signal-to-noise ratio dataset, training set, test set, and validation set.

[0017] The feature extraction module is used to perform synchronous compressed wavelet transform on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data is then subjected to time-frequency domain transformation and energy suppression processing to obtain the target waveform features. At the same time, high-frequency information is extracted from the multi-level decomposed data and superimposed. The superimposed high-frequency information is then subjected to time-frequency graph transformation to obtain the target spectrogram features.

[0018] The feature fusion module is used to fuse the target waveform features and the target spectrogram features, and optimize the feature fusion weights through three-dimensional parallel PSO dynamic weight search to obtain the adaptive wind noise fused features.

[0019] The model optimization module is used to call the YOLO network and optimize the network structure based on the fused features to obtain the feature extraction network model and the object detection network model.

[0020] The model training and validation module is used to train the feature extraction network model and the object detection network model using the training set and the test set, save the model parameters, and use the validation set to verify the effectiveness of the model.

[0021] The model deployment module is used to deploy the feature extraction network model and the target detection network model to mobile devices or ground stations after the model validity verification is passed, to detect low, slow and small targets, and output the target detection results.

[0022] In summary, the beneficial effects of the present invention are as follows:

[0023] This invention provides a rapid detection method for low-speed, small targets under adaptive wind noise. It constructs a high signal-to-noise ratio (SNR) dataset, training set, test set, and validation set to provide a foundation for model robustness. Then, it performs synchronous compressed wavelet transform on the high SNR dataset to separate high- and low-frequency information, extracting dual-modal features related to the target waveform and spectrogram, thus overcoming the limitations of single-feature representation and enhancing the saliency of target features under wind noise. Furthermore, during feature fusion, it incorporates three-dimensional parallel PSO to dynamically optimize the fusion weights, achieving adaptive wind noise compensation and significantly improving the anti-interference capability of the fused features. Simultaneously, it optimizes the YOLO network structure based on the fused features to obtain a feature extraction network model and a target detection network model. The optimized model is then validated for effectiveness. After successful validation, the model can be flexibly deployed on mobile devices or ground stations to achieve accurate detection and recognition of low-speed, small targets under dynamic wind noise. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.

[0025] Figure 1 This is a flowchart of a method for rapid detection of small, slow targets with adaptive wind noise according to the present invention.

[0026] Figure 2 This is a flowchart of the feature extraction process of the present invention;

[0027] Figure 3 This is a schematic diagram of the fusion feature weight trainer structure of the present invention;

[0028] Figure 4 This is a flowchart illustrating the feature fusion process of the present invention;

[0029] Figure 5 This is a comparative schematic diagram of the structural optimization of the YOLO network of this invention;

[0030] Figure 6 This is a schematic diagram of the optimized feature detection network structure of the present invention;

[0031] Figure 7 This is a flowchart of the rapid detection process for small, slow targets according to the present invention;

[0032] Figure 8 This is a schematic diagram illustrating the application process of the method of the present invention;

[0033] Figure 9 This is a functional block diagram of a rapid detection system for low-speed, small targets with adaptive wind noise according to the present invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Unless otherwise specified, the present invention and the various features in the embodiments can be combined with each other, all of which are within the protection scope of the present invention.

[0035] Currently, technologies for UAV detection based on spatial scanning of radar electromagnetic signals and target echo analysis have the following main drawbacks:

[0036] (1) The reliability of radar signal and echo analysis detection and identification technology is low.

[0037] Radar signal detection and identification methods are mainly divided into low-resolution radar and high-resolution radar. Low-resolution radar can quickly detect and identify targets within the monitoring range, while high-resolution radar can analyze the detailed information of targets through high-resolution Doppler data. However, due to the similarity of radar echoes for small, slow-moving targets, low-resolution radar often struggles to distinguish them. High-resolution radar, on the other hand, is typically expensive, requires large amounts of data to process, and demands significant processing resources. Conventional monitoring methods generally do not employ high-resolution radar for target detection and identification. Furthermore, the analysis of high-resolution radar signal echoes requires computational power and time, and it usually lacks the capability for rapid on-site target analysis.

[0038] (2) Exposure characteristics of detection and identification technologies for radar signals and echo analysis

[0039] Low-speed-small detection and identification technology based on radar signal and echo analysis is an active detection technology. Non-cooperative intruding targets can easily detect radar waves and conduct electromagnetic interference, thereby causing radar echo distortion and preventing the detection and analysis of intruding targets.

[0040] (3) Limitations of conventional acoustic detection and identification technologies

[0041] For the detection and identification of silent electromagnetic targets and illegal drone targets with radar countermeasures, there is currently passive acoustic detection technology that can detect and identify low, slow and small targets. However, it is extremely sensitive to wind noise, can only perform short-range detection and identification, has low reliability, and is mostly used as a supplement to radar detection and identification technology.

[0042] In summary, addressing the issues of poor reliability and high vulnerability in radar signal detection and identification of low, slow, and small targets, this invention employs an acoustic detection-based approach. Acoustic detection primarily relies on the sound source signals actively emitted by the target, thus offering strong concealment for airspace target detection. Secondly, detection methods based on active emission signals also include the detection of radiation source signals, which also provides strong concealment. However, UAVs can fly autonomously in silent mode, making it difficult to effectively detect and identify radiation source signals from silent targets. Therefore, this invention utilizes an acoustic detection-based approach to detect and identify low, slow, and small targets. Detailed implementation details are provided in the following embodiments.

[0043] Example 1: Refer to Figure 1 As shown, Figure 1 This is a schematic flowchart of a fast detection method for low-speed, small targets with adaptive wind noise according to an embodiment of the present invention. The method of this embodiment includes:

[0044] Acoustic data from drones and bird acoustic datasets were collected to construct a high signal-to-noise ratio dataset, training set, test set, and validation set.

[0045] Synchronous compressed wavelet transform is performed on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data is then subjected to time-frequency domain transformation and energy suppression processing to obtain the target waveform features.

[0046] High-frequency information is extracted from multi-layer decomposed data and superimposed. Then, time-frequency transformation is performed on the superimposed high-frequency information to obtain the target spectrogram features.

[0047] Feature fusion is performed on the target waveform features and target spectrogram features, and the feature fusion weights are optimized by three-dimensional parallel PSO dynamic weight search to obtain adaptive wind noise fused features;

[0048] By calling the YOLO network and optimizing the network structure based on fused features, a feature extraction network model and an object detection network model are obtained.

[0049] The feature extraction network model and the object detection network model are trained using the training set and the test set, and the model parameters are saved. At the same time, the effectiveness of the model is verified using the validation set.

[0050] After the model's effectiveness is verified, the feature extraction network model and the target detection network model are deployed to mobile devices or ground stations to detect low, slow, and small targets and output the target detection results.

[0051] In this embodiment, the method first collects a robust dataset to ensure model training. Specifically, it includes real drone acoustic datasets and bird acoustic datasets. Based on the two acoustic datasets, a high signal-to-noise ratio dataset, training set, test set, and validation set with multiple signal-to-noise ratios are constructed. This mixes data with accurate signal-to-noise ratios to ensure the robustness of the model.

[0052] Secondly, the method performs multi-level data decomposition on the acoustic signals in the high signal-to-noise ratio dataset and extracts target waveform features and target spectrogram features from the decomposed data, overcoming the shortcomings of single-modal features and further significantly improving the ability to adapt to wind noise. Subsequently, a dynamic fusion mechanism for multi-modal features is established, and feature fusion is performed on the target waveform features and target spectrogram features to further enhance the utilization of the advantages of multi-modal features. This enables the fused features to adaptively adjust the multi-modal feature weights and compensate for wind noise feature loss, allowing for rapid adaptation to wind noise changes and exhibiting strong adaptability, anti-interference, and robustness. Then, based on the image information features and structure contained in this fused feature, the YOLO network structure is improved and optimized to construct a multi-channel fusion feature extraction network model and a target detection network model. These two models are used for feature extraction and feature fusion respectively, effectively improving the processing speed and timeliness in practical applications without increasing the complexity of feature extraction or the number of model parameters. Finally, the method conducts comprehensive verification of the optimized model for wind noise detection to ensure its effectiveness. The model will be deployed to mobile key or large ground stations to achieve high-precision detection and identification of low, slow and small targets in complex environments, ensuring the safety of public airspace, information security of municipal platforms, and public safety.

[0053] In this embodiment, acoustic data from drones is collected and a bird acoustic dataset is retrieved to construct a high signal-to-noise ratio dataset. Simultaneously, training, testing, and validation sets are constructed, including:

[0054] The acoustic signals of various drones during flight were collected and labeled, and the collected acoustic signals were denoised to obtain a drone acoustic dataset.

[0055] The bird sound signal dataset is retrieved from the bird data database. Noise reduction is performed on each bird recording in the bird sound signal dataset. The bird recordings are then divided into multiple acoustic data segments according to the concentrated areas of bird calls to obtain the bird acoustic dataset.

[0056] High signal-to-noise ratio datasets, training sets, test sets, and validation sets were constructed by extracting drone acoustic data and acoustic data segments from drone acoustic datasets and bird acoustic datasets, respectively.

[0057] Specifically, the establishment of the UAV acoustic dataset can be based on data collected by real UAV equipment. This dataset undergoes professional pre-processing, signal-to-noise ratio processing, and adjustable noise superposition to form a UAV acoustic dataset superimposed with real-world environmental noise. In the specific data collection process, this embodiment uses a mobile phone microphone and two multi-rotor UAVs to collect and label the dataset. Pre-processing noise reduction is performed on the data, and 1000 acoustic data points are generated for each UAV as needed, resulting in the UAV acoustic dataset.

[0058] Secondly, this embodiment utilizes publicly available datasets from bird databases to establish a bird acoustic dataset. This data still employs professional methods to extract bird sounds (the bird sounds in publicly available datasets are somewhat unclear), and controllable real-world environmental noise is superimposed to form a bird acoustic dataset with superimposed real-world environmental noise. Specifically, the publicly available dataset can be the one from "Fudan University Social Science Data Platform / Data Resources / 2022 Scientific Data Management Course / Fifth Group of Documents / Birdsong". By segmenting each recording in the publicly available dataset into 3 to 4 data segments based on the concentrated bird calls region, a total of 1000 bird acoustic data points are obtained.

[0059] After obtaining the UAV acoustic dataset and the bird acoustic dataset, a high signal-to-noise ratio dataset was further constructed. Specifically, it was obtained by extracting UAV acoustic data and combining acoustic data segments from the UAV acoustic dataset and the bird acoustic dataset. The high signal-to-noise ratio dataset contains 2,000 multi-rotor UAV acoustic data points and 1,000 bird sound data points.

[0060] To construct the training and testing sets for the UAV detection network, 1600 acoustic data points of multi-rotor UAV technology and 800 bird sound data points were extracted to construct the training and testing sets, with a data ratio of 3:1.

[0061] Constructing a validation dataset for the drone detection network: Extract the remaining sound data from the drone acoustic dataset and the bird acoustic dataset, and mix in wind noise with adjustable intensity to obtain 400×4 multi-rotor drone acoustic data (containing 2 drones, 4 levels of mixed noise data, 20dB, 10dB, 0dB, -10dB) and 200×4 bird acoustic data (containing 4 levels of mixed noise data, 20dB, 10dB, 0dB, -10dB). This validation set is used to verify the reliability of the present invention in resisting wind noise when detecting drones.

[0062] In this embodiment, to achieve effective detection of low-speed, small targets, a robust, anti-interference, and significantly target-feature-rich acoustic multi-level, multi-mode anti-dynamic wind noise feature extraction method is proposed, effectively splitting a single data set into multi-modal features. This feature extraction method first employs energy-suppressed wavelet transform based on multi-level decomposition for acoustic target waveform feature extraction. This method can extract stationary features from non-stationary feature states, significantly enhancing target features and stability, and exhibiting strong anti-interference capabilities. Secondly, spectrogram transform is performed using data based on multi-level decomposition. This feature extraction method can extract hierarchical high-frequency features and filter low-frequency background noise, resulting in target image features with anti-wind noise characteristics. By extracting two modal features based on splitting, the drawbacks of using a single feature can be compensated for, significantly improving the reliability of low-speed, small target detection based on acoustic features under conditions of unavoidable wind noise.

[0063] Based on the above feature extraction method, this embodiment performs synchronous compressed wavelet transform on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The process of performing time-frequency domain transformation and energy suppression processing on the low-frequency information of the multi-level decomposed data to obtain the target waveform features specifically includes:

[0064] The acoustic signals in the high signal-to-noise ratio dataset are sampled into discrete signals, and synchronous compressed wavelet transform is performed on the discrete signals to decompose the discrete signals into multi-level decomposed data.

[0065] Low-frequency and high-frequency information of each layer of decomposed data are separated by low-pass and high-pass filters; time-frequency domain transformation is performed on the low-frequency information of each layer to obtain the time-frequency domain characteristics of the low-frequency information of each layer.

[0066] Calculate the energy proportion of low-frequency information in each layer, and suppress the energy of the time-frequency domain features of each layer according to the energy proportion to obtain the target waveform features.

[0067] Specifically, in signal research, radio frequency fingerprint features based on time-frequency domain transformation can retain fixed fingerprint features in steady-state signals, while also retaining nonlinear mutation features in transient signals, and have strong anti-interference capabilities. It is also applicable in the study of waveform features of acoustic signals. In this embodiment, in order to align the waveform features with the time-frequency spectrum features extracted later, a three-level decomposition of synchronous compressed wavelet transform is used for signal transformation, which is then aligned with the three-channel parameters in the image.

[0068] Reference Figure 2 As shown, when extracting target waveform features, the continuous sound signal is sampled as a discrete signal:

[0069] ;

[0070] Basic formula for synchronous compressed wavelet transform:

[0071] ;

[0072] ;

[0073] in, , Represents the time axis variable in time-domain information. Represents the mother wavelet function; Indicates the scale factor; This represents the time shift factor. Represents the cosine function. This represents the angle value after the time-domain variable has been transformed into the frequency-domain variable. Representing discrete variables, This indicates the corresponding signal amplitude. This represents the expression after the signal has been discretized. This represents the angle conversion value after synchronous compression. express The value of the i-th discrete variable (i is within the interval 0-N, and N is the length of the discrete signal). This represents the output value that is synchronously compressed and changed. This represents a complex exponential basis function with parameter b.

[0074] The sound signal is continuously decomposed using synchronous compressed wavelet transform. The specific decomposition formula is as follows:

[0075] ;

[0076] ;

[0077] ;

[0078] ;

[0079] ;

[0080] ;

[0081] in, This represents the first layer of synchronous compressed wavelet transform; This represents the low-frequency information of the first layer of the synchronous compressed wavelet transform; Indicates a low-pass filter; This represents the high-frequency information of the first layer of synchronous compressed wavelet transform; Indicates a high-pass filter; This indicates the second layer of synchronous compressed wavelet transform; This represents the low-frequency information of the second layer of the synchronous compressed wavelet transform; This represents the high-frequency information of the second layer of the synchronous compressed wavelet transform; This indicates the third layer of synchronous compressed wavelet transform; This represents the low-frequency information of the third layer of the synchronous compressed wavelet transform; This represents the high-frequency information of the third layer of the synchronous compressed wavelet transform.

[0082] The above abbreviation:

[0083] ;

[0084] ;

[0085] ;

[0086] in, This represents the low-frequency information of the first layer of the synchronous compressed wavelet transform. abbreviated identifier; similarly, and These are abbreviations representing the low-frequency information of the second and third layers of the synchronous compressed wavelet transform, respectively. This represents the high-frequency information of the first layer of the synchronous compressed wavelet transform. abbreviated identifier; similarly, and This is a shorthand identifier representing the high-frequency information of the second and third layers of the synchronous compressed wavelet transform.

[0087] Since the synchronous compressed wavelet transform does not reduce the length of the original signal in each decomposition, only the low-frequency similarity decomposition features are used as the cumulative expression, resulting in the triple cumulative transform formula, as shown in the following equation:

[0088] ;

[0089] in, This represents the triple cumulative eigenvalue obtained based on the above values ​​obtained from the synchronous compression transform.

[0090] Furthermore, following the principle of time-frequency domain cross-transformation, and using a simplified form:

[0091] ;

[0092] in, For Fast Fourier Transform, This represents the triple cumulative time-frequency domain features obtained based on triple cumulative eigenvalues; Indicates based on The first low-frequency time-frequency domain eigenvalues ​​after Fourier transform of the features; express The second layer of low-frequency time-frequency domain eigenvalues ​​after Fourier transform; express The features are the third-level low-frequency time-frequency domain features after Fourier transform.

[0093] The continuous discrete decomposition transformation of the time-domain wavelet to the time-frequency domain reference plane is specifically as follows:

[0094] ;

[0095] in, This represents the complex exponential basis function.

[0096] Since the energy value of synchronous compressed wavelet transform increases rapidly with the increase of decomposition level, an energy suppression step is performed to balance the weight values ​​of multi-level decomposition:

[0097] ;

[0098] ;

[0099] in, This is represented as calculating the low-frequency decomposition value of the three-layer feature. , , The total energy; Let i represent the low-frequency feature value of the i-th layer decomposition, where i ranges from 1 to 3; , , These are the weight values ​​for balancing the first, second, and third levels of decomposition, respectively.

[0100] The energy calculation formula E is specifically expressed as follows:

[0101] ;

[0102] in, Indicates input After that, I received The corresponding energy value.

[0103] The final expression for synchronous compressed wavelet transform multi-level decomposition energy suppression feature extraction is:

[0104] ;

[0105] in, This represents the final extracted waveform feature value of the sound signal.

[0106] At this point, the wavelet features obtained from the 3-level decomposition approximation are aligned with the 3-channel RGB image features. The weight allocation suppresses the energy growth problem of the multi-level decomposition and has excellent representational properties of voiceprint features.

[0107] In this embodiment, high-frequency information is extracted from multi-layer decomposition data and superimposed, and time-frequency transformation is performed on the superimposed high-frequency information to obtain target spectrogram features, including:

[0108] The high-frequency information separated from the multi-layer decomposed data through a high-pass filter is obtained and superimposed. The superimposed high-frequency information is subjected to continuous short-time Fourier transform to convert it into the corresponding time-frequency map. Based on the time-frequency map, a 3-channel target spectrogram feature is generated.

[0109] In one embodiment, refer to Figure 2 As shown, the specific process of extracting target features from anti-interference spectrograms based on multi-level decomposition is as follows:

[0110] First, extract the decomposition results obtained in the above embodiments. , , This information is high-frequency information from synchronous compression transformation and multi-layer decomposition. For audio signals, the transformation and jitter of high-frequency information are richer, and low-frequency background noise has been eliminated. It is more reasonable to convert it into a time-frequency diagram as a voiceprint feature.

[0111] Secondly, the above high-frequency information is directly superimposed:

[0112] ;

[0113] in, Indicates the first High-frequency information of the layer This represents the superposition value of three high-frequency features in the synchronous compressed wavelet transform eigenvalue decomposition.

[0114] Finally, Directly convert the time-frequency graph. The time-frequency graph transformation is a continuous short-time Fourier transform process:

[0115] ;

[0116] in, Represents the short-time Fourier transform. For the signal that needs to be transformed, For continuous transformation windowing functions, As the center point of the time domain, The center frequency value, The complex exponential basis function is calculated by adding 2πk / N.

[0117] Therefore, the high-frequency anti-interference audio signal based on multi-layer decomposition is:

[0118] ;

[0119] The above results in the anti-interference acoustic spectrogram target features based on multi-level decomposition. This image is a 3-channel color spectrogram matrix, and the matrix is ​​related to... Shape alignment. Among them, It is a three-channel graphical feature vector. This represents the graph eigenvector of channel 1. This represents the graph eigenvector of channel 2. This represents the graph feature vector of channel 3. Ultimately, we obtain the following: Figure 2 The multi-modal anti-dynamic interference characteristics of the multi-layer decomposition are shown.

[0120] In this embodiment, before constructing the feature extraction network model, the two modal features are first effectively fused. This allows the fused features to express the advantages of both waveform and graphic features, and the final fused features are effectively applicable to the subsequent detection network. Therefore, in the above feature extraction process, the semantic morphology of the two modal features is pre-aligned. Furthermore, fusing these two modal features into an effective 3-channel image not only improves the anti-interference ability and feature expression ability of a single modal feature, but also provides a method for multi-modal feature fusion.

[0121] Secondly, to adapt to dynamic wind noise, this embodiment proposes a fusion feature weight trainer and a three-dimensional parallel PSO dynamic weight search algorithm, which adaptively and dynamically adjusts the weight values ​​involved in the dual-modal feature fusion to adapt to wind noise, and provides feature compensation values ​​that can be adaptively adjusted due to wind noise loss in dynamic fusion, which greatly improves the confidence of fusion features and robustness in complex environments.

[0122] In addition, this embodiment also constructs a fast target detection network based on deep learning for the 3-channel image information features. According to the final fusion form of the modal features, the network utilizes and optimizes part of the general YOLO network structure to complete the deep extraction of image features with smaller network parameters. It also combines a small-parameter convolutional network and a dense network to extract and train the fused features, which can output detection results faster than conventional image detection networks.

[0123] In this embodiment, refer to Figure 4 As shown, feature fusion is performed on the target waveform features and target spectrogram features, and the feature fusion weights are optimized through three-dimensional parallel PSO dynamic weight search to obtain adaptive wind noise fused features, specifically including:

[0124] The target waveform features and target spectrogram features, along with the initial waveform feature weights, initial image feature weights, and bias terms, are input into a preset fusion feature weight trainer. Based on the initial waveform feature weights, initial image feature weights, and initial bias terms, a three-dimensional parallel PSO dynamic weight search is performed to obtain the globally optimal waveform feature weights, optimal image feature weights, and optimal bias terms.

[0125] The target waveform features and target spectrogram features are normalized by a preset fusion feature weight trainer. The normalized target waveform features and target spectrogram features are then linearly combined according to the optimal waveform feature weight and the optimal image feature weight. Finally, a compensation term calculated based on the optimal bias term is superimposed to obtain the adaptive wind noise fusion feature.

[0126] Specifically, the dual-mode fission feature fusion process for adaptive dynamic wind noise is as follows:

[0127] First, obtain the target waveform features by acquiring the multilayer suppression energy decomposition:

[0128] ;

[0129] Then obtain the multi-layer high-frequency energy decomposition target spectral characteristics:

[0130] ;

[0131] The above steps have aligned the morphological features of the two features. Here, the two features are fused, specifically as follows:

[0132] ;

[0133] in, As the initial waveform feature weights, As the initial image feature weights, Indicates fusion features, norm The expression represents the 255 quantization elements for aligning image channels, specifically as follows:

[0134] ;

[0135] in, This represents the variable value of the input Norm; This represents the minimum value of data; express The maximum value.

[0136] At this point, the effective fusion of the dual-modal features of the fission is completed, and the three-channel image elements of RGB are aligned to facilitate the construction of the subsequent fast target detection network.

[0137] The weight values ​​of the two features in the above situation under high signal-to-noise ratio data , All weights are set to 0.5. This weight is an empirical value under ideal data. However, it is difficult to obtain a fixed empirical value based on the dynamic wind noise weight value. It needs to be readjusted and optimized according to the actual situation.

[0138] In this embodiment, the training process of the preset fusion feature weight trainer is as follows:

[0139] A dual-channel residual network structure is used to construct a fusion feature weight trainer. The initial waveform feature weights, initial image feature weights, initial bias terms, target waveform features, and target spectrogram features are set as inputs to the trainer, and the class probability distribution of low, slow, and small targets is used as the output.

[0140] Initial waveform feature weights and initial image feature weights Set to 0.5, initial bias term Set to 0, and train the fusion feature weight trainer using high signal-to-noise ratio data. Save the trainer parameters after training. An initial bias term is added to compensate for the target feature loss caused by dynamic interference wind noise. .

[0141] Among them, reference Figure 3 As shown, the fusion feature weight trainer specifically includes: a fusion feature layer, which is used to directly superimpose the dual-channel output values.

[0142] The probability distribution layer, also known as the softmax layer, is a regular output layer in deep learning networks, representing the probability distribution of three targets in the output.

[0143] The target decision layer is used to output specific low, small, and slow targets, which include UAV target 1, UAV target 2, and bird target.

[0144] Specifically, this embodiment uses the PSO algorithm to perform a three-dimensional parallel PSO dynamic weight search. The PSO algorithm is an automatic search algorithm that simulates the social behavior of flocks of birds or schools of fish. It motivates a swarm of particles, causing each particle to move within the search space to find a local optimum, and then uses the particle swarm behavior to find the global optimum.

[0145] The typical iterative formula for the PSO algorithm is as follows:

[0146] ;

[0147] in, For the new iteration value, Represents the value of the previous iteration. Represents the inertia weighting factor. , Represents learning factors, , Represents a random number, used to increase the randomness of ion searches. This represents a local optimum. Represents the globally optimal solution;

[0148] The specific process of PSO location update and search completion is shown in the following formula:

[0149] ;

[0150] in, The position parameter of the last search. The position shift obtained from the iterative formula, These are the position parameters obtained after one search.

[0151] Each particle will go through the above iterative process and obtain its own local optimum. If any particle Compared to the current global optimal solution If it's better, then update. .

[0152] The overall formula input value for the PSO automatic search algorithm is: The output value is a single particle. and particle swarm .

[0153] The PSO algorithm suffers from getting stuck in local optima and failing to find the global optimum in a short number of iterations. Therefore, optimization is necessary. This embodiment proposes a three-dimensional parallel PSO dynamic weight search algorithm to better adapt to dynamic wind noise and automatically compensate for feature loss. Leveraging the PSO algorithm's faster speed than deep learning models in dynamically deciding weights, the algorithm strategically optimizes the weights and biases to be adjusted in this embodiment. This avoids the particle swarm getting stuck in local optima and allows the same 60 particles to quickly find the global optimum in fewer iterations.

[0154] Input values ​​for the search algorithm in this embodiment:

[0155] Weight , and bias It can be viewed as a three-dimensional coordinate system. , , );

[0156] The search algorithm in this embodiment outputs the following value in a single iteration:

[0157] =-1×characteristic variance + standard deviation of characteristic probability distribution

[0158] Based on a set obtained from each iteration of the search algorithm ( , , ) and the extracted target waveform features and target graphic features of the acoustic signal under dynamic wind noise interference, through Figure 3 The trained fusion feature trainer yields a fusion feature layer, which is a dense layer with 128 bits. The output value of the fusion feature layer is represented as... .based on =0.5, =0.5, =0, following the above process, another set of fused feature layers is obtained, and its output value is expressed as Therefore, the characteristic variation value can be calculated, and the specific calculation formula is as follows:

[0159] ;

[0160] in, Indicates based on search weight ( , , The i-th output value of the corresponding 128-dimensional dense layer, where i ranges from 1 to 128; This represents the i-th output value of the 128-dimensional dense layer corresponding to the initial weights (0.5, 0.5, 0).

[0161] Standard deviation of characteristic probability distribution:

[0162] The standard deviation of the feature probability distribution is based on ( , , ) Dense layer of fused features under weighted inference The next layer probability distribution value The distribution represents the probability of outputting the three targets, so we have:

[0163] Standard deviation of feature probability distribution = ;

[0164] in, Indicates the first i The probability distribution value of the layer.

[0165] In this embodiment, a three-dimensional parallel PSO dynamic weight search is performed based on the initial waveform feature weights, initial image feature weights, and initial bias term to obtain the globally optimal waveform feature weights, optimal image feature weights, and optimal bias term, including:

[0166] Six particle swarms are initialized, each containing 10 particles, and the initial waveform feature weights, initial image feature weights, and initial bias terms are used as the three-dimensional coordinates. , , The initial positions of the particle swarm are set to (0.5, 0.5, 0), and the inertia weighting factor is... and learning factors , All values ​​are set to 0.5, and different initial position offsets are set for each particle swarm in different search directions.

[0167] Each particle swarm is assigned an independent iterative direction vector, and the six particle swarms perform the optimal value search in parallel. The local optimal solution of each particle and the global optimal solution of the particle swarm are determined by calculating the feature mutation value and the standard deviation of the feature probability distribution.

[0168] If the optimization magnitude of the global optimal solution of a certain particle swarm is less than 0.1 for three consecutive iterations, the particle swarm is judged to be trapped in a local optimum. Local optimum decoupling is performed according to the preset decoupling logic until the best waveform feature weight, best image feature weight and best bias term of the global optimum are obtained.

[0169] Specifically, an offset is set for each particle group, so that the six particle groups search from the beginning based on six different directions. The initial position offsets of the six particle groups are as follows: Particle group 1 (0.55,0.5,0), Particle group 2 (0.5,0.55,0), Particle group 3 (0.5,0.5,0.05), Particle group 4 (0.55,0.55,0), Particle group 5 (0.5,0.55,0.05), and Particle group 6 (0.55,0.5,0.05).

[0170] In this study, each particle swarm is independently searched using the PSO iterative formula and search algorithm, resulting in six independent and parallel search behaviors for the six particle swarms. These six independent search behaviors are configured with six different iteration directions: Particle swarm 1 (1, s11, s12); Particle swarm 2 (s21, 1, s22); Particle swarm 3 (s31, s32, 1); Particle swarm 4 (1, 1, s4); Particle swarm 5 (s5, 1, 1); and Particle swarm 6 (1, s6, 1).

[0171] Taking particle swarm 1 as an example, the specific search behavior is as follows:

[0172] (1) Initialization:

[0173] =(0.5,0.5,0)+(0.05,0,0)=(0.55,0.5,0);

[0174] Where (0.55, 0.5, 0) is the initial position of particle swarm 1. The initial position for the group is (0.5, 0.5, 0). , The offset position of particle swarm 1 along the x-axis is (0.05, 0, 0).

[0175] Value range constraint: Restricting iterative updates ( , () represents a value between [0,1] containing two decimal places; The value is limited to a range of [0, 0.5] with two decimal places.

[0176] Iterative displacement update:

[0177] ;

[0178] in, This is the new iteration value, i.e., the position shift obtained from the iteration formula; Represents the value of the previous iteration; Represents the inertia weighting factor; , Represents learning factors; , Represents a random number, used to increase the randomness of ion searches; This represents a local optimal solution (the calculation process has been explained in the iterative output value of the search algorithm of this invention mentioned above). This represents the globally optimal solution.

[0179] Value range constraint: Inertia weighting factor w and learning factors , All are set to 0.5; , For each round of iteration, a random number is generated.

[0180] Search algorithm location update:

[0181] ;

[0182] in, The position parameter of the last search. The position shift obtained from the iterative formula, The position parameters are obtained after one search. (1, s11, s12) is the iteration direction vector of particle swarm 1, representing that the particle swarm's main iteration direction vector is along the x-axis. s11 and s12 are random values ​​between [0, 0.1] with two decimal places. Their function is to calculate... In the formula , similar.

[0183] The other five particle swarms simultaneously perform parallel and independent searches according to the search behavior of particle swarm 1.

[0184] At this point, the six particle swarms, starting from their collective initial position (0.5, 0.5, 0) plus six initial offsets in different directions, perform parallel searches for their respective swarm's optimal value along their own search directions. These six independent search behaviors in different directions allow the algorithm's search range to sufficiently cover the entire three-dimensional coordinate system, significantly improving the search breadth and thus increasing search efficiency.

[0185] In this embodiment, local optimal decoupling is performed according to preset decoupling logic, including:

[0186] Determine if the current global optimum of the particle swarm is also the global optimum of all six particle swarms. If it is, retain the global optimum and reduce the inertia factor of the current particle swarm. Increase learning factor and ;

[0187] If it is not the global optimal solution of the six particle swarms, then the iteration direction vector of the current particle swarm is modified based on the position of the global optimal solution of the six particle swarms, so that the iteration direction vector of the current particle swarm evolves toward the direction vector of the global optimal solution of the six particle swarms.

[0188] In this embodiment, refer to Figure 5 As shown, the YOLO network is invoked and its structure is optimized based on fused features to obtain a feature extraction network model and an object detection network model, including:

[0189] Access the YOLO network. Figure 5 In the YOLO network, the network structure consists of a backbone network, a neck network, a head network, and a feature detection network connected in sequence.

[0190] Three convolutional structures of size 32×3×3 are used as the three backbone networks of the YOLO network, and one head network is retained. The head network is connected to the neck network with two convolutional structures of size 32×3×3 to obtain the feature extraction network model.

[0191] The YOLO network's feature detection network is divided into multiple parallel branches, and a feature concatenation layer is set to connect the outputs of multiple parallel branches. At the same time, a pooling layer, a lightweight Dense layer, and a classification decision layer are added after the feature concatenation layer to obtain the object detection network model. Each parallel branch adopts a stacked structure of convolutional layer + BN layer + ReLU activation function + Dense layer.

[0192] Specifically, this embodiment constructs a YOLO fusion feature extraction network and a fast target detection network based on fission dual-mode fusion features. The feature extraction network is designed for fission dual-mode feature fusion features, which are represented as RGB image features. Therefore, it is constructed based on a general YOLO network. However, in order to quickly and accurately extract the key feature information of the high- and low-frequency superposition features extracted from acoustic non-stationary signals, the general YOLO network structure is optimized and simplified here. The specific structural optimization is as follows: Figure 5 As shown, and as follows:

[0193] Feature optimization based on the backbone network: The original YOLO network uses a backbone network for initial image feature extraction, while the two fission features in this invention are RGB image features based on multi-layer decomposition feature recombination. Therefore, the 3-channel features are extracted independently first, and then merged into the neck network for the next step of feature extraction. (Generally, YOLO uses the relatively lightweight backbone model ResNet-18 with 33 million parameters. Here, the backbone network is a 32×3×3 structure with 20,000 parameters, while the three backbone networks have a total of 60,000 parameters, which is only 0.2% of the parameters of ResNet-18).

[0194] Feature optimization based on the head network: The original YOLO network splits the features of the neck network into two streams. However, the ultimate goal of this invention is only to detect the target and the target category. Therefore, only one head network is used for category feature extraction. Based on the previous multi-channel independent feature extraction and merging, YOLO fusion features are extracted here. (The design idea of ​​the head network in the general YOLO structure is adopted here to connect with the output features of the neck network and design two 32×3×3 overall networks).

[0195] A fast target detection network based on a feature detection network: Based on the multi-level decomposition features of this invention, the YOLO fusion features are still a further feature representation of the three-channel multi-level decomposition features, therefore, as shown here... Figure 6 As shown, the feature detection network is split and then fused, and further hierarchical abstraction and fusion unified expression processes are performed to increase the final category detection accuracy (because the underlying logic of the general YOLO detection network is still binary classification detection between 0 and 1, while the method of this invention can be further extended to target detection and multi-class recognition. This part increases the network size and parameters compared to the general YOLO, which is used to improve the practical application possibility and scalability of the detection network of this invention).

[0196] In summary, the YOLO fusion feature extraction and fast detection network construction of this invention has a total parameter scale of only 0.3% of that of a typical YOLO detection network. It effectively extracts abstract features from image information by leveraging the backbone, neck, and head network structure design of the YOLO network, while still maintaining the effectiveness of the hierarchical abstract features of the fission multimodal features.

[0197] In this embodiment, the search behavior itself ensures a global competitive relationship among the six particle groups in six directions, making it less likely for the overall search behavior to get trapped in local optima. Secondly, communication relationships are established between the six particle groups to decouple the local optima problem of a single particle group. Taking particle group 1 as an example again.

[0198] When the global optimum of this particle swarm Value (6 independent search actions result in 6 global optima) If the difference between the output values ​​in three consecutive iterations is less than 0.1, it is considered to be trapped in a local optimum, and is handled in the following two ways:

[0199] like If the current optimal value is already found for the six particle swarms, then retain the current optimal value and reduce the inertia factor of that particle swarm. Increase learning factor and This improves search speed and search displacement, allowing for a quick escape from local values.

[0200] like If it is not the optimal value for the 6 particle swarms, then modify the direction values ​​(1, s11, s12): At this point, the optimal value for the 6 particle swarms is located... ; This represents the x-axis coordinate corresponding to the location of the optimal value. This represents the y-coordinate corresponding to the location of the optimal value. This represents the z-axis coordinate corresponding to the optimal value position.

[0201] Particle swarm 1 is currently trapped in a local optimum. Location ;in, This represents the x-axis coordinate corresponding to the location of the local optimum. This represents the y-coordinate corresponding to the location of the local optimum. This represents the z-axis coordinate corresponding to the local optimum; the iteration direction vector (1, s11, s12) of particle swarm 1 is changed to evolve towards the direction vectors of the six particle swarm optima.

[0202] ;

[0203] During the search execution in the above embodiments, if a single particle swarm is locally optimal, the decoupling method described above is used to improve information exchange between particle swarms. This ensures both the breadth of the six parallel search behaviors and information exchange between the six parallel searches, significantly improving the search efficiency of this search algorithm.

[0204] In summary, the steps of the above embodiments complete the setup and execution flow of this search algorithm. It is worth mentioning that, based on the three-dimensional parallel PSO dynamic weight search algorithm of this embodiment, under the premise of 60 particle swarms, the ordinary PSO algorithm needs to calculate the global optimum through 50 iterations, and may encounter the phenomenon of local optima not being decoupled after 10 iterations; while the three-dimensional parallel PSO dynamic weight search algorithm of this embodiment can obtain the global optimum after 5 parallel iterations, significantly improving the search efficiency, thereby improving the adjustment speed of adaptive wind noise fusion mode weights and feature compensation.

[0205] Specifically, based on the three-dimensional parallel PSO dynamic weight search algorithm, the following is obtained: The adaptive wind noise fusion features are obtained as follows:

[0206] ;

[0207] in, These are the optimal weights and feature loss bias supplementary values ​​obtained through the weight trainer and the three-dimensional parallel PSO dynamic weight search, respectively. The matrix corresponding to the target waveform features extracted in the above embodiments is a 3-channel matrix. The matrix corresponding to the target graphic features extracted in the above embodiments is also a 3-channel matrix. These are commonly used normalization formulas for graphical matrices (these three expressions have been described and reasoned previously). It is a 3-channel matrix of all ones, used to match the bias term. The operation.

[0208] Specifically, in this embodiment, after obtaining the feature extraction network model and the object detection network model, the feature extraction network model and the object detection network model are trained based on the training set and the test set, and the model parameters are saved. At the same time, the validity of the model is verified using the validation set.

[0209] The training process of the model is as follows:

[0210] 1) The training and testing datasets of the previously constructed UAV detection network are used: including 1600 multi-rotor UAV acoustic data (including 2 UAVs) and 800 bird sound datasets, labeled as UAV 1, UAV 2 and birds (the training and testing ratio is 3:1, and the datasets are all 30dB high signal-to-noise ratio data).

[0211] 2) The YOLO fusion feature extraction network and the fast object detection network were trained based on the training set, with training parameters of learning_rate=0.01, batch_size=64, and epoch=500.

[0212] 3) Once the loss values ​​of both the training and test sets have stabilized, save the model parameters with the highest accuracy on the test set.

[0213] The effectiveness verification includes: 1) an ablation experiment between the adaptive wind noise fusion feature method and single-modal detection to verify the effectiveness of the proposed method in utilizing multimodal feature fission and adaptive wind noise fusion; 2) a reliability comparison with mainstream acoustic feature detection methods for low, slow, and small targets under different signal-to-noise ratios to verify the wind noise resistance performance of the proposed method; and 3) a comparison with the timeliness of mainstream acoustic target detection networks to verify the computational efficiency of the proposed method.

[0214] like Figure 7 As shown, the overall process for detecting small, slow targets in this invention includes:

[0215] 1) Construct the dataset;

[0216] 2) Extract multi-level stationary features, i.e., multi-level data decomposition;

[0217] 3) Extract modal feature 1, i.e., target waveform features, which have anti-interference ability and stability characteristics;

[0218] 4) Extract modal feature 2, i.e. target graphic features, which have anti-interference ability and spectral characteristics;

[0219] 5) Adaptive wind noise fusion of modal feature 1 and modal feature 2, and provision of feature loss compensation;

[0220] 6) Construct a feature extraction network model;

[0221] 7) Construct an object detection network model;

[0222] 8) Train the feature extraction network model and the object detection network model, and obtain the final model parameters;

[0223] 9) Deploy feature extraction network models and target detection network models to achieve detection of small, slow targets.

[0224] Specifically, the ablation verification process for the validity of the dual-mode features based on acoustic data fission in this embodiment is as follows:

[0225] This embodiment splits a single dataset into two types of model features and performs adaptive fusion with wind noise to determine if it significantly improves performance. The verification method is based solely on the effectiveness of the model features. A general convolutional network with a combination of conv(3,3) is used as the target detection network. The training and test sets of the target detection network from step three are used to save the network training set parameters. Then, a verification dataset (containing multiple signal-to-noise ratio levels) is used to compare the effectiveness of the features.

[0226] As shown in Table 1 below, the overall accuracy of adaptive fusion features is the highest. Secondly, the accuracy of fixed fusion features with empirical values ​​and no feature compensation drops sharply as the signal-to-noise ratio decreases, but it is still higher than that of single modal features. The detection accuracy of several single modal features decreases significantly as the signal-to-noise ratio decreases, especially at the 0dB stage. However, the accuracy of adaptive fusion features only decreases by 6%, modal feature 1 decreases by 7%, and modal feature 2 decreases by 14%.

[0227] Table 1 Comparison of Ablation Experiments

[0228] Therefore, the adaptive fusion feature not only has strong anti-interference ability under the overall mixed signal-to-noise ratio, but its superior anti-interference ability is even more significant under the condition of sharp deterioration of signal-to-noise ratio. This shows the effectiveness of the weight trainer and parallel PSO weight search algorithm proposed in this invention. Secondly, the performance of the fixed fusion feature still shows that multimodal feature fusion can effectively make up for the defects of two single features, bringing about the enhancement of anti-interference performance and overall accuracy.

[0229] This embodiment has been compared and comprehensively verified for effectiveness with mainstream feature extraction methods. The specific process is as follows:

[0230] 1) Comparison of feature extraction methods:

[0231] This section only tests the effectiveness of feature extraction. Therefore, the experiment adopts the feature extraction method of the above embodiment to test the effectiveness of features, uses high signal-to-noise ratio data to train the model, and uses mixed signal-to-noise ratio data to verify the model.

[0232] The following comparison methods were adopted:

[0233] Method 1 Feature Extraction: EEMD-ICA feature extraction method (Source identification of gasoline engine noise based on continuous wavelet transform and EEMD–RobustICA), a sound signal enhancement method used for separating sound signals from noise;

[0234] Method 2 Feature Extraction: The fast-ICA feature extraction method (Acoustic UAV detection method based on blind source separation framework) uses feature extraction to achieve blind signal separation, and is also used for target detection.

[0235] Mainstream feature extraction methods: MFCC, GTCC, commonly used audio signal feature extraction methods based on mainstream feature analysis.

[0236] The results are shown in Table 2. Compared with the modal 1 features (target waveform features) and adaptive fusion features of the present invention, the above method shows that the features extracted by the present invention have a greater advantage in mixed signal-to-noise ratio, whether in the comparison of single features or adaptive fusion features. This demonstrates that the feature extraction algorithm of the present invention has strong anti-interference ability and the adaptive fusion algorithm enables the fusion features to have strong flexibility and adaptability in harsh environments.

[0237] Table 2 Comparison of Feature Extraction Experiments

[0238]

[0239] This embodiment verifies the effectiveness of the target detection network and the overall method steps. The verification process is as follows:

[0240] This embodiment tests the effectiveness of the target detection network and the overall method steps of this embodiment. Therefore, the experiment adopts the feature extraction method of the above embodiment to test the feature effectiveness, uses high signal-to-noise ratio data to train the model, and uses mixed signal-to-noise ratio data to verify the model.

[0241] Adaptive fusion features are used as common features, and the following object detection network is employed for validation:

[0242] Mainstream detection networks: SVM, KNN; basic detection networks used through feature extraction;

[0243] Method 1 Detection Network: YAMNet (Analysis of Distance and Environmental Impact on UAV)

[0244] Acoustic Detection); reducing the signal-to-noise ratio by increasing the distance is also a detection method under interference.

[0245] Method 2 uses the HRL (UAV Detection Using Reinforcement Learning) network, which is inspired by the radio frequency fingerprint detection network.

[0246] As shown in Table 3, when the adaptive fusion feature of this invention is compared with other target detection networks, its effectiveness is similar to that of using a simple convolutional network; while the effectiveness of using mainstream detection networks SVN and KNN is significantly reduced, which means that the effectiveness of SVN and KNN, which are commonly used in acoustic experiments, is significantly lower than that of simple CNN networks; while the method of using the adaptive fusion feature + target detection network of this invention can achieve a detection success rate of 94% even at -10dB.

[0247] Although this invention describes the feature extraction network and the object detection network as two separate networks, it is actually similar to the comparison methods YAMNET and HRL. The object detection network built based on deep learning is a combination of deep learning object feature extraction network and final decision network. Therefore, the comprehensive comparison method in Table 3 is effective.

[0248] Table 3 Comparative Experiments of Target Detection Networks

[0249]

[0250] This embodiment also verifies the timeliness of the model. The specific verification process is as follows:

[0251] In the timeliness comparison, this embodiment adopts a comprehensive application of feature extraction, feature fusion, and object detection networks from the above embodiments, rather than just model application verification. Inference time refers to the processing time for feature processing and model inference of all validation set data, used to measure the timeliness of the method of this invention. Both SVM and KNN adopt the feature extraction method described in the paper "Acoustic UAV detection method based on blind source separation framework". The specific verification results are shown in Table 4 below.

[0252] Table 4 Comparison of Object Detection Networks

[0253]

[0254] Based on Table 4 above, the target detection network in this embodiment requires 0.8s. This is because the feature extraction algorithm it uses has a complexity of oNlog(N), which is relatively lightweight in terms of time complexity for feature processing in non-stationary signals. The adaptive PSO search algorithm is a parallel 10-particle scale with a very low number of iterations, and the inference detection model has a normal deep learning parameter scale. Therefore, the search and inference time is negligible, and the total target inference time is mainly the time consumed by feature processing. Among other methods, the time complexity of the methods used by SVM and KNN is olog(N^3). The feature extraction method of HRL is not disclosed, so the feature extraction method of this invention is used. Since HRL's inference process is different from that of general deep learning, its inference time is about 4s. The feature extraction method used by YAMNet includes 160 features, mainly MFCC. Its noise reduction effect is very limited. Its processing time is similar to that of this invention, but the final effect is very different. In Table 2, compared with the same network, the accuracy of MFCC features under low signal-to-noise ratio is only 30%, while the accuracy of the present invention is about 90%.

[0255] The target detection method in this embodiment has strong anti-interference ability, can adapt to different degrees of wind noise, has great adaptability to complex environments, and due to its low-complexity feature processing algorithm and the construction of a small-scale detection network, it has high timeliness in target analysis and detection, and can quickly process low, slow and small targets that intrude into the airspace.

[0256] In this embodiment, after the model validity verification is passed, the feature extraction network model and the target detection network model are deployed to the mobile terminal or ground station to detect low, slow and small targets and output the target detection results.

[0257] Specifically, such as Figure 8 As shown, in the early stages of low-profile, small, and slow-moving target detection, data reception and preliminary processing are achieved through radio equipment and ground-based software. Data feature construction and model training, including feature extraction network models and target detection network models, are then performed using GPU servers. After completing these preliminary preparations, there are two main application approaches:

[0258] The first method involves saving the model parameters and porting them to an AI computing embedded board. By configuring a real-time acoustic receiver, the device is loaded onto a mobile terminal for the detection and recognition of low, slow, and small targets. The mobile terminal can be a ground-based mobile patrol vehicle.

[0259] The second method involves directly configuring the model onto the server, and then managing and monitoring airspace safety directly through the server device on the control tower server.

[0260] The model in this embodiment has demonstrated strong support for the efficiency and reliability of real-time data processing and inference in preliminary verification. It can be readily applied to AI computing embedded boards and mounted on the mobile terminal shown in the illustration. The device terminal weighs no more than 1.5kg, making it highly flexible. Furthermore, it can directly perform large-scale data training and airspace management via fixed ground stations in the detection and recognition of low-speed, small targets in complex scenarios. Its high timeliness gives it excellent emergency handling capabilities in complex situations.

[0261] The application principles of this invention are not limited to the mobile terminal and ground tower monitoring equipment shown in the figure. Based on the lightweight advantage of the loading equipment and the timeliness of the algorithm model, the model of this invention can be widely applied to small embedded terminals, air-space-ground mobile devices, large fixed ground stations, etc., to perform integrated networking for low-speed, small target detection and recognition, thereby improving the ability to detect and recognize low-speed, small targets in extremely complex scenarios.

[0262] Example 2: Refer to Figure 9 As shown, based on Embodiment 1 above, this embodiment also provides an adaptive wind noise low-speed small target rapid detection system, the system comprising:

[0263] The dataset construction module is used to collect drone acoustic data and retrieve bird acoustic datasets to build a high signal-to-noise ratio dataset, training set, test set, and validation set.

[0264] The feature extraction module is used to perform synchronous compressed wavelet transform on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data is then subjected to time-frequency domain transformation and energy suppression processing to obtain the target waveform features. At the same time, high-frequency information is extracted from the multi-level decomposed data and superimposed. The superimposed high-frequency information is then subjected to time-frequency graph transformation to obtain the target spectrogram features.

[0265] The feature fusion module is used to fuse the target waveform features and the target spectrogram features, and optimize the feature fusion weights through three-dimensional parallel PSO dynamic weight search to obtain the adaptive wind noise fused features.

[0266] The model optimization module is used to call the YOLO network and optimize the network structure based on the fused features to obtain the feature extraction network model and the object detection network model.

[0267] The model training and validation module is used to train the feature extraction network model and the object detection network model using the training set and the test set, save the model parameters, and use the validation set to verify the effectiveness of the model.

[0268] The model deployment module is used to deploy the feature extraction network model and the target detection network model to mobile devices or ground stations after the model validity verification is passed, to detect low, slow and small targets, and output the target detection results.

[0269] Specifically, the system in this embodiment has the following technical advantages:

[0270] 1. This invention first decomposes non-stationary acoustic signals into multi-level features to extract detailed features and reduce interference. Then, through modal decomposition, the acoustic signals are decomposed into graphic data and waveform data to increase data modes and make up for the problem that the waveform features are not well matched with deep learning networks (most of which are derived from graphic extraction networks), as well as the problem that graphic data loses some detailed features (the original form of acoustics is waveform).

[0271] 2. This invention provides a method for adaptive fusion of two modal features under different wind noise conditions, effectively combining the advantages of different modal features. It also innovatively proposes an adaptive wind noise weight training model and weight search algorithm to quickly and efficiently complete the fusion of dual modal features. Furthermore, it provides an adaptively adjustable wind noise feature loss compensation value to quickly and efficiently solve the problem of feature performance degradation under dynamic wind noise. In addition, it provides an improved YOLO extraction model that fits the fused features and a fast target detection network for subsequent decision-making, which greatly improves feature utilization efficiency and target detection accuracy.

[0272] 3. This invention designs a verification and comparison experiment method with current mainstream methods. The main evaluation is carried out through ablation experiments, feature extraction comparison, detection network comparison, and time complexity comparison. In the ablation experiment, the adaptive fusion feature extraction has a 28% higher detection accuracy than the single feature extraction at low signal-to-noise ratio. The adaptive fusion feature extraction has a 10% higher detection accuracy than the fixed fusion feature extraction at low signal-to-noise ratio, which shows the effectiveness of the dual-mode feature extraction and fusion, weight training, and parallel PSO dynamic weight search algorithm. In the comparison of feature extraction methods, this patent outperforms mainstream methods such as MFCC by nearly 60% at low signal-to-noise ratios (SNR), and outperforms the latest methods by nearly 40% at low SNR, demonstrating the advanced nature of its feature extraction. In the comparison of target detection networks, based on the feature extraction method of this embodiment, the target detection network of this embodiment still outperforms conventional methods by 15% and the latest methods by 5% at low SNR, demonstrating the advanced nature of the deep learning detection model proposed in this patent. Regarding time complexity, this embodiment, combining feature extraction and target detection efficiency under the same data processing conditions, also exhibits significant advantages in response time and resource space compared to current mainstream and newer methods.

[0273] 4. This invention designs a way to apply the method to embedded applications on mobile terminals in complex environments and to applications on large ground station servers, solving the problem of moving the algorithm from theory to practical application.

[0274] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for rapid detection of small, slow-moving targets with adaptive wind noise, characterized in that, include: Acoustic data from drones and bird acoustic datasets were collected to construct a high signal-to-noise ratio dataset, training set, test set, and validation set. Synchronous compressed wavelet transform is performed on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data is then subjected to time-frequency domain transformation and energy suppression processing to obtain the target waveform features. High-frequency information is extracted from multi-layer decomposed data and superimposed. Then, time-frequency transformation is performed on the superimposed high-frequency information to obtain the target spectrogram features. Feature fusion is performed on the target waveform features and target spectrogram features, and the feature fusion weights are optimized by three-dimensional parallel PSO dynamic weight search to obtain adaptive wind noise fusion features; the three-dimensional parallel PSO specifically refers to: using the PSO automatic search algorithm, the initial waveform feature weights, initial image feature weights and initial bias terms as three-dimensional search coordinates, and performing dynamic weight search in parallel through 6 independent particle swarms. By calling the YOLO network and optimizing the network structure based on fused features, a feature extraction network model and an object detection network model are obtained. The feature extraction network model and the object detection network model are trained using the training set and the test set, and the model parameters are saved. At the same time, the effectiveness of the model is verified using the validation set. After the model validity verification is passed, the feature extraction network model and the target detection network model are deployed to mobile devices or ground stations to detect drones and output target detection results.

2. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The process involves collecting acoustic data from drones and retrieving bird acoustic datasets to construct a high signal-to-noise ratio dataset. Simultaneously, training, testing, and validation sets are constructed, including: The acoustic signals of various drones during flight were collected and labeled, and the collected acoustic signals were denoised to obtain a drone acoustic dataset. The bird sound signal dataset is retrieved from the bird data database. Noise reduction is performed on each bird recording in the bird sound signal dataset. The bird recordings are then divided into multiple acoustic data segments according to the concentrated areas of bird calls to obtain the bird acoustic dataset. High signal-to-noise ratio datasets, training sets, test sets, and validation sets were constructed by extracting drone acoustic data and acoustic data segments from drone acoustic datasets and bird acoustic datasets, respectively.

3. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The synchronous compressed wavelet transform of the acoustic signals in the high signal-to-noise ratio dataset is performed to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data undergoes time-frequency domain transformation and energy suppression processing to obtain the target waveform features, including: The acoustic signals in the high signal-to-noise ratio dataset are sampled into discrete signals, and synchronous compressed wavelet transform is performed on the discrete signals to decompose the discrete signals into multi-level decomposed data. Low-frequency and high-frequency information of each layer of decomposed data are separated by low-pass and high-pass filters; time-frequency domain transformation is performed on the low-frequency information of each layer to obtain the time-frequency domain characteristics of the low-frequency information of each layer. Calculate the energy proportion of low-frequency information in each layer, and suppress the energy of the time-frequency domain features of each layer according to the energy proportion to obtain the target waveform features.

4. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The process of extracting high-frequency information from multi-layer decomposition data, superimposing it, and performing time-frequency transformation on the superimposed high-frequency information to obtain target spectrogram features includes: The high-frequency information separated from the multi-layer decomposed data through a high-pass filter is obtained and superimposed. The superimposed high-frequency information is subjected to continuous short-time Fourier transform to convert it into the corresponding time-frequency map. Based on the time-frequency map, a 3-channel target spectrogram feature is generated.

5. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The process of fusing target waveform features and target spectrogram features, and optimizing the feature fusion weights through three-dimensional parallel PSO dynamic weight search to obtain adaptive wind noise fused features includes: The target waveform features and target spectrogram features, along with the initial waveform feature weights, initial image feature weights, and bias terms, are input into a preset fusion feature weight trainer. Based on the initial waveform feature weights, initial image feature weights, and initial bias terms, a three-dimensional parallel PSO dynamic weight search is performed to obtain the globally optimal waveform feature weights, optimal image feature weights, and optimal bias terms. The target waveform features and target spectrogram features are normalized by a preset fusion feature weight trainer. The normalized target waveform features and target spectrogram features are then linearly combined according to the optimal waveform feature weight and the optimal image feature weight. Finally, a compensation term calculated based on the optimal bias term is superimposed to obtain the adaptive wind noise fusion feature.

6. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The training process of the preset fusion feature weight trainer is as follows: A dual-channel residual network structure is used to construct a fusion feature weight trainer. The initial waveform feature weights, initial image feature weights, initial bias terms, target waveform features, and target spectrogram features are set as inputs to the trainer, and the UAV category probability distribution is used as the output. Set the initial waveform feature weights and initial image feature weights to 0.5, and the initial bias term to 0. Train the fusion feature weight trainer using high signal-to-noise ratio data, and save the trainer parameters after training.

7. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The three-dimensional parallel PSO dynamic weight search based on initial waveform feature weights, initial image feature weights, and initial bias term to obtain the globally optimal waveform feature weights, optimal image feature weights, and optimal bias term includes: Six particle swarms are initialized, each containing 10 particles. The initial waveform feature weights, initial image feature weights, and initial bias terms are used as the three-dimensional coordinates. The initial position of the particle swarm is set to (0.5, 0.5, 0). The inertia factor w and the learning factors c1 and c2 are both set to 0.

5. At the same time, different initial position offsets are set for each particle swarm in different search directions. Each particle swarm is assigned an independent iterative direction vector, and the six particle swarms perform the optimal value search in parallel. The local optimal solution of each particle and the global optimal solution of the particle swarm are determined by calculating the feature mutation value and the standard deviation of the feature probability distribution. If the optimization magnitude of the global optimal solution of a certain particle swarm is less than 0.1 for three consecutive iterations, the particle swarm is judged to be trapped in a local optimum. Local optimum decoupling is performed according to the preset decoupling logic until the best waveform feature weight, best image feature weight and best bias term of the global optimum are obtained.

8. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The step of performing local optimal decoupling according to preset decoupling logic includes: Determine whether the global optimal solution of the current particle swarm is the global optimal solution of all 6 particle swarms. If it is the global optimal solution of all 6 particle swarms, retain the global optimal solution and reduce the inertia factor w of the current particle swarm, while increasing the learning factors c1 and c2. If it is not the global optimal solution of the six particle swarms, then the iteration direction vector of the current particle swarm is modified based on the position of the global optimal solution of the six particle swarms, so that the iteration direction vector of the current particle swarm evolves toward the direction vector of the global optimal solution of the six particle swarms.

9. The adaptive wind noise-based fast detection method for small, slow targets, characterized in that, The process of calling the YOLO network and optimizing the network structure based on fused features to obtain a feature extraction network model and an object detection network model includes: The YOLO network is retrieved, and the network structure of the YOLO network includes a backbone network, a neck network, a head network, and a feature detection network connected in sequence. Three convolutional structures of size 32×3×3 are used as the three backbone networks of the YOLO network, and one head network is retained. The head network is connected to the neck network with two convolutional structures of size 32×3×3 to obtain the feature extraction network model. The YOLO network's feature detection network is divided into multiple parallel branches, and a feature concatenation layer is set to connect the outputs of multiple parallel branches. At the same time, a pooling layer, a lightweight Dense layer, and a classification decision layer are added after the feature concatenation layer to obtain the object detection network model. Each parallel branch adopts a stacked structure of convolutional layer + BN layer + ReLU activation function + Dense layer.

10. A rapid detection system for low-speed, small targets with adaptive wind noise, characterized in that, include: The dataset construction module is used to collect drone acoustic data and retrieve bird acoustic datasets to build a high signal-to-noise ratio dataset, training set, test set, and validation set. The feature extraction module is used to perform synchronous compressed wavelet transform on the acoustic signals in the high signal-to-noise ratio dataset to obtain multi-level decomposed data. The low-frequency information of the multi-level decomposed data is then subjected to time-frequency domain transformation and energy suppression processing to obtain the target waveform features. At the same time, high-frequency information is extracted from the multi-level decomposed data and superimposed. The superimposed high-frequency information is then subjected to time-frequency graph transformation to obtain the target spectrogram features. The feature fusion module is used to fuse the target waveform features and the target spectrogram features, and optimize the feature fusion weights through three-dimensional parallel PSO dynamic weight search to obtain adaptive wind noise fused features. The three-dimensional parallel PSO specifically uses the PSO automatic search algorithm, with the initial waveform feature weights, initial image feature weights and initial bias terms as three-dimensional search coordinates, and performs dynamic weight search in parallel through 6 independent particle swarms. The model optimization module is used to call the YOLO network and optimize the network structure based on the fused features to obtain the feature extraction network model and the object detection network model. The model training and validation module is used to train the feature extraction network model and the object detection network model using the training set and the test set, save the model parameters, and use the validation set to verify the effectiveness of the model. The model deployment module is used to deploy the feature extraction network model and the target detection network model to mobile devices or ground stations after the model validity verification is passed, so as to detect drones and output target detection results.

Citation Information

Patent Citations

  • CN119626256A

  • CN121092985A