Unmanned aerial vehicle classification method and system based on noise perception model
Through the combination of noise perception model and convolutional neural network, the accuracy and resource limitation of drone recognition technology in complex environments is solved, and a high-precision drone classification and supervision solution is realized.
Patent Information
- Application Number
- CN202510847652.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing drone identification technology is difficult to achieve high-precision classification in complex environments, especially in scenarios where light is insufficient or the target distance is far away, the accuracy and reliability of visual identification technology have significantly decreased, and it is difficult to widely use in scenarios with limited resources.
The noise perception model is used to combine the convolutional neural network, and noise quantization, noise reduction processing and feature enhancement are performed by acquiring the drone audio data, generating a Mel spectrum, and inputting the convolutional neural network for the probability output of the drone category.
Implement high-precision drone classification in complex environments, reduce dependence on visual recognition technology, and is suitable for resource-constrained scenarios, providing new drone supervision and air traffic management solutions.
Smart Images

Figure CN120375863A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of sound processing and unmanned aerial vehicle (UAV) technology, and particularly to a UAV classification method and system based on a noise perception model. Background Art
[0002] With the rapid development of UAV technology, UAVs are increasingly widely used in fields such as aerial photography, logistics, agricultural monitoring, and disaster relief. However, most of the existing UAV recognition technologies rely on visual or radar signals, and these methods have significant limitations in complex environments (such as at night, in haze, rain, snow, etc.) or when the line of sight is blocked, making it difficult to achieve high-precision UAV classification.
[0003] Traditional UAV classification methods mainly rely on image recognition technology, but in scenarios with insufficient light or a long target distance, the accuracy and reliability of image recognition significantly decrease. In addition, visual recognition technology has high requirements for computing resources and hardware devices and is difficult to be widely applied in resource-constrained scenarios.
[0004] UAVs generate specific sound signals during flight, and these sound signals are closely related to the type, structure, working state, etc. of the UAV. Compared with visual signals, sound signals have the advantage of being not restricted by light and line of sight and can work stably in complex environments.
[0005] The present invention proposes a UAV classification method based on sound processing by combining a noise perception model and a convolutional neural network. This method can not only effectively handle classification tasks in complex noise environments but also reduce the dependence on visual recognition technology, providing a new solution for UAV supervision, air traffic management, and safety protection. The present invention has broad application prospects and can be applied to fields such as military reconnaissance, border monitoring, and urban security. Summary of the Invention
[0006] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a UAV classification method and system based on a noise perception model, which solve the problems mentioned in the background art.
[0007] To achieve the above object, the present invention provides the following technical solution: A UAV classification method based on a noise perception model, comprising the following steps: Step S1: Obtain UAV audio data, process the UAV audio data to obtain mono UAV audio data, resample the mono UAV audio data to obtain resampled audio data, and calculate the resampled audio data to obtain the audio data energy; Step S2: Process the energy of the audio data to obtain time-domain quantization metrics and frequency-domain quantization metrics, and process the time-domain quantization metrics and frequency-domain quantization metrics to obtain the denoised time-domain signal and the denoised frequency-domain signal. Then, fuse the time-domain quantization metrics and frequency-domain quantization metrics with the denoised time-domain signal and the denoised frequency-domain signal to obtain the enhanced signal; Step S3: Optimize the enhanced signal using the Mel spectrum transform to obtain the low-frequency band and the high-frequency band of the enhanced signal. Analyze the low-frequency band and the high-frequency band of the enhanced signal to generate a mixed-resolution spectrum; Optimize the low-frequency band and the high-frequency band of the enhanced signal through the differential Mel filter bank to obtain the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient. Process the mixed-resolution spectrum and the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient to generate a Mel spectrogram; Step S4: Input the Mel spectrogram into a convolutional neural network for processing and output the probability of the UAV category.
[0008] Furthermore, the specific process of obtaining the enhanced signal is as follows: Perform frame division on the input audio data energy and then calculate through the short-time Fourier transform to respectively obtain the energy of the k-th frame and the estimated value of the noise energy of the k-th frame , and use the minimum statistical noise tracking algorithm to perform dynamic updates on the energy of the k-th frame and the estimated value of the noise energy of the k-th frame to obtain the time-domain quantization metric , which is expressed as: ; In the formula, represents the total number of frames of the audio data energy division; represents the mean value of the energy of the k-th frame; represents the standard deviation of the energy of the k-th frame; represents the hyperbolic tangent function; represents the adaptive adjustment factor; ; SNR represents the signal-to-noise ratio; log represents the logarithmic operation; Extract the audio data energy through the Mel filter bank to respectively obtain the signal spectrum components of the v-th frame and the noise spectrum components of the v-th frame , and use the noise floor estimation method to calculate the extracted signal spectrum components of the v-th frame and the noise spectrum components of the v-th frame to obtain the frequency-domain quantization metric , which is expressed as: ; Wherein, represents the total number of channels of the filter bank; represents the slope parameter of the Sigmoid function; represents the energy threshold; e represents the natural constant; Normalize the signal-to-noise ratio value range to [0,1] to obtain the normalized signal-to-noise ratio ; ; represents the maximum value of the signal-to-noise ratio; represents the minimum value of the signal-to-noise ratio; Through the normalized signal-to-noise ratio Calculate the parameters p and q, ; ; p represents the signal amplitude spectrum power adjustment parameter; q represents the dynamic noise suppression coefficient; Based on p and q, calculate at the frequency point The spectral attention mask value at, which means: ; Wherein, represents the spectral attention mask value at the frequency point ; represents the amplitude spectrum of the signal at the frequency point ; represents the average amplitude spectrum of the dynamic noise in the frequency band ; represents the adaptive scaling exponent Adjust through the frequency domain quantization index ; represents the balance factor; wherein, = ; Perform noise reduction calculation on the spectral attention mask value at the frequency point to obtain the denoised time-domain signal and the denoised frequency-domain signal ; represents the inverse Fourier transform; Through dynamic weight fusion, the denoised time-domain signal and the denoised frequency-domain signal and the time-domain quantization index and the frequency-domain quantization index , to obtain the enhanced signal , which means: .
[0009] Furthermore, the specific process of generating the Mel spectrogram is as follows: According to the noise environment, based on the time-domain quantization index and the noise spectrum entropy, calculate the resolution adjustment coefficient , based on the resolution adjustment coefficient , the time-frequency resolution of the Mel spectrum transform is used to dynamically adjust and optimize the enhanced signal to obtain the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal ; The resolution adjustment coefficient , which means: ; In the formula, sigmoid represents the activation function; Hnoise represents the noise spectrum entropy; w1 represents the time-domain quantization index weight; w2 represents the noise spectrum entropy weight; Analyze the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal to generate a mixed-resolution spectrum, which means: ; In the formula, represents the mixed-resolution spectrum; represents the low-frequency band spectrum of the enhanced signal ; represents the high-frequency band spectrum of the enhanced signal ; The differentiable Mel filter bank, based on the resolution adjustment coefficient , dynamically fuses the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal , which means: ; In the formula, represents the frequency band boundary of the th differentiable Mel filter bank under the control of the resolution adjustment coefficient ; represents the low-frequency band filter boundary of the th enhanced signal ; represents the high-frequency band filter boundary of the th enhanced signal ; Based on the frequency band boundary of the th differentiable Mel filter bank under the control of the resolution adjustment coefficient , define the weight coefficient of the th Mel filter at the frequency point ; Based on the mixed-resolution spectrum and the th Mel filter at the frequency point The weight coefficient at generates a Mel spectrogram, indicating: ; In the formula, represents the Mel spectrogram; represents the number of points of the fast Fourier transform FFT; represents the th Mel filter at the frequency point The weight coefficient at; represents the minimum value.
[0010] Furthermore, the convolutional neural network is composed of four convolutional blocks connected in sequence, and each convolutional block is composed of a convolutional layer, a ReLU activation layer, and a pooling layer connected in sequence; the Mel spectrogram is input into the convolutional neural network and passes through four convolutional blocks in sequence to output the category probability of the drone.
[0011] Furthermore, the specific process of obtaining the energy of the audio data is as follows: Convert the acquired drone audio data into a unified format to obtain multi-channel drone audio data in a unified format; Convert the multi-channel drone audio data into mono drone audio data, indicating: ; In the formula, represents the mono drone audio data; represents the total number of channels of the multi-channel drone audio data; represents the drone audio data of the cth channel; Resample the mono drone audio data to obtain resampled audio data, indicating: ; In the formula, represents the resampled audio data; represents the resampling operation; represents the original sampling rate of the mono drone audio data; represents the target sampling rate; Calculate the resampled audio data to obtain the audio data energy, indicating: ; In the formula, represents the audio data energy; represents the total number of resampling points; represents the th sampling point of the resampled audio data.
[0012] A drone classification system based on a noise perception model, applying the described drone classification method based on a noise perception model, includes: The acquisition module is used to obtain the audio data of the drone, process the audio data of the drone to obtain the mono drone audio data, resample the mono drone audio data to obtain the resampled audio data, and calculate the resampled audio data to obtain the audio data energy; The processing module is used to process the audio data energy to obtain the time-domain quantization index and the frequency-domain quantization index, and process the time-domain quantization index and the frequency-domain quantization index to obtain the denoised time-domain signal and the denoised frequency-domain signal, and fuse the time-domain quantization index and the frequency-domain quantization index with the denoised time-domain signal and the denoised frequency-domain signal to obtain the enhanced signal; The optimization module is used to optimize the enhanced signal by using the Mel spectrum transform to obtain the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal, analyze the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal to generate a mixed-resolution spectrum; optimize the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal through the differential Mel filter bank to obtain the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient, and process the mixed-resolution spectrum and the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient to generate a Mel spectrogram; The output module is used to input the Mel spectrogram into a convolutional neural network for processing and output the probability of the drone category.
[0013] An electronic device includes a processor, a memory, and a bus. The processor and the memory are connected through the bus. Among them, the memory is used to store a set of program codes, and the processor is used to call the program codes stored in the memory to execute the described drone classification method based on the noise perception model.
[0014] A non-volatile computer storage medium stores computer-executable instructions, and the computer-executable instructions are the described drone classification method based on the noise perception model.
[0015] Compared with the existing technology, the present invention has the following beneficial effects: (1) The present invention performs noise quantization, noise reduction processing, and feature enhancement on the drone audio data through the noise perception model, and combines a convolutional neural network to achieve high-precision drone classification; the noise perception model can effectively extract the essential features of the drone audio data, significantly improving the classification accuracy.
[0016] (2) The present invention is based on audio data processing technology, is not limited by light and line of sight, and can work stably in complex environments such as night, haze, rain, and snow, enhancing the environmental adaptability of drone recognition.
[0017] (3) The method of the present invention is simple and efficient, and is applicable to real-time UAV classification tasks; through the combination of the noise perception model and the convolutional neural network, it can be widely applied in resource-constrained scenarios, providing a new solution for UAV supervision, air traffic management, and security protection. Description of the Drawings
[0018] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0019] As Figure 1 shown, the present invention provides a technical solution: a UAV classification method based on a noise perception model, including the following steps: Step S1: Obtain UAV audio data, process the UAV audio data to obtain mono UAV audio data, resample the mono UAV audio data to obtain resampled audio data, and calculate the resampled audio data to obtain audio data energy; Step S2: Process the audio data energy to obtain time-domain quantization indexes and frequency-domain quantization indexes, and process the time-domain quantization indexes and frequency-domain quantization indexes to obtain a denoised time-domain signal and a denoised frequency-domain signal, and fuse the time-domain quantization indexes and frequency-domain quantization indexes with the denoised time-domain signal and denoised frequency-domain signal to obtain an enhanced signal; Step S3: Optimize the enhanced signal by using the Mel spectrogram transform to obtain the low-frequency band and high-frequency band of the enhanced signal, analyze the low-frequency band and high-frequency band of the enhanced signal to generate a mixed-resolution spectrum; optimize the low-frequency band and high-frequency band of the enhanced signal by using a differential Mel filter bank to obtain the frequency band boundaries of the differential Mel filter bank under the control of a resolution adjustment coefficient, and process the mixed-resolution spectrum and the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient to generate a Mel spectrogram; Step S4: Input the Mel spectrogram into a convolutional neural network for processing, and output the probability of the UAV category.
[0020] Among them, the convolutional neural network is composed of four convolutional blocks connected in sequence, and each convolutional block is composed of a convolutional layer, a ReLU activation layer, and a pooling layer connected in sequence; input the Mel spectrogram into the convolutional neural network and pass through the four convolutional blocks in sequence to output the category probability of the UAV; The operation of the convolutional layer is expressed as: ; In the formula, represents the value of the output enhanced feature map of the convolutional layer at the position ; represents the convolutional kernel size; represents the convolutional kernel at the position Weight value; Represents the input Mel spectrogram; Represents the bias term; ReLU activation layer operation, indicating: ; In the formula, Represents the output of the ReLU activation function; Represents the function to take the maximum value; Represents the value of the enhanced feature map output by the convolutional layer at position ; Pooling layer operation, downsampling the output of the ReLU activation function through the max pooling layer, indicating: ; In the formula, Represents the value of the feature map output by the pooling layer at position ; R represents the range of the pooling window; Represents the output of the ReLU activation function; Flatten the value of the feature map output by the pooling layer at position into a one-dimensional vector and input it into the fully connected layer, calculating the probability of the drone category through the Softmax function, indicating: ; In the formula, Represents the probability of the i-th drone category; e represents the natural constant; Represents the output value of the fully connected layer for the i-th drone category; Represents the output value of the fully connected layer for the -th drone category; Represents the total number of categories output by the fully connected layer.
[0021] Among them, the specific process of obtaining the energy of the audio data is as follows: Uniformly convert the acquired drone audio data into the wav format to obtain multi-channel drone audio data in the wav format; Convert the multi-channel drone audio data into mono-channel drone audio data, indicating: ; In the formula, Represents the mono-channel drone audio data; Represents the total number of channels of the multi-channel drone audio data; Represents the drone audio data of the c-th channel; Resample the mono-channel drone audio data to obtain the resampled audio data, indicating: ; In the formula, represents the resampled audio data; represents the resampling operation; represents the original sampling rate of the monaural drone audio data; represents the target sampling rate; Calculate the resampled audio data to obtain the audio data energy, which is expressed as: ; In the formula, represents the audio data energy; represents the total number of resampling points; represents the th sampling point of the resampled audio data.
[0022] Among them, the specific process of obtaining the enhanced signal is as follows: Time-frequency dual-stream noise quantization. An adaptive dual-stream noise perception algorithm (ADNPA) that fuses time-frequency domain features is used to quantize the audio data energy from two levels in the time domain and the frequency domain respectively, obtaining the time-domain quantization index and the frequency-domain quantization index ; The specific process of the time-domain quantization index is as follows: The input audio data energy is frame-divided and then calculated through the short-time Fourier transform (STFT) to obtain the energy of the kth frame and the noise energy estimation value of the kth frame. The minimum statistical noise tracking algorithm (MSNR) is used to dynamically update the energy of the kth frame and the noise energy estimation value of the kth frame to obtain the time-domain quantization index , which is expressed as: ; In the formula, represents the total number of frames of the audio data energy; represents the mean value of the energy of the kth frame; represents the standard deviation of the energy of the kth frame; represents the hyperbolic tangent function; represents the adaptive adjustment factor; ; SNR represents the signal-to-noise ratio; log represents the logarithmic operation; The specific process of the frequency-domain quantization index is as follows: Extract the audio data energy through the Mel filter bank to obtain the signal spectral component of the vth frame and the noise spectral component , the signal spectral components of the v-th frame extracted by using the noise floor estimation method (NBE) and the noise spectral components of the v-th frame , to obtain the frequency-domain quantization index , which is expressed as: ; In the formula, represents the filter channel index ; represents the total number of channels of the filter bank; represents the slope parameter of the Sigmoid function, controlling the steepness of the transition region; represents the energy threshold; β = 0.8; τ = 0.6; Normalize the signal-to-noise ratio value range to [0, 1] to obtain the normalized signal-to-noise ratio ; ; represents the maximum value of the signal-to-noise ratio; represents the minimum value of the signal-to-noise ratio; Through the normalized signal-to-noise ratio calculate the parameters p and q, ; ; p represents the signal amplitude spectrum power adjustment parameter; q represents the dynamic noise suppression coefficient; Based on p and q, calculate the spectral attention mask value at the frequency point , which is expressed as: ; In the formula, represents the spectral attention mask value at the frequency point ; represents the amplitude spectrum of the signal at the frequency point ; represents the average amplitude spectrum of the dynamic noise in the frequency band ; represents the adaptive scaling exponent Adjusted by the frequency-domain quantization index ; represents the balance factor; among them, = ; Perform noise reduction calculation on the spectral attention mask value at the frequency point to obtain the denoised time-domain signal and the denoised frequency-domain signal ; represents the inverse Fourier transform; Through dynamic weight fusion, the denoised time-domain signal and the denoised frequency-domain signal and the time-domain quantization index and frequency-domain quantization metrics , to obtain an enhanced signal , which means: .
[0023] Among them, the specific process of generating the Mel spectrogram is as follows: First, according to the noise environment, based on the time-domain quantization metrics and the noise spectrum entropy, calculate the resolution adjustment coefficient , based on the resolution adjustment coefficient adopt the time-frequency resolution of the Mel spectrum transformation to dynamically adjust and optimize the enhanced signal , obtain the low-frequency band (harmonic components) of the enhanced signal and the high-frequency band (transient noise) of the enhanced signal , to adapt to the feature extraction requirements of different frequency bands; The resolution adjustment coefficient , which means: ; In the formula, sigmoid represents the activation function; Hnoise represents the noise spectrum entropy (range 0-10), which characterizes the noise complexity; w1 represents the time-domain quantization metrics weight; w2 represents the noise spectrum entropy weight; α ∈ (0,1); when α → 1, the low-frequency high-resolution weight of the low-frequency band of the enhanced signal increases, and when α → 0, the high-frequency time-resolution weight of the high-frequency band of the enhanced signal increases; Dynamic adaptability: The traditional Mel spectrum transformation uses a fixed window, while the Mel spectrum transformation of the present invention adjusts the resolution adjustment coefficient in real time through a noise perception model , for example, when the time-domain quantization metrics is low, the high-frequency time-resolution weight increases; when the time-domain quantization metrics is high, the low-frequency high-resolution weight increases to extract the characteristics of harmonic components, and directly uses the time-domain quantization metrics and the noise spectrum entropy to form a closed-loop optimization, which is not achieved by the traditional method; Analyze the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal ; In the formula, represents the mixed-resolution spectrum; represents the enhanced signal The spectrum of the low-frequency band (0 - 2 kHz) uses the short-time Fourier transform STFT with a long time window (window length of 50 ms, number of points for fast Fourier transform FFT is 1024, frequency resolution is 10 Hz / bin); represents the enhanced signal The spectrum of the high-frequency band (greater than 2 kHz) uses the short-time Fourier transform STFT with a short time window (window length of 10 ms, number of points for fast Fourier transform FFT is 256, time resolution is 1 ms / bin); kHz is kilohertz; Hz / bin is frequency resolution; ms / bin is time resolution; The low-frequency harmonics of the drone rotor require high low-frequency resolution, and the high-frequency noise of the motor requires high high-frequency time resolution; the traditional Mel spectrogram with a fixed window (such as 25 ms) cannot balance both. The above analysis of the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal can balance the feature expression capabilities of the low-frequency band and the high-frequency band of the enhanced signal ; The differentiable Mel filter bank, based on the resolution adjustment coefficient , dynamically fuses the low-frequency band and the high-frequency band of the enhanced signal , expressed as: ; In the formula, ; where represents the frequency band boundary of the th differentiable Mel filter bank under the control of the resolution adjustment coefficient ; represents the filter boundary of the low-frequency band (0 - 2 kHz) of the th enhanced signal , designed to be densely spaced (one filter per 50 Hz); represents the filter boundary of the high-frequency band (greater than 2 kHz) of the th enhanced signal , designed to be sparsely spaced (one filter per 200 Hz); Based on the frequency band boundary of the th differentiable Mel filter bank under the control of the resolution adjustment coefficient , define the weight coefficient of the th Mel filter at the frequency point ; The frequency bands of the traditional Mel filter bank are fixed (such as 40 channels evenly distributed), while the adaptive multi-resolution time-frequency encoder AMR-TFE, through differentiable design, allows dynamic adjustment of the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal The high-frequency band filter boundary adapts to the physical characteristics of UAV audio data; Based on the mixed-resolution spectrum and the weight coefficient of the th Mel filter at the frequency point, generate a Mel spectrogram, which represents: ; In the formula, represents the Mel spectrogram; represents the number of points of the fast Fourier transform FFT; represents the th Mel filter at the frequency point ; represents the minimum value.
[0024] A UAV classification system based on a noise perception model, which is applied to the above-mentioned UAV classification method based on a noise perception model, includes: An acquisition module, which is used to acquire UAV audio data, process the UAV audio data to obtain mono UAV audio data, resample the mono UAV audio data to obtain resampled audio data, and calculate the resampled audio data to obtain audio data energy; A processing module, which is used to process the audio data energy to obtain time-domain quantization indexes and frequency-domain quantization indexes, and process the time-domain quantization indexes and frequency-domain quantization indexes to obtain a denoised time-domain signal and a denoised frequency-domain signal, and fuse the time-domain quantization indexes and frequency-domain quantization indexes with the denoised time-domain signal and the denoised frequency-domain signal to obtain an enhanced signal; An optimization module, which is used to optimize the enhanced signal by using Mel spectrum transformation to obtain the low-frequency band and high-frequency band of the enhanced signal, analyze the low-frequency band and high-frequency band of the enhanced signal to generate a mixed-resolution spectrum; optimize the low-frequency band and high-frequency band of the enhanced signal by using a differential Mel filter bank to obtain the frequency band boundary of the differential Mel filter bank under the control of a resolution adjustment coefficient, and process the mixed-resolution spectrum and the frequency band boundary of the differential Mel filter bank under the control of the resolution adjustment coefficient to generate a Mel spectrogram; An output module, which is used to input the Mel spectrogram into a convolutional neural network for processing and output the probability of the UAV category.
[0025] An electronic device includes a processor, a memory and a bus, the processor is connected to the memory through the bus, wherein, the memory is used to store a set of program codes, and the processor is used to call the program codes stored in the memory to execute the above-mentioned UAV classification method based on a noise perception model.
[0026] A non-volatile computer storage medium stores computer-executable instructions that can execute the described method for classifying drones based on a noise perception model.
[0027] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for classifying drones based on a noise perception model, characterized in that, It includes the following steps: Step S1: Obtain the UAV audio data, process the UAV audio data to obtain the mono UAV audio data, resample the mono UAV audio data to obtain the resampled audio data, calculate the resampled audio data to obtain the audio data energy; Step S2: Process the audio data energy to obtain the time-domain quantization index and the frequency-domain quantization index, and process the time-domain quantization index and the frequency-domain quantization index to obtain the denoised time-domain signal and the denoised frequency-domain signal, and fuse the time-domain quantization index and the frequency-domain quantization index with the denoised time-domain signal and the denoised frequency-domain signal to obtain the enhanced signal; Step S3: Optimize the enhanced signal by using the Mel spectrum transform to obtain the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal, analyze the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal to generate the mixed-resolution spectrum; optimize the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal through the differential Mel filter bank to obtain the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient, and process the mixed-resolution spectrum and the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient to generate the Mel spectrogram; Step S4: Input the Mel spectrogram into the convolutional neural network for processing and output the probability of the UAV category.
2. The drone classification method based on a noise perception model according to claim 1, characterized in that: The specific process of obtaining the enhanced signal is as follows: After performing frame processing on the energy of the input audio data, calculate it through the short-time Fourier transform to obtain the energy of the k-th frame respectively and the estimated value of the noise energy of the k-th frame , adopt the minimum statistical noise tracking algorithm for the energy of the k-th frame and the estimated value of the noise energy of the k-th frame to perform dynamic update and obtain the time-domain quantization index , which means: ; In the formula, represents the total number of framed audio data energy; represents the mean value of the energy of the k-th frame; represents the standard deviation of the energy of the k-th frame; represents the hyperbolic tangent function; represents the adaptive adjustment factor; ; SNR represents the signal-to-noise ratio; log represents the logarithmic operation; Extract the energy of audio data through the Mel filter bank to obtain the signal spectrum components of the v-th frame respectively and the noise spectrum components of the v-th frame , and use the noise floor estimation method to calculate the signal spectrum components of the extracted v-th frame and the noise spectrum components of the v-th frame , and obtain the frequency-domain quantization index , which means: ; In the formula, represents the total number of channels of the filter bank; represents the slope parameter of the Sigmoid function; represents the energy threshold; e represents the natural constant; Normalize the signal-to-noise ratio value range to [0, 1] to obtain the normalized signal-to-noise ratio ; ; represents the maximum value of the signal-to-noise ratio; represents the minimum value of the signal-to-noise ratio; By normalizing the signal-to-noise ratio calculate the parameters p and q, ; ; p represents the signal amplitude spectrum power adjustment parameter; q represents the dynamic noise suppression coefficient; Calculate the spectral attention mask value at the frequency point based on p and q, which means: ; In the formula, represents the spectral attention mask value at the frequency point ; represents the amplitude spectrum of the signal at the frequency point ; represents the average amplitude spectrum of the dynamic noise in the frequency band ; represents the adaptive scaling exponent adjusted by the frequency-domain quantization index ; represents the balance factor; where = ; Perform noise reduction calculation on the spectral attention mask value at the frequency point to obtain the time-domain signal after noise reduction and the frequency-domain signal after noise reduction respectively; denotes the inverse Fourier transform; The time-domain signal after noise reduction through dynamic weight fusion and the frequency-domain signal after noise reduction along with the time-domain quantization index and the frequency-domain quantization index , to obtain the enhanced signal , which means: 。 3. The drone classification method based on a noise perception model according to claim 2, wherein: The specific process of generating the Mel spectrogram is as follows: Based on the noise environment, calculate the resolution adjustment coefficient based on the time-domain quantization index and the noise spectrum entropy . Based on the resolution adjustment coefficient , use the time-frequency resolution of the Mel spectrum transform to dynamically adjust and optimize the enhanced signal to obtain the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal ; Resolution adjustment coefficient , which means: ; Wherein, sigmoid represents an activation function; Hnoise represents the noise spectrum entropy; w1 represents the time-domain quantization index weight; w2 represents the noise spectrum entropy weight; For the enhanced signal in the low-frequency band and the enhanced signal in the high-frequency band are analyzed to generate a mixed-resolution spectrum, indicating: ; In the formula, represents the mixed-resolution spectrum; represents the low-frequency band spectrum of the enhanced signal ; represents the high-frequency band spectrum of the enhanced signal ; Differentiable Mel filter bank based on resolution adjustment coefficient , dynamically fuse and enhance the signal in the low-frequency band and the enhanced signal in the high-frequency band, expressed as: ; In the formula, represents the frequency band boundary of the th differentiable Mel filter bank under the control of the resolution adjustment coefficient ; represents the low-frequency band filter boundary of the th enhanced signal ; represents the high-frequency band filter boundary of the th enhanced signal ; Based on the frequency band boundaries of the first differentiable Mel filter bank under the control of the resolution adjustment coefficient, define the weight coefficient of the ith Mel filter at the frequency point; Based on the weights of the hybrid-resolution spectrum and the th Mel filter at the frequency point to generate a Mel spectrogram, which represents: ; In the formula, represents the Mel spectrogram; represents the number of points of the fast Fourier transform FFT; represents the th Mel filter at the frequency point weight coefficient at; represents the minimum value.
4. A method for classifying drones based on a noise perception model according to claim 1, characterized in that: The convolutional neural network consists of four convolutional blocks connected in sequence, and each convolutional block consists of a convolutional layer, a ReLU activation layer, and a pooling layer connected in sequence; input the Mel spectrogram into the convolutional neural network and pass through the four convolutional blocks in sequence to output the category probability of the UAV.
5. The drone classification method based on a noise perception model according to claim 3, characterized in that: The specific process of obtaining the audio data energy is as follows: Convert the obtained UAV audio data into a unified format to obtain the multi-channel UAV audio data in the unified format; Convert the multi-channel UAV audio data into mono UAV audio data, denoted as: ; In the formula, represents the monaural drone audio data; represents the total number of channels of the multi-channel drone audio data; represents the drone audio data of the c-th channel; Resample the mono UAV audio data to obtain the resampled audio data, denoted as: ; In the formula, represents the resampled audio data; represents the resampling operation; represents the original sampling rate of the mono UAV audio data; represents the target sampling rate; Calculate the resampled audio data to obtain the audio data energy, denoted as: ; In the formula, represents the energy of the audio data; represents the total number of resampling points; represents the th sampling point of the resampled audio data.
6. A UAV classification system based on a noise perception model, which is applied to the UAV classification method based on a noise perception model according to any one of claims 1-5, characterized in that, It includes: An acquisition module, which is used to obtain the UAV audio data, process the UAV audio data to obtain the mono UAV audio data, resample the mono UAV audio data to obtain the resampled audio data, calculate the resampled audio data to obtain the audio data energy; A processing module, which is used to process the audio data energy to obtain the time-domain quantization index and the frequency-domain quantization index, and process the time-domain quantization index and the frequency-domain quantization index to obtain the denoised time-domain signal and the denoised frequency-domain signal, and fuse the time-domain quantization index and the frequency-domain quantization index with the denoised time-domain signal and the denoised frequency-domain signal to obtain the enhanced signal; An optimization module, which is used to optimize the enhanced signal by using Mel spectrum transformation to obtain the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal, analyze the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal to generate a mixed-resolution spectrum; optimize the low-frequency band of the enhanced signal and the high-frequency band of the enhanced signal through a differential Mel filter bank to obtain the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient, and process the mixed-resolution spectrum and the frequency band boundaries of the differential Mel filter bank under the control of the resolution adjustment coefficient to generate a Mel spectrogram; An output module, which is used to input the Mel spectrogram into a convolutional neural network for processing and output the probability of the drone category.
7. An electronic device, characterized in that, It includes a processor, a memory and a bus. The processor and the memory are connected through the bus. Among them, the memory is used to store a set of program codes, and the processor is used to call the program codes stored in the memory to execute a drone classification method according to any one of claims 1-5.
8. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions execute a drone classification method according to any one of claims 1-5.
Citation Information
Patent Citations
Audio signal encoding method, decoding method, encoding equipment and decoding equipment
CN113593586A
Speech synthesis method and system based on bone conduction signal and lip image fusion
CN116343793A
Environmental sound classification method and system based on metric learning
CN119296550A
Unmanned aerial vehicle identification method based on ResNet deep learning network
CN119513717A
Audio processing method, electronic equipment and storage medium
CN119946506A
Cited By
Audio processing system of photographic camera
CN120600041A