Bird sound noise reduction method and system for adaptive sub-band and acoustic index consistency loss facing bird sound production

By employing an adaptive frequency band division and acoustic index consistency loss method for bird sound noise reduction, the problems of fixed frequency band division and insufficient noise robustness in field ecological acoustic monitoring are solved, achieving high signal-to-noise ratio bird sound signal reconstruction and improving the accuracy of ecological monitoring.

CN121725804APending Publication Date: 2026-03-24GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies are ill-suited to the diverse spectral structures of bird calls in field ecological acoustic monitoring. Fixed frequency band divisions lead to uneven information aggregation, insufficient decoupling between frequency band modeling and temporal modeling, inadequate noise robustness, and insufficient spectral enhancement resolution, all of which affect the accuracy of bird call detection and acoustic index calculation.

Method used

A bird sound noise reduction method with adaptive frequency band division and acoustic index consistency loss is adopted. Through a learnable frequency band prototype dynamic soft band division mechanism, combined with time Transformer and frequency Transformer blocks, time-frequency feature fusion is performed, and acoustic index consistency loss is introduced to achieve adaptive optimization of spectrum division and noise suppression.

Benefits of technology

It improves the signal-to-noise ratio of outdoor soundscapes, preserves bird call details, enhances the accuracy and interpretability of ecological monitoring, adapts to complex noise environments, and ensures the consistency of acoustic indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725804A_ABST
    Figure CN121725804A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of bird sound noise reduction, and discloses a bird sound noise reduction method and system for adaptive sub-band and acoustic index consistency loss of bird sound production, and the method comprises the steps: S1, carrying out the superposition synthesis of a clean bird sound sample and a noise sample, and generating a training sample; s2, processing the training sample to obtain a frequency band feature tensor; s3, performing time dimension and frequency dimension feature fusion processing on the frequency band feature tensor to obtain a fusion feature; and S4, performing calculation based on the fusion features, and reconstructing a de-noised bird sound audio signal. The high signal-to-noise ratio improvement, the distortion rate reduction and the target bird chirp audibility improvement are realized in the complex field sound scene, the accuracy of subsequent automatic identification and acoustic index calculation can be remarkably improved, and the method has important application value for ecological monitoring, bird protection and habitat evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bird noise reduction, and more particularly to a bird noise reduction method and system that adapts to frequency division and acoustic index consistency loss for bird vocalization. Background Technology

[0002] With the rapid development of ecological acoustic monitoring technology, bird calls, as an important acoustic indicator of ecosystem status, are widely used in research fields such as biodiversity assessment and habitat quality monitoring. Birds, as a sensitive group in ecosystems, not only reflect population dynamics and behavioral rhythms through their vocalizations, but also characterize the health of ecosystems and environmental change trends. Long-term automated acoustic monitoring can construct time-series features of ecological soundscapes, providing quantitative evidence for ecological protection and environmental management. However, in real-world environments, soundscape recordings are often interfered with by various noise sources, including wind, rain, traffic, human voices, construction machinery, and other animal vocalizations. These non-target sound sources exhibit significant non-stationarity and spectral complexity, often overlapping with bird calls in the time and frequency domains, leading to a significant decrease in the signal-to-noise ratio and severely affecting the detection, identification, and calculation accuracy of acoustic indicators for bird calls. Especially in forest, wetland, and coastal ecosystems, due to drastic changes in background noise, dense sound sources, and complex reflections, traditional signal enhancement algorithms struggle to effectively distinguish target calls from background interference. Furthermore, different bird species exhibit significant differences in the frequency distribution, duration, modulation patterns, and energy structure of their calls. Some species with high-frequency soft or short calls are easily masked in noisy environments, while species with low-frequency calls are easily disturbed by low-frequency environmental noise. This cross-species frequency mobility and spectral diversity pose significant challenges to existing fixed-band or full-domain noise reduction algorithms. Furthermore, speech enhancement models focus only on subjective or objective audio quality indicators, neglecting the ability to restore the ecological soundscape. In field ecological monitoring, acoustic indices (such as ACI, ADI, and BI) are key indicators for assessing bird activity and biodiversity; however, these indicators are easily affected by background noise and inconsistencies in audio noise reduction processing, leading to biased or even erroneous monitoring results. Therefore, while existing methods improve audio audibility or objective indicators, they often fail to reliably preserve the true characteristics of the ecological soundscape. Moreover, the massive scale of field acoustic data and the significant differences in recording quality place higher demands on the adaptability, generalization, and real-time performance of noise reduction algorithms.

[0003] In the fields of bird sound noise reduction and ecological acoustic monitoring, deep learning models (such as CNN, Conformer, BSRNN, etc.) have been widely used for bird sound noise reduction tasks. However, when faced with complex noise backgrounds in the wild, spectral differences among bird species, and the high dynamism of ecological data, existing technologies still have the following core problems: (1) The frequency band division is fixed, making it difficult to adapt to the diverse spectral structure of bird sounds.

[0004] Traditional frequency domain noise reduction models typically divide the input signal into fixed frequency bands (such as 32 or 64 bands) according to a linear or Mel scale, and then model each band independently. However, the frequency distribution of bird calls varies significantly among different bird species, with some high-frequency calls concentrated in the 8–12 kHz range, while low-frequency bird calls are mainly distributed in the 1–3 kHz range. Fixed band division leads to uneven information aggregation: for narrowband calls, important features are dispersed and diluted; for broadband calls, intra-band details are severely lost, resulting in masking residues and frequency artifacts, which limit the model's generalization and adaptability.

[0005] (2) Frequency band modeling and time modeling are not fully decoupled, and RNN and other structures do not make sufficient use of context information.

[0006] Most existing noise reduction methods only perform unidirectional modeling in the time or frequency dimensions, making it difficult to capture the dependencies between the time and frequency domains simultaneously. For example, while RNN structures can transmit information in the time dimension, their memory length is limited, gradients are prone to decay, and they are difficult to model dynamic changes in long-term contexts; at the same time, they lack the ability to perceive energy transfer across different frequency bands. In bird calls, rhythmic variations are often accompanied by cross-band resonance and energy transfer, and simple temporal modeling is insufficient to restore the consistency of the acoustic spectrum structure. Because the frequency and time domain features are not effectively decoupled and fused, existing systems often suffer from problems such as "temporal domain smoothing transition" or "frequency domain pseudo-enhancement," resulting in the loss or distortion of output signal details.

[0007] (3) Insufficient noise robustness, unstable performance under dynamic conditions, and difficulty in reflecting real sound scene conditions.

[0008] Background noise in field ecological recordings varies dramatically with time, climate, and topography, including non-stationary components such as wind, rain, insect sounds, and mechanical noise. Existing deep noise reduction models are mostly trained based on fixed noise statistical characteristics, resulting in poor adaptability to changes in noise type and intensity. When encountering unfamiliar noise conditions, these models often exhibit oversuppression (falsely eliminating the target signal) or noise leakage (enhancement failure), thus reducing the sensitivity of bird call detection in ecological monitoring. Furthermore, noise reduction models do not consider the influence of acoustic indices and other ecological monitoring methods on the noise reduction results. Acoustic indices are important indicators characterizing changes in species diversity and the ecological state of the soundscape; their calculation is highly susceptible to interference from noise reduction errors, further reducing the explanatory power of the ecological soundscape structure, increasing the bias in biodiversity estimation, and distorting species activity rhythms and soundscape dynamics. Therefore, current technologies cannot guarantee the authenticity and scientific monitoring value of the noise-reduced ecological soundscape.

[0009] (4) The spectrum enhancement resolution is insufficient, resulting in low sound detail reproduction.

[0010] Existing methods typically only perform coarse-grained amplitude adjustments to the overall spectrum during the enhancement process, ignoring subtle differences between adjacent frequency points. This results in insufficient restoration of details in the high-frequency and transient parts of the enhanced bird sounds, leading to timbre distortion or a sense of blurriness and poor time-frequency structure integrity. Summary of the Invention

[0011] The purpose of this invention is to disclose a bird sound noise reduction method and system based on adaptive frequency division and acoustic index consistency loss for bird sounds, thereby solving the technical problems mentioned in the background art.

[0012] To achieve the above objectives, the present invention provides the following technical solution: On one hand, this invention provides a bird sound noise reduction method based on adaptive frequency band division and acoustic index consistency loss for bird vocalizations, comprising: S1, combine clean bird sound samples and noise samples to generate training samples; S2, process the training samples to obtain the frequency band feature tensor; S3, perform feature fusion processing on the frequency band feature tensor in the time dimension and frequency dimension to obtain the fused features; S4, based on the fusion features, performs calculations to reconstruct the denoised bird sound audio signal.

[0013] Preferably, S1 includes: S10: Download bird recording samples related to the target area from a public bird sound database, process the bird recording samples, and obtain clean bird sound sample segments. S11, Obtain field noise samples, slice the field noise samples to obtain field noise sample fragments; S12, superimpose clean bird sound sample segments and field noise sample segments to obtain training samples.

[0014] Preferably, the bird recording samples are processed to obtain clean bird sound sample segments, including: Bird recording samples were converted to WAV format and resampled to 32,000 Hz to obtain intermediate samples; The intermediate samples are sliced ​​into segments of 5 seconds each to obtain multiple intermediate sample segments. Intermediate sample segments with a duration of less than 2 seconds are deleted. The intermediate sample segments are filtered out, and intermediate sample segments with noise ratios greater than the set threshold or with signal distortion are deleted, thereby obtaining clean bird sound sample segments.

[0015] Preferably, the types of outdoor noise samples include wind noise, rain noise, vehicle noise, construction noise, cicada chirping, and cricket chirping.

[0016] Preferably, the field noise sample is sliced, which includes: slicing the field noise sample, dividing the field noise sample into segments of 5 seconds in length to obtain multiple field noise sample segments, and deleting field sample segments with a duration of less than 2 seconds.

[0017] Preferably, S2 includes: S20, perform STFT transformation on the training samples to obtain the complex spectrum X; S21 maps the complex spectrum X to high-dimensional features. :

[0018] Where b is the batch size, f is the frequency index, and t is the time frame index; This represents mapping the concatenated vector of real and imaginary parts onto the embedding space inside the model; and Let X be the real part and the imaginary part, respectively. S22, Define K learnable frequency band prototype vectors:

[0019] Represents the set of learnable frequency band prototype vectors, each It represents a typical feature of an ideal frequency band center, which is dynamically updated during training to adapt to the distribution of bird acoustic structures. Calculate the similarity score between the high-dimensional features corresponding to each frequency point and the typical features of the ideal frequency band center:

[0020] The similarity score is calculated using temperature-scaled Softmax to obtain the probability of each frequency point belonging to each frequency band:

[0021] It is a temperature coefficient; the smaller the value, the sharper the distribution, approaching a hard division; the larger the value, the smoother the distribution. It is the weight of continuous frequency bands; S23, based on Will Aggregate to K frequency bands to obtain the frequency band feature tensor :

[0022] Where F represents the number of frequency bins of the input STFT, T represents the number of time frames of the STFT, and N represents the subband feature dimension.

[0023] Preferably, S3 includes: S30, for Perform a dimensional transformation to obtain ; BK represents the aggregation of the two dimensions B and K; S31, will Input to the time Transformer block, get ,right Perform dimensional restoration to obtain ; S32, for Perform a dimensional transformation to obtain ,Will The input is fed into a frequency Transformer block, and its output is obtained by convolving the frequency band plot. After undergoing dimensionality transformation and temporal cross-attention of the input frequency band, the fused features are finally obtained. .

[0024] BT represents the aggregation of the two dimensions B and T.

[0025] Preferably, S4 includes: S40, Obtain the gain mask corresponding to the fused feature:

[0026] in It's the batch size. It refers to the number of frequency bands. It is a time frame. It is the number of channels in the complex spectrum. The number of frames; S41, based on Mapping G back to the original STFT frequency dimension:

[0027] in It is a frequency index, where i and j are the indices of the three frames and the real and virtual channels; The sub-band feature prediction gain mask is represented by M, and M represents the complex mask after re-projection. S42, based on The complex spectrum X is enhanced to obtain the enhanced complex spectrum S:

[0028]

[0029]

[0030]

[0031] in, This represents the enhanced complex STFT spectrum of the first frame at t=0; This represents the complex mask coefficients predicted by the model in frame 0 and used to process the current frame; This represents the complex STFT spectrum of the input audio at frame 0; This represents the complex mask coefficients predicted by the model in frame 0, which are used to process the next frame; This represents the complex STFT spectrum of the input audio in the first frame, serving as a future frame reference for the first frame enhancement, used to compensate for the lack of forward information in the first frame; This represents the complex STFT spectrum after the intermediate frame enhancement; This represents the complex mask coefficients predicted by the model in frame t, used to process the previous frame; This represents the complex STFT spectrum of the input audio in the previous frame, used as historical information for enhancement in the current frame; This represents the complex mask coefficients predicted by the model in frame t and used to process the current frame; This represents the complex STFT spectrum of the input audio at frame t, i.e., the time-frequency representation of the current frame; This represents the complex mask coefficients predicted by the model in frame t and used to process the next frame; This represents the complex STFT spectrum of the input audio in the next frame, used as future reference information for current frame enhancement; This represents the enhanced complex STFT spectrum of the last frame; This represents the complex mask coefficients predicted by the model in the last frame and used to process the previous frame; This represents the complex STFT spectrum of the input audio in the penultimate frame t=t−2, which serves as the forward information input for the last frame; This represents the complex mask coefficients predicted by the model in the last frame and used to process the current frame; This represents the complex STFT spectrum of the last frame t=t−1, i.e., the current spectrum of the last frame; Represents the complex STFT spectrum after three frames are fused and enhanced; S43 performs an inverse Fourier transform on S to reconstruct the denoised bird sound audio signal.

[0032] On the other hand, the present invention provides a bird sound noise reduction system with adaptive frequency division and acoustic index consistency loss for bird sounds, including a superposition module, a first processing module, a second processing module and a reconstruction module. The overlay module is used to overlay clean bird sound samples and noise samples to generate training samples; The first processing module is used to process the training samples and obtain the frequency band feature tensor; The second processing module is used to perform time-dimensional and frequency-dimensional feature fusion processing on the frequency band feature tensor to obtain fused features; The reconstruction module is used to perform calculations based on the fusion features to reconstruct the denoised bird sound audio signal.

[0033] Beneficial effects: This invention introduces a dynamic soft-banding mechanism with a learnable frequency band prototype, enabling adaptive optimization of spectrum allocation based on the time-frequency variation characteristics of bird calls. This fully extracts the rhythmic structure and frequency band coupling features of bird calls, enhancing the accuracy of acoustic signal reconstruction and maintaining the temporal continuity and harmonic structure integrity of the calls. Thus, while effectively suppressing non-stationary noises such as wind, rain, and insect sounds, it maximizes the preservation of bird vocal details, improving the naturalness and ecological interpretability of the noise-reduced audio. Based on these technological innovations, this invention achieves a significantly improved signal-to-noise ratio, reduced distortion rate, and improved audibility of target bird calls in complex outdoor soundscapes. It can significantly improve the accuracy of subsequent automatic identification and acoustic index calculation, and has important application value for ecological monitoring, bird conservation, and habitat assessment. Attached Figure Description

[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of the bird noise reduction method of the present invention.

[0036] Figure 2 This is a schematic diagram of the structure of the Transformer network of the present invention.

[0037] Figure 3 This is a schematic diagram of the time-frequency feature interaction fusion mechanism of the present invention.

[0038] Figure 4a This is a schematic diagram of the bird sound signal before noise reduction.

[0039] Figure 4b This is a schematic diagram of the bird sound signal after noise reduction. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] refer to Figure 1 This invention provides a bird sound noise reduction method based on adaptive frequency band division and acoustic index consistency loss for bird vocalizations, comprising: S1, combine clean bird sound samples and noise samples to generate training samples; S2, process the training samples to obtain the frequency band feature tensor; S3, perform feature fusion processing on the frequency band feature tensor in the time dimension and frequency dimension to obtain the fused features; S4, based on the fusion features, performs calculations to reconstruct the denoised bird sound audio signal.

[0042] Preferably, S1 includes: S10: Download bird recording samples related to the target area from a public bird sound database, process the bird recording samples to obtain clean bird sound sample segments.

[0043] Specifically, bird recording samples related to the target area were downloaded from the public bird sound database Xeno-Canto (https: / / xeno-canto.org / ), covering a variety of ecological types and representative species.

[0044] S11, Obtain field noise samples, slice the field noise samples to obtain field noise sample fragments; S12, superimpose clean bird sound sample segments and field noise sample segments to obtain training samples.

[0045] To simulate multi-signal-to-noise ratio (SNR) scenarios in real-world environments, this invention employs a dynamic mixing strategy to superimpose and synthesize clean bird calls and noise samples. Furthermore, data augmentation strategies are used to ensure the model's generalization ability, generating multi-signal-to-noise ratio (SNR) training samples. Specifically, the SNR is set to [-5, 0, 5, 10, 15, 20] dB, with distribution ratios of [20%, 20%, 20%, 15%, 15%, 10%], reflecting the high frequency of low SNR samples in natural environments. The generation of the mixed audio follows the principles of energy matching and temporal smoothing. Frame-by-frame windowing and overlapping addition methods are used to handle splicing boundaries, ensuring the continuity and naturalness of the generated audio.

[0046] Preferably, the bird recording samples are processed to obtain clean bird sound sample segments, including: Bird recording samples were converted to WAV format and resampled to 32,000 Hz to obtain intermediate samples.

[0047] Considering that the original recordings came from different devices and recording conditions, the data contained issues such as differences in sound quality, background noise interference, and labeling errors. Therefore, systematic cleaning and preprocessing were necessary. All recording files were first uniformly converted to WAV format and resampled to 32,000 Hz to ensure the consistency of the model input.

[0048] The intermediate samples are sliced ​​into segments of 5 seconds each, resulting in multiple intermediate sample fragments. Intermediate sample fragments with a duration of less than 2 seconds are deleted.

[0049] Noise detection and time slicing were performed on the recordings, dividing the audio into 5-second segments. Short segments of less than 2 seconds and samples without valid bird calls were removed. For audio segments with significant background interference, waveform and spectrograms were manually compared and verified, and samples with excessive noise or signal distortion were deleted to ensure the purity and representativeness of the training samples.

[0050] The intermediate sample segments are filtered out, and intermediate sample segments with noise ratios greater than the set threshold or with signal distortion are deleted, thereby obtaining clean bird sound sample segments.

[0051] Preferably, the types of outdoor noise samples include wind noise, rain noise, vehicle noise, construction noise, cicada chirping, and cricket chirping.

[0052] Some noise samples were taken from real-world scenarios, while others came from the open-source platform Freesound (https: / / freesound.org / ) to ensure sample diversity and realistic distribution characteristics. Each type of noise was segmented into 5-second segments to facilitate matching with bird sound samples.

[0053] Preferably, the field noise sample is sliced, which includes: slicing the field noise sample, dividing the field noise sample into segments of 5 seconds in length to obtain multiple field noise sample segments, and deleting field sample segments with a duration of less than 2 seconds.

[0054] After preprocessing the audio signal, this invention employs an adaptive dynamic frequency banding strategy. The core idea of ​​this module is to use learnable band prototypes to allow the model to automatically learn the feature distribution of different frequency regions during training, and dynamically determine the frequency band division method based on the time-frequency structure of the input signal, thereby replacing the traditional fixed-boundary frequency band division scheme.

[0055] Preferably, S2 includes: S20, perform STFT transformation on the training samples to obtain the complex spectrum X; S21 maps the complex spectrum X to high-dimensional features. :

[0056] Where b is the batch size, f is the frequency index, and t is the time frame index; This represents mapping the concatenated vector of real and imaginary parts onto the embedding space inside the model; and Let X be the real part and the imaginary part, respectively. Bird sounds are typically concentrated in the 2–8 kHz range; low and ultra-high frequencies are often noise (wind, rain, motor noise). Calculate the actual frequency value corresponding to each frequency point F, assigning higher weights to the target range (e.g., 2–8 kHz) and lower weights to non-target ranges, generating a weight vector:

[0057] Weights are injected into the similarity score through logarithmic transformation. This forces the model to focus on key frequencies.

[0058] Bird sounds typically exhibit localized energy peaks in the frequency spectrum; the amplitude is calculated as follows:

[0059] Local average: Kurtosis index:

[0060] The scores are also injected in logarithmic form. This enables dynamic resolution enhancement of the peak region of the spectrum, preventing key peak information from being masked by redundant frequencies.

[0061] S22, Define K learnable frequency band prototype vectors:

[0062] Represents the set of learnable frequency band prototype vectors, each It represents a typical feature of an ideal frequency band center, which is dynamically updated during training to adapt to the distribution of bird acoustic structures. Calculate the similarity score between the high-dimensional features corresponding to each frequency point and the typical features of the ideal frequency band center:

[0063] The similarity score is calculated using temperature-scaled Softmax to obtain the probability of each frequency point belonging to each frequency band:

[0064] It is a temperature coefficient; the smaller the value, the sharper the distribution, approaching a hard division; the larger the value, the smoother the distribution. It is the weight of continuous frequency bands; S23, based on Will Aggregate to K frequency bands to obtain the frequency band feature tensor :

[0065] Where F represents the number of frequency bins of the input STFT, T represents the number of time frames of the STFT, and N represents the subband feature dimension.

[0066] The cross-dimensional Transformer feature interaction fusion module consists of four parts: temporal Transformer block, frequency Transformer block, graph structure enhancement, and cross-axis attention fusion.

[0067] The bird sound noise reduction model's sequence modeling module includes Transformer modules for modeling temporal and frequency-dimensional correlations, namely the temporal Transformer block and the frequency Transformer block. These, along with cross-dimensional interaction enhancement, break down dimensional isolation and achieve information complementarity. The Transformer leverages multi-head attention mechanisms to capture long-distance dependencies, demonstrating powerful sequence modeling capabilities.

[0068] The sequence modeling module is divided into a temporal Transformer block and a frequency Transformer block, which can sequentially capture the correlation of global and local information in the time and frequency dimensions. Graph structure enhancement dynamically constructs an inter-band similarity map based on the average context information of Z after the frequency band axis transformation, and performs frequency band graph convolution (GCN) to inject frequency band topological priors. Cross-axis attention fusion fuses information between the time axis and the frequency band axis. The time-frequency feature interaction fusion mechanism is as follows... Figure 3 As shown.

[0069] Preferably, S3 includes: S30, for Perform a dimensional transformation to obtain ; BK represents the aggregation of the two dimensions B and K; S31, will Input to the time Transformer block, get ,right Perform dimensional restoration to obtain ; Dimensional transformation and dimensional restoration are existing techniques, such as reshape and transpose, which are commonly used dimensional transformation methods in deep learning.

[0070] S32, for Perform a dimensional transformation to obtain ,Will The input is fed into a frequency Transformer block, and its output is obtained by convolving the frequency band plot. After undergoing dimensionality transformation and temporal cross-attention of the input frequency band, the fused features are finally obtained. .

[0071] BT represents the aggregation of the two dimensions B and T.

[0072] The structure of a Transformer network is as follows: Figure 2 As shown, the multi-head self-attention mechanism is defined as follows:

[0073] Where Q, K, and V are the values ​​of the input after linear mapping, T represents the transpose, and d represents the dimension of K.

[0074] Preferably, S4 includes: S40, Obtain the gain mask corresponding to the fused feature:

[0075] in It's the batch size. It refers to the number of frequency bands. It is a time frame. It is the number of channels in the complex spectrum. The number of frames; The input to S40 comes from the sub-band features of the dual Transformer output. ; S41, based on Mapping G back to the original STFT frequency dimension:

[0076] in It is the frequency index, i,j are the indices of the three frames and the real and virtual channels; here, soft weights are used to ensure mask smoothness and preservation of frequency band overlap information, avoiding structural breaks caused by traditional hard banding.

[0077] The sub-band feature prediction gain mask is represented by M, and M represents the complex mask after re-projection. S41 mainly utilizes the soft allocation weights generated by dynamic frequency division to map the sub-band mask back to the original STFT frequency dimension.

[0078] S42, based on The complex spectrum X is enhanced to obtain the enhanced complex spectrum S:

[0079]

[0080]

[0081]

[0082] in, This represents the enhanced complex STFT spectrum at t=0 for the first frame. This represents the complex mask coefficients predicted by the model in frame 0 and used to process the current frame; This represents the complex STFT spectrum of the input audio at frame 0; This represents the complex mask coefficients predicted by the model in frame 0, which are used to process the next frame; This represents the complex STFT spectrum of the input audio in the first frame, serving as a future frame reference for the first frame enhancement, used to compensate for the lack of forward information in the first frame; This represents the complex STFT spectrum after the intermediate frame enhancement; This represents the complex mask coefficients predicted by the model in frame t, used to process the previous frame; This represents the complex STFT spectrum of the input audio in the previous frame, used as historical information for enhancement in the current frame; This represents the complex mask coefficients predicted by the model in frame t and used to process the current frame; This represents the complex STFT spectrum of the input audio at frame t, i.e., the time-frequency representation of the current frame; This represents the complex mask coefficients predicted by the model in frame t, which are used to process the next frame; This represents the complex STFT spectrum of the input audio in the next frame, used as future reference information for current frame enhancement; This represents the enhanced complex STFT spectrum of the last frame; This represents the complex mask coefficients predicted by the model in the last frame and used to process the previous frame; This represents the complex STFT spectrum of the input audio in the penultimate frame t=t−2, which serves as the forward information input for the last frame; This represents the complex mask coefficients predicted by the model in the last frame and used to process the current frame; This represents the complex STFT spectrum of the last frame t=t−1, i.e., the current spectrum of the last frame; Represents the complex STFT spectrum after three frames are fused and enhanced; S43 performs an inverse Fourier transform on S to reconstruct the denoised bird sound audio signal.

[0083] The denoised bird sound audio signal was recovered by ISTFT (Inverse Short Time Fourier Transform). The bird sound signals before and after denoising are as follows: Figure 4a and Figure 4b As shown.

[0084] Figure 4a This is a mixed signal of vehicle noise and bird calls, with overlapping bird noise components. Figure 4b This is the bird sound signal after noise reduction.

[0085] To enhance the practicality of noise-reduced audio in acoustic ecological monitoring, this invention designs an Index Consistency Loss, which ensures that the ecological characteristics of bird calls in the noise-reduced audio remain consistent with the reference signal in terms of acoustic indices. This loss function integrates ecological information from three dimensions: temporal dynamics, biological activity levels, and frequency distribution diversity of bird calls. The form of the loss function is as follows:

[0086] in , These are weighting coefficients used to balance the various losses.

[0087] The ACI acoustic complexity consistency loss is defined as follows:

[0088]

[0089] The BI Bioacoustic Index Consistency Loss is defined as follows: Background noise is calculated using soft minimum:

[0090] ReLU smooth substitution:

[0091]

[0092] ADI defines acoustic diversity consistency loss as follows:

[0093]

[0094]

[0095] To avoid numerical instability or gradient explosion caused by excessively small denominator values, a small positive number is introduced during the calculation process. Perform smoothing processing. To compare and evaluate the performance of the various models in this study, the following evaluation metrics were used: .

[0096] On the other hand, the present invention provides a bird sound noise reduction system with adaptive frequency division and acoustic index consistency loss for bird sounds, including a superposition module, a first processing module, a second processing module and a reconstruction module. The overlay module is used to overlay clean bird sound samples and noise samples to generate training samples; The first processing module is used to process the training samples and obtain the frequency band feature tensor; The second processing module is used to perform time-dimensional and frequency-dimensional feature fusion processing on the frequency band feature tensor to obtain fused features; The reconstruction module is used to perform calculations based on the fusion features to reconstruct the denoised bird sound audio signal.

[0097] The innovation of this invention lies in: 1. An adaptive frequency band division method based on a learnable frequency band prototype of bird call energy frequency distribution is proposed: In the spectral feature processing stage, prior weights for bird calls and dynamic peak perception enhancement are added. A learnable band prototype parameter matrix is ​​introduced. By calculating the similarity between each frequency feature and the band prototype, and combining the temperature-controllable Sftmax function, the assignment weights of each frequency point to each band are generated, thus achieving "soft banding". The assignment weights are dynamically adjusted according to the time-frequency characteristics of the input bird call audio, so that the band division has the adaptive capability of the time dimension.

[0098] It overcomes the problems of broken harmonic structure and fragmented energy distribution of bird calls caused by traditional fixed frequency band division (such as Mel filtering), and adapts to the complex and time-varying time-frequency characteristics of bird calls (2-8kHz core frequency band).

[0099] The original spectral features are mapped to a fixed number of sub-band features, which compresses the feature dimension, improves feature compactness, and enhances robustness to noise. This provides structurally stable and information-complete input features for subsequent modeling, ensuring the effective preservation of key acoustic features of bird calls (such as call rhythm and harmonic intervals).

[0100] 2. A cross-dimensional Transformer feature interaction fusion structure is proposed: The dual-axis alternating modeling structure achieves comprehensive capture of the joint time-frequency context, solving the problem that single-dimensional modeling (such as traditional Transformer) cannot effectively balance local (frequency band) and global (time) dependencies when processing high-dimensional time-frequency data, and solving the problem that the features of weak bird sounds are submerged in the background of strong noise (such as wind, rain, and insect sounds).

[0101] By injecting structural priors through GCN and achieving bidirectional flow of information across two axes through cross-attention, the distinguishability of weak target acoustic signals in complex background noise is further improved.

[0102] 3. A loss constraint method based on acoustic exponential consistency is proposed: During the model training phase, acoustic index consistency constraints are introduced. By calculating key ecological acoustic indices such as ACI, BI, and ADI for bird sounds before and after noise reduction, and penalizing the differences in these indices in the loss function, the model can maintain the ecological characteristics and rhythmic structure of bird sound signals during the noise reduction process.

[0103] To avoid harmonic distortion and loss of time-frequency structure caused by excessive suppression, ensure the reliability and interpretability of the noise-reduced signal in subsequent biodiversity assessment, bird identification, and soundscape analysis.

[0104] 4. A time-frequency reconstruction method based on sub-band mask prediction and three-frame fusion is proposed: Gain prediction is performed in the sub-bands, and the signal is then dynamically weighted back to the original frequency to achieve joint enhancement of the complex spectrum amplitude and phase. This ensures the continuity of the harmonic structure and the naturalness of the audio, effectively suppressing non-stationary noise such as wind, rain, and insect sounds.

[0105] By fusing information from adjacent frames during complex spectrum enhancement, inter-frame structural smoothing and signal detail recovery are achieved. This significantly reduces noise reduction artifacts and structural jumps, improving listening quality and ecological interpretability.

[0106] 5. An end-to-end bird noise reduction system architecture for complex ecological scenarios is proposed: The system integrates the "bird adaptive frequency band division module", the "cross-dimensional Transformer feature interaction fusion module", and the "sub-band mask prediction and three-frame fusion reconstruction module" to form an end-to-end bird sound noise reduction system. The system has built-in bird sound frequency band prior (2-8kHz core frequency band weight enhancement) and peak perception gating (dynamically adjusting the band resolution based on the spectral peak) to adapt to complex ecological scenarios in the wild (such as bird sound characteristics at different altitudes and in different climates).

[0107] This system overcomes the limitations of traditional noise reduction systems (which require multi-module, step-by-step processing and rely on manual parameter tuning) for long-term field monitoring tasks, addressing the need for automation and high precision in bird call noise reduction in complex ecological scenarios. It supports the automatic deployment and long-term operation of field monitoring equipment, and the noise-reduced bird call signals can be directly used for bird call identification and acoustic index inversion (such as species richness calculation); improving the reliability of ecological assessment and species conservation data while reducing manual data processing costs.

[0108] In summary, compared with existing bird sound noise reduction methods, this invention has significant advantages in modeling mechanism and engineering performance. Firstly, at the frequency band processing level, existing technologies generally use fixed frequency band division, which easily leads to the breakage of key harmonic structures in bird sounds (especially the 2-8kHz core frequency band) and the fragmentation of weak signal energy. This invention, however, by using the frequency band distribution characteristics of bird sound energy, a learnable frequency band prototype, and a dynamic soft-banding mechanism, can adjust the frequency point assignment weights in real time according to the time-frequency characteristics of bird sounds. This preserves the complete acoustic signature structure and improves the ability to capture weak bird sound signals, solving the problem of poor adaptability of fixed band division. Secondly, this invention uses dual-dimensional Transformer joint modeling, combining cross-dimensional cross-attention and frequency band graph convolution to achieve synergistic enhancement of time-frequency information. It can simultaneously capture bird sound rhythm information and cross-frequency band coupling structures, achieving more accurate acoustic signature representation and noise suppression effects, significantly superior to existing single-dimensional feature modeling methods. Furthermore, this invention employs a sub-band masking and three-frame fusion complex spectrum reconstruction mechanism, which maintains temporal continuity and harmonic structure integrity while achieving strong noise suppression, resulting in higher sound quality and greater reliability in ecological analysis of the output audio. This invention also incorporates an Acoustic Indices Consistency Loss, ensuring that the bird call ecological characteristics of the noise-reduced frequency are consistent with the reference signal in terms of acoustic indices. This method possesses end-to-end trainability and high computational efficiency, allowing for seamless integration with ecological monitoring equipment and adaptability to various complex soundscape environments. In summary, this invention significantly outperforms the best existing technologies in terms of adaptive modeling capabilities, signal reconstruction quality, noise robustness, and practical monitoring value, demonstrating outstanding innovation and practicality. Through dual innovation at both the generation and classification ends, this invention achieves highly robust and sensitive identification of rare birds in complex noise environments in the wild, exhibiting stronger stability and ecological practical value compared to existing methods, and can be widely applied to biodiversity monitoring and ecological protection assessment.

[0109] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A bird sound noise reduction method based on adaptive frequency band division and acoustic index consistency loss for bird sounds, characterized in that, include: S1, combine clean bird sound samples and noise samples to generate training samples; S2, process the training samples to obtain the frequency band feature tensor; S3, perform feature fusion processing on the frequency band feature tensor in the time dimension and frequency dimension to obtain the fused features; S4, based on the fusion features, performs calculations to reconstruct the denoised bird sound audio signal.

2. The method according to claim 1, characterized in that, S1 includes: S10: Download bird recording samples related to the target area from a public bird sound database, process the bird recording samples, and obtain clean bird sound sample segments. S11, Obtain field noise samples, slice the field noise samples to obtain field noise sample fragments; S12, superimpose clean bird sound sample segments and field noise sample segments to obtain training samples.

3. The method according to claim 2, characterized in that, The bird recording samples were processed to obtain clean bird sound sample segments, including: Bird recording samples were converted to WAV format and resampled to 32,000 Hz to obtain intermediate samples; The intermediate samples are sliced ​​into segments of 5 seconds each to obtain multiple intermediate sample segments. Intermediate sample segments with a duration of less than 2 seconds are deleted. The intermediate sample segments are filtered out, and intermediate sample segments with noise ratios greater than the set threshold or with signal distortion are deleted, thereby obtaining clean bird sound sample segments.

4. The method according to claim 2, characterized in that, The types of outdoor noise samples include wind noise, rain noise, vehicle noise, construction noise, cicada chirping, and cricket chirping.

5. The method according to claim 2, characterized in that, The field noise sample is sliced, including: the field noise sample is sliced ​​into segments of 5 seconds in length to obtain multiple field noise sample segments, and field sample segments with a duration of less than 2 seconds are deleted.

6. The method according to claim 1, characterized in that, S2 include: S20, perform STFT transformation on the training samples to obtain the complex spectrum X; S21 maps the complex spectrum X to high-dimensional features. : Where b is the batch size, f is the frequency index, and t is the time frame index; This represents mapping the concatenated vector of real and imaginary parts onto the embedding space inside the model; and Let X be the real part and the imaginary part, respectively. S22, Define K learnable frequency band prototype vectors: Represents the set of learnable frequency band prototype vectors, each It represents a typical feature of an ideal frequency band center, which is dynamically updated during training to adapt to the distribution of bird acoustic structures. Calculate the similarity score between the high-dimensional features corresponding to each frequency point and the typical features of the ideal frequency band center: The similarity score is calculated using temperature-scaled Softmax to obtain the probability of each frequency point belonging to each frequency band: It is a temperature coefficient; the smaller the value, the sharper the distribution, approaching a hard division; the larger the value, the smoother the distribution. It is the weight of continuous frequency bands; S23, based on Will Aggregate to K frequency bands to obtain the frequency band feature tensor : Where F represents the number of frequency bins of the input STFT, T represents the number of time frames of the STFT, and N represents the subband feature dimension.

7. The method according to claim 6, characterized in that, S3 include: S30, for Perform a dimensional transformation to obtain ; BK represents the aggregation of the two dimensions B and K; S31, will Input to the time Transformer block, get ,right Perform dimensional restoration to obtain ; S32, for Perform a dimensional transformation to obtain ,Will The input is fed into a frequency Transformer block, and its output is obtained by convolving the frequency band plot. After undergoing dimensionality transformation and temporal cross-attention of the input frequency band, the fused features are finally obtained. . BT represents the aggregation of the two dimensions B and T.

8. The method according to claim 7, characterized in that, S4 include: S40, Obtain the gain mask corresponding to the fused feature: in It's the batch size. It refers to the number of frequency bands. It is a time frame. It is the number of channels in the complex spectrum. The number of frames; S41, based on Mapping G back to the original STFT frequency dimension: in It is a frequency index, where i and j are the indices of the three frames and the real and virtual channels; The sub-band feature prediction gain mask is represented by M, and M represents the complex mask after re-projection. S42, based on The complex spectrum X is enhanced to obtain the enhanced complex spectrum S: in, This represents the enhanced complex STFT spectrum of the first frame at t=0; This represents the complex mask coefficients predicted by the model in frame 0 and used to process the current frame; This represents the complex STFT spectrum of the input audio at frame 0; This represents the complex mask coefficients predicted by the model in frame 0, which are used to process the next frame; This represents the complex STFT spectrum of the input audio in the first frame, serving as a future frame reference for the first frame enhancement, used to compensate for the lack of forward information in the first frame; This represents the complex STFT spectrum after the intermediate frame enhancement; This represents the complex mask coefficients predicted by the model in frame t, used to process the previous frame; This represents the complex STFT spectrum of the input audio in the previous frame, used as historical information for enhancement in the current frame; This represents the complex mask coefficients predicted by the model in frame t and used to process the current frame; This represents the complex STFT spectrum of the input audio at frame t, i.e., the time-frequency representation of the current frame; This represents the complex mask coefficients predicted by the model in frame t, which are used to process the next frame; This represents the complex STFT spectrum of the input audio in the next frame, used as future reference information for current frame enhancement; This represents the enhanced complex STFT spectrum of the last frame; This represents the complex mask coefficients predicted by the model in the last frame and used to process the previous frame; This represents the complex STFT spectrum of the input audio in the penultimate frame t=t−2, which serves as the forward information input for the last frame; This represents the complex mask coefficients predicted by the model in the last frame and used to process the current frame; This represents the complex STFT spectrum of the last frame t=t−1, i.e., the current spectrum of the last frame; Represents the complex STFT spectrum after three frames are fused and enhanced; S43 performs an inverse Fourier transform on S to reconstruct the denoised bird sound audio signal.

9. A bird sound noise reduction system with adaptive frequency band division and acoustic index consistency loss for bird sounds, characterized in that, It includes an overlay module, a first processing module, a second processing module, and a reconstruction module; The overlay module is used to overlay clean bird sound samples and noise samples to generate training samples; The first processing module is used to process the training samples and obtain the frequency band feature tensor; The second processing module is used to perform time-dimensional and frequency-dimensional feature fusion processing on the frequency band feature tensor to obtain fused features; The reconstruction module is used to perform calculations based on the fusion features to reconstruct the denoised bird sound audio signal.