A method for urban ecological noise elimination in bird sound recording and related equipment
Patent Information
- Application Number
- CN202610866549.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]综上,相关技术存在以下问题:1、基于信号处理的传统降噪方法依赖先验条件,无法对非平稳噪声进行建模,会产生音乐噪声
[0019]本申请实施例至少包括以下有益效果:本申请提供一种用于鸟声录音中的城市生态噪声消除方法及相关设备,该方案收集多类城市噪声数据,构建干净鸟声数据集合噪声数据集,进而划分得到用于模型训练的训练集、验证集和测试集;根据所述训练集、验证集和测试集,训练得到鸟声降噪网络;通过短时傅里叶方法对输入的现场录音信号进行信号提取,得到所述现场录音信号的实部和虚部;将所述实部和所述虚部输入所述鸟声降噪网络进行计算,得到复数理想比率掩码;将所述复数理想比率掩码作用于所述现场录音信号的实部和虚部,得到现场录音信号经过去噪后的傅里叶变换结果;将去噪后的傅里叶变换结果进行傅里叶逆变换得到现场录音信号的去噪结果。本申请能够提高模型的泛化性,能够促进编码器的信息流动,为解码器提供更丰富的信息,进而提升了噪声消除的效果。
Smart Images

Figure CN122821971A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a method and related equipment for eliminating urban ecological noise in bird sound recordings. Background Technology
[0002] With rapid urbanization, urban ecosystems face immense pressure. Birds are a crucial component of these ecosystems, and their presence helps assess ecosystem stability. In urban ecology, birds serve as indicator species for environmental change; their species, numbers, and behaviors effectively reflect the state of the urban ecosystem. In recent years, passive acoustic monitoring technology has been widely applied in urban bird monitoring and conservation. It overcomes the limitations of visual observation, continuously collecting environmental sound data over long periods and on large spatial scales. This method can reduce the damage to the ecological environment caused by human activities while lowering monitoring costs. Long-term, automated audio recordings are used to analyze and assess the species richness and spatiotemporal distribution of birds in urban green spaces, parks, and residential areas, thereby evaluating the ecological environment. However, most bird acoustic recordings are made in open outdoor environments. In actual passive acoustic recordings, target bird signals are often severely contaminated by various environmental noises. These noise sources include natural wind and rain, traffic or mechanical noise from human activities, and calls from other non-target organisms (such as high-intensity cicada choruses and katydids). These noises coexist with bird calls in the time domain and partially or completely overlap in the frequency domain. A low signal-to-noise ratio can affect the accuracy of automated tasks such as bird call detection, segmentation, recognition, and classification, thereby impacting the assessment of the ecological environment.
[0003] In bird call noise reduction, previous work typically utilized traditional signal processing methods. The most common methods are spectral subtraction and short-time spectral amplitude estimation based on the minimum mean square error criterion (MMSE-STSA). These methods are simple to implement and suitable for stationary noise scenarios, but struggle with non-stationary urban ecological noise. They are also prone to producing musical noise artifacts during processing, disrupting the fine structure of bird calls. In recent years, data-driven deep learning techniques have been introduced into the field of bird call noise reduction. Neural networks, trained on large-scale datasets, can model non-stationary noise, thereby achieving complex nonlinear noise reduction mappings.
[0004] Deep learning-based monophonic bird sound denoising techniques can currently be divided into two main types. The first type uses image segmentation to denoise the audio. Recorded audio is converted into a spectrogram through a short-time Fourier transform. The model learns the acoustic edges of the bird sounds and segments the image to separate the bird sounds from the noise, selectively preserving the bird sounds and discarding the noise, thus achieving denoising. The second type uses an encoder-modeling-decoder structure. The encoder converts the noisy bird sounds into high-dimensional features, the modeling module models the bird sounds, and the decoder decodes the modeled data to recover the original bird sounds. The model learns a mask that maps noisy bird sounds to clean bird sounds.
[0005] In current denoising methods, most networks use short-time Fourier transform to convert the noisy time-domain waveform into another feature representation as network input, such as amplitude spectrum or complex spectrum. The encoder and decoder typically consist of convolutional neural networks (deconvolutional neural networks), normalization layers, and activation functions. The bottleneck layer usually employs recurrent neural networks (RNNs) and their variants (Long Short-Term Memory networks (LSTM), gated recurrent units (GRUs)), temporal convolutional modules (TCMs), and Transformers and Mamba. The network output is then reconstructed using an inverse short-time Fourier transform to reconstruct the time-domain waveform.
[0006] In summary, the relevant technologies have the following problems: 1. Traditional noise reduction methods based on signal processing rely on prior conditions and cannot model non-stationary noise, resulting in musical noise. 2. Current deep learning methods only target single or a few types of noise, lacking research on noise in multi-category and complex scenarios, and have insufficient model generalization. Summary of the Invention
[0007] The main objective of this application is to propose a method and related equipment for eliminating urban ecological noise in bird sound recordings, which can improve the generalization of the model.
[0008] To achieve the above objectives, one aspect of this application proposes a method for eliminating urban ecological noise in bird call recordings, comprising: Collect various types of urban noise data, construct a clean bird sound dataset, and then divide it into training, validation, and test sets for model training. The bird sound noise reduction network is trained based on the training set, validation set, and test set. The input on-site recording signal is extracted using the short-time Fourier transform method to obtain the real and imaginary parts of the on-site recording signal; The real part and the imaginary part are input into the bird sound noise reduction network for calculation to obtain the complex ideal ratio mask; The complex ideal ratio mask is applied to the real and imaginary parts of the on-site recording signal to obtain the Fourier transform result of the on-site recording signal after denoising. The denoised Fourier transform result is then subjected to inverse Fourier transform to obtain the denoised result of the on-site recording signal.
[0009] In some embodiments, the collection of multiple types of urban noise data to construct a clean bird sound dataset includes: Record various types of noise data within the target geographic area; Collect various types of raw bird sound data within the target geographic area; The noise data and the original bird sound data are processed in a unified manner, and configured into a unified audio format, sampling frequency, quantization accuracy and number of audio channels; Birdsong data with low background noise are filtered from the original bird calls data and corresponding noise reduction processing is performed to obtain clean bird calls data. The noise data of multiple types is sliced, and the duration of a single noise sample is processed to 10 seconds so that each type of noise contains a preset number of training audios. The processed clean bird sound data and noise data are divided into training set, validation set and test set according to a preset ratio.
[0010] In some embodiments, the method further includes: during the training process, mixing clean bird sound data and various noise data according to different discrete signal-to-noise ratio segments to obtain a training set and a validation set, this step including: For each training sample, one clean bird sound data and one noise data are randomly selected from all the clean bird sound data and noise data in the training set. From the selected noise data, a noise segment with a duration of 4 seconds is randomly cut out; Random sampling is performed within a preset discrete signal-to-noise ratio range; Based on the discrete signal-to-noise ratio range, the selected clean bird sound data is mixed with the cropped noise segments to synthesize noisy bird sound data, thereby constructing the training set and validation set.
[0011] In some embodiments, the step of extracting the signal from the input on-site recording signal using the short-time Fourier method to obtain the real and imaginary parts of the on-site recording signal includes: Perform a short-time Fourier transform on the input on-site recording signal to obtain a complex spectrum; The complex spectrum is decomposed into a real matrix and an imaginary matrix using Euler's formula.
[0012] In some embodiments, the bird noise reduction network includes an encoder, a bottleneck layer, and a decoder; The encoder is used to extract multi-scale time-frequency features of bird sounds layer by layer and compress the time dimension; The bottleneck layer is used for global modeling of the long-range dependence of bird sound signals and the correlation between frequency subbands; The decoder is used to filter out irrelevant noise information while preserving and enhancing the amplitude consistency of bird calls.
[0013] In some embodiments, the encoder includes six cascaded complex coding units, each layer of complex coding units comprising complex convolution, complex batch normalization, and parameterized linear rectified unit activation functions; The bottleneck layer includes a complex feedforward neural network module, a complex multi-head attention mechanism module with relative position encoding, and a complex convolution enhancement module; The complex feedforward neural network module consists of two complex linear transformation layers sandwiching a complex Swish activation function. A Dropout mechanism is introduced after each linear transformation layer to prevent overfitting. The intermediate layer dimension is expanded to four times the input dimension to enhance nonlinear expressive power. The processing steps of the complex multi-head attention mechanism module include: generating corresponding rotation factors for the frequency index and time index in the complex domain; combining the two rotation factors through an outer product to obtain a joint positional encoding; injecting the joint positional encoding into the query and key, and performing linear projection to generate the corresponding query, key, and value; calculating scaling points and attention based on the query, key, and value; and constructing a complex convolution enhancement module and a complex feedforward network to complete the construction process of the complex multi-head attention mechanism module. The complex convolution enhancement module includes complex layer normalization, complex point convolution, complex gated linear units, complex depthwise convolution, complex batch normalization, complex Swish activation, and a second complex point convolution; the complex convolution enhancement module also includes cross-layer residual connections, used to add the input features to the convolutional features.
[0014] In some embodiments, the decoder construction process includes: A frequency-aware mechanism is applied to combine the skip-gating link between the decoder and the encoder; Based on the output features of the encoder, attention weights are generated in the frequency dimension; The generated attention weights are used to perform element-wise weighted processing on the encoder features; The weighted encoder features and decoder features are fused together to complete the construction of the decoder; The process of fusing the weighted encoder features and decoder features includes: concatenating the weighted encoder features and decoder features along the channel dimension and fusing them through a convolutional layer to generate a gated signal; and multiplying the weighted encoder features element-wise with the gated signal and then passing them through a convolutional layer to obtain the final output features.
[0015] Another aspect of this application provides an urban ecological noise cancellation device for bird call recording, comprising: The first module is used to collect various types of urban noise data, construct a clean bird sound data set (noise dataset), and then divide it into training set, validation set and test set for model training. The second module is used to train a bird sound noise reduction network based on the training set, validation set, and test set. The third module is used to extract the signal from the input on-site recording signal using the short-time Fourier method, and obtain the real part and imaginary part of the on-site recording signal. The fourth module is used to input the real part and the imaginary part into the bird sound noise reduction network for calculation to obtain a complex ideal ratio mask; The fifth module is used to apply the complex ideal ratio mask to the real and imaginary parts of the on-site recording signal to obtain the Fourier transform result of the on-site recording signal after denoising. The sixth module is used to perform an inverse Fourier transform on the denoised Fourier transform result to obtain the denoised result of the on-site recording signal.
[0016] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0017] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0018] This application also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.
[0019] The embodiments of this application include at least the following beneficial effects: This application provides a method and related equipment for urban ecological noise cancellation in bird sound recordings. This scheme collects multiple types of urban noise data, constructs a clean bird sound data set (noise dataset), and then divides it into a training set, a validation set, and a test set for model training. Based on the training set, validation set, and test set, a bird sound denoising network is trained. The input field recording signal is extracted using the short-time Fourier transform method to obtain the real and imaginary parts of the field recording signal. The real and imaginary parts are input into the bird sound denoising network for calculation to obtain a complex ideal ratio mask. The complex ideal ratio mask is applied to the real and imaginary parts of the field recording signal to obtain the Fourier transform result of the denoised field recording signal. The denoised Fourier transform result is then subjected to an inverse Fourier transform to obtain the denoised result of the field recording signal. This application can improve the generalization of the model, promote information flow in the encoder, provide richer information to the decoder, and thus improve the noise cancellation effect. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application; Figure 2 This is a flowchart of the overall steps provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the implementation process in a specific scenario provided in the embodiments of this application; Figure 4 This is an architecture diagram of the complex number encoding unit provided in the embodiments of this application; Figure 5 This is a structural diagram of the complex feedforward neural network module provided in the embodiments of this application; Figure 6 This is an architecture diagram of a complex Conformer module with relative position encoding provided in an embodiment of this application; Figure 7 This is an architecture diagram of the complex convolution module provided in an embodiment of this application; Figure 8 This is an architecture diagram of the decoder module provided in an embodiment of this application; Figure 9 This is a schematic diagram of a clean bird sound signal provided in an embodiment of this application; Figure 10 This is a schematic diagram of a noisy bird sound signal provided in an embodiment of this application; Figure 11 This is a schematic diagram of the noise-reduced bird sound signal provided in an embodiment of this application; Figure 12 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0022] It is understood that the terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0025] Before providing a detailed description of the embodiments of this application, some related technologies involved in the embodiments of this application will be described first, as follows: In bird call noise reduction techniques, previous work typically utilized traditional signal processing methods. The most common methods are spectral subtraction and short-time spectral amplitude estimation based on the minimum mean square error criterion (MMSE-STSA). These methods are simple to implement and suitable for stationary noise scenarios, but struggle with non-stationary urban ecological noise. They are also prone to generating musical noise artifacts during processing, disrupting the fine structure of bird calls. In recent years, data-driven deep learning techniques have been introduced into the field of bird call noise reduction. Neural networks, trained on large-scale datasets, can model non-stationary noise, thereby achieving complex nonlinear noise reduction mappings.
[0026] Deep learning-based monophonic bird sound denoising techniques can currently be divided into two main types. The first type uses image segmentation to denoise the audio. Recorded audio is converted into a spectrogram through a short-time Fourier transform. The model learns the acoustic edges of the bird sounds and segments the image to separate the bird sounds from the noise, selectively preserving the bird sounds and discarding the noise, thus achieving denoising. The second type uses an encoder-modeling-decoder structure. The encoder converts the noisy bird sounds into high-dimensional features, the modeling module models the bird sounds, and the decoder decodes the modeled data to recover the original bird sounds. The model learns a mask that maps noisy bird sounds to clean bird sounds.
[0027] In current denoising methods, most networks use short-time Fourier transform to convert the noisy time-domain waveform into another feature representation as network input, such as amplitude spectrum or complex spectrum. The encoder and decoder typically consist of convolutional neural networks (deconvolutional neural networks), normalization layers, and activation functions. The bottleneck layer usually employs recurrent neural networks (RNNs) and their variants (Long Short-Term Memory networks (LSTM), gated recurrent units (GRUs)), temporal convolutional modules (TCMs), and Transformers and Mamba. The network output is then reconstructed using an inverse short-time Fourier transform to reconstruct the time-domain waveform.
[0028] In view of this, this application provides a method and related equipment for eliminating urban ecological noise in bird sound recordings. The scheme proposes a bird sound noise reduction network for urban ecological scenarios. The network is trained using various types of noise and bird sounds, which improves the generalization of the model. A network based on a complex encoding and decoding structure is constructed, a complex Conformer is designed as the bottleneck layer for feature extraction, and a frequency domain gating attention mechanism is designed to promote the information flow of the encoder and provide richer information to the decoder.
[0029] The urban ecological noise cancellation method and related equipment for bird sound recording provided in this application relates to the field of deep learning technology. The urban ecological noise cancellation method for bird sound recording provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited thereto; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the urban ecological noise cancellation method for bird sound recording, but is not limited to the above forms.
[0030] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0031] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0032] like Figure 1The diagram shown is a schematic representation of an implementation environment provided in an embodiment of this application. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0033] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0034] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0035] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc. It can also be a vehicle-mounted terminal of the various device types described above, but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment does not impose any limitations.
[0036] Exemplary based on Figure 1 The implementation environment shown in this application embodiment provides a method for eliminating urban ecological noise in bird sound recordings. The following description uses the application of this method for eliminating urban ecological noise in bird sound recordings in server 101 as an example. It can be understood that this method can also be applied to terminal 102.
[0037] Reference Figure 2 , Figure 2 The flowchart illustrates a method for urban ecological noise cancellation in bird call recordings applied to a server, as provided in this application embodiment. The execution entity of this method can be any of the aforementioned computer devices (including a server or terminal). (Refer to...) Figure 2 The method may include the following steps: S210. Collect various types of urban noise data, construct a clean bird sound data set (noise dataset), and then divide it into training set, validation set, and test set for model training. S220. Based on the training set, validation set, and test set, a bird sound noise reduction network is trained. S230. The input on-site recording signal is extracted using the short-time Fourier method to obtain the real and imaginary parts of the on-site recording signal; S240. Input the real part and the imaginary part into the bird sound noise reduction network for calculation to obtain the complex ideal ratio mask; S250. Apply the complex ideal ratio mask to the real and imaginary parts of the on-site recording signal to obtain the Fourier transform result of the on-site recording signal after denoising. S260. Perform an inverse Fourier transform on the denoised Fourier transform to obtain the denoised result of the on-site recording signal.
[0038] In some embodiments, the collection of multiple types of urban noise data to construct a clean bird sound dataset includes: Record various types of noise data within the target geographic area; Collect various types of raw bird sound data within the target geographic area; The noise data and the original bird sound data are processed in a unified manner, and configured into a unified audio format, sampling frequency, quantization accuracy and number of audio channels; Birdsong data with low background noise are filtered from the original bird calls data and corresponding noise reduction processing is performed to obtain clean bird calls data. The noise data of multiple types is sliced, and the duration of a single noise sample is processed to 10 seconds so that each type of noise contains a preset number of training audios. The processed clean bird sound data and noise data are divided into training set, validation set and test set according to a preset ratio.
[0039] In some embodiments, the method further includes: during the training process, mixing clean bird sound data and various noise data according to different discrete signal-to-noise ratio segments to obtain a training set and a validation set, this step including: For each training sample, one clean bird sound data and one noise data are randomly selected from all the clean bird sound data and noise data in the training set. From the selected noise data, a noise segment with a duration of 4 seconds is randomly cut out; Random sampling is performed within a preset discrete signal-to-noise ratio range; Based on the discrete signal-to-noise ratio range, the selected clean bird sound data is mixed with the cropped noise segments to synthesize noisy bird sound data, thereby constructing the training set and validation set.
[0040] In some embodiments, the step of extracting the signal from the input on-site recording signal using the short-time Fourier method to obtain the real and imaginary parts of the on-site recording signal includes: Perform a short-time Fourier transform on the input on-site recording signal to obtain a complex spectrum; The complex spectrum is decomposed into a real matrix and an imaginary matrix using Euler's formula.
[0041] In some embodiments, the bird noise reduction network includes an encoder, a bottleneck layer, and a decoder; The encoder is used to extract multi-scale time-frequency features of bird sounds layer by layer and compress the time dimension; The bottleneck layer is used for global modeling of the long-range dependence of bird sound signals and the correlation between frequency subbands; The decoder is used to filter out irrelevant noise information while preserving and enhancing the amplitude consistency of bird calls.
[0042] In some embodiments, the encoder includes six cascaded complex coding units, each complex coding unit comprising complex convolution, complex batch normalization, and parameterized linear rectifier unit activation functions; The bottleneck layer includes a complex feedforward neural network module, a complex multi-head attention mechanism module with relative position encoding, and a complex convolution enhancement module; The complex feedforward neural network module consists of two complex linear transformation layers sandwiching a complex Swish activation function. A Dropout mechanism is introduced after each linear transformation layer to prevent overfitting. The intermediate layer dimension is expanded to four times the input dimension to enhance nonlinear expressive power. The processing steps of the complex multi-head attention mechanism module include: generating corresponding rotation factors for the frequency index and time index in the complex domain; combining the two rotation factors through an outer product to obtain a joint positional encoding; injecting the joint positional encoding into the query and key, and performing linear projection to generate the corresponding query, key, and value; calculating scaling points and attention based on the query, key, and value; and constructing a complex convolution enhancement module and a complex feedforward network to complete the construction process of the complex multi-head attention mechanism module. The complex convolution enhancement module includes complex layer normalization, complex point convolution, complex gated linear units, complex depthwise convolution, complex batch normalization, complex Swish activation, and a second complex point convolution; the complex convolution enhancement module also includes cross-layer residual connections, used to add the input features to the convolutional features.
[0043] In some embodiments, the decoder construction process includes: A frequency-aware mechanism is applied to combine the skip-gating link between the decoder and the encoder; Based on the output features of the encoder, attention weights are generated in the frequency dimension; The generated attention weights are used to perform element-wise weighted processing on the encoder features; The weighted encoder features and decoder features are fused together to complete the construction of the decoder; The process of fusing the weighted encoder features and decoder features includes: concatenating the weighted encoder features and decoder features along the channel dimension and fusing them through a convolutional layer to generate a gated signal; and multiplying the weighted encoder features element-wise with the gated signal and then passing them through a convolutional layer to obtain the final output features.
[0044] The following, with reference to the accompanying drawings, details the specific implementation process of this application using a particular application scenario as an example: refer to Figure 3 The training and testing algorithm for a deep learning-based bird sound noise reduction model for urban ecology proposed in this application includes the following steps: S1. Construct a clean bird sound dataset and a noisy dataset, divide them into training set, validation set and test set, and design a dynamic hybrid loading strategy; S2. Use the short-time Fourier transform method to extract the real and imaginary parts of the signal and construct the network input and output; S3. Construct a bird sound noise reduction network, namely an encoder, a bottleneck layer, and a decoder; S4. Construct the training loss function and the validation set evaluation function; S5. Set network parameters such as optimizer, learning rate, batch size, and number of training iterations; S6. Train the model and test its performance.
[0045] S7, Real-world audio applications.
[0046] The specific implementation process of steps S1-S7 above will be explained in detail below: 1. The specific details of S1 are as follows: (1) Collect various types of urban noise data. Record various types of noise signals within the target geographical area. The noise signals may include cicada chirping, frog croaking, grasshopper chirping, wind noise, rain noise, thunder noise, airplane noise, car horn noise, pile driver noise, lawn mower noise, waterfall noise, and several other common urban ecological noises.
[0047] (2) Collect bird call data for training and validation. The geographical location of the bird species is not limited, and the number of bird calls is Ns. The bird species used for testing exist in the geographical area of interest, and the number of bird call data is Nt. The frequencies of the above bird calls are distributed in the range of 200Hz-12000Hz. Each category of data includes several audio files.
[0048] (3) The above data format is standardized: audio format: wav, sampling frequency: 32000Hz, quantization precision: 16 bits, number of audio channels: mono.
[0049] (4) Select bird sound data with low background noise, and then use audio software to reduce noise to obtain clean bird sound data. The signal-to-noise ratio of the obtained clean bird sound data should not be less than 30dB. Use a speech activity detection signal processing algorithm to remove silences and ensure that the time interval between two adjacent bird sound signals in each audio is less than 4 seconds. Then manually check to ensure that the bird sound signals are clean. The duration of a single audio for each category of processed bird sound data is 4s, and there are at least 200 training audios for each category of bird sound.
[0050] (5) Slice the noise data of various types. The duration of a single noise sample after processing is 10 seconds. There are at least 400 training audios for each type of noise.
[0051] (6) Divide the clean bird sounds and noise audio that have been processed above into training set, validation set and test set in a ratio of 8:1:1.
[0052] (7) In each training phase, the clean bird sound data and various noise data of the training set and validation set are mixed according to different discrete signal-to-noise ratio segments to obtain the training set and validation set. In the evaluation phase, in order to ensure the stability of the evaluation results, the signal-to-noise ratio is fixed at 5dB. The test set ensures that each type of noise is mixed in each signal-to-noise ratio segment.
[0053] (8) The mixing process of the training set, validation set, and test set is as follows: the duration of a single clean bird sound sample is 4 seconds, and the duration of a noise sample is 10 seconds. During the training phase, for each training sample, a clean bird sound is randomly selected from the clean bird sounds and noise samples in the training set. And noise. From the selected noise, randomly cut out a 4-second noise clip. Random samples are taken from a preset discrete signal-to-noise ratio (SNR) range [-5, 0, 5, 10, 15] dB. Then, based on this SNR... and Mix and synthesize noisy bird sounds. The calculation formula is as follows: , , in, and These represent the power of the clean bird call and the noise signal, respectively. During the verification and testing phase, a 4-second interval was selected from the start of the noise data as... .
[0054] 2. The specific details of step S2 are as follows: Extracting the real and imaginary parts of a signal using the short-time Fourier transform: Let the input time-domain signal be... Perform a short-time Fourier transform on the signal: , in, Represents a predefined window function. Represents the time frame index. Represents the frequency frame index. It represents a plural number.
[0055] Then we use Euler's formula: , The above complex spectrum Decomposed into real part matrix and imaginary part matrix .
[0056] The feature map of the input data of this network is as follows: Where B is the batch number, C is the channel number (composed of real and imaginary parts), F is the number of frequency points in the complex spectrum, and T is the total number of frames. The network input and output dimensions are the same, and the final output is a complex ratio mask, which, after being applied to noisy bird sounds, undergoes inverse short-time Fourier transform to reconstruct the time-domain waveform.
[0057] 3. The specific details of step S3 are as follows: The bird sound noise reduction network structure consists of three parts: an encoder, a bottleneck layer, and a decoder, which need to be built sequentially. (1) Constructing the encoder of the bird sound denoising network. The encoder consists of 6 cascaded complex encoder blocks, used to extract multi-scale time-frequency features of bird sounds layer by layer and compress the time dimension. Each layer contains complex convolution (ComplexConv2d), complex batch normalization (Complex Batch Normalization), and parameterized linear rectified unit (PRelu) activation functions. The architecture is shown in [reference needed]. Figure 4 The input complex spectrum characteristics of the encoder can be expressed as: Two-dimensional convolutional filter Then the entire complex number encoding unit operation can be represented as , in, , This represents the real and imaginary parts of the input to the complex convolutional layer; , Represents the real and imaginary parts of the complex convolution kernel; Normalization of the number of plurals; Represents the PReLU activation function; Represents the output of the convolutional layer; Representing the Layer coding layer; Represents the convolution operation; This is the imaginary part identifier. After 6 encoding modules, the number of channels expands sequentially to {32, 64, 128, 256, 256, 256}. Finally, the time dimension is compressed to 1 / 64 of the original.
[0058] (2) Constructing the bottleneck layer of the bird sound noise reduction network. A complex Conformer module is constructed between the encoder and decoder as the bottleneck layer to globally model the long-range dependencies of the bird sound signal and the correlation between frequency subbands. This is particularly suitable for suppressing non-stationary sudden noise in urban ecological environments. The construction of the bottleneck layer can be divided into the following three modules: 1) Complex Feedforward Neural Network Module: like Figure 5 As shown, the complex feedforward neural network module consists of two layers of complex linear transformation sandwiching a complex Swish activation function. A Dropout mechanism is introduced after each layer of linear transformation to prevent overfitting. The intermediate layer dimension is expanded to four times the input dimension to enhance nonlinear expressive power.
[0059] 2) Complex multi-head attention mechanism module with relative position encoding: like Figure 6 As shown, to address the issues of phase information loss and insufficient modeling of long sequence dependencies in traditional Transformers when processing complex spectra, Joint Complex Rotation Position Encoding (RoPE) is introduced at the bottleneck layer. The feature map output by the encoder is shown. Reshaped into input feature maps that conform to the bottleneck layer In the complex field, we index the frequencies respectively. and time index Generate rotation factor and And combine them through an outer product: , in, This is a hyperparameter that controls the rotation frequency. Before calculating the attention, we inject this joint positional encoding into the query and key: , To strictly adhere to the rules of complex number operations, we use the complex linear projection generated above to generate the query. ,key Sum Perform scaling dot product attention calculation: , in, This represents the conjugate transpose. The specific calculation of the real and imaginary parts involves interaction terms of four components, thus enabling the capture of subtle phase correlations in bird calls: , , Following the multi-head attention mechanism, a complex convolution enhancement module and a complex feedforward network are constructed to form a complete complex Conformer structure, further refining the local time-frequency features.
[0060] 3) Complex Convolution Enhancement Module: like Figure 7 As shown, the complex convolution adopts a sandwich structure, which includes complex layer normalization, complex point convolution, complex gated linear units, complex depthwise convolution, complex batch normalization, complex Swish activation, and a second complex point convolution in sequence. In addition, this module also includes a cross-layer residual connection, which adds the input features to the features processed by the convolution to alleviate the gradient vanishing problem of deep networks and preserve the original information.
[0061] (3) Decoder Construction. The construction of the decoder needs to be combined with the skip-gated link of the encoder. This skip link applies a frequency-aware Squeeze-and-Excitation (SE) mechanism, such as... Figure 8 Assuming encoder features This mechanism generates attention weights in the frequency dimension through the following steps: , in, This indicates that average pooling is performed only along the time dimension, thus preserving information in the frequency dimension. , Two 1x1 convolutional layers are used for feature transformation and dimensionality reduction. The encoder features are then element-wise weighted using the generated attention weights. , Through this frequency-aware weighting, it can prioritize and enhance specific frequency channels containing birdcall energy while suppressing noise-dominated bands. To achieve adaptive interaction between the encoder and decoder, we weight the encoder features... With decoder features The signals are then fused. First, they are concatenated along the channel dimension and then fused through a convolutional layer to generate a soft-gated signal. , Final output Through weighted encoder features With gate signal Element-wise multiplication followed by a convolutional layer yields: , This gated fusion mechanism can adaptively adjust the contribution of encoder features according to the current decoding stage requirements, effectively filtering out irrelevant noise information while preserving and enhancing the amplitude consistency of bird sounds, providing more robust feature fusion in complex acoustic environments.
[0062] 4. The specific details of step S4 are as follows: (1) During the training process, we used a joint loss function in the time domain and frequency domain: , in, It is 0.01. It is 0.3. To calculate the scale-invariant signal-to-noise ratio loss of the signal, To calculate the mean square error loss of the amplitude, To calculate the mean square error loss for the real and imaginary parts.
[0063] (2) The validation process uses scale-invariant signal-to-noise ratio (SI-SNR) to evaluate model performance. The calculation process of SI-SNR is as follows: , , , in, This represents the estimated audio signal. This indicates the audio signal of the tag.
[0064] 5. The specific details of step S5 are as follows: (1) Set the number of epochs and the batch size to 100 and 8 respectively; (2) Optimizer selection. The Adam optimizer was used as the stochastic batch gradient descent optimizer with a learning rate of 0.001, a weight decay of 0.01, and momentum parameters of (0.9, 0.999).
[0065] (3) The learning rate scheduling adopts ReduceLROnPlateau. When the validation set loss has not improved for 5 consecutive epochs, the learning rate is multiplied by 0.7, and the lower limit of the learning rate is set to 1×10. -6 .
[0066] 6. The specific details of step S6 are as follows: (1) The model was trained using a self-built clean bird sound dataset and a bird sound dataset with various types of noise; (2) During the training process, the model training is completed when the loss value of the validation set does not decrease after 12 epochs.
[0067] (3) Select the model that performs best on the validation set as the best model and use the test set to test its performance. If the test performance is not good, fine-tune the model, such as by modifying the learning rate.
[0068] 7. The specific details of step S7 are as follows: (1) Perform Fourier transform on the on-site recording signal, with 640 Fourier transform points, a window length of 320 points, a frame shift of 640 points, and a Hamming window. Calculate the real and imaginary parts of the transformed result using Euler's formula.
[0069] (2) Input the real and imaginary parts calculated above into the optimal model weights to calculate the complex ideal ratio mask.
[0070] (3) Apply the complex ideal ratio mask calculated above to the real and imaginary parts obtained in step (1) to obtain the Fourier transform result of the field recording signal after denoising.
[0071] (4) Perform inverse Fourier transform on the denoised Fourier transform obtained in step (3) to obtain the denoised result of the on-site recording signal.
[0072] The noise reduction effect of the method proposed in this application is compared with that of relevant prior art, and the results are shown in Table 1 below: Table 1. Noise reduction performance of different noise reduction models, with original signal-to-noise ratio -5dB.
[0073]
[0074] ΔSI-SNR represents the improvement in scale-invariant signal-to-noise ratio, which is an indicator of bird noise reduction performance; the higher the better.
[0075] In addition, this embodiment demonstrates Figure 9 , Figure 10 and Figure 11 ,in, Figure 9 The time-frequency distribution of a clean bird call signal is shown. Figure 10 The time-frequency distribution of the noisy bird sound signal obtained by superimposing urban ecological noise onto a clean bird sound signal is shown. Figure 11 The noise reduction results after processing with the urban ecological noise cancellation method described in this application are shown. The horizontal axis represents time, the vertical axis represents frequency, and the energy intensity in the figure is used to characterize the acoustic signal distribution at different time and frequency locations. Specifically, Figure 9The clean bird calls signal is mainly characterized by a relatively clear, continuous or discrete bird call pattern structure within a local time period, while the non-target energy in the background area is low. Figure 10 In this study, noisy bird calls, after being superimposed with urban ecological noises such as aircraft noise, pile driver noise, and katydid noise, exhibit continuous or diffuse noise energy distribution over a wide frequency range. Furthermore, some noise energy overlaps with the bird call signal in both time and frequency dimensions, resulting in the masking of the bird call signal's edge structure, main frequency band structure, and local time-frequency texture. Further reference... Figure 11 After processing by the method of this application, Figure 10 Background noise energy distributed over medium to large areas, low-frequency continuous noise energy, and non-target broadband noise energy are significantly reduced; meanwhile, Figure 9 The corresponding bird call frequency band, harmonic structure, frequency modulation trajectory, and main time-frequency texture are in Figure 11 The data was well preserved. Therefore, through comparison... Figure 9 , Figure 10 and Figure 11 It can be confirmed that: Figure 10 Compared to Figure 9 It significantly increases urban ecological noise interference and reduces the discernibility of target bird calls; Figure 11 Compared to Figure 10 It removed a large amount of background noise and non-target noise components, and Figure 11 Bird sound patterns and Figure 9 The clean bird call spectrum in the data maintains high consistency in appearance time, main frequency range, and local structure. This result demonstrates that the method described in this application can effectively improve the signal-to-noise ratio of noisy bird call signals, enhance the target bird call components, and reduce the interference of urban ecological noise on downstream tasks such as bird call detection, identification, or classification. It can be seen that this application extracts the real and imaginary parts of the field recording signal through short-time Fourier transform, estimates a complex ideal ratio mask using a bird call denoising network, and then applies this complex ideal ratio mask to the complex spectrum of the noisy signal. This can suppress non-target noise components in the complex time-frequency domain while preserving the main time-frequency structure of the target bird call signal.
[0076] Compared with existing technologies, this application constructs a model architecture based on complex number encoding and decoding structure, designs a complex number Conformer as the bottleneck layer for feature extraction, and as described in step S3 above, this application designs a frequency domain gated attention mechanism based on SE. This application achieves better noise reduction effect than other models on various types of noise, and the denoised spectrogram is cleaner, which is more conducive to improving the performance of downstream tasks.
[0077] Another aspect of this application provides an urban ecological noise cancellation device for bird call recording, comprising: The first module is used to collect various types of urban noise data, construct a clean bird sound data set (noise dataset), and then divide it into training set, validation set and test set for model training. The second module is used to train a bird sound noise reduction network based on the training set, validation set, and test set. The third module is used to extract the signal from the input on-site recording signal using the short-time Fourier method, and obtain the real part and imaginary part of the on-site recording signal. The fourth module is used to input the real part and the imaginary part into the bird sound noise reduction network for calculation to obtain a complex ideal ratio mask; The fifth module is used to apply the complex ideal ratio mask to the real and imaginary parts of the on-site recording signal to obtain the Fourier transform result of the on-site recording signal after denoising. The sixth module is used to perform an inverse Fourier transform on the denoised Fourier transform result to obtain the denoised result of the on-site recording signal.
[0078] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0079] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for eliminating urban ecological noise in bird call recording. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0080] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0081] Please see Figure 12 , Figure 12 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1201 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1202 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1202 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1202 and is called and executed by the processor 1201 to execute the urban ecological noise cancellation method for bird sound recording according to the embodiments of this application. The input / output interface 1203 is used to implement information input and output; The communication interface 1204 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1205 transmits information between various components of the device (e.g., processor 1201, memory 1202, input / output interface 1203, and communication interface 1204); The processor 1201, memory 1202, input / output interface 1203 and communication interface 1204 are connected to each other within the device via bus 1205.
[0082] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for eliminating urban ecological noise in bird sound recordings.
[0083] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0084] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0085] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0086] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0087] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0088] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0089] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0090] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0091] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0092] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for eliminating urban ecological noise in bird call recordings, characterized in that, include: Collect various types of urban noise data, construct a clean bird sound dataset, and then divide it into training, validation, and test sets for model training. The bird sound noise reduction network is trained based on the training set, validation set, and test set. The input on-site recording signal is extracted using the short-time Fourier transform method to obtain the real and imaginary parts of the on-site recording signal; The real part and the imaginary part are input into the bird sound noise reduction network for calculation to obtain the complex ideal ratio mask; The complex ideal ratio mask is applied to the real and imaginary parts of the on-site recording signal to obtain the Fourier transform result of the on-site recording signal after denoising. The denoised Fourier transform result is then subjected to inverse Fourier transform to obtain the denoised result of the on-site recording signal.
2. The method for eliminating urban ecological noise in bird sound recordings according to claim 1, characterized in that, The process of collecting various types of urban noise data and constructing a clean bird sound dataset includes: Record various types of noise data within the target geographic area; Collect various types of raw bird sound data within the target geographic area; The noise data and the original bird sound data are processed in a unified manner, and configured into a unified audio format, sampling frequency, quantization accuracy and number of audio channels; Birdsong data with low background noise are filtered from the original bird calls data and corresponding noise reduction processing is performed to obtain clean bird calls data. The noise data of multiple types is sliced, and the duration of a single noise sample is processed to 10 seconds so that each type of noise contains a preset number of training audios. The processed clean bird sound data and noise data are divided into training set, validation set and test set according to a preset ratio.
3. The method for eliminating urban ecological noise in bird sound recordings according to claim 2, characterized in that, The method further includes: during the training process, mixing clean bird sound data and various noise data according to different discrete signal-to-noise ratio segments to obtain a training set and a validation set. This step includes: For each training sample, one clean bird sound data and one noise data are randomly selected from all the clean bird sound data and noise data in the training set. From the selected noise data, a noise segment with a duration of 4 seconds is randomly cut out; Random sampling is performed within a preset discrete signal-to-noise ratio range; Based on the discrete signal-to-noise ratio range, the selected clean bird sound data is mixed with the cropped noise segments to synthesize noisy bird sound data, thereby constructing the training set and validation set.
4. The method for eliminating urban ecological noise in bird call recordings according to claim 1, characterized in that, The step of extracting the real and imaginary parts of the input on-site recording signal using the short-time Fourier transform method includes: Perform a short-time Fourier transform on the input on-site recording signal to obtain a complex spectrum; The complex spectrum is decomposed into a real matrix and an imaginary matrix using Euler's formula.
5. A method for eliminating urban ecological noise in bird call recordings according to claim 1, characterized in that, The bird noise reduction network includes an encoder, a bottleneck layer, and a decoder; The encoder is used to extract multi-scale time-frequency features of bird sounds layer by layer and compress the time dimension; The bottleneck layer is used for global modeling of the long-range dependence of bird sound signals and the correlation between frequency subbands; The decoder is used to filter out irrelevant noise information while preserving and enhancing the amplitude consistency of bird calls.
6. A method for eliminating urban ecological noise in bird sound recordings according to claim 5, characterized in that, The encoder includes six cascaded complex coding units, each layer of which contains complex convolution, complex batch normalization, and parameterized linear rectifier unit activation functions; The bottleneck layer includes a complex feedforward neural network module, a complex multi-head attention mechanism module with relative position encoding, and a complex convolution enhancement module; The complex feedforward neural network module consists of two complex linear transformation layers sandwiching a complex Swish activation function. A Dropout mechanism is introduced after each linear transformation layer to prevent overfitting. The intermediate layer dimension is expanded to four times the input dimension to enhance nonlinear expressive power. The processing steps of the complex multi-head attention mechanism module include: generating corresponding rotation factors for the frequency index and time index in the complex domain; combining the two rotation factors through an outer product to obtain a joint positional encoding; injecting the joint positional encoding into the query and key, and performing linear projection to generate the corresponding query, key, and value; calculating scaling points and attention based on the query, key, and value; and constructing a complex convolution enhancement module and a complex feedforward network to complete the construction process of the complex multi-head attention mechanism module. The complex convolution enhancement module includes complex layer normalization, complex point convolution, complex gated linear units, complex depthwise convolution, complex batch normalization, complex Swish activation, and a second complex point convolution; the complex convolution enhancement module also includes cross-layer residual connections, used to add the input features to the convolutional features.
7. A method for eliminating urban ecological noise in bird sound recordings according to claim 5, characterized in that, The construction process of the decoder includes: A frequency-aware mechanism is applied to combine the skip-gating link between the decoder and the encoder; Based on the output features of the encoder, attention weights are generated in the frequency dimension; The generated attention weights are used to perform element-wise weighted processing on the encoder features; The weighted encoder features and decoder features are fused together to complete the construction of the decoder; The process of fusing the weighted encoder features and decoder features includes: concatenating the weighted encoder features and decoder features along the channel dimension and fusing them through a convolutional layer to generate a gated signal; and multiplying the weighted encoder features element-wise with the gated signal and then passing them through a convolutional layer to obtain the final output features.
8. A device for eliminating urban ecological noise in bird call recording, characterized in that, include: The first module is used to collect various types of urban noise data, construct a clean bird sound data set (noise dataset), and then divide it into training set, validation set and test set for model training. The second module is used to train a bird sound noise reduction network based on the training set, validation set, and test set. The third module is used to extract the signal from the input on-site recording signal using the short-time Fourier method, and obtain the real part and imaginary part of the on-site recording signal. The fourth module is used to input the real part and the imaginary part into the bird sound noise reduction network for calculation to obtain a complex ideal ratio mask; The fifth module is used to apply the complex ideal ratio mask to the real and imaginary parts of the on-site recording signal to obtain the Fourier transform result of the on-site recording signal after denoising. The sixth module is used to perform an inverse Fourier transform on the denoised Fourier transform result to obtain the denoised result of the on-site recording signal.
9. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.