Automatic Acoustic Identification System and Method for Birds in Wetland Environments

By employing a four-step process of multi-source noise modeling, blind source separation, and multi-label recognition, combined with an acoustic dictionary and a hybrid multi-label classification network, the problem of bird call recognition in wetland environments with overlapping multiple sound sources and noisy backgrounds has been solved. This has enabled efficient and accurate automatic acoustic recognition of birds, supporting ecological monitoring and protection decisions.

CN120220702BActive Publication Date: 2025-10-28云南省林业调查规划院(云南省森林和草原资源监测中心、云南省自然保护地研究监测中心)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510471442.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-10-28
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively separate and accurately identify multiple bird calls in wetland environments, especially when multiple sound sources overlap and the background is noisy, resulting in a significant drop in accuracy. Furthermore, traditional models are prone to confusion or missed detection of calls from new species that are not recorded or labeled, leading to low efficiency and accuracy in automatic monitoring.

Method used

The method employs a four-step process: multi-source noise modeling, blind source separation and noise reduction, multi-label recognition, and multi-modal fusion. By constructing an acoustic dictionary, a blind source separation model, a hybrid multi-label classification network, and an adaptive update strategy, combined with CNN and Transformer structures, it achieves the decomposition of multiple sound sources and the precise recognition of bird call features.

Benefits of technology

Maintaining high recognition accuracy in noisy, multi-species chorus scenarios, reducing misclassification and omissions, supporting long-term ecological monitoring and conservation decision-making, and providing an efficient automatic bird acoustic recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220702B_ABST
    Figure CN120220702B_ABST
Patent Text Reader

Abstract

This invention discloses an automatic acoustic identification system and method for birds in wetland environments, relating to the fields of ecological monitoring and acoustic signal recognition technology. Addressing the complex acoustic environment of wetlands, it proposes a four-step process: multi-source noise modeling, blind source separation and noise reduction, multi-label identification, and multi-modal fusion feedback. Step one involves constructing an acoustic dictionary and mixing noise and birdsong using contribution coefficients. Step two uses a self-supervised method to demix multiple sound sources, outputting a quasi-clean track. Step three combines multi-label classification with dynamic spectrum enhancement time and frequency enhancement factors to identify and label overlapping birdsong. Step four uses microphone array and external environmental parameter fusion to determine the noise type and updates the noise and birdsong templates to form a closed-loop iteration, maintaining high recognition accuracy even in noisy, multi-species chorus scenarios, providing efficient support for wetland ecological monitoring and protection decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological monitoring and acoustic signal recognition technology, specifically to an automatic acoustic recognition system and method for birds in wetland environments. Background Technology

[0002] In wetland ecosystems, birds emit various types of calls at different times of day to meet their survival and reproductive needs. These acoustic signals not only reflect species characteristics but also environmental changes and ecological health. Wetland environments typically contain rich and varied background noise sources, such as wind, water flow, frog croaks, and insect calls, whose intensity and frequency distribution dynamically change with weather, season, and topography. Furthermore, multiple bird species may inhabit wetlands simultaneously, and the common dawn "chorus" phenomenon results in a high degree of overlap in the spectrum of various bird calls. Faced with such a complex acoustic environment with multiple sound sources, manual segment labeling and filtering is time-consuming and laborious, making it difficult to meet the needs of large-scale, long-term monitoring. Therefore, to achieve long-term, automated collection and analysis of wetland bird diversity and behavioral patterns, it is necessary to construct a highly robust acoustic processing system adaptable to multiple scenarios and noise types to obtain timely bird population information and support ecological conservation decisions.

[0003] However, existing technologies often suffer from a significant drop in recognition accuracy or fail to effectively separate multiple bird calls when faced with chorusing birds in wetlands and noisy backgrounds. Since most traditional algorithms or models assume a single, clear species' sound as input, they struggle to accurately demix and classify multiple overlapping sound sources. Furthermore, the amplitude and frequency range of various noises in wetlands vary greatly; wind and water sounds exhibit randomness and time-varying characteristics, and background noises such as frog croaks and cicada chirps may overlap with the target bird calls in frequency bands or time periods, rendering simple filtering or fixed-bandwidth detection ineffective.

[0004] When new species calls that are not recorded or labeled are discovered, traditional models are more likely to cause confusion or miss detection. In extreme environments or under fault conditions, blind source separation and identification will also fail, requiring professionals to repeatedly listen to and manually screen a large number of audio clips, which directly reduces the efficiency and accuracy of automatic monitoring and limits the subsequent in-depth analysis and scientific application of bird ecological data.

[0005] Therefore, the present invention provides an automatic acoustic identification system and method for birds in wetland environments. Summary of the Invention

[0006] (a) Technical problems to be solved

[0007] To address the shortcomings of existing technologies, this invention provides an automatic acoustic identification system and method for birds in wetland environments. By proposing a four-step process of multi-source noise modeling, blind source separation and noise reduction, multi-label identification, and multi-modal fusion feedback, it maintains high identification accuracy in noisy, multi-species chorus scenarios, providing efficient support for wetland ecological monitoring and protection decisions; and solves the technical problems described in the background art.

[0008] (II) Technical Solution

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] Automatic acoustic identification methods for birds in wetland environments, including:

[0011] When initial processing of wetland audio is detected, background sound R is collected over multiple time periods. bg With pure birdsong R sing And construct an acoustic dictionary D wetland Combined with the controllable contribution coefficient α b and β s Noise and birdsong are simulated and preliminarily segmented and anomaly removed to reduce the risk of data contamination. The acquisition time and scene description are recorded synchronously in the metadata, and then the labeled and effective target audio recordings are output. target With abnormal or unavailable fragments R drop ;

[0012] If multi-source audio source decomposition is required to effectively target the recording R target Call the blind source separation model Φ sep Iterative training is performed using a loss function, along with matching degree. The mixed signal track was broken down and denoised one by one to preserve acoustic details for weak bird calls, improve the signal-to-noise ratio to avoid confusion caused by multiple sound source overlap, and finally generate a near-clean track.

[0013] When the orbit is nearly pure When generating and determining the bird species included in each track, the data is input into a hybrid multi-label classification network Φ. class We combine CNN and Transformer structures for deep representation, use Renyi entropy to improve loss and enhance multi-species discrimination, and then use enhancement factors to dynamically weight the dawn or main frequency band to highlight weak bird calls. This produces an information set containing multi-species labels and confidence scores, and labels the corresponding time segments.

[0014] When the recognition results and environmental monitoring requirements trigger multimodal fusion, multi-channel audio data M is collected. qUsing the spatial positioning vector G(τ) and external environmental parameters E(τ), the system identifies dominant noise sources such as strong winds or water flow through a spatial positioning function and an adaptive update strategy. Upon detecting false positives, missed detections, or the emergence of new species, the system feeds back the corresponding orbital data and labeled logs to previous steps to supplement the noise template Φ. b (t,f) and bird song template Ψ s (t,f).

[0015] Preferably, a diverse dataset of raw wetland audio recordings is collected at different time intervals. Taw Background sound R bg Pure birdsong sing ;

[0016] For the audio recording dataset R raw Key bird song segments appearing in the data were tagged with species and recorded along with the corresponding background noise types for the same time periods, forming a preliminary tagging table L. stagel Audio segments with quality defects will be removed or separately recorded as fault records. fault ,

[0017] Preferably, an acoustic dictionary D is constructed based on the obtained multi-source data. wetland Among them: several background noise templates Φ b (t,f)(b=1,…,B) and single-species bird song template Ψ s (t,f), s=1,…,S, and construct the entire wetland environment combination model W(t,f);

[0018] By combining different contribution coefficients α b and contribution coefficient β s The values ​​of R are used to synthesize various complex scenes to obtain simulated synthesized soundscapes. simu The original recording R raw Combined with preliminary labeling table L stagel Perform more detailed segmentation and anomaly removal;

[0019] For audio segments that are determined to be invalid or subject to particularly severe interference, abnormal or unusable segments (R) can be filtered according to preset filtering rules. drop And indicate its type;

[0020] For segments that can be partially utilized, time-series segmentation and extraction function F are used. segment Divide the bird song segment R containing the main bird songs target and background segment R containing only background bgOnly ;

[0021] Preferably, construct the training sample set R. mix The training sample set R mix It can be directly from actual recordings, or from the acoustic dictionary D.wetland Perform random weight combination;

[0022] Let the depth model be Φ sep Its input is a single-channel mixed signal x(t) or its corresponding time-frequency representation X(t,f), and its output is M separate tracks. Introducing a self-supervised approach to define a comprehensive loss function Using the above loss function and training sample set R mix For the depth model Φ sep Perform iterative training;

[0023] Preferably, for the trained separation network Φ sep Inference is performed, with actual wetland recordings as input, and the separation network Φ is used. sep Output M initial separation tracks

[0024] With the help of the acoustic dictionary D wetland Medium background template Φ b (t,f) and separation results The degree of matching of the construction

[0025] Matching degree For orbits exceeding the threshold θ, a nonlinear noise reduction filtering operator γ is applied. m Noise reduction processing is performed to create a near-clean orbital array. Preserve the time-frequency distribution corresponding to each track and noise confidence labels;

[0026] Preferred, with quasi-pure orbit Using the input, construct a hybrid multi-label classification network Φ class ;

[0027] For each quasi-pure orbit It is segmented into several short time windows and time-frequency features are extracted, and then applied to a hybrid multi-label classification network Φ. class Batch processing is performed to obtain the corresponding predicted species distribution P. m =(p m,1 ,p m,2 ,…,p m,s ), where p m,s This represents an estimate of the probability that the orbit contains species s;

[0028] Constructing an improved multi-label classification loss based on the Renyi entropy concept And designed a single-class penalty function Φ R (p m,s ,α);

[0029] When the probability estimate pm,s Deviation from true label y m,s When, the single-class penalty function Φ R (p m,s ,α) generates a non-linearly increasing penalty, improving the differentiation of easily confused species in multi-species scenarios; if multiple single orbitals appear in the same time period Then, according to the hybrid multi-label classification network Φ class Output probability estimate p m,s

[0030] Preferably, a dynamic spectral enhancement function Ω(t,f) is constructed and aligned with the pure orbit. Perform targeted weighting to obtain the time-frequency distribution. And then construct the time enhancement factor Γ time and frequency enhancement factor Γ freq Time-frequency distribution For re-feeding into the hybrid multi-label classification network φ class Perform a re-evaluation and output time segment markers;

[0031] Preferably, when monitoring at fixed points in wetlands or during drone patrols, microphone arrays are deployed at multiple spatial locations to generate multi-channel audio data {M1, M2, ..., M...}. Q}; and make multi-channel audio data M q (t) and quasi-pure orbit Corresponding to step size or sampling rate;

[0032] Collect external environmental parameters E(τ) that are strongly correlated with the wetland scene; perform GPS / BeiDou positioning on drones or fixed points to generate spatial positioning vector G(τ);

[0033] Preferably, based on the acquired multi-channel audio data {M1, M2, ..., M... Q} and the spatial positioning vector G(τ), which can be used to estimate the spatial location of the target bird's call using techniques such as time of arrival or coherent processing;

[0034] Constructing the positioning function The estimated coordinates R of the output sound source in three-dimensional space m (τ) is used to distinguish birdsong from different directions. Combined with species tags and time segment markers, it is used to analyze and obtain information on which birds are making calls in which places and when.

[0035] Define adaptive adjustment function When an abnormal noise distribution or new interference is detected, the contribution coefficient α is locally corrected based on a known background template or newly collected noise samples. b and β s wait;

[0036] If rare bird species are observed at a specific spatiotemporal location even under extreme environmental parameter E(τ), then the quasi-pure orbit will be re-tested. Dynamic spectrum enhancement and re-determination are performed to add emphasis markers to the multi-species label output, along with the estimated coordinates R. m (τ) is derived;

[0037] Preferably, in the event of false detection or missed detection, the relevant audio track and its time period, as well as the species tag information, are recorded. If the actual bird species does not match the identification tag, the difference is recorded in the error log H. error ;

[0038] For new or previously unrecorded noise types, write them to the noise log H. noise ;

[0039] If noise log H noise If new noise samples or new noise combinations are recorded, then the background template Φ... b (t,f) can be expanded or the training dataset R can be modified. simu Make incremental additions;

[0040] If error log H error This indicates that some species have a high rate of identification confusion. Increasing the pure bird song R of that species can help. sing Or build more hybrid scenarios α b and β s Combining these technologies improves the ability to identify related confused species;

[0041] Based on the feedback and updates generated in the first two steps, the monitoring strategy will be adjusted in subsequent deployments:

[0042] Increase sampling frequency for periods or locations where missed detections frequently occur, deploy denser microphone arrays in key wetland areas, and if it is confirmed that noise is extremely severe during certain periods or that drone recording fails under specific weather conditions, then the workflow can flexibly skip those periods or adopt alternative data sources to improve overall efficiency and accuracy.

[0043] An automatic acoustic identification system for birds in wetland environments, including:

[0044] The audio data processing module, when it detects that initial processing of wetland audio is required, acquires background audio data from multiple time periods. bg With pure birdsong R sing And construct an acoustic dictionary D wetland Combined with the controllable contribution coefficient α b and β s Noise and birdsong are simulated and preliminarily segmented and anomaly removed to reduce the risk of data contamination. The acquisition time and scene description are recorded synchronously in the metadata, and then the labeled and effective target audio recordings are output. targetWith abnormal or unavailable fragments R drop ;

[0045] The separation module, if it is necessary to decompose the effective target recording from multiple sound sources, R target Call the blind source separation model Φ sep Iterative training is performed using a loss function, along with matching degree. The mixed signal track was broken down and denoised one by one to preserve acoustic details for weak bird calls, improve the signal-to-noise ratio to avoid confusion caused by multiple sound source overlap, and finally generate a near-clean track.

[0046] The marking module, when the orbit is quasi-pure. When generating and determining the bird species included in each track, the data is input into a hybrid multi-label classification network Φ. class We combine CNN and Transformer structures for deep representation, use Renyi entropy to improve loss and enhance multi-species discrimination, and then use enhancement factors to dynamically weight the dawn or main frequency band to highlight weak bird calls. This produces an information set containing multi-species labels and confidence scores, and labels the corresponding time segments.

[0047] The adaptive update module acquires multi-channel audio data M when the recognition results and environmental monitoring requirements trigger multi-modal fusion. q Using the spatial positioning vector G(τ) and external environmental parameters E(τ), the system identifies dominant noise sources such as strong winds or water flow through a spatial positioning function and an adaptive update strategy. Upon detecting false positives, missed detections, or the emergence of new species, the system feeds back the corresponding orbital data and labeled logs to previous steps to supplement the noise template Φ. b (t,f) and bird song template Ψ s (t,f).

[0048] (III) Beneficial Effects

[0049] This invention provides an automatic acoustic identification system and method for birds in wetland environments, which has the following beneficial effects:

[0050] By establishing a flexibly schedulable acoustic dictionary D wetland And using the contribution coefficient α b With β s For background template Φ b (t,f) and the bird song template Ψ s Controllable mixing of (t,f) effectively improves the realism of simulating multiple noise sources and multiple bird chorus scenes in wetlands, and helps the subsequent model remain robust under extreme disturbance conditions;

[0051] Using a self-supervised blind source separation network Φ sep and matching degree Multi-track splitting of real or synthetic mixed audio allows for the separation and fine-tuning of previously superimposed bird calls and noise. Without relying on a clean bird call benchmark, by adjusting the filtering intensity and similarity threshold of each track, misclassification and omission are effectively reduced, while preserving sufficient acoustic details for subsequent detection of weak or rare species.

[0052] A multi-label classification network combines a hybrid structure of CNN and Transformer, improves species discrimination accuracy by enhancing the loss function through Renyi entropy, and utilizes a time-enhancing factor Γ. time With frequency enhancement factor Γ freq Dynamic weighting of dawn or specific frequency bands can be used to enhance weak calls in a targeted manner. At the same time, it supports the parallel discrimination of multiple species within the same orbit, thus taking into account the wetland chorus phenomenon. This dynamic enhancement mechanism can make full use of the activity patterns of wetland birds in time and frequency, and also avoid missing rare or weak bird species.

[0053] Step four introduces multimodal fusion, utilizing spatial positioning information and environmental parameters such as wind speed and water flow to help identify strong wind noise and other unconventional signals, and adaptively updates the noise template or blind source separation model Φ. sep The parameters enable long-term model evolution. The feedback mechanism for false positives and false negatives can also incorporate newly added species or novel noises into the acoustic dictionary D. wetland It enables continuous learning and maintains high reliability. When facing dynamic changes in the ecological environment, it can maximize the detection efficiency of multiple bird species and significantly reduce the burden of manual verification, providing more efficient and reliable data support for wetland protection and biodiversity research. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the automatic acoustic identification method for birds in wetland environments according to the present invention;

[0055] Figure 2 This is a schematic diagram of the automatic acoustic identification system for birds in wetland environments according to the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Please see Figure 1 This invention provides an automatic acoustic identification method for birds in wetland environments, including,

[0058] Step 1: When initial processing of wetland audio is detected, background sound R is collected over multiple time periods. bg With pure birdsong R sing And construct an acoustic dictionary D wetland Combined with the controllable contribution coefficient α b and β s Noise and birdsong are simulated and preliminarily segmented and anomaly removed to reduce the risk of data contamination. The acquisition time and scene description are recorded synchronously in the metadata, and then the labeled and effective target audio recordings are output. target With abnormal or unavailable fragments R drop ;

[0059] Step one includes the following:

[0060] Step 101: Multi-source wetland audio acquisition and time-segmented labeling

[0061] Based on methods such as fixed monitoring points and drone patrols, diverse raw wetland audio data were collected at different times. The data object was named Audio Dataset R. Taw For typical background noises such as water flow, wind, insect chirping, and amphibian calls, separate pure background segments were recorded, and these audio segments were uniformly named "Background Sound R". bg .

[0062] For common bird species, efforts are made to obtain individual bird calls with high signal-to-noise ratios and clear species identification, which are named "Pure Bird Call R". sing ;

[0063] With manual or semi-automatic assistance, the audio dataset R... raw Key bird song segments appearing in the data were tagged with species and recorded along with the corresponding background noise types for the same time periods, forming a preliminary tagging table L. stagel ;

[0064] Considering the temporal and spatial variability of wetland environments, repeated data collection is necessary under different weather and seasonal conditions to obtain a more comprehensive understanding of scene differences.

[0065] Audio segments with quality defects (such as microphone malfunction, extremely weak recording signal, etc.) will be removed or separately marked as fault records. fault This is to ensure the robustness of subsequent model training and analysis;

[0066] When in use, various types of pure background sound R are collected by separation. bg It can precisely characterize common noise sources in wetlands, reduce uncertainties in subsequent processing, and collect pure bird song data. singSpecies labeling is performed to enable the construction of biologically meaningful reference templates, which has higher distinguishability in multi-bird mixed scenes. Faulty or extremely poor quality recordings are identified and removed separately to avoid invalid information from polluting the training process of subsequent recognition models.

[0067] Step 102: Construction of environmental acoustic dictionary and preprocessing of target audio segments

[0068] An acoustic dictionary D is constructed based on the obtained multi-source data. wetland Among them: several background noise templates Φ b (t,f)b(=1,…,B) and single-species bird song template Ψ s (t,f), s=1,…,S, and the general expression of the entire wetland environment combination model W(t,f) is defined using the following formula:

[0069]

[0070] Where: Φ b (t,f), Ψ s (t, f) represent the background sound R. bg And pure birdsong R sing The extracted time-frequency template function is W(t,f), which is the composite environment time-frequency distribution function, i.e., the wetland environment combination model, used to simulate the phenomenon of multiple sound sources superimposed in wetlands.

[0071] α b Background noise template Φ b The contribution coefficient of (t,f) is a variable whose larger value indicates a stronger noise component, and can be taken in the range [0,1].

[0072] β s Template Ψ for the song of a certain species s The contribution coefficient of (t,f) is used to control the relative loudness of the species in the mixed scene, and can be taken in the range [0,1]; B and S are the upper limits of the number of background noise and bird song templates;

[0073] t and f are time and frequency coordinates used to characterize the distribution of sound changes in the short-time Fourier transform or other time-frequency transform domains;

[0074] Among them, the background sound R is usually used bgEach background sound segment is processed using a Short-Time Fourier Transform (STFT) or other time-frequency transforms (such as CQT, wavelet transform, etc.) to obtain a two-dimensional distribution of the amplitude spectrum evolving over time. For the STFT implementation, a window length (e.g., 1024 points) and a step size (e.g., 256 points) can be set to frame the audio signal and calculate the complex spectrum. After extracting the time-frequency representations of several background sound segments, an average or representative template can be formed in the following ways: directly select a relatively stable background interval without sudden interference as the template; statistically aggregate the amplitude spectra of multiple background segments (e.g., take the average or median), and appropriately sample the time dimension to obtain a relatively stable and representative time-frequency distribution function; if it is necessary to subdivide different background types (wind sound, water flow, etc.), a corresponding noise template Φ can be generated for each type. b (t,f); Birdsong template Ψ s The construction method of (t,f) is similar;

[0075] By combining the above formulas, different contribution coefficients α can be obtained. b and contribution coefficient β s The value of can be used to synthesize various complex scenes (including backgrounds of strong winds and rapid currents, choruses of multiple species, etc.) in the laboratory or during the algorithm training phase, to obtain important simulated synthesized soundscapes R for training and validating subsequent source separation models. simu The original recording R raw Combined with preliminary labeling table L stagel Perform more detailed segmentation and anomaly removal;

[0076] For audio segments that are determined to be invalid or subject to particularly severe interference, abnormal or unusable segments (R) can be filtered according to preset filtering rules. drop The audio segment should be labeled with its type. The rules can be as follows: if the audio has no valid signal for most of the time (e.g., amplitude too low, almost zero level, or extreme overload), or if fault indicators such as device power failure or microphone malfunction appear in the metadata record, this audio segment should be classified as faulty or invalid. Audio recordings may contain fragmented files that are too short or not fully recorded due to faults; these can also be merged into the abnormal or unusable segment category R. drop ;

[0077] For segments that can be partially utilized, time-series segmentation and extraction function F are used. segment Divide the bird song segment R containing the main bird songs target and background segment R containing only background bgOnly Temporal segmentation and extraction function F segmentIt receives input audio data (such as filtered candidate audio) and outputs several sub-segments, where each sub-segment is either a core segment containing the main bird calls or an auxiliary control segment without bird calls or containing only background; that is, it outputs and labels classification results through multi-level threshold detection, template or feature matching, time sequence merging and boundary correction.

[0078] In this process, the background template Φ in the wetland environment composite model W(t,f) constructed in this step can be used. b (t,f) performs matching and verification of noise features to better identify non-birdsong-dominated segments;

[0079] The final acoustic dictionary D wetland Its core consists of various background templates Φ b Birdsong templates Ψ s and the corresponding set of contribution coefficients {α} b}、{β s The effective target recording R after segmentation is composed of [data / structure]. target Recording with optional background contrast R bgOnly Along with abnormal or unavailable fragments R with clear reasons drop The above data and labels will be imported into the subsequent multi-source separation and recognition process;

[0080] When using it, the contribution coefficient α b With β s The dynamic allocation can produce a variety of realistic synthesized soundscapes, which can simulate the actual noisy and multi-bird overlapping scenes to a large extent. The accurate background template and target call template provide a traceable and adjustable foundation for subsequent self-supervised separation and multi-label recognition, making the algorithm more adaptable and generalizable. Based on different noise intensity and mixing degree cases synthesized by the dictionary model, the error-prone points of the algorithm in extreme environments can be discovered more quickly and targeted improvements can be made.

[0081] Step 2: If multiple sound sources are required, analyze the effective target recording R. target Call the blind source separation model Φ sep Iterative training is performed using a loss function, along with matching degree. The mixed signal track was broken down and denoised one by one to preserve acoustic details for weak bird calls, improve the signal-to-noise ratio to avoid confusion caused by multiple sound source overlap, and finally generate a near-clean track.

[0082] Step two includes the following:

[0083] Step 201: Self-supervised network training and sound source separation structure construction

[0084] Based on the bird song segment R output in step one target Background section RbgOnly and analog synthesized soundscapes R simu Construct a representative training sample set R mix , where R mix It can be directly from a real recording (containing multiple birds and background noise), or from the acoustic dictionary D in step one. wetland Perform random weight combination (using contribution coefficient α) b and β s Synthesized to cover a wide range of scenarios with varying numbers of species and different noise intensities;

[0085] Let the deep model (separation network) be Φ sep Its input is a single-channel mixed signal x(t) or its corresponding time-frequency representation X(t,f), and its output is M separate tracks. In practice, M can be dynamically set according to the maximum number of sound sources to be extracted; the separation network Φ sep It is generally composed of convolutional layers, attention mechanisms or other neural network modules. To ensure compatibility with subsequent noise reduction processes, an intermediate layer that can uniformly process multiple sound sources needs to be retained in the network structure.

[0086] Traditional separation methods often rely on a clean target reference, but it is difficult to obtain a clean ground truth in the mixed bird environment of wetlands. Therefore, a self-supervised approach is introduced: it only requires that the linear or energy superposition of all output orbitals matches the original mixed input, while minimizing cross-contamination between orbitals. A comprehensive loss function can be defined as follows.

[0087]

[0088] In the formula: R represents the distribution of the m-th separated orbit output by the network in the time-frequency domain; mix (t,f) represents the time-frequency domain representation of the mixed input; ∥·∥ p It is the generalized Minkowski norm (p>1), which aims to provide a higher-order measure of the overall reconstruction error;

[0089] Γ(·,·) is the cross-discrimination measure function between orbits. Renyi divergence, improved coherence distance or other higher-order measures can be used to ensure that cross-contamination between different orbits is minimized.

[0090] λ is a balance coefficient that can be adjusted within the range of 0 to 1 to control the relative weight of reconstruction accuracy and track discrimination.

[0091] Using the above loss function, the deep model (separation network) Φ sep Iterative training was conducted, with the training set consisting of real wetland mixed audio and the synthesis contribution coefficient α. b ,βs Different values ​​yield the simulated synthesized soundscape R simu ;

[0092] The separation effect is continuously monitored using a validation set (which may also include actual wetland recordings from different times and locations). If insufficient separation is found for specific noise types or particular bird calls, further steps can be taken in the acoustic dictionary D from step one. wetland Add corresponding background templates Φ b (t,f) or bird song template Ψ s (t,f) to enrich the training data;

[0093] When used, network training can be completed without obtaining clean recordings of birdsong one by one, which greatly reduces the cost of data annotation. By using high-order metrics and orbital discrimination Γ, the accuracy of mixed sound reconstruction is guaranteed while effectively reducing the problem of mutual contamination of birdsong in the same frequency band. With the help of the controllable synthesized soundscape provided in the first step, the model can be exposed to a wide range of sounds during the training phase, improving its adaptability to extreme noise and multiple bird superposition.

[0094] Step 202: Noise Reduction and Quasi-Pure Orbit Generation

[0095] For the trained separation network Φ sep Inference is performed, with the input being actual wetland recordings (filtered R). target Or new real-time data), separate network Φ sep Output M initial separation tracks

[0096] To identify which tracks are the main sources of noise, an acoustic dictionary D can be used. wetland Medium background template Φ b (t,f) and separation results The similarity is determined using the following matching degree.

[0097]

[0098] Where: ζ can be defined as a certain spectral morphology matching degree (such as generalized correlation or adaptive kernel contrast). The higher the value, the closer the track is to the known background noise pattern, and the higher the confidence level of judging it as a noise track.

[0099] Matching degree For orbits exceeding the threshold θ, a set of nonlinear noise reduction filtering operators γ are applied. m For example, multi-scale wavelet thresholding or sparse mapping filters can reduce the amplitude of noise components;

[0100] If some bird calls are mixed with noise in the same track, the intensity of this filtering operator can be appropriately reduced to prevent excessive attenuation of bird call details. The parameter controlling this intensity can be defined as γ.m ∈[0,1], determined by the background similarity and the prior bird call energy ratio;

[0101] After noise reduction processing, the main noise in each track is suppressed, and the bird call signal is more prominent, forming a near-clean track set. This result will be used for multi-label bird species identification in step three. During the output process, the time-frequency distribution corresponding to each track must be preserved. And noise confidence labels, so that step three can selectively further enhance the target frequency band or fuse other modal information.

[0102] In practice, the mixed signal is first decomposed through blind source separation, and then adaptive noise reduction is performed based on the background template matching degree, forming a two-level suppression mode for complex wetland noise. This is achieved through an adjustable nonlinear noise reduction filtering operator γ. m To avoid damaging weak bird calls when excessively suppressing noise, it is important to maintain the sensitivity of subsequent recognition and improve the matching accuracy of noise. Explicit retention facilitates reference during label recognition, improving the accuracy of classification networks in noisy residual environments.

[0103] Step 3, when the orbit is nearly pure When generating and determining the bird species included in each track, the data is input into a hybrid multi-label classification network Φ. class We combine CNN and Transformer structures for deep representation, use Renyi entropy to improve loss and enhance multi-species discrimination, and then use enhancement factors to dynamically weight the dawn or main frequency band to highlight weak bird calls. This produces an information set containing multi-species labels and confidence scores, and labels the corresponding time segments.

[0104] Step three includes the following:

[0105] Step 301: Multi-label bird species identification and advanced interactive classification model

[0106] With a near-pure orbit Using the input, construct a hybrid multi-label classification network Φ class The network can combine CNN to extract local time-frequency features through convolution, and then capture long-range dependencies through Transformer or other attention mechanisms; for each quasi-pure orbit... It can be segmented into several short time windows and time-frequency features (such as logarithmic amplitude spectrum or Mel-frequency cepstral spectrum) can be extracted, and then applied to a hybrid multi-label classification network Φ. class Batch processing is performed to obtain the corresponding predicted species distribution P. m =(p m,1 ,p m,2 ,…,p m,S ), where pm,s This represents an estimate of the probability that the orbit contains species s;

[0107] Considering that multiple bird species may appear on the same track in wetlands simultaneously, to avoid the mutual exclusion constraints caused by simple independent binary cross-entropy, a multi-label classification loss based on the Renyi entropy concept is constructed.

[0108]

[0109] In the formula: Y m =(y m,1 ,…,y m, s) is the multi-label vector corresponding to orbit m during the training or labeling phase, y m,s ∈

[0110] {0,1} indicates whether species s actually exists;

[0111] Φ R (·,α) is a single-class penalty function designed using the Rényi entropy concept, which can be defined as:

[0112]

[0113] Where α>1 is the adjustment index, which is used to impose different degrees of penalty on the predicted probability in different confidence intervals;

[0114] κ is the global scaling factor, which takes a value greater than 0 and is used to balance the overall loss magnitude.

[0115] When the probability estimate p m,s Deviation from true label y m,s When, the single-class penalty function Φ R (p m,s ,α) will produce a non-linearly increasing penalty, thereby improving the ability to distinguish easily confused species in multi-species scenarios;

[0116] During the reasoning phase, multiple single tracks appearing within the same time period are considered. Based on a hybrid multi-label classification network Φ class Output probability estimate p m,s Furthermore, by incorporating time overlap information between orbits, species predictions can be further merged or revised.

[0117] Specifically, if multiple tracks give a high confidence level to species s within the same time range, the overall probability of species s' existence during that time period can be increased. For weak bird song tracks, if their recognition probability is relatively low but they have a clear temporal correlation with other tracks, they can also be moderately increased through probability compensation strategies to avoid being ignored due to weak signals.

[0118] When using it, multi-label loss is employed. Even when multiple bird calls exist on a single track, the system can still correctly identify each species, rather than forcing each track to have only one label. The CNN combined with the Transformer structure can extract subtle bird call details and capture the continuous or overlapping phenomena of multiple species calls in the time domain, which greatly improves the recognition accuracy. Through the track merging strategy, low-energy or temporally similar bird calls can be cross-verified, thereby reducing the false negative rate.

[0119] Step 302: Dynamic Spectrum Enhancement and Target Bird Call Highlighting

[0120] In the process of identifying multiple bird calls in wetlands, some residual noise or cross-track interference may still interfere with the classification network. To address this, a dynamic spectrum enhancement function Ω(t,f) is constructed. Its main principle is to align the frequency band activity and time series peaks (such as dawn) of typical wetland birds with clean orbits. Perform targeted weighting:

[0121]

[0122] In the formula: The time-frequency distribution after dynamic weighting is used for subsequent final output or to assist in duplicate detection;

[0123] The dynamic spectral enhancement function Ω(t,f) can be obtained by multiplying two parts:

[0124] Ω(t,f)=Γ time (t)×Γf req (f)

[0125] By utilizing the peak characteristics of wetland bird calls during dawn chorus and dusk calls, a time-series-based weighting function can be set up, such as constructing a time enhancement factor Γ. time :

[0126] Γ time (t)=1+ρ dawn exp{-ω dawn (t-t0) 2}

[0127] In the formula: t0 represents the dawn time or other peak center, ω dawn >0 is a parameter that controls the width of the time distribution, ρ dawn >0 represents the peak amplitude coefficient; the larger the value, the stronger the enhancement during the dawn period.

[0128] For certain dominant frequency bands where common birds congregate (such as the 2-8 kHz range), bandpass enhancement in the following form can be introduced, such as a frequency enhancement factor Γ. freq :

[0129] Γfreq (f)=1+ρ band Θ(f;f1,f2)

[0130] In the formula: Θ(f; f1, f2) represents a function that increases or maintains a high value within the interval [f1, f2], and can take piecewise linear or other smooth transition forms to highlight the active frequency range of birdsong; ρ band To enhance the amplitude coefficient, the value should be greater than or equal to 0, and the adjustment should be made according to the dominant frequency range of the target bird species.

[0131] Time-frequency distribution after weighting by dynamic spectrum enhancement function Ω(t,f) This step is mainly used for:

[0132] The data is then fed back into a hybrid multi-label classification network. class A second assessment is conducted to improve the final recognition accuracy in cases of weak bird calls or residual noise.

[0133] Output time segment marking: High-energy or high-confidence segments in time are marked for use in the next step of multimodal fusion and continuous feedback optimization, so as to be jointly analyzed with spatial positioning or meteorological data;

[0134] When used, it makes full use of the biological patterns of wetland birdsong being active in specific time periods or frequency bands, effectively improving the detectability of target calls in noisy environments. Combined with interactive multi-label judgment, it can highlight weak bird calls that may be overlooked, reducing the missed detection of rare or faint species. The enhanced time-frequency segments output can be matched with the spatial information or meteorological data of the sensor array in the fourth step, further improving the stability and dynamic adaptability of the identification.

[0135] Through the advanced multi-label recognition strategy in step 301 and the dynamic time-frequency enhancement in step 302, accurate species detection and effective feature highlighting of separated tracks in wetland multi-bird chorus environment are achieved. This provides high-confidence multi-species identification results and enhanced audio data for spatiotemporal correlation analysis for the final multimodal fusion, thereby promoting the overall performance improvement of wetland bird monitoring system in complex and multi-species overlapping scenarios.

[0136] Step 4: When the recognition results trigger multimodal fusion based on environmental monitoring needs, acquire multi-channel audio data M. q Using the spatial positioning vector G(τ) and external environmental parameters E(τ), the system identifies dominant noise sources such as strong winds or water flow through a spatial positioning function and an adaptive update strategy. Upon detecting false positives, missed detections, or the emergence of new species, the system feeds back the corresponding orbital data and labeled logs to previous steps to supplement the noise template Φ. b (t,f) and bird song template Ψ s (t,f);

[0137] Step four includes the following:

[0138] Step 401: Multimodal data acquisition and synchronization alignment

[0139] When monitoring at fixed points in wetlands or during drone patrols, microphone arrays are deployed at multiple spatial locations to generate multi-channel audio data {M1, M2, ..., M...}. Q}, where Q represents the number of microphone channels;

[0140] Each channel recording needs to be synchronized with the single track identified in step three on the timeline. Ensure proper alignment of timestamps t (i.e., use the same reference time base) to guarantee that subsequent sound localization or coherence analysis can be performed based on spatial differences;

[0141] Synchronization identifier: Define a global time reference τ so that the multi-channel audio data M q (t) and quasi-pure orbit Corresponding to step size or sampling rate;

[0142] Collect non-audio modal data that are strongly correlated with wetland scenarios, such as wind speed v wind Water flow rate u flow Temperature T air And so on, and uniformly map them to the same time index τ, and denote these data as a vector form of the external environment parameter E(τ):

[0143] E(τ)=(v wind (τ),u flow (τ),T air (τ),...)

[0144] GPS / BeiDou positioning is performed on drones or fixed points to generate a spatial positioning vector G(τ)=(lat(τ),lng(τ),alt(τ));

[0145] By combining the relative positions of the microphone array in three-dimensional space, the location of multiple sound sources can be determined and the bird calls from near and far or from different habitats can be better distinguished.

[0146] When in use, by using multi-channel microphones and external environmental parameters, the correlation between bird sound propagation and external factors such as meteorology and hydrology can be further revealed, breaking the limitations of single sound source processing and forming a unified time axis in the entire system, providing a basis for subsequent analysis (such as noise component determination and spatial positioning). Since the wetland environment varies greatly with meteorology and hydrology, collecting and aligning this information can significantly enhance the system's stability against sudden situations or seasonal changes.

[0147] Step 402: Multimodal fusion and dynamic parameter adjustment

[0148] Based on the acquired multi-channel audio data {M1,M2,…,M… Q The spatial location vector G(τ) can be used to estimate the spatial location of the target bird's call using techniques such as time of arrival or coherent processing.

[0149] Define positioning function The estimated coordinates R of the output sound source in three-dimensional space m (τ) is used to distinguish bird calls from different directions. This information will be associated with the species tags and time segment markers produced in step three to determine which bird species is calling from which location and when. The recorded external environmental parameter E(τ) can help identify certain strong wind or water flow noise scenarios. For example, when the wind speed v wind (τ) is greater than the preset wind speed threshold θ wind If the estimated location of the sound source is concentrated above or in an open area, it can be determined that the noise is likely dominated by wind. If a new noise source is confirmed to be dominant, the blind source separation model Φ in step two (blind source separation and noise reduction) can be modified. sep The corresponding noise template Φ b (t,f) performs incremental updates, or in the noise identification operator Λ noise Adjust the matching parameters to better suppress similar noise;

[0150] Define adaptive adjustment function When an abnormal noise distribution or new interference is detected, the contribution coefficient α is locally corrected based on a known background template or newly collected noise samples. b and β s (See the weighting coefficients in steps one and two);

[0151] Where the positioning function The approach is as follows:

[0152] Acquire multi-channel audio data M q With orbit Y′ m Synchronization time interval; Extract TDOA: For the time window corresponding to τ, calculate the time delay difference Δt using cross-correlation methods such as GCC-PHAT. ij Coordinate calculation: based on the microphone's relative position (x...) i ,y i ,z i ), speed of sound C and Δt ij The sound source location is solved using geometric or least squares methods; matching determination: if the solved sound source coordinates match the coordinates within the time window... If the features are highly consistent, then the estimated coordinates R will be... m (τ) Record it, otherwise mark it as a mismatch or noise source;

[0153] Output estimated coordinates R m(τ), and the estimated coordinates R available in subsequent steps. m (τ) Assists in determining noise distribution, bird flock location, or performing closed-loop updates.

[0154] Adaptive adjustment function The approach is as follows:

[0155] Obtain external environment parameters E(τ) and estimate coordinates R m (τ), Quasi-pure orbit After detecting an anomaly or new situation in step four, the system calls Θ. update ;

[0156] Noise type identification: Based on meteorological and hydrological data and location information, it was initially determined to be strong wind, water flow noise, or other interference;

[0157] Comparison with background / birdsong template: Residual noise characteristics of this track with acoustic dictionary D wetland Perform a similarity assessment. If the matching degree is low and there is sufficient confidence that this is a new scene or a new species, then proceed to the update branch.

[0158] Generate update instructions: Adjust key parameters of blind source separation or multi-label model and expand noise template Φ according to the judgment results. b (t,f) or bird song template Ψ s Templates in (t,f);

[0159] Record and execute: Based on this instruction, the model is automatically updated and the dataset is expanded in the next training cycle or during offline retraining to complete the adaptive closed loop.

[0160] If, in step 401, a rare bird species is observed at a specific spatiotemporal location under extreme environmental parameter E(τ), then step three can be notified via a rare event pipeline to re-test the quasi-pure orbit. Dynamic spectrum enhancement and re-judgment are performed to improve the accuracy of rare species identification. This is achieved by adding emphasis markers (such as the appearance of rare bird species) to the multi-species label output, along with the estimated coordinates R. m (τ) is exported to facilitate subsequent protection or monitoring actions;

[0161] When in use, by combining spatial positioning and environmental parameters, it can more quickly and accurately distinguish between the source of natural noise and the source of target sounds, reduce the problem of over-adaptation of blind source separation to rare noise scenarios, and promptly expand templates or parameters for new noise types under extreme conditions such as strong winds and floods, maintain the long-term stability of the system, increase attention to weak or rare birds in key time and space areas, and significantly improve the value and sensitivity of ecological monitoring.

[0162] By combining noise source localization and environmental parameter-guided noise dominance factor inference, the template weights for blind source separation are dynamically updated, forming an active learning-based noise suppression process. Based on environmental data and spatiotemporal distribution information, it triggers further enhancement or classification re-judgment, providing more focused detection capabilities for rare bird species.

[0163] Step 403: Continuous Feedback Loop and Step-by-Step Iterative Optimization

[0164] In step three, during the secondary localization process, if false detections (misclassifying non-target sounds as birdsong) or missed detections (birdsong is detected but not classified), the relevant audio tracks, their time periods, and species labels must be recorded. If, during manual review, a discrepancy is found between the actual bird species and the identification label, this difference should also be noted and written into the error log H. error ;

[0165] For new or previously unrecorded noise types, also write them to the noise log H. noise So that it can be updated to the acoustic dictionary D later. wetland and blind source separation model Φ sep ;

[0166] If noise log H noise If new noise samples or new noise combinations are recorded, then in step one, a background template Φ can be used. b (t,f) can be expanded or the training dataset R can be modified. simu Make incremental additions;

[0167] If error log H error This indicates that some species have a high confusion rate. Therefore, in the joint training of step two and step three (or subsequent offline retraining), the pure bird calls of that species are increased. sing Or build more hybrid scenarios α b and β s Combining these technologies improves the ability to identify related confused species;

[0168] Based on the feedback and updates generated in the first two steps, the monitoring strategy will be adjusted in subsequent deployments:

[0169] Increase sampling frequency for periods or locations where missed detections frequently occur, deploy denser microphone arrays in key wetland areas, and if it is confirmed that noise is extremely severe during certain periods or that drone recording fails under specific weather conditions, then the workflow can flexibly skip those periods or adopt alternative data sources to improve overall efficiency and accuracy.

[0170] In use, by accumulating error and noise logs, it can be continuously updated across monitoring cycles of several months or years, rapidly adapting to new environments and bird species, and gradually reducing false positives and false negatives in the iterative process, providing more accurate long-term data for subsequent wetland ecological decisions; for multi-channel audio data Mq Acquisition and synchronization with environmental parameter E(τ) to align with a clean orbit In addition to the species identification period, step 402 makes up for the limitations of relying solely on single-channel audio analysis. By utilizing spatial positioning and external environmental information, the noise identification and rare species detection of blind source separation are further optimized, and the sound and environmental information are integrated more deeply.

[0171] Please see Figure 2 This invention provides an automatic acoustic identification system for birds in wetland environments, comprising:

[0172] The audio data processing module, when it detects that initial processing of wetland audio is required, acquires background audio data from multiple time periods. bg With pure birdsong R sing And construct an acoustic dictionary D wetland Combined with the controllable contribution coefficient α b and β s Noise and birdsong are simulated and preliminarily segmented and anomaly removed to reduce the risk of data contamination. The acquisition time and scene description are recorded synchronously in the metadata, and then the labeled and effective target audio recordings are output. target With abnormal or unavailable fragments R drop ;

[0173] The separation module, if it is necessary to decompose the effective target recording from multiple sound sources, R target Call the blind source separation model Φ sep Iterative training is performed using a loss function, along with matching degree. The mixed signal track was broken down and denoised one by one to preserve acoustic details for weak bird calls, improve the signal-to-noise ratio to avoid confusion caused by multiple sound source overlap, and finally generate a near-clean track.

[0174] The marking module, when the orbit is quasi-pure. When generating and determining the bird species included in each track, the data is input into a hybrid multi-label classification network Φ. class We combine CNN and Transformer structures for deep representation, use Renyi entropy to improve loss and enhance multi-species discrimination, and then use enhancement factors to dynamically weight the dawn or main frequency band to highlight weak bird calls. This produces an information set containing multi-species labels and confidence scores, and labels the corresponding time segments.

[0175] The adaptive update module acquires multi-channel audio data M when the recognition results and environmental monitoring requirements trigger multi-modal fusion. qUsing the spatial positioning vector G(τ) and external environmental parameters E(τ), the system identifies dominant noise sources such as strong winds or water flow through a spatial positioning function and an adaptive update strategy. Upon detecting false positives, missed detections, or the emergence of new species, the system feeds back the corresponding orbital data and labeled logs to previous steps to supplement the noise template Φ. b (t,f) and bird song template Ψ s (t,f).

[0176] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0178] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0179] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0180] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An automatic acoustic identification method for birds in wetland environments, characterized by: include, By collecting background sounds and pure birdsong from multiple time periods and constructing an acoustic dictionary, and combining controllable contribution coefficients, noise and birdsong are simulated and preliminary segmentation and anomaly removal are completed. The labeled valid target recordings and abnormal or unusable segments are output. If it is necessary to extract effective target recordings from multiple sound sources, a loss function is used to iteratively train the blind source separation model, and the matching degree is used to decompose and denoise the mixed signal track to generate a quasi-clean track; among these, an acoustic dictionary is used. Medium background template With separation results The similarity is determined using the following matching degree. : ;in: For spectral morphology matching degree; The quasi-pure orbit is input into a hybrid multi-label classification network, which combines CNN and Transformer structures for deep representation. Renyi entropy and enhancement factors are used to strengthen the weak bird song features, producing an information set containing multi-species labels and confidence scores. When the identification results and environmental monitoring needs trigger multimodal fusion, the dominant noise source is identified through spatial localization function and adaptive update strategy. When false positives or false negatives or new species are detected, the corresponding orbital data and annotation logs are sent back to the previous steps to supplement the noise template and bird song template.

2. The method for automatic acoustic identification of birds in wetland environments according to claim 1, characterized in that: Collect diverse raw wetland audio datasets, background sounds, and pure birdsong at different times; The bird song segments appearing in the audio recording dataset are labeled with species, and their corresponding background noise types are recorded to form a preliminary labeling table. Audio segments with quality defects are removed or separately marked as fault records.

3. The automatic acoustic identification method for birds in wetland environments according to claim 2, characterized in that: An acoustic dictionary was constructed based on the obtained multi-source sound data, which included several background noise templates and single-species bird song templates. A combined model of the entire wetland environment was also constructed. By combining different contribution coefficients, various complex scenes were synthesized to obtain simulated synthesized soundscapes. The original recordings were combined with a preliminary labeling table for segmentation and anomaly removal. Bird song segments and background segments were divided by temporal segmentation and extraction functions.

4. The automatic acoustic identification method for birds in wetland environments according to claim 1, characterized in that: The training sample set is constructed from real recordings or from an acoustic dictionary with random weight combinations. A self-supervised approach is introduced to define a comprehensive loss function. The comprehensive loss function and the training sample set are used to iteratively train the constructed deep model. The trained separation network is then used for inference. The input is actual wetland recordings, and the separation network outputs preliminary separation tracks.

5. The automatic acoustic identification method for birds in wetland environments according to claim 4, characterized in that: A matching degree is constructed using background templates from the acoustic dictionary and the separation results; For orbitals with a matching degree exceeding the matching threshold, a nonlinear noise reduction filtering operator is applied to perform noise reduction processing, forming a quasi-clean orbital set, while retaining the time-frequency distribution and noise confidence label corresponding to each orbital.

6. The automatic acoustic identification method for birds in wetland environments according to claim 5, characterized in that: For each quasi-pure orbit, it is divided into several short time windows and time-frequency features are extracted. These features are then processed in batches in the constructed hybrid multi-label classification network to obtain the corresponding species prediction distribution. Construct a multi-label classification loss and a single-class penalty function based on Renyi entropy; When the probability estimate deviates from the true label, the single-class penalty function generates a non-linearly increasing penalty, which improves the discrimination of easily confused species in multi-species scenarios; if multiple single tracks appear in the same time period, the probability estimate is output according to the hybrid multi-label classification network.

7. The automatic acoustic identification method for birds in wetland environments according to claim 6, characterized in that: A dynamic spectrum enhancement function is constructed to obtain the time-frequency distribution after targeted weighting of the quasi-pure orbit. Then, time enhancement factors and frequency enhancement factors are constructed. The time-frequency distribution is used to feed the hybrid multi-label classification network for re-judgment and output time segment labels.

8. The automatic acoustic identification method for birds in wetland environments according to claim 7, characterized in that: When monitoring fixed points in wetlands or when drones are patrolling, microphone arrays are arranged in multiple spatial locations to form multi-channel audio data, and the multi-channel audio data corresponds to the quasi-clean orbit in terms of step size or sampling rate. After collecting external environmental parameters that are strongly correlated with the wetland scene, the drone or fixed point is located to generate a spatial positioning vector.

9. The automatic acoustic identification method for birds in wetland environments according to claim 8, characterized in that: Based on the acquired multi-channel audio data and spatial positioning vectors, the constructed positioning function outputs the estimated coordinates of the sound source in three-dimensional space to distinguish birdsong from different directions. Combined with species tags and time segment markers, the system analyzes and obtains information on which birds are singing, where, and when. When abnormal noise distribution or new interference is detected, an adaptive adjustment function is defined to locally correct the contribution coefficient based on the known background template or newly collected noise samples.

10. The method for automatic acoustic identification of birds in wetland environments according to claim 9, characterized in that: If a rare bird species is observed at a specific spatiotemporal location under extreme environmental parameter conditions, dynamic spectrum enhancement and re-judgment are performed again on the quasi-pure orbit, thereby adding emphasis markers to the multi-species label output; If false detections or missed detections occur, record the relevant audio tracks, their time periods, and species tag information; If the actual bird species does not match the identification tag, the difference in production is recorded in the error log; if it corresponds to a new or unrecorded noise type, it is recorded in the noise log.

11. The method for automatic acoustic identification of birds in wetland environments according to claim 10, characterized in that: If new noise samples or new noise combinations are recorded in the noise log, the background template is expanded or the training dataset is incrementally supplemented. If the error log indicates that certain species have a high confusion rate, increase the pure bird calls of that species or build more mixed scene combinations.

12. An automatic acoustic identification system for birds in wetland environments, employing the automatic identification method according to any one of claims 1 to 11, characterized in that: include, The sound data processing module collects background sounds and pure birdsong from multiple time periods and constructs an acoustic dictionary. It then uses a controllable contribution coefficient to simulate noise and birdsong, completes preliminary segmentation and anomaly removal, and outputs labeled valid target recordings and abnormal or unusable segments. The separation module, if it is necessary to decompose the effective target recording from multiple sound sources, uses a loss function to iteratively train the blind source separation model, and uses the matching degree to decompose and denoise the mixed signal track to generate a quasi-clean track; The labeling module inputs the quasi-pure orbital into a hybrid multi-label classification network, combines CNN and Transformer structures for deep representation, uses Renyi entropy and enhancement factors to strengthen weak bird song features, and produces an information set containing multi-species labels and confidence scores. The adaptive update module identifies the dominant noise source through spatial localization function and adaptive update strategy when the identification results and environmental monitoring needs trigger multimodal fusion. When false positives or false negatives or new species are detected, the corresponding orbital data and annotation logs are sent back to the previous steps to supplement the noise template and bird song template.

Citation Information

Patent Citations

  • Aliasing twitter separation method based on deep learning

    CN117789746A

  • Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device

    CN119027775A