Acoustic automatic recognition system and method for birds in wetland environment
By proposing acoustic automatic identification methods of multi-step processes in wetland environments, including multi-source noise modeling, blind source separation, multi-label recognition and multi-modal fusion, the problem of low recognition accuracy in noisy wetland environments is solved, and efficient automatic identification of birds is achieved.
Patent Information
- Application Number
- CN202510471442.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art is difficult to achieve high-accurate automatic bird acoustic recognition in wetland multi-bird chorus and noisy backgrounds, especially when multiple sound sources overlap and background noise are complex.
A multi-step process method is proposed, including multi-source noise modeling, blind source separation and noise reduction, multi-label recognition and multi-modal fusion feedback, through these steps, maintaining high recognition accuracy in noisy environments.
It achieves high recognition accuracy in noisy and multi-species chorus scenes, provides efficient support for decisions on ecological monitoring and protection of wetlands, and reduces the burden of manual review.
Smart Images

Figure CN120220702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ecological monitoring and acoustic signal recognition, and specifically to an automatic bird acoustic recognition system and method in wetland environments. Background Art
[0002] In wetland ecological environments, birds often emit various types of vocalizations at different times due to survival and reproduction needs. These acoustic signals not only reflect species characteristics but also environmental changes and ecological health status. Wetland environments usually contain rich and variable background noise sources, such as wind sounds, water flow sounds, frog calls, and insect chirps, whose sound intensity and frequency band distribution change dynamically with factors such as weather, season, and terrain. In addition, multiple bird species may inhabit the wetland simultaneously, and the common "chorus" phenomenon during dawn makes the frequency spectra of multiple bird calls highly superimposed. Faced with such a complex acoustic environment with multiple sound sources coexisting, manual segment-by-segment annotation and filtering are time-consuming and laborious, and it is difficult to meet the needs of large-scale long-term monitoring. Therefore, to achieve long-term, automated acquisition and analysis of wetland bird diversity and behavior patterns, it is necessary to construct a highly robust acoustic processing system that adapts to multiple scenarios and multiple noise types in order to timely obtain bird population information and support ecological protection decision-making.
[0003] However, existing technologies often experience a significant decrease in recognition accuracy or are unable to effectively separate multiple bird call signals when faced with multiple bird choruses and noisy backgrounds in wetlands. Since most traditional algorithms or models default the input to be a single clear species sound, it is difficult to perform precise unmixing and classification when multiple sound sources overlap; in addition, the amplitude and frequency range of various types of noise in wetlands vary greatly, and wind sounds and water flow sounds have random and time-varying characteristics, and backgrounds such as frog calls and cicada chirps may also overlap with target bird calls in frequency bands or time periods, resulting in less than ideal results for simple filtering or fixed bandwidth detection.
[0004] Once a new species call that is not included or annotated appears, traditional models are more likely to be confused or missed; and in extreme environments or fault conditions, blind source separation and recognition will also fail, requiring professionals to repeatedly listen to and manually screen a large number of audio segments, directly reducing the efficiency and accuracy of automatic monitoring, and also restricting the in-depth analysis and scientific application of subsequent bird ecological data.
[0005] Therefore, the present invention provides an automatic bird acoustic recognition system and method in wetland environments. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] In view of the deficiencies of the prior art, the present invention provides an automatic bird acoustic recognition system and method in wetland environments. By proposing a four-step process of multi-source noise modeling, blind source separation and noise reduction, multi-label recognition, and multi-modal fusion feedback, it maintains high recognition accuracy in noisy and multi-species chorus scenarios, providing efficient support for wetland ecological monitoring and protection decision-making; solving the technical problems recorded in the background art.
[0008] (II) Technical Solution
[0009] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0010] An automatic bird acoustic recognition method in wetland environments, including,
[0011] When it is detected that initial processing of wetland audio is required, background sounds R bg and pure bird songs R sing are collected over multiple time periods and an acoustic dictionary D wetland is constructed. Combining the controllable contribution coefficients α b and β s to simulate noise and bird songs and complete preliminary segmentation and anomaly elimination to reduce the risk of data contamination, and synchronously record the collection time and scene description in the metadata, and then output the labeled effective target recording R target and abnormal or unavailable segments R drop ;
[0012] If multi-source disassembly of the effective target recording R target is required, the blind source separation model Φ sep is called to perform iterative training using the loss function, and the matching degree is used to disassemble and denoise each mixed signal track one by one, so as to retain acoustic details for weak bird songs, improve the signal-to-noise ratio to avoid confusion caused by multi-source overlap, and finally generate a quasi-pure track
[0013] When the quasi-pure track is generated and it is necessary to determine the bird species contained in each track, it is input into the hybrid multi-label classification network Φ class to perform deep characterization by combining the CNN and Transformer structures, and use the Renyi entropy to improve the loss to strengthen the multi-species discrimination. Then, dynamic energy weighting is performed on the dawn or main frequency band through the enhancement factor to highlight the characteristics of weak bird songs, outputting an information set containing multi-species labels and confidence levels, and marking the corresponding time segments;
[0014] When the recognition result and the environmental monitoring requirements trigger multi-modal fusion, multi-channel audio data M q(t) and external environmental parameters E(τ) as well as spatial positioning vector G(τ), identify noise-dominated sources such as strong winds or water currents through the spatial positioning function and the adaptive update strategy, and when false detections or missed detections or the emergence of new species are detected, transmit the corresponding track data and annotation logs back to the previous steps to supplement the noise template Φ b (t,f) and bird song template Ψ s (t,f).
[0015] Preferably, collect diverse original wetland recording data sets R Taw 、background sound R bg 、pure bird song R sing ;
[0016] Perform species label annotation on the key bird song segments that appear in the recording data set R raw and record them together with the background noise types in the corresponding time periods to form a preliminary marking table L stagel ; Eliminate audio segments with quality defects or separately stockpile and identify them as fault records R fault ,
[0017] Preferably, construct an acoustic dictionary D wetland , which contains several background noise templates Φ b (t,f) (b = 1,..., B) and single-species bird song templates Ψ s (t,f), s = 1,..., S, and construct the entire wetland environment combination model W(t,f);
[0018] By combining different contribution coefficients α b and contribution coefficient β s synthesize various complex scenarios to obtain a simulated synthetic soundscape R simu , combine the original recording R raw with the preliminary marking table L stagel for more detailed segmentation and abnormal elimination;
[0019] For audio segments determined to be invalid or with particularly severe interference, filter them according to the preset filtering rules to form abnormal or unavailable segments R drop and label their types;
[0020] For paragraphs that can be partially utilized, divide them into bird song segments R segment containing the main bird songs and background segments R target containing only the background through the time series segmentation and extraction function F bgOnly ;
[0021] Preferably, construct a training sample set R mix , where the training sample set R mix can be directly from real recordings or from the acoustic dictionary Dwetland Perform random weight combination;
[0022] Let the deep model be Φ sep , whose input is the single-channel mixed signal x(t) or its corresponding time-frequency representation X(t,f), and the output is M separated tracks Introduce the self-supervised idea to define a comprehensive loss function Use the above loss function and the training sample set R mix Perform iterative training on the deep model Φ sep ;
[0023] Preferably, perform inference on the trained separation network Φ sep , with the input being the actual wetland recording, and the separation network Φ sep outputs M preliminary separated tracks
[0024] With the help of the acoustic dictionary D wetland the background template Φ b (t,f) and the separation result of the constructed matching degree
[0025] For the matching degree apply the non-linear noise reduction filtering operator γ to the tracks whose matching degree exceeds the threshold θ m perform noise reduction processing to form a quasi-pure track set Retain the time-frequency distribution corresponding to each track and the noise confidence label;
[0026] Preferably, use the quasi-pure track as the input to construct a hybrid multi-label classification network Φ class ;
[0027] For each quasi-pure track divide it into several short-time windows and extract time-frequency features, and then perform batch processing in the hybrid multi-label classification network Φ class to obtain the corresponding species prediction distribution P m =(p m,1 ,p m,2 ,…,p m,s ), where p m,s represents the probability estimate that the track contains the species s;
[0028] Construct a multi-label classification loss improved based on the concept of Renyi entropy and design a one-class penalty function Φ R (p m,s ,α);
[0029] When the probability estimate pm,s Deviation from the true label y m,s When this occurs, the single-class penalty function Φ R (p m,s , α) generates a non-linearly increasing penalty, which improves the discrimination of easily confused species in a multi-species scenario; if multiple single orbits appear within the same time period Then, according to the hybrid multi-label classification network Φ class Output the probability estimate p m,s
[0030] Preferably, construct a dynamic spectrum enhancement function Ω(t, f), perform targeted weighting on the pure orbit to obtain the time-frequency distribution and then construct a time enhancement factor Γ time and a frequency enhancement factor Γ freq ; The time-frequency distribution is used to be sent into the hybrid multi-label classification network φ again class for rejudgment and output the time segment label;
[0031] Preferably, when at a fixed wetland monitoring point or during drone cruising, arrange microphone arrays at multiple spatial positions to form multi-channel audio data {M1, M2, …, M Q}; And make the multi-channel audio data M q (t) correspond to the quasi-pure orbit in terms of step size or sampling rate;
[0032] Collect external environmental parameters E(τ) that are strongly related to the wetland scenario; perform positioning on the drone or fixed point using GPS / Beidou, etc., to generate a spatial positioning vector G(τ);
[0033] Preferably, based on the obtained multi-channel audio data {M1, M2, …, M Q} and the spatial positioning vector G(τ), techniques such as time of arrival or coherent processing can be used to estimate the azimuth of the target bird call in space;
[0034] Construct a positioning function Output the estimated coordinates R m (τ) of the sound source in three-dimensional space to distinguish bird calls from different azimuths, and combine the species label and time segment label to analyze what kind of bird makes a sound where and when;
[0035] Define an adaptive adjustment function When abnormal noise distribution or new interference is detected, locally correct the contribution coefficients α b and β s etc.;
[0036] If a rare bird species is observed at a specific spatio-temporal location under extreme environmental parameters E(τ), then the quasi-pure orbit is re-evaluated Perform dynamic spectrum enhancement and re-verification, thereby adding key markings to the multi-species label output and deriving it along with the estimated coordinates R m (τ);
[0037] Preferably, if misdetection or missed detection occurs, record the relevant audio track, its time period, and species label information. If the actual bird species does not match the identification label, write the produced difference annotation into the error log H error ;
[0038] For new or unrecorded noise types, write them into the noise log H noise ;
[0039] If the noise log H noise records new noise samples or new noise combination methods, then expand the background template Φ b (t,f) or perform incremental supplementation on the training dataset R simu ;
[0040] If the error log H error indicates that the recognition confusion rate of certain species is relatively high, improve the recognition performance of related confused species by increasing the pure bird calls R sing of this species or constructing more mixed scenarios α b and β s combinations;
[0041] Based on the feedback and updates formed in the first two steps, adjust the monitoring strategy in subsequent deployments:
[0042] Increase the sampling frequency for time periods or locations where missed detection often occurs, deploy a denser microphone array in key wetland areas. If it is confirmed that the noise is extremely severe during certain time periods or the recording of the drone fails under specific meteorological conditions, then the corresponding time period can be skipped flexibly in the workflow or alternative data sources can be adopted to improve the overall efficiency and accuracy.
[0043] An automatic bird acoustic recognition system in a wetland environment includes
[0044] A sound data processing module. When it is detected that the wetland audio needs to be initially processed, collect the background sound R bg and pure bird calls R sing in multiple time periods and construct an acoustic dictionary D wetland , combine the controllable contribution coefficients α b and β s to simulate noise and bird calls and complete preliminary segmentation and anomaly elimination to reduce the risk of data contamination, and synchronously record the acquisition time and scene description in the metadata, and then output the labeled effective target recording R targetWith abnormal or unavailable segment R drop ;
[0045] Separation module, if multi-source disassembly of the effective target recording R is required target , call the blind source separation model Φ sep Through iterative training using the loss function and the matching degree Disassemble and denoise each mixed signal track one by one, so as to retain the acoustic details for weak bird calls, improve the signal-to-noise ratio to avoid confusion caused by multi-source overlap, and finally generate a quasi-pure track
[0046] Labeling module, when the quasi-pure track is generated and the bird species contained in each track needs to be determined, input it into the hybrid multi-label classification network Φ class Combine the CNN and Transformer structures for deep representation, use the Renyi entropy to improve the loss to enhance the multi-species discrimination, and then perform dynamic energy weighting on the dawn or main frequency band through the enhancement factor to highlight the characteristics of weak bird calls, output an information set containing multi-species labels and confidence levels, and mark the corresponding time segments;
[0047] Adaptive update module, when the recognition result triggers multi-modal fusion with the environmental monitoring requirements, collect multi-channel audio data M q (t) and external environmental parameters E(τ) and spatial positioning vector G(τ), identify the noise-dominated sources such as strong wind or water flow through the spatial positioning function and the adaptive update strategy, and when misdetection or missed detection or the appearance of a new species is detected, transmit the corresponding track data and annotation logs back to the previous steps to supplement the noise template Φ b (t,f) and the bird call template Ψ s (t,f).
[0048] (III) Beneficial effects
[0049] The present invention provides an automatic bird acoustic recognition system and method in a wetland environment, having the following beneficial effects:
[0050] By establishing a flexibly schedulable acoustic dictionary D wetland , and using the contribution coefficients α b and β s to controllably mix the background template Φ b (t,f) and the bird call template Ψ s (t,f), effectively improving the authenticity of the simulated wetland multi-noise source and multi-bird chorus scenarios, and helping the subsequent model to remain robust under extreme interference conditions;
[0051] Utilize the self-supervised blind source separation network Φ sep and the matching degree Split real or synthetic mixed audio into multiple tracks, so that the originally superimposed bird calls and noise can be separated and finely noise-reduced. Without relying on a clean bird call baseline, by adjusting the filtering strength and similarity threshold of each track, it effectively reduces misclassification and omission, while retaining sufficient acoustic details for subsequent detection of weak or rare species;
[0052] The multi-label classification network combines the hybrid structure of CNN and Transformer, improves the accuracy of species discrimination by improving the loss function through Renyi entropy, and uses the time enhancement factor Γ time and the frequency enhancement factor Γ freq Dynamic weighting is performed at dawn or in specific frequency bands to enhance weak calls in a targeted manner. At the same time, it supports the parallel identification of multiple species in the same track, thus taking into account the phenomenon of wetland chorus. This dynamic enhancement mechanism not only makes full use of the active patterns of wetland birds in time and frequency, but also avoids missing rare or weak-sounding bird species.
[0053] Step 4 introduces multimodal fusion, using spatial positioning information and environmental parameters such as wind speed and water flow to help identify strong wind noise and other unconventional signals, and adaptively update the noise template or blind source separation model Φ sep The feedback mechanism for false detection and missed detection can also add new species or new noises to the acoustic dictionary D wetland , achieving continuous learning and maintaining high reliability. In the face of dynamic changes in the ecological environment, it can maximize the detection efficiency of multiple bird species and significantly reduce the burden of manual review, providing more efficient and reliable data support for wetland protection and biodiversity research. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic diagram of the process flow of the method for automatic acoustic identification of birds in a wetland environment of the present invention;
[0055] Figure 2 It is a schematic diagram of the structure of the automatic acoustic identification system for birds in a wetland environment of the present invention. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] See also Figure 1 The present invention provides a method for automatic acoustic identification of birds in a wetland environment, comprising:
[0058] Step 1: When it is detected that the wetland audio needs to be initially processed, collect the background sound R at multiple time periods bg and the pure bird song R sing and construct the acoustic dictionary D wetland , combined with the controllable contribution coefficients α b and β s simulate the noise and bird song and complete the preliminary segmentation and anomaly elimination to reduce the risk of data contamination, and synchronously record the collection time and scene description in the metadata, and then output the labeled effective target recording R target and the abnormal or unavailable segments R drop ;
[0059] The first step includes the following contents:
[0060] Step 101: Multi-source wetland audio collection and time-period annotation
[0061] Based on methods such as fixed monitoring points and drone cruising, collect diverse wetland raw recording data at different time periods, and name the data object as the recording dataset R Taw ; For typical background noises such as water flow, wind sound, insect chirping, and amphibian calls, record their pure background segments separately, and uniformly name these audio segments as the background sound R bg .
[0062] For common bird species, try to obtain separate bird songs with high signal-to-noise ratio and clear species, and name them as pure bird song R sing ;
[0063] Under manual or semi-automatic assistance, label the species tags for the key bird song sections in the recording dataset R raw , and record them together with the corresponding background noise types in different time periods to form a preliminary marking table L stagel ;
[0064] Considering the spatio-temporal variability of the wetland environment, it is necessary to repeat the collection at different weather and season time periods to obtain more comprehensive scene differences;
[0065] Eliminate or separately stock and mark as a fault record R the audio segments with quality defects (such as microphone failures, extremely weak recording signals, etc.) fault to ensure the robustness of subsequent model training and analysis;
[0066] When in use, by separating and collecting various pure background sounds R bg , the common noise sources in the wetland can be finely characterized, reducing the uncertainty in the subsequent processing process, and collecting pure bird song R singAnd perform species annotation to enable the subsequent construction of a reference template with biological significance, which has a higher discrimination degree for multi-bird mixed scenarios. Identify and eliminate recordings with faults or extremely poor quality separately, which can avoid contaminating the training process of the subsequent recognition model with invalid information.
[0067] Step 102, Construction of environmental acoustic dictionary and preprocessing of target audio segmentation
[0068] Construct an acoustic dictionary D based on the obtained multi-source data wetland , where: it contains several background noise templates Φ b (t,f)b(=1,…,B) and single-species bird song templates Ψ s (t,f), s = 1,…,S, and use the following formula to define the general expression of the combined model W(t,f) of the entire wetland environment:
[0069]
[0070] In the formula: Φ b (t,f), Ψ s (t,f) are the time-frequency template functions obtained by extraction from the background sound R bg and the pure bird song R sing respectively; W(t,f) is the composite environmental time-frequency distribution function, that is, the combined model of the wetland environment, which is used to simulate the multi-source sound superposition phenomenon in the wetland;
[0071] α b is the contribution coefficient of the background noise template Φ b (t,f), and the larger the value, the stronger the noise component, and the value can be taken in the interval [0,1];
[0072] β s is the contribution coefficient of a certain species of bird song template Ψ s (t,f), which is used to control the relative loudness of the species in the mixed scenario, and the value can be taken in the interval [0,1]; B, S are the upper limits of the number of background noise and bird song templates;
[0073] t, f are the time and frequency coordinates, which are used to describe the change distribution of the sound in the short-time Fourier transform or other time-frequency transform domains;
[0074] Among them, usually for the background sound R bgEach background sound in the image is transformed by short-time Fourier transform (STFT) or other time-frequency transform (such as CQT, wavelet transform, etc.) to obtain the two-dimensional distribution of the corresponding amplitude spectrum evolving over time. For STFT implementation, a window length (such as 1024 points) and a step (such as 256 points) can be set to frame the audio signal and calculate the complex spectrum. After extracting the time-frequency representation of several background sounds, an average or representative template can be formed in the following ways: directly select a relatively stable background interval without sudden interference as a template; statistically aggregate the amplitude spectra of multiple backgrounds (such as taking the average or median), and appropriately sample the time dimension to obtain a relatively stable and representative time-frequency distribution function; if different background types (wind, water flow, etc.) need to be subdivided, the corresponding noise template Φ can be generated for each type. b (t,f); bird song template Ψ s (t,f) is constructed in a similar way;
[0075] By combining different contribution coefficients α through the above formula b and contribution coefficient β s The value of can be used to synthesize a variety of complex scenes (including backgrounds of strong winds and rapid water, chorus of multiple species, etc.) in the laboratory or algorithm training stage to obtain the important simulated synthetic soundscape R used to train and verify the subsequent source separation model. simu , the original recording R raw Combined with the preliminary marking table L stagel Perform more detailed segmentation and exception elimination;
[0076] For audio segments that are determined to be invalid or have serious interference, they can be filtered according to preset filtering rules to form abnormal or unusable segments R drop The rules can be: if there is no valid signal in most of the audio period (such as too low amplitude, almost zero level or extreme overload), or if there are fault signs such as power failure or poor microphone contact in the metadata record, the audio will be listed as a faulty or invalid segment. The audio recording may contain fragmented files that are too short or not fully recorded due to a fault, which can also be merged into abnormal or unavailable segments R drop ;
[0077] For the partially usable paragraphs, the time series segmentation and extraction function F segment Divide the bird song section R containing the main bird songs target and the background segment R containing only the background bgOnly ; Time series segmentation and extraction function F segmentReceive the input audio data (such as the filtered candidate audio) and output several sub - segments, where each sub - segment is either a core segment containing the main bird song or an auxiliary control segment without bird song or only containing background noise; that is, through multi - level threshold detection, template or feature matching, and timely sequence merging and boundary correction, the classification results are output and marked;
[0078] During this process, the background template Φ b (t,f) in the wetland environment combination model W(t,f) constructed in this step can be used to match and verify the noise characteristics to better identify the sections not dominated by bird song;
[0079] The finally formed acoustic dictionary D wetland , whose core consists of each background template Φ b , each bird song template Ψ s , and the corresponding contribution coefficient sets {α b}, {β s}, and the effective target recording R target after segmentation, the optional background control recording R bgOnly , together with the abnormal or unavailable segments R drop with clear reasons, the above - mentioned data and labels will be imported into the subsequent multi - source separation and recognition process;
[0080] When in use, through the dynamic allocation of the contribution coefficients α b and β s , a variety of realistic synthetic soundscapes can be obtained, which can simulate the actual noisy and multi - bird overlapping scenes to a large extent. The accurate background template and target bird song template provide a traceable and adjustable basis for subsequent self - supervised separation and multi - label recognition, making the adaptability and generalization ability of the algorithm stronger. According to the cases of different noise intensities and mixing degrees synthesized by the dictionary model, the error - prone points of the algorithm in extreme environments can be found faster and targeted improvements can be made.
[0081] Step 2: If it is necessary to disassemble the effective target recording R target , call the blind source separation model Φ sep to perform iterative training through the use of the loss function pair, and the matching degree to disassemble and denoise each mixed signal track one by one, so as to retain the acoustic details for weak bird songs, improve the signal - to - noise ratio to avoid confusion caused by multi - source overlap, and finally generate a quasi - pure track
[0082] The above - mentioned Step 2 includes the following content:
[0083] Step 201: Self - supervised network training and construction of the sound source separation structure
[0084] According to the bird song section R target and background section RbgOnly and the simulated synthetic soundscape R simu , construct a representative training sample set R mix , where R mix can be directly from real recordings (including the mixture of multiple birds and background noise), or can be synthesized by randomly combining the weights of the acoustic dictionary D in step one wetland (using the contribution coefficients α b and β s ) to cover rich scenarios with different numbers of species and different noise intensities;
[0085] Let the deep model (separation network) be Φ sep , whose input is a single-channel mixed signal x(t) or its corresponding time-frequency representation X(t,f), and the output is M separated tracks In actual operation, M can be dynamically set according to the maximum number of sound sources expected to be extracted; the separation network Φ sep is generally composed of convolutional layers, attention mechanisms or other neural network modules. To ensure compatibility with the subsequent noise reduction process, an intermediate layer that can uniformly process multiple sound sources needs to be retained in the network structure;
[0086] Traditional separation methods often rely on a pure target reference, which is difficult to obtain a clean ground truth in a multi-bird mixed wetland environment. Therefore, a self-supervised idea is introduced: only requiring the linear or energy superposition of all output tracks to match the original mixed input, and at the same time minimizing the mutual contamination between tracks as much as possible, the following comprehensive loss function can be defined
[0087]
[0088] In the formula: represents the distribution of the m-th separated track output by the network in the time-frequency domain; R mix (t,f) represents the time-frequency domain representation of the mixed input; ∥·∥ p is the generalized Minkowski norm (p>1), aiming to perform a high-order measurement on the overall reconstruction error;
[0089] Γ(·,·) is the cross-discrimination measure function between tracks, and Renyi divergence, improved coherence distance or other high-order metrics can be used to ensure the minimum mutual contamination between different tracks;
[0090] λ is the balance coefficient, which can be adjusted in the range of 0 to 1 to control the relative weights of the reconstruction accuracy and the track discrimination;
[0091] Use the above loss function to iteratively train the deep model (separation network) Φ sep , and the training set includes real wetland mixed audio and the synthetic contribution coefficients α b , βs Simulated synthetic soundscapes R obtained for different values simu ;
[0092] Continuously monitor the separation effect through the validation set (actual wetland recordings at different times and locations can also be adopted). If it is found that the separation of specific noise types or special bird calls is insufficient, corresponding background templates Φ wetland in the acoustic dictionary D in step one can be further increased b (t, f) or bird call template Ψ s (t, f) to enrich the training data;
[0093] During use, the network training can be completed without obtaining the clean recordings of bird calls one by one, greatly reducing the data annotation cost. Through the high-order metric and track distinguishability Γ, while ensuring the reconstruction accuracy of the mixed sound, the problem of mutual contamination of bird calls in the same frequency band is effectively reduced. With the help of the controllable synthetic soundscape provided in the first step, the model can be well-informed during the training stage and improve its adaptability to extreme noise and multi-bird superposition.
[0094] Step 202, Noise reduction and quasi-pure track generation
[0095] Perform inference on the trained separation network Φ sep with the input being the actual wetland recording (R obtained after filtering target or new real-time data), and the separation network Φ sep outputs M preliminary separation tracks
[0096] To identify which tracks are the main noises, the similarity between the background template Φ wetland in the acoustic dictionary D b (t, f) and the separation result can be used, and the following matching degree
[0097]
[0098] where: ζ can be defined as a certain spectral form matching degree (such as generalized correlation or adaptive kernel contrast), and the higher the value, the closer the track is to the known background noise pattern, and the higher the confidence level of being determined as a noise track;
[0099] For the tracks with the matching degree exceeding the threshold θ, apply a set of non-linear noise reduction filtering operators γ m , such as multi-scale wavelet threshold or sparse mapping filter, so that the amplitude of the noise component is weakened;
[0100] If there are some bird calls mixed with the noise in the same track, the intensity of this filtering operator can be appropriately reduced to prevent excessive weakening of the bird call details. The parameter controlling this intensity can be defined as γm ∈[0,1], which is jointly determined by the background similarity and the prior bird song energy ratio;
[0101] After noise reduction processing, the main noise in each track is suppressed, and the bird song signal becomes more prominent, forming a quasi-pure track set This result will be called for multi-label bird species recognition in Step 3. During the output process, the time-frequency distribution corresponding to each track needs to be retained and the noise confidence label, so that in Step 3, the target frequency band can be selectively further enhanced or other modality information can be fused.
[0102] When in use, first decompose the mixed signal through blind source separation, and then perform adaptive noise reduction based on the background template matching degree to form a two-stage suppression mode for the complex wetland noise. Through the adjustable non-linear noise reduction filtering operator γ m , try to avoid damaging the weak signal bird song when over-suppressing the noise, which can maintain the sensitivity of subsequent recognition. The matching degree of noise matching is explicitly retained for reference during label recognition, improving the accuracy of the classification network in a noisy residual environment.
[0103] Step 3. When the quasi-pure track is generated and it is necessary to determine the bird species contained in each track, input it into the hybrid multi-label classification network Φ class Combine the CNN and Transformer structures for deep representation, and use the Renyi entropy to improve the loss to strengthen the multi-species discrimination. Then, dynamically weight the energy of the dawn or main frequency band through the enhancement factor to highlight the weak bird song features, output an information set containing multi-species labels and confidence levels, and mark the corresponding time segments;
[0104] The said Step 3 includes the following contents:
[0105] Step 301. Multi-label bird species recognition and advanced interactive classification model
[0106] Using the quasi-pure track as the input, construct the hybrid multi-label classification network Φ class , where this network can combine the CNN to perform convolutional extraction of local time-frequency features, and then capture long-range dependencies through the Transformer or other attention mechanisms; for each quasi-pure track it can be segmented into several short-time windows and time-frequency features (such as logarithmic amplitude spectrum or Mel cepstrum) are extracted, and then batch processed in the hybrid multi-label classification network Φ class to obtain the corresponding species prediction distribution P m =(p m,1 , p m,2 ,…, p m,S ), where pm,s represents the probability estimate that the track contains species s;
[0107] Considering that multiple birds in the wetland may appear in the same track simultaneously, to avoid the mutual exclusion constraint brought by the simple independent binary cross-entropy, a multi-label classification loss improved based on the concept of Renyi entropy is constructed
[0108]
[0109] In the formula: Y m =(y m,1 ,…,y m, s) is the multi-label vector corresponding to track m in the training or annotation stage, and y m,s ∈
[0110] {0,1} indicates whether species s truly exists;
[0111] Φ R (·,α) is a one-class penalty function designed using the Renyi entropy idea and can be defined as:
[0112]
[0113] where α>1 is the adjustment exponent used to impose different degrees of penalty on the predicted probability in different confidence intervals;
[0114] κ is the global scaling coefficient, with a value greater than 0, used to balance the overall loss amplitude;
[0115] When the probability estimate p m,s deviates from the true label y m,s , the one-class penalty function Φ R (p m,s ,α) will produce a non-linearly increasing penalty, thereby improving the discrimination of easily confused species in the multi-species scenario;
[0116] In the inference stage, for multiple single tracks that appear within the same time period The probability estimate p class can be output according to the hybrid multi-label classification network Φ m,s And combined with the time overlap information between tracks to further merge or correct species predictions.
[0117] Specifically, if multiple tracks give a high confidence level for species s within the same time range, the overall presence probability of species s in that time period can be increased; for weak bird call tracks, if its recognition probability is relatively low but there is an obvious temporal association with other tracks, it can also be moderately increased through a probability compensation strategy to avoid being ignored due to weak signals.
[0118] When used, through multi-label loss When there are multiple bird calls in a single track, each species can still be correctly identified, rather than forcing each track to have only one label. The CNN combined with the Transformer structure can not only extract tiny details of bird calls, but also capture the continuity or overlap of calls of multiple species in the time domain, greatly improving recognition accuracy. Through the inter-track merging strategy, low-energy or similar-time bird calls can be mutually verified, thereby reducing the missed detection rate.
[0119] Step 302: Dynamic spectrum enhancement and target bird song highlighting
[0120] In the process of multi-bird song recognition in wetlands, some residual noise or cross-track interference may still interfere with the classification network. For this reason, a dynamic spectrum enhancement function Ω(t,f) is constructed. Its main principle is to align the pure track according to the frequency band activity and time sequence peak (such as dawn) of typical wetland birds. Carry out targeted weighting:
[0121]
[0122] Where: It is the time-frequency distribution after dynamic weighting, which is used for subsequent final output or auxiliary repetition determination;
[0123] The dynamic spectrum enhancement function Ω(t,f) can be obtained by multiplying two parts:
[0124] Ω(t,f)=Γ time (t)×Γf req (f)
[0125] Using the peak characteristics of wetland birds at dawn chorus, dusk singing, etc., a time series-based weighting function can be set, such as constructing a time enhancement factor Γ time :
[0126] Γ time (t)=1+ρ dawn exp{-ω dawn (t-t0) 2}
[0127] Where: t0 represents the dawn time or other peak center, ω dawn >0 is the parameter that controls the width of the time distribution, ρ dawn >0 is the peak amplitude coefficient, and the larger the value, the stronger the enhancement during the dawn period;
[0128] For some main frequency bands where common birds gather (such as the 2-8kHz range), the following form of bandpass enhancement can be introduced, such as the frequency enhancement factor Γ freq :
[0129] Γfreq $(f)=1 + \rho$ band $\Theta(f; f_1, f_2)$
[0130] Where: $\Theta(f; f_1, f_2)$ represents a function that increases or remains high within the interval $[f_1, f_2]$, and can take a piecewise linear or other smoothly transitioning form to highlight the active frequency band of bird calls; $\rho$ band is the enhancement amplitude coefficient, with a value greater than or equal to 0, and is adjusted according to the main frequency range of the target bird species.
[0131] The time - frequency distribution weighted by the dynamic spectrum enhancement function $\Omega(t, f)$ is mainly used in this step for:
[0132] being sent back into the hybrid multi - label classification network $\Phi$ class for re - judgment to improve the final recognition accuracy in the case of weak bird calls or residual noise;
[0133] Outputting time segment markers: Marking the high - energy segments or high - confidence segments in time for the next - step multi - modal fusion and continuous feedback optimization calls, so as to conduct joint analysis with spatial positioning or meteorological data;
[0134] When in use, make full use of the biological law that wetland bird calls are active in specific time periods or specific frequency bands, effectively improve the detectability of target calls in a noisy environment, combine interactive multi - label determination, re - highlight the weak bird calls that may be ignored, reduce the missed detection of rare or weak species, and the enhanced time - frequency segments output can be matched with the spatial information or meteorological data of the sensor array in the fourth step to further improve the stability and dynamic adaptability of recognition.
[0135] Through the advanced multi - label recognition strategy in step 301 and the dynamic time - frequency enhancement in step 302, accurate species detection and effective feature highlighting of separated tracks in the multi - bird chorus environment of wetlands are achieved, providing high - confidence multi - species recognition results for the final multi - modal fusion and enhanced audio data for spatio - temporal correlation analysis, thus promoting the overall performance improvement of the wetland bird monitoring system in scenarios with increased noise and multi - species overlap.
[0136] Step Four: When the recognition result triggers multi - modal fusion with environmental monitoring requirements, collect multi - channel audio data $M$ q $(t)$ and external environmental parameters $E(\tau)$ as well as spatial positioning vector $G(\tau)$, identify the noise - dominant sources such as strong wind or water flow through the spatial positioning function and adaptive update strategy, and when misdetection, missed detection or the appearance of new species is detected, send the corresponding track data and annotation logs back to the previous steps to supplement the noise template $\Phi$ b $(t, f)$ and the bird call template $\Psi$ s $(t, f)$;
[0137] Step 4 includes the following contents:
[0138] Step 401, Multi-modal data collection and synchronous alignment
[0139] When at a fixed wetland monitoring point or during an unmanned aerial vehicle (UAV) cruise, microphone arrays are arranged at multiple spatial positions to form multi-channel audio data {M1, M2, …, M Q} where Q represents the number of microphone channels;
[0140] The recording of each channel needs to be aligned with the single track identified in Step 3 by making the time stamp t aligned (i.e., uniformly adopting the same reference time base), ensuring that subsequent sound localization or coherence analysis can be carried out based on spatial differences;
[0141] Synchronization identifier: Define a set of global time reference τ to make the multi-channel audio data M q (t) correspond to the quasi-pure track in terms of step size or sampling rate;
[0142] Collect non-audio modal data strongly related to the wetland scene, such as wind speed v wind , water flow u flow , air temperature T air etc., and uniformly map them to the same time index τ, and record these data in the vector form of the external environmental parameter E(τ):
[0143] E(τ) = (v wind (τ), u flow (τ), T air (τ), …)
[0144] Locate the UAV or fixed point using GPS / Beidou, etc., to generate a spatial positioning vector G(τ) = (lat(τ), lng(τ), alt(τ));
[0145] Combined with the relative position of the microphone array in three-dimensional space, subsequent azimuth determination of multiple sound sources and better differentiation of the distance or bird songs in different habitats can be carried out;
[0146] When in use, through the multi-channel microphone and external environmental parameters, the association between bird sound propagation and external factors such as meteorology and hydrology is further revealed, breaking the limitation of single sound source processing, forming a unified time axis in the whole system, and providing a basis for subsequent analysis (such as noise component determination and spatial positioning). Since the wetland environment varies greatly with meteorology and hydrology, collecting and aligning this information can significantly enhance the stability of the system against sudden situations or seasonal changes.
[0147] Step 402, Multi-modal fusion and dynamic parameter adjustment
[0148] Based on the acquired multi-channel audio data {M1, M2, …, M Q}, and the spatial positioning vector G(τ), techniques such as time of arrival or coherent processing can be used to estimate the azimuth of the target bird song in space.
[0149] Define the positioning function Output the estimated coordinates R m (τ) of the sound source in three-dimensional space to distinguish bird songs in different azimuths. This information will be associated with the species label and time segment marker produced in step three, so as to know which bird makes a sound where and when; the recorded external environmental parameter E(τ) can help identify some strong wind or water flow noise scenarios. For example, when the wind speed v wind (τ) is greater than the preset wind speed threshold θ wind and the sound source azimuth estimation is concentrated above or in an open area, it can be determined that the noise may be mainly wind noise at this time; if it is confirmed that the new noise dominates, the corresponding noise template Φ sep in the blind source separation model Φ b (t, f) in step two (blind source separation and noise reduction) can be incrementally updated, or the matching parameters can be adjusted in the noise recognition operator Λ noise to better suppress similar noises;
[0150] Define the adaptive adjustment function When abnormal noise distribution or new interference is detected, locally correct the contribution coefficients α b and β s etc. (see the weighting coefficients in steps one and two);
[0151] Among them, the positioning function The idea is as follows:
[0152] Obtain the multi-channel audio data M q and the synchronization time interval of the track Y′ m ; Extract TDOA: In the time window corresponding to τ, calculate the time delay difference Δt using cross-correlation methods such as GCC-PHAT ij ; Coordinate solution: Based on the relative positions of the microphones (x i , y i , z i ), the speed of sound C and Δt ij , use geometric or least squares methods to solve the sound source position; Matching determination: If the calculated sound source coordinates are consistent with the characteristic height within this time window, record the estimated coordinates R m (τ), otherwise mark it as unmatched or a noise source;
[0153] Output the estimated coordinates R m(τ), and the estimated coordinates R available in subsequent steps m (τ) is used to assist in determining the noise distribution, the position of the bird flock, or performing operations such as closed-loop updates.
[0154] Adaptive adjustment function The idea is as follows:
[0155] Obtain the external environment parameters E(τ), the estimated coordinates R m (τ), and the quasi-pure orbit After detecting an anomaly or a new situation in step four, the system calls Θ update ;
[0156] Noise type judgment: Based on meteorological & hydrological data + positioning information, initially determine whether it is strong wind, water flow sound, or other interferences;
[0157] Compare with the background / bird call template: Compare the remaining noise characteristics of this orbit with the acoustic dictionary D wetland Perform a similarity determination. If the matching degree is low and there is sufficient confidence to believe that this is a new scenario or a new species, then enter the update branch;
[0158] Generate update instructions: Adjust the key parameters of blind source separation or multi-label models according to the judgment results, and expand the noise template Φ b (t,f) or the template in the bird call template Ψ s (t,f);
[0159] Record and execute: Automatically update the model and expand the dataset based on this instruction during the next training cycle or offline retraining to complete the adaptive closed loop.
[0160] If a rare bird species is observed at a specific spatio-temporal position in the extreme case of the environmental parameter E(τ) in step 401, a rare event pipeline can be used to notify step three to perform dynamic spectrum enhancement and re-judgment on the said quasi-pure orbit to improve the accuracy of rare species identification, thereby adding a key mark (such as the appearance of a rare bird species) to the multi-species label output, and exporting it together with the estimated coordinates R m (τ) for subsequent protection or monitoring actions;
[0161] When in use, by combining spatial positioning and environmental parameters, it is possible to more quickly and accurately distinguish the source of natural noise and target bird calls, reduce the over-fitting problem of blind source separation to rare noise scenarios, and can timely perform template or parameter expansion for new noise types in extreme situations such as strong wind and floods, maintain the long-term stability of the system, increase the attention to weak or rare birds in key spatio-temporal regions, and significantly improve the value and sensitivity of ecological monitoring.
[0162] Infer the noise dominant factor by combining sound source localization and environmental parameter guidance, dynamically update the template weights of blind source separation, form an active learning-based noise suppression process, trigger re-enhancement or classification re-judgment according to environmental data and spatio-temporal distribution information, and provide more focused detection capabilities for rare bird species.
[0163] Step 403: Continuous feedback loop and stepwise iterative optimization
[0164] In the recognition output of Step 3 or the secondary localization of this step, if there is a misdetection (judging non-target sound as bird song) or a missed detection (there is bird song in the actual measurement but the classification fails), the relevant audio track and its time period and species label information need to be recorded. If it is found in the manual assisted review that the actual bird species does not match the recognition label, this difference will also be marked and written into the error log H error ;
[0165] For new or uncollected noise types, they are also written into the noise log H noise , so as to update to the acoustic dictionary D wetland and the blind source separation model Φ sep ;
[0166] If the noise log H noise records new noise samples or new noise combination methods, then in Step 1, the background template Φ b (t,f) can be expanded or the training data set R simu can be incrementally supplemented;
[0167] If the error log H error indicates that the recognition confusion rate of some species is relatively high, then in the joint training of Step 2 and Step 3 (or subsequent offline retraining), by increasing the pure bird song R sing of this species or constructing more mixed scenarios α b and β s combinations, the recognition performance for relevant confused species can be improved;
[0168] Based on the feedback and update formed in the previous two steps, adjust the monitoring strategy in the subsequent deployment:
[0169] Increase the sampling frequency for the time periods or locations where missed detections often occur, deploy a more dense microphone array for key wetland areas. If it is confirmed that the noise is extremely severe in certain time periods or the recording of the drone fails under specific meteorological conditions, then the time period can be skipped flexibly in the work process or alternative data sources can be adopted to improve the overall efficiency and accuracy.
[0170] When in use, through the accumulation of the error and noise logs, it can be continuously updated across monitoring cycles of several months or years, quickly adapt to new environments and new bird species, gradually reduce misdetections and missed detections in the iteration, and provide more accurate long-term data for subsequent wetland ecological decision-making; for multi-channel audio data Mq Collection and synchronization with the environmental parameter E(τ), aligning with the pure orbit During the species identification period, it makes up for the limitation of relying only on single-channel audio analysis. In step 402, spatial positioning and external environmental information are used to further optimize the noise identification and rare species detection of blind source separation, and the sound and environmental information are more deeply integrated.
[0171] Please refer to Figure 2 , the present invention provides an automatic bird acoustic recognition system in a wetland environment, including,
[0172] A sound data processing module, when it detects that the wetland audio needs to be initially processed, collects the background sound R bg and pure bird calls R sing at multiple time periods and constructs an acoustic dictionary D wetland , combines the controllable contribution coefficients α b and β s to simulate noise and bird calls and complete preliminary segmentation and anomaly elimination to reduce the risk of data contamination, and synchronously record the acquisition time and scene description in the metadata, and then output the labeled effective target recording R target and abnormal or unavailable segments R drop ;
[0173] A separation module, if it is necessary to disassemble the effective target recording R target with multiple sound sources, calls the blind source separation model Φ sep to perform iterative training through the use of a loss function, and the matching degree to disassemble and denoise each mixed signal orbit one by one, so as to retain acoustic details for weak bird calls, improve the signal-to-noise ratio to avoid confusion caused by multi-source overlap, and finally generate a quasi-pure orbit
[0174] A marking module, when the quasi-pure orbit is generated and it is necessary to determine the bird species contained in each orbit, input it into the hybrid multi-label classification network Φ class to perform deep characterization by combining the CNN and Transformer structures, use the Renyi entropy to improve the loss to strengthen the multi-species discrimination, and then perform dynamic energy weighting on the dawn or main frequency band through the enhancement factor to highlight the weak bird call characteristics, output an information set containing multi-species labels and confidence levels, and mark the corresponding time segments;
[0175] An adaptive update module, when the recognition result triggers multi-modal fusion with the environmental monitoring requirements, collects multi-channel audio data M q(t) and external environment parameter E(τ) as well as spatial positioning vector G(τ), identify noise-dominated sources such as strong winds or water currents through the spatial positioning function and the adaptive update strategy, and when misdetection or missed detection is detected or a new species appears, transmit the corresponding trajectory data and annotation logs back to the previous steps to supplement the noise template Φ b (t,f) and birdcall template Ψ s (t,f).
[0176] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0177] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0178] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0179] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for automatic acoustic identification of birds in a wetland environment, characterized by: include, By collecting background sounds and pure bird calls from multiple time periods and constructing an acoustic dictionary, the noise and bird calls are simulated with controllable contribution coefficients, and preliminary segmentation and abnormal elimination are completed, and the annotated valid target recordings and abnormal or unusable segments are output; If multiple sound sources need to be disassembled into valid target recordings, the blind source separation model is called to iteratively train the mixed signal track by using the loss function and matching degree to disassemble and reduce noise, generating a quasi-pure track; The quasi-pure tracks are input into a hybrid multi-label classification network, which combines CNN and Transformer structures for deep representation. The Renyi entropy and enhancement factor are used to enhance the weak bird song features, and an information set containing multi-species labels and confidence levels is generated. When the recognition results and environmental monitoring needs trigger multimodal fusion, the dominant source of noise is identified through spatial positioning functions and adaptive update strategies. When false detections, missed detections, or the appearance of new species are detected, the corresponding track data and annotation logs are fed back to the previous steps to supplement the noise template and bird song template.
2. The method for automatic acoustic identification of birds in a wetland environment according to claim 1, characterized in that: Collect original recording data sets, background sounds and pure bird calls from various wetlands at different time periods; The bird song segments that appear in the recording dataset are labeled with species labels, and the background noise types of the corresponding time periods are recorded to form a preliminary marking table. The audio segments with quality defects are removed or separately stored as faulty records.
3. The method for automatic acoustic identification of birds in a wetland environment according to claim 2, characterized in that: Based on the obtained multi-source sound data, an acoustic dictionary is constructed, which includes several background noise templates and single-species bird song templates, and a combined model of the entire wetland environment is constructed; a variety of complex scenes are synthesized by combining different contribution coefficients to obtain a simulated synthetic soundscape, and the original recording is segmented and anomalies are eliminated in combination with a preliminary marking table, and the bird song segment and background segment are divided through time series segmentation and extraction functions.
4. The method for automatic acoustic identification of birds in a wetland environment according to claim 1, characterized in that: The training sample set is constructed by random weight combination of real recordings or acoustic dictionaries; A self-supervised approach is introduced to define a comprehensive loss function. The constructed deep model is iteratively trained using the comprehensive loss function and training sample set. The trained separation network is inferred. The input is actual wetland recordings, and the separation network outputs preliminary separation tracks.
5. The method for automatic acoustic identification of birds in a wetland environment according to claim 4, characterized in that: Use the background template in the acoustic dictionary to build a match with the separation result; For tracks whose matching degree exceeds the matching threshold, a nonlinear noise reduction filter operator is applied to perform noise reduction processing to form a quasi-pure track set, retaining the time-frequency distribution and noise confidence label corresponding to each track.
6. The method for automatic acoustic identification of birds in a wetland environment according to claim 5, characterized in that: For each quasi-pure track, it is divided into several short time windows and the time-frequency features are extracted. The tracks are batch processed in the constructed hybrid multi-label classification network to obtain the corresponding species prediction distribution. Construct multi-label classification loss and single-class penalty function based on Renyi entropy improvement; When the probability estimate deviates from the true label, the single-class penalty function produces a nonlinearly increasing penalty, which improves the discrimination of easily confused species in multi-species scenarios; if multiple single tracks appear in the same period, the probability estimate is output according to the hybrid multi-label classification network.
7. The method for automatic acoustic identification of birds in a wetland environment according to claim 6, characterized in that: A dynamic spectrum enhancement function is constructed to perform targeted weighting on the quasi-pure track to obtain the time-frequency distribution, and then the time enhancement factor and frequency enhancement factor are constructed. The time-frequency distribution is used to be sent to the hybrid multi-label classification network for re-judgment and output the time segment label.
8. The method for automatic acoustic identification of birds in a wetland environment according to claim 7, characterized in that: When the wetland is monitored at a fixed point or when the drone is cruising, microphone arrays are arranged at multiple spatial locations to form multi-channel audio data, and the multi-channel audio data is made to correspond to the quasi-pure track in terms of step size or sampling rate; After collecting external environmental parameters that are strongly related to the wetland scene, the UAV or fixed point is positioned to generate a spatial positioning vector.
9. The method for automatic acoustic identification of birds in a wetland environment according to claim 8, characterized in that: Based on the acquired multi-channel audio data and spatial positioning vector, the constructed positioning function outputs the estimated coordinates of the sound source in three-dimensional space to distinguish bird calls from different directions. Combined with species labels and time segment markers, it is analyzed to obtain which bird species sings where and when. When abnormal noise distribution or new interference is detected, the local correction contribution coefficient of the adaptive adjustment function is defined according to the known background template or the newly collected noise samples.
10. The method for automatic acoustic identification of birds in a wetland environment according to claim 9, characterized in that: If rare bird species are observed at a specific spatiotemporal location under extreme environmental parameters, dynamic spectrum enhancement and re-judgment are performed again on the quasi-pure track, thereby adding focus marks in the multi-species label output; If there is a false detection or missed detection, record the relevant audio track and its time period and species label information; If the actual bird species does not match the identification tag, the difference will be noted in the error log; for new or unrecorded noise types, the noise log will be written.
11. The method for automatic acoustic identification of birds in a wetland environment according to claim 10, characterized in that: If new noise samples or new noise combinations are recorded in the noise log, the background template is expanded or the training data set is incrementally supplemented; If the error log indicates that the confusion rate of identifying certain species is high, increase the pure bird song of that species or construct more mixed scene combinations.
12. An automatic acoustic identification system for birds in a wetland environment, characterized by: include, The sound data processing module collects background sounds and pure bird calls from multiple time periods and constructs an acoustic dictionary. It simulates noise and bird calls with controllable contribution coefficients and completes preliminary segmentation and abnormal elimination, outputting the annotated valid target recordings and abnormal or unusable segments. Separation module: If multiple sound sources need to be disassembled into valid target recordings, the blind source separation model is called to iteratively train the mixed signal track by using the loss function and matching degree to disassemble and reduce noise, and generate a quasi-pure track; The labeling module inputs the quasi-pure tracks into a hybrid multi-label classification network, combines CNN and Transformer structures for deep representation, uses Renyi entropy and enhancement factors to enhance weak bird song features, and produces an information set containing multi-species labels and confidence levels; The adaptive update module, when the recognition results and environmental monitoring needs trigger multimodal fusion, identifies the dominant source of noise through spatial positioning functions and adaptive update strategies, and when false detections, missed detections or the appearance of new species are detected, the corresponding track data and annotation logs are fed back to the previous steps to supplement the noise template and bird song template.
Citation Information
Patent Citations
Bird whistling classification and identification method and device
CN115762533A
Aliasing twitter separation method based on deep learning
CN117789746A
Bird sound feature extraction and recognition method and system for complex sound scene
CN117831544A
Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device
CN119027775A
Bird sound event detection method and system oriented to national key protection bird monitoring
CN119296548A
Cited By
Natural sound recognition device of online monitor
CN121122292A
Valuable and rare bird recognition method by means of voiceprint recognition technology
CN121122293A
A rare bird identification method by means of voiceprint recognition technology
CN121122293B
Urban street bird activity distribution monitoring method based on bird chirp intensity
CN122090876A