Bird voiceprint collection system and device based on AI voiceprint recognition

By using AI voiceprint recognition technology, combined with acoustic and environmental data acquisition, quantization, and adaptive correction units, and dynamically correcting the soundscape channel operator, the problem of decreased recognition accuracy in complex environments of traditional systems is solved, achieving highly robust and stable bird voiceprint acquisition and recognition.

CN120932655BActive Publication Date: 2026-04-28JIANGSU TIANNING ECOLOGICAL GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU TIANNING ECOLOGICAL GRP CO LTD
Filing Date
2025-08-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional bird voiceprint acquisition systems struggle to filter out noise in complex field environments and cannot dynamically adjust model parameters, leading to decreased recognition accuracy and recognition failure.

Method used

A bird voiceprint acquisition system based on AI voiceprint recognition is adopted. Through the collaborative work of multiple units, including acoustic and environmental data acquisition, environmental state quantification, sound source signal decoupling, voiceprint recognition, and adaptive correction units, acoustic signals and environmental parameters are acquired in real time, environmental stress factors are quantified, and soundscape channel operators are dynamically corrected to adapt to environmental changes.

Benefits of technology

It achieves highly robust bird voiceprint acquisition and recognition, improves resistance to environmental interference, ensures the stability and accuracy of recognition results, and can quickly recover performance to avoid recognition failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932655B_ABST
    Figure CN120932655B_ABST
Patent Text Reader

Abstract

The application discloses a kind of bird voiceprint collection systems and devices based on AI voiceprint identification, it is related to bird voiceprint collection technical field, comprising: acoustic and environmental data acquisition unit, for obtaining original acoustic signal and synchronous collection environmental physical parameter;Environment state quantization unit, for quantifying to generate environmental stress factor;Sound source signal decoupling unit is configured to apply correctable soundscape channel operator, generate decoupling reliability factor;Voiceprint identification unit is used to analyze pure sound source signal, and output species identification confidence;Self-adaptive correction unit is used to determine correction learning rate, calculate channel model loss, and carry out closed-loop correction to soundscape channel operator, to adapt to the interference of environmental change to acoustic signal.The application realizes high robustness bird voiceprint collection, can adapt to environmental change, recover performance quickly through hierarchical correction and historical database, and improve identification reliability by combining multi-factor calculation loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bird voiceprint acquisition technology, specifically to a bird voiceprint acquisition system and device based on AI voiceprint recognition. Background Technology

[0002] Bird voiceprint recognition is a crucial tool in ecological monitoring, enabling species identification and behavioral studies through the analysis of bird calls. With the development of artificial intelligence, voiceprint recognition technology is increasingly being applied to bird monitoring alongside AI. However, the complexities of the wild environment and the susceptibility of acoustic signals to interference from environmental factors mean that traditional bird voiceprint collection methods often employ fixed models to analyze acoustic signals, failing to effectively adapt to environmental changes that could disrupt the signal.

[0003] Existing bird voiceprint acquisition systems have significant shortcomings in complex field environments. On the one hand, environmental changes such as fluctuations in temperature, humidity, and vegetation cover can interfere with the propagation of acoustic signals, making it difficult for traditional bird voiceprint acquisition systems to effectively filter out noise, resulting in a decrease in recognition accuracy. On the other hand, traditional systems lack adaptive correction mechanisms, making it impossible to dynamically adjust model parameters when the environment changes abruptly, and the reliability assessment of the signal decoupling process is lacking, making them prone to recognition failure due to model mismatch. Summary of the Invention

[0004] The purpose of this invention is to provide a bird voiceprint acquisition system and device based on AI voiceprint recognition, which solves the problems existing in the background technology.

[0005] To address the aforementioned technical problems, this invention provides a bird voiceprint acquisition system based on AI voiceprint recognition, comprising: an acoustic and environmental data acquisition unit, used to acquire in real time the original acoustic signal containing the call of the target bird, and simultaneously acquire the environmental physical parameters of the location where the original acoustic signal is generated;

[0006] An environmental state quantification unit is used to compare the environmental physical parameters with a preset environmental historical baseline and quantify and generate environmental stress factors that characterize the degree of deviation of the current environment from the normal state.

[0007] The sound source signal decoupling unit is configured to apply a modifiable sound scene channel operator to separate the pure sound source signal from the original acoustic signal, and generate a decoupling reliability factor to evaluate the reliability of the process based on the degree of change of the original acoustic signal caused by the separation process; the sound scene channel operator generates the signal by mapping the environmental physical parameters through environmental mapping parameters.

[0008] A voiceprint recognition unit is used to analyze the pure sound source signal and output the species identification confidence level;

[0009] The adaptive correction unit is used for:

[0010] Based on the environmental stress factor, a correction learning rate is determined for modifying the soundscape channel operator;

[0011] Based on the species identification confidence level and the decoupling reliability factor, calculate the channel model loss that characterizes the current accuracy of the soundscape channel operator;

[0012] By combining the channel model loss and the correction learning rate, the acoustic scene channel operator is subjected to closed-loop correction to adapt to the interference of environmental changes on the acoustic signal.

[0013] Preferably, the environmental physical parameters include the current ambient temperature, ambient humidity, and vegetation cover density index; the environmental state quantification unit is specifically used to compare the current ambient temperature, ambient humidity, and vegetation cover density index with their corresponding historical average values, generate their respective deviation values, and perform a weighted summation of the deviation values ​​to constitute the environmental stress factor.

[0014] Preferably, the adaptive correction unit is configured to determine the correction learning rate based on the environmental stress factor and a preset stress sensitivity coefficient through a preset nonlinear positive correlation, such that the larger the value of the environmental stress factor, the higher the correction learning rate.

[0015] Preferably, the decoupling reliability factor includes:

[0016] Extract the original acoustic features of the original acoustic signal and the pure acoustic features of the pure sound source signal;

[0017] Calculate the distortion between the original acoustic feature and the pure acoustic feature;

[0018] Based on the distortion degree, the decoupling reliability factor is generated through a preset negative correlation function relationship. The greater the distortion degree, the lower the decoupling reliability factor.

[0019] Preferably, before calculating the channel model loss, the adaptive correction unit further includes:

[0020] The species identification confidence level is multiplied by the decoupling reliability factor to obtain the reliability-adjusted confidence level.

[0021] The channel model loss is calculated based on the confidence level after reliability adjustment.

[0022] Preferably, the voiceprint recognition unit is further configured to calculate the voiceprint complexity factor based on the acoustic information entropy and bandwidth of the pure sound source signal; the calculation of the channel model loss is further based on the voiceprint complexity factor, and when the confidence level after reliability adjustment is the same, the higher the voiceprint complexity factor, the lower the channel model loss.

[0023] Preferably, the step of the adaptive correction unit performing closed-loop correction on the soundscape channel operator includes:

[0024] Multiply the correction learning rate by the channel model loss to obtain the correction amount;

[0025] Subtracting the correction amount from the standard value yields the correction coefficient;

[0026] The soundscape channel operator is multiplied by the correction coefficient to generate the corrected soundscape channel operator.

[0027] Preferably, the adaptive correction unit further includes:

[0028] The channel model loss is compared with a first preset threshold and a second preset threshold, wherein the second preset threshold is greater than the first preset threshold;

[0029] If the channel model loss is greater than the first preset threshold but not greater than the second preset threshold, then the environment mapping parameters are fine-tuned.

[0030] If the channel model loss is greater than the second preset threshold, then the sound scene channel operator that was successfully identified in a similar environment in the historical database is used for replacement.

[0031] A bird voiceprint acquisition device based on AI voiceprint recognition is also provided, comprising:

[0032] The acoustic and environmental data acquisition module is equipped with a microphone array and environmental sensors to collect the raw acoustic signals of the target bird calls in real time, and simultaneously acquire physical parameters such as ambient temperature, humidity and vegetation cover density index at the location where the signal is generated.

[0033] The environmental state quantification module includes a data processing unit, which compares the collected environmental physical parameters with a preset historical baseline and generates an environmental stress factor by weighted summation of deviation values ​​to characterize the degree to which the environment deviates from the normal state.

[0034] The sound source signal decoupling module integrates a sound scene channel operator calculation unit. It generates operators through environmental mapping parameters to separate the pure sound source signal from the original acoustic signal, and calculates the decoupling reliability factor based on the degree of signal change.

[0035] The voiceprint recognition module, equipped with an AI processing chip, analyzes pure sound source signals, outputs species identification confidence, and calculates the voiceprint complexity factor based on signal entropy and bandwidth.

[0036] The adaptive correction module, including the control processor, is configured as follows:

[0037] The correction learning rate is determined through a nonlinear relationship based on environmental stress factors and preset sensitivity coefficients.

[0038] The channel model loss is calculated by combining species identification confidence, decoupling reliability factor, and voiceprint complexity factor.

[0039] The soundscape channel operator is closed-loop corrected based on the learning rate and loss value. Specifically, the correction coefficient is calculated through the correction amount to realize the operator update.

[0040] Compare the channel model loss with a preset threshold. When the loss exceeds the first threshold, fine-tune the environment mapping parameters. When it exceeds the second threshold, call the operator in the historical database under similar conditions for replacement.

[0041] The storage module is used to save historical environmental baselines, soundscape channel operators, and successfully identified voiceprint data.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] 1. Through multi-unit collaborative work, it achieves highly robust acquisition and recognition of bird voiceprints, can acquire acoustic and environmental data in real time, quantify environmental stress factors, separate pure sound source signals and evaluate decoupling reliability, and combine multiple factors to dynamically modify the soundscape channel operator, effectively adapting to the interference of environmental changes on acoustic signals and improving the ability to resist environmental interference.

[0044] 2. By adopting a graded correction strategy and historical database, when the channel model loss exceeds different thresholds, parameter fine-tuning or operator replacement is performed respectively to ensure that the system can quickly recover performance under different degrees of deviation, avoid recognition failure due to severe model mismatch, and improve the stability and robustness of the system.

[0045] 3. By comprehensively considering species identification confidence, decoupling reliability factor and voiceprint complexity factor, the channel model loss is calculated, making the loss signal more accurately point to the channel model problem, avoiding misjudgment of voiceprint signals with suspicious or complex sources, and improving the reliability and accuracy of the identification results. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a logic block diagram of the system of the present invention;

[0048] Figure 2This is a structural diagram of a bird voiceprint acquisition device based on AI voiceprint recognition; Detailed Implementation

[0049] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Example 1

[0050] Please see Figure 1 This invention provides a bird voiceprint acquisition system based on AI voiceprint recognition, comprising:

[0051] The acoustic and environmental data acquisition unit is used to acquire raw acoustic signals containing the calls of target birds in real time, and simultaneously acquire environmental physical parameters of the location where the raw acoustic signals are generated; the environmental state quantification unit is used to compare the environmental physical parameters with a preset environmental historical baseline, and quantify and generate an environmental stress factor characterizing the degree of deviation of the current environment from the norm; the sound source signal decoupling unit is configured to apply a correctable soundscape channel operator to separate the pure sound source signal from the raw acoustic signal, and generate a decoupling reliability factor to evaluate the reliability of the process based on the degree of change of the raw acoustic signal during the separation process; the soundscape channel operator maps the environmental physical parameters through a set of environmental mapping parameters; the voiceprint recognition unit is used to analyze the pure sound source signal and output the species identification confidence level; the adaptive correction unit is used to: determine the correction learning rate for correcting the soundscape channel operator based on the environmental stress factor; calculate the channel model loss characterizing the current accuracy of the soundscape channel operator based on the species identification confidence level and the decoupling reliability factor; and perform closed-loop correction of the soundscape channel operator by combining the channel model loss and the correction learning rate to make the system adapt to the interference of environmental changes on the acoustic signal;

[0052] This invention provides a bird voiceprint acquisition system based on AI voiceprint recognition. Through the collaborative work of multiple units, it achieves highly robust acquisition and recognition of bird voiceprints. When the system is working, the acoustic and environmental data acquisition unit is deployed in the field monitoring scene, such as forest or wetland, and is responsible for capturing composite acoustic signals containing the target bird's call in real time, and simultaneously recording environmental physical parameters such as temperature, humidity and vegetation cover at the location.

[0053] The environmental state quantification unit receives these environmental parameters, compares them with the historical environmental baseline established by long-term observation, and generates a standardized environmental stress factor to numerically represent the degree of anomaly in the current environment.

[0054] The sound source signal decoupling unit uses a dynamically modifiable soundscape channel operator to process the received composite acoustic signal, aiming to filter out attenuation and noise interference caused by environmental propagation and restore the pure bird song signal; the voiceprint recognition unit performs deep learning analysis on the pure signal and outputs the specific bird species identification result and its confidence level.

[0055] As the central hub of the system, the adaptive correction unit integrates environmental stress factors, species recognition confidence, and reliability assessment results of the sound source signal decoupling process. It dynamically calculates the channel model loss and, combined with the correction learning rate determined by the degree of environmental stress, performs closed-loop feedback correction on the sound scene channel operator. This design enables the entire system to actively adapt to dynamic changes in the environment, such as changes in sound propagation paths caused by weather changes, thereby continuously optimizing the accuracy of signal processing and ensuring stable and reliable bird recognition results even in complex and ever-changing field environments.

[0056] In addition, the system also includes a historical database, which stores historical sound scene channel operators that have been successfully identified under various known combinations of environmental parameters. This historical database provides the basis for subsequent secondary correction strategies, ensuring that when the system detects a serious mismatch in the current channel operator, it can query and call successful operators under similar historical conditions for rapid replacement, thereby ensuring the robustness and rapid recovery capability of the system. Example 2

[0057] Environmental physical parameters include the current ambient temperature, ambient humidity, and vegetation cover density index; the environmental state quantification unit is specifically used to compare the current ambient temperature, ambient humidity, and vegetation cover density index with their corresponding historical averages, generate their respective deviation values, and perform a weighted summation of the deviation values ​​to constitute the environmental stress factor.

[0058] The adaptive correction unit is configured to determine the correction learning rate based on the environmental stress factor and a preset stress sensitivity coefficient through a preset nonlinear positive correlation, so that the larger the value of the environmental stress factor, the higher the correction learning rate.

[0059] ;

[0060] in: Indicates environmental stress factors; The current ambient temperature. Given the current ambient humidity, This represents the current vegetation cover density index. This represents the historical average baseline value of the ambient temperature recorded at this monitoring point over a long period of time. This represents the historical average baseline value of environmental humidity recorded at this monitoring point over a long period of time. This represents the historical average baseline value of the vegetation cover density index recorded over a long period of time at this monitoring point; The environmental temperature parameter weights are set based on prior knowledge; The environmental humidity parameter weights are set based on prior knowledge; The vegetation cover density index parameter weights are set based on prior knowledge, and the preset weight values ​​are set according to the physical principles of sound wave propagation.

[0061] The adaptive correction unit then adjusts according to the environmental stress factor. The corrective learning rate is determined through a pre-defined nonlinear positive correlation, designed as an exponential function to amplify the effects of high stress.

[0062] ;

[0063] in: This represents the correction learning rate for the i-th segment of the signal; This represents the preset base learning rate, which indicates the system's fine-tuning rate under the most stable conditions. This represents the preset stress sensitivity coefficient, used to control the rate at which the learning rate increases with environmental stress. This represents the environmental stress factor corresponding to the current signal segment.

[0064] The overall effect of this design is that when the system detects drastic changes in the environment, such as a sudden drop in humidity due to heavy rain... Sudden rise, The value of will increase significantly, which in turn leads to an exponential increase in the learning rate. The system is rapidly improved, enabling quick and significant model corrections to adapt to new acoustic environments and ensure the continuity and accuracy of recognition. Example 3

[0065] The decoupling reliability factor includes: extracting the original acoustic features of the original acoustic signal and the pure acoustic features of the pure sound source signal; calculating the distortion degree between the original acoustic features and the pure acoustic features; and generating the decoupling reliability factor based on the distortion degree through a preset negative correlation function relationship, wherein the greater the distortion degree, the lower the decoupling reliability factor.

[0066] After separating the pure acoustic source signal, the sound source signal decoupling unit further quantifies the reliability of its operation. This unit first extracts acoustic feature vectors from the received original acoustic signal and the estimated pure acoustic source signal after decoupling using a feature extraction operator, such as a Mel-frequency cepstral coefficient extractor. Then, by calculating and normalizing the Euclidean distance between these two feature vectors, the decoupling distortion degree is obtained.

[0067] ;

[0068] in: This represents the degree of decoupling distortion of the i-th signal segment; Indicates the feature extraction operator; This represents the received raw acoustic signal; This represents the estimated clean sound source signal after decoupling; This indicates the calculation of the vector magnitude; this distortion degree It reflects the extent to which the decoupling process modifies the original signal;

[0069] Finally, this unit converts distortion into a decoupling reliability factor based on a pre-defined negative correlation function:

[0070] ;

[0071] in: This represents the decoupling reliability factor; This represents a preset distortion sensitivity coefficient;

[0072] The preset value of this coefficient is based on prior knowledge of the scenario. For example, in an open area with sparse vegetation, the channel itself is stable, and any large signal distortion may be due to model inaccuracy, so it will be set to a large value. Conversely, in a dense forest, the channel itself will cause huge attenuation, requiring strong decoupling, so it will be set to a small value to tolerate necessary signal distortion. The effect of this mechanism is that when the system makes drastic modifications to the restored signal, the reliability of the result will be judged to be low, thus providing key uncertainty information for subsequent decisions and preventing the system from making an overconfident judgment on a signal that may be overprocessed. Example 4

[0073] Before calculating the channel model loss, the adaptive correction unit also includes: multiplying the species identification confidence score by the decoupling reliability factor to obtain a reliability-adjusted confidence score; the channel model loss is calculated based on the reliability-adjusted confidence score.

[0074] The voiceprint recognition unit is also used to calculate a voiceprint complexity factor based on the acoustic information entropy and bandwidth of the pure sound source signal; specifically, the voiceprint recognition unit is equipped with a built-in signal analysis module, which can analyze the input pure sound source signal. Processing involves two aspects: First, by performing time-frequency analysis methods such as Fast Fourier Transform (FFT), the signal's frequency domain distribution range is determined, thereby calculating its effective bandwidth. On the other hand, by statistically analyzing the amplitude of the signal or the probability distribution of the encoded symbols, its information entropy is calculated to quantify the amount of information and uncertainty contained in the signal itself. These two calculated features provide direct input for the generation of the voiceprint complexity factor. The calculation of the channel model loss is further based on the voiceprint complexity factor. When the confidence level after reliability adjustment is the same, the higher the voiceprint complexity factor, the lower the channel model loss.

[0075] The adaptive correction unit and the speaker recognition unit work together to calculate a more refined channel model loss; first, the speaker recognition unit analyzes the clean sound source signal and calculates its speaker complexity factor:

[0076] ;

[0077] in: Indicates the voiceprint complexity factor; and These represent the estimated pure sound source signal. Before being substituted into the formula, the information entropy and bandwidth are normalized to make them dimensionless values. This indicates the preset weighting coefficients;

[0078] The adaptive correction unit converts the original confidence level output by the species identification unit. Decoupling reliability factor from the sound source signal decoupling unit Multiply by this to obtain the confidence level adjusted for reliability. Ultimately, the channel model loss is calculated based on the adjusted confidence level and complexity factor:

[0079] ;

[0080] in: The channel model loss represents the i-th segment of the signal; Indicates the original confidence level; Indicates the adjusted confidence level; This represents the decoupling reliability factor from the sound source signal decoupling unit;

[0081] This method produces highly robust results: if the AI ​​model has a high confidence level in identifying a signal... The signal is very high, but it was obtained through a very drastic decoupling process. It's very low, so the adjusted confidence level This will be significantly reduced, thus leading to the final loss. Maintaining a high level, while also considering the complexity of a voiceprint itself (z iEven with low confidence in recognizing high-pitched bird calls, the system will calculate a relatively low loss value, reflecting tolerance for the inherent difficulty of the recognition task. This allows the loss signal to more accurately point to problems in the channel model itself, rather than the limits of AI recognition capabilities. Example 5

[0082] The steps of the adaptive correction unit to perform closed-loop correction on the sound scene channel operator include: multiplying the correction learning rate by the channel model loss to obtain a correction amount; subtracting the correction amount from a standard value to obtain a correction coefficient; and multiplying the sound scene channel operator by the correction coefficient to generate the corrected sound scene channel operator.

[0083] The adaptive correction unit performs closed-loop correction as follows: it applies the dynamically calculated correction learning rate and channel model loss to the update of the soundscape channel operator:

[0084] ;

[0085] in: The new soundscape channel operator generated after correction; The current soundscape channel operator to be corrected; It is the corrected learning rate determined by environmental stress factors; It is a channel model loss that integrates recognition confidence, decoupling reliability, and voiceprint complexity;

[0086] During this correction process This constitutes the correction amount, and This is the correction coefficient; the effect of this step is to achieve a dynamic, on-demand, and refined adjustment; for example, in a stable environment ( Low) and the system identifies accurately ( In the case of low (low) conditions, the correction amount is extremely small, and the soundscape channel operator Maintaining stability avoids unnecessary disturbances; however, once the environment changes abruptly ( High) or decreased recognition performance ( (High), the correction amount will increase significantly, leading to... A significant adjustment is made to enable it to quickly adapt to the new environment or correct model biases, thereby ensuring the high performance and high adaptability of the entire system during continuous operation. Example 6

[0087] The adaptive correction unit further includes: comparing the channel model loss with a first preset threshold and a second preset threshold, wherein the second preset threshold is greater than the first preset threshold; if the channel model loss is greater than the first preset threshold but not greater than the second preset threshold, then fine-tuning the environment mapping parameters; if the channel model loss is greater than the second preset threshold, then calling the sound scene channel operator that was successfully identified in a similar environment in the historical database for replacement.

[0088] The adaptive correction unit is configured to employ a tiered correction strategy to address different levels of system bias; the unit internally presets two loss thresholds, the first of which is... Second preset threshold ,and The calculated channel model loss during each correction period It will be compared with these two loss thresholds, triggering different correction behaviors:

[0089] like The system determines the current soundscape channel operator. A severe mismatch has occurred, and conventional correction methods are no longer sufficient. At this point, the system uses the current environmental parameters as query conditions to retrieve and invoke a soundscape channel operator that has been successfully recognized under similar conditions from the historical database. Perform a complete replacement and send a maintenance request to the backend system, recommending calibration of the physical sensors; this is the "emergency refactoring" level, designed to quickly restore the system's basic functionality.

[0090] like If the system determines that there is a moderate deviation, it will trigger the "online fine-tuning" strategy; at this time, the system will not replace the entire operator, but will instead adjust the generated... The environmental mapping parameters are slightly adaptively adjusted to optimize the mapping relationship between environmental parameters and channel model.

[0091] when The system considers the model to be accurate, the deviation to be within an acceptable range, the conventional correction amount to be minimal, the operator to remain stable, and the system to operate normally.

[0092] The effectiveness of this hierarchical strategy lies in providing the system with solutions that match the severity of the problem, achieving a balance between robustness and stability. In the face of sudden and severe environmental events, the system can decisively replace the model and achieve rapid recovery. For gradual and minor environmental changes, it adapts through gentle parameter fine-tuning, avoiding new instability factors that may be introduced by drastic adjustments, and ensuring the long-term stable operation of the system. Example 7

[0093] Please see Figure 1 and Figure 2 A bird voiceprint acquisition device based on AI voiceprint recognition, comprising:

[0094] The acoustic and environmental data acquisition module is equipped with a microphone array and environmental sensors to collect the raw acoustic signals of the target bird calls in real time, and simultaneously acquire physical parameters such as ambient temperature, humidity and vegetation cover density index at the location where the signal is generated.

[0095] The environmental state quantification module includes a data processing unit, which compares the collected environmental physical parameters with a preset historical baseline and generates an environmental stress factor by weighted summation of deviation values ​​to characterize the degree to which the environment deviates from the normal state.

[0096] The sound source signal decoupling module integrates a sound scene channel operator calculation unit. It generates operators through environmental mapping parameters to separate the pure sound source signal from the original acoustic signal, and calculates the decoupling reliability factor based on the degree of signal change.

[0097] The voiceprint recognition module, equipped with an AI processing chip, analyzes pure sound source signals, outputs species identification confidence, and calculates the voiceprint complexity factor based on signal entropy and bandwidth.

[0098] The adaptive correction module, including the control processor, is configured as follows:

[0099] The correction learning rate is determined through a nonlinear relationship based on environmental stress factors and preset sensitivity coefficients.

[0100] The channel model loss is calculated by combining species identification confidence, decoupling reliability factor, and voiceprint complexity factor.

[0101] The soundscape channel operator is closed-loop corrected based on the learning rate and loss value. Specifically, the correction coefficient is calculated through the correction amount to realize the operator update.

[0102] Compare the channel model loss with a preset threshold. When the loss exceeds the first threshold, fine-tune the environment mapping parameters. When it exceeds the second threshold, call the operator in the historical database under similar conditions for replacement.

[0103] The storage module is used to save historical environmental baselines, soundscape channel operators, and successfully identified voiceprint data;

[0104] Specific scenario setting: A bird voiceprint collection device based on AI voiceprint recognition is designed as an independent, solar-powered monitoring device, installed in the core area of ​​a wetland nature reserve for long-term, unattended monitoring and identification of the activities of endangered waterbirds.

[0105] Device composition and working process:

[0106] Acoustic and environmental data acquisition module: An array of eight microphones is installed on the top of the monitoring pile to capture the surrounding 360-degree sound with high fidelity and to preliminarily determine the direction of the sound source; the pile also integrates a high-precision environmental sensor to measure the ambient temperature and humidity at the location of the device in real time, as well as a miniature camera to assess vegetation density. The vegetation cover density index is obtained through image analysis. These data are collected synchronously and timestamped.

[0107] Environmental state quantification module: The built-in data processing unit of the module calls the historical environmental baseline data in the storage module; it compares the real-time collected temperature and humidity with the baseline, and calculates an environmental stress factor as shown in Example 2. This factor quantifies the potential impact of the current environment on sound propagation.

[0108] Sound source signal decoupling module: When the microphone array captures the original audio signal containing the target waterbird's call, the sound scene channel operator calculation unit in this module is immediately activated; it uses the environmental stress factor generated by the environmental state quantization module to generate a sound scene channel operator in real time; this operator is applied to the original audio to filter out wind noise, water flow noise, and specific frequency band attenuation caused by high temperature and high humidity environment, thereby separating a purer bird call signal. As described in Example 3, it calculates a decoupling reliability factor by comparing the Mel frequency cepstral coefficient characteristics of the original signal and the pure signal, which is used to evaluate the credibility of this signal separation operation.

[0109] Voiceprint recognition module: The core of this module is an AI processing chip; it receives the pure bird song signal output by the decoupling module, runs a pre-trained deep learning model, analyzes the voiceprint, and finally outputs the species identification result; in addition, this module will also calculate a voiceprint complexity factor based on the information entropy and bandwidth of the signal, which is used to characterize the recognition difficulty of the bird song itself.

[0110] Adaptive correction module: This is the core of the device, executed by a control processor; it integrates various data to make decisions.

[0111] Calculate the model loss: Multiply the 95% confidence level output by the voiceprint recognition module by the reliability factor output by the sound source signal decoupling module, and then combine it with the voiceprint complexity factor to calculate the final channel model loss according to the formula in Example 4. This loss value reflects the accuracy of the current state of the entire system.

[0112] Closed-loop correction: The learning rate is dynamically adjusted according to the environmental stress factor. The learning rate is multiplied by the channel model loss to obtain a correction amount. This correction amount is used to fine-tune the current soundscape channel operator to make it more adaptable to environmental changes.

[0113] Tiered strategy: If the calculated channel model loss exceeds a preset lower threshold, the module will automatically fine-tune the environmental mapping parameters used to generate the channel operator; if the loss exceeds a higher threshold due to sudden rainstorms or other reasons, the module will determine that the current operator has failed severely, and will query the storage module to retrieve a sound scene channel operator that was successfully identified in a similar rainstorm environment in the past to replace the current operator, so as to achieve rapid system recovery and send a maintenance alarm to the background.

[0114] Storage module: The device has a built-in industrial-grade solid-state drive for long-term storage of critical data: historical environmental parameters of the protected area, soundscape channel operators and corresponding environmental parameters that have been successfully identified under various weather conditions, and records of all successfully identified bird voiceprints; this data is not only used for adaptive correction, but also provides valuable information for subsequent ecological research.

[0115] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A bird voiceprint acquisition system based on AI voiceprint recognition, characterized in that, include: The acoustic and environmental data acquisition unit is used to acquire raw acoustic signals containing the calls of target birds in real time, and simultaneously acquire environmental physical parameters of the location where the raw acoustic signals are generated. An environmental state quantification unit is used to compare the environmental physical parameters with a preset environmental historical baseline and quantify and generate environmental stress factors that characterize the degree of deviation of the current environment from the normal state. The sound source signal decoupling unit is configured to apply a modifiable sound scene channel operator to separate the pure sound source signal from the original acoustic signal, and generate a decoupling reliability factor to evaluate the reliability of the process based on the degree of change of the original acoustic signal caused by the process of separating the pure sound source signal; the sound scene channel operator generates the environmental physical parameters by mapping the environmental mapping parameters. A voiceprint recognition unit is used to analyze the pure sound source signal and output the species identification confidence level; The adaptive correction unit is used for: Based on the environmental stress factor, a correction learning rate is determined for modifying the soundscape channel operator; Based on the species identification confidence level and the decoupling reliability factor, calculate the channel model loss that characterizes the current accuracy of the soundscape channel operator; By combining the channel model loss and the correction learning rate, the acoustic scene channel operator is subjected to closed-loop correction to adapt to the interference of environmental changes on the acoustic signal; The step of the adaptive correction unit performing closed-loop correction on the soundscape channel operator includes: Multiply the correction learning rate by the channel model loss to obtain the correction amount; Subtracting the correction amount from the standard value yields the correction coefficient; The soundscape channel operator is multiplied by the correction coefficient to generate the corrected soundscape channel operator.

2. The bird voiceprint acquisition system based on AI voiceprint recognition according to claim 1, characterized in that, The environmental physical parameters include the current ambient temperature, ambient humidity, and vegetation cover density index; the environmental state quantification unit is specifically used to compare the current ambient temperature, ambient humidity, and vegetation cover density index with their corresponding historical average values, generate their respective deviation values, and perform a weighted summation of the deviation values ​​to constitute the environmental stress factor.

3. The bird voiceprint acquisition system based on AI voiceprint recognition according to claim 1, characterized in that, The adaptive correction unit is configured to determine the correction learning rate based on the environmental stress factor and a preset stress sensitivity coefficient through a preset nonlinear positive correlation, such that the larger the value of the environmental stress factor, the higher the correction learning rate.

4. The bird voiceprint acquisition system based on AI voiceprint recognition according to claim 1, characterized in that, The decoupling reliability factor includes: Extract the original acoustic features of the original acoustic signal and the pure acoustic features of the pure sound source signal; Calculate the distortion between the original acoustic feature and the pure acoustic feature; Based on the distortion degree, the decoupling reliability factor is generated through a preset negative correlation function relationship. The greater the distortion degree, the lower the decoupling reliability factor.

5. A bird voiceprint acquisition system based on AI voiceprint recognition according to claim 1, characterized in that, Before calculating the channel model loss, the adaptive correction unit further includes: The species identification confidence level is multiplied by the decoupling reliability factor to obtain the reliability-adjusted confidence level. The channel model loss is calculated based on the confidence level after reliability adjustment.

6. A bird voiceprint acquisition system based on AI voiceprint recognition according to claim 5, characterized in that, The voiceprint recognition unit is also used to calculate the voiceprint complexity factor based on the acoustic information entropy and bandwidth of the pure sound source signal; the calculation of the channel model loss is further based on the voiceprint complexity factor. When the confidence level after reliability adjustment is the same, the higher the voiceprint complexity factor, the lower the channel model loss.

7. A bird voiceprint acquisition system based on AI voiceprint recognition according to claim 1, characterized in that, The adaptive correction unit further includes: The channel model loss is compared with a first preset threshold and a second preset threshold, wherein the second preset threshold is greater than the first preset threshold; If the channel model loss is greater than the first preset threshold but not greater than the second preset threshold, then the environment mapping parameters are fine-tuned. If the channel model loss is greater than the second preset threshold, then the sound scene channel operator that was successfully identified in a similar environment in the historical database is used for replacement.

8. A bird voiceprint acquisition device based on AI voiceprint recognition, applied to a bird voiceprint acquisition system based on AI voiceprint recognition as described in any one of claims 1 to 7, characterized in that, include: The acoustic and environmental data acquisition module is equipped with a microphone array and environmental sensors to collect the raw acoustic signals of the target bird calls in real time, and simultaneously acquire physical parameters such as ambient temperature, humidity and vegetation cover density index at the location where the signal is generated. The environmental state quantification module includes a data processing unit, which compares the collected environmental physical parameters with a preset historical baseline and generates an environmental stress factor by weighted summation of deviation values ​​to characterize the degree to which the environment deviates from the normal state. The sound source signal decoupling module integrates a sound scene channel operator calculation unit. It generates operators through environmental mapping parameters to separate the pure sound source signal from the original acoustic signal, and calculates the decoupling reliability factor based on the degree of signal change. The voiceprint recognition module, equipped with an AI processing chip, analyzes pure sound source signals, outputs species identification confidence, and calculates the voiceprint complexity factor based on signal entropy and bandwidth. The adaptive correction module, including the control processor, is configured as follows: The correction learning rate is determined through a nonlinear relationship based on environmental stress factors and preset sensitivity coefficients. The channel model loss is calculated by combining species identification confidence, decoupling reliability factor, and voiceprint complexity factor. The soundscape channel operator is closed-loop corrected based on the learning rate and loss value. Specifically, the correction coefficient is calculated through the correction amount to realize the operator update. Compare the channel model loss with a preset threshold. When the loss exceeds the first threshold, fine-tune the environment mapping parameters. When it exceeds the second threshold, call the operator in the historical database under similar conditions for replacement. The storage module is used to save historical environmental baselines, soundscape channel operators, and successfully identified voiceprint data.

Citation Information

Patent Citations

  • Bird trajectory prediction system based on bird detection radar and big data model analysis

    CN119148132A

  • Robust neural network acoustic model with side task prediction of reference signals

    US10147442B1