Intelligent deep learning method and system for animal audio voiceprint recognition

Through intelligent deep learning methods, multimodal feature extraction and voiceprint modeling are used to use distributed acquisition terminals and hybrid deep learning models, solving the problems of low recognition efficiency and low accuracy in traditional methods, and achieving more efficient and accurate animal voiceprint recognition.

CN120148525AInactive Publication Date: 2025-06-13XINTONG CONSTR TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510622265.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional animal voiceprint recognition methods are inefficient, have low recognition accuracy, and lack effective voiceprint feature evaluation and screening mechanisms, which affect the recognition effect.

Method used

Using intelligent deep learning methods, animal audio is collected through distributed acquisition terminals, multimodal feature extraction and confidence evaluation are carried out, hybrid deep learning model is constructed for voiceprint modeling, and voiceprint feature library is generated to perform target voiceprint matching and context semantic completion.

Benefits of technology

It improves the accuracy and efficiency of animal voiceprint recognition, enhances the characterization ability of voiceprint features, and improves the generalization ability and recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148525A_ABST
    Figure CN120148525A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent deep learning method and system for animal audio voiceprint recognition, and relates to the technical field of audio recognition. Animal audios of a plurality of monitoring nodes in a target area are collected, the animal audios are processed into standardized audio information packets, non-target sound sources are filtered out, and target sound sources are extracted; the method comprises the steps of performing multi-modal feature extraction on a target sound source, performing confidence evaluation after obtaining multi-modal voiceprint features, obtaining the multi-modal voiceprint features meeting a confidence screening threshold as a modeling data set, constructing a mixed deep learning model, inputting the modeling data set, performing voiceprint modeling, and generating a voiceprint feature library and a target voiceprint template. And inputting animal audio voiceprints needing to be recognized into the voiceprint feature library, calculating voiceprint similarity, marking out a target voiceprint fragment conforming to a target voiceprint template, judging whether a fuzzy fragment region exists or not, determining whether context semantic completion is performed or not according to a judgment result, and outputting a recognition result of final complete voiceprint information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio recognition, and specifically to an intelligent deep learning method and system for animal audio voiceprint recognition. Background Art

[0002] In the field of wild animal monitoring and protection, the recognition and analysis of animal voiceprints is an important technology. Traditional animal voiceprint recognition methods mainly rely on manual recognition and simple signal processing techniques, and these methods have the following limitations: 1. Manual recognition is inefficient, easily affected by subjective factors, and cannot process a large amount of audio data; 2. Simple signal processing techniques are difficult to extract the complex features of animal voiceprints, resulting in low recognition accuracy; 3. Lack of an effective voiceprint feature evaluation and screening mechanism, making the quality of the modeling data set not high and affecting the final recognition effect. Summary of the Invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide an intelligent deep learning method and system for animal audio voiceprint recognition.

[0004] The purpose of the present invention can be achieved through the following technical solutions: An intelligent deep learning method for animal audio voiceprint recognition includes the following steps: Step S1: Set up distributed acquisition terminals, collect animal audio at multiple monitoring nodes in the target area, process the animal audio into standardized audio information packets, filter out non-target sound sources in the standardized audio information packets, and extract the target sound sources; Step S2: Perform multi-modal feature extraction on the target sound sources to obtain corresponding multi-modal voiceprint features, evaluate the confidence of the multi-modal voiceprint features, and obtain all multi-modal voiceprint features that meet the confidence screening threshold as the modeling data set; Step S3: Build a hybrid deep learning model, input the modeling data set for voiceprint modeling, and then generate a voiceprint feature library. Set the target voiceprint template of the voiceprint feature library, input the animal audio voiceprint to be recognized into the voiceprint feature library, calculate the voiceprint similarity, and mark the target voiceprint segments that meet the target voiceprint template; Step S4: Determine whether there is a fuzzy segment area in the target voiceprint segment, and decide whether to perform context semantic completion according to the judgment result, and output the recognition result of the final complete voiceprint information.

[0005] Furthermore, the process of setting up distributed acquisition terminals, collecting animal audio at multiple monitoring nodes in the target area, and processing the animal audio into standardized audio information packets includes: Set up a distributed acquisition terminal, which consists of several terminal recording devices. Select multiple monitoring nodes within the target area and deploy a terminal recording device at each monitoring node to collect animal audio through the terminal recording device at each monitoring node; Integrate the animal audio at all monitoring nodes into a preset blank information packet, convert the blank information packet into an audio information packet corresponding to the stored animal audio, and convert it into a standardized audio information packet after audio noise reduction and format standardization for the audio information packet.

[0006] Furthermore, filtering out non-target sound sources from the standardized audio information packet, the process of extracting target sound sources includes: Construct a target sound source feature parameter library and a non-target sound source feature template library; Obtain the sound source spectra of the target sound source and non-target sound sources respectively, perform probability modeling for the target sound source and non-target sound sources based on the Gaussian mixture model, and generate their respective spectral analysis models; Perform multi-channel blind source separation on the standardized audio information packet, apply the independent component analysis algorithm to decompose the standardized audio information packet into several independent sound sources, calculate the matching degree between each independent sound source and the target sound source feature parameter library through the spectral analysis model of the target sound source, and record it as S-match. Set a matching degree threshold and record it as T-match; When S-match ≥ T-match, determine the corresponding independent sound source as the target sound source, and extract the current target sound source in the standardized audio information packet; When S-match < T-match, determine the corresponding independent sound source as a non-target sound source, and filter out the current non-target sound source in the standardized audio information packet.

[0007] Furthermore, the process of extracting multi-modal features of the target sound source to obtain the corresponding multi-modal voiceprint features includes: The multi-modal feature extraction performed on the target sound source includes time-domain feature extraction, frequency-domain feature extraction, and semantic feature extraction; Obtain the time-domain waveform information corresponding to the target sound source through time-domain feature extraction, and synchronously construct a time-domain waveform diagram corresponding to the target sound source according to the time-domain waveform information. The time-domain waveform diagram is used to record the voiceprint time-domain features of the target sound source; Obtain the frequency-domain spectral line information corresponding to the target sound source through frequency-domain feature extraction, and synchronously construct a frequency-domain spectral line diagram corresponding to the target sound source according to the frequency-domain spectral line information. The frequency-domain spectral line diagram is used to record the voiceprint frequency-domain features of the target sound source; Obtain the semantic segment features corresponding to the target sound source through semantic feature extraction; Summarize the voiceprint time-domain features, voiceprint frequency-domain features, and semantic segment features of the target sound source as the multi-modal voiceprint features of the target sound source.

[0008] Furthermore, the process of performing confidence evaluation on the multimodal voiceprint features and obtaining all multimodal voiceprint features that meet the confidence screening threshold as a modeling data set includes: Define the confidence evaluation indexes of the voiceprint time domain features, voiceprint frequency domain features and semantic segment features, generate the corresponding confidence coefficients according to the confidence evaluation indexes, perform weighted fusion on the confidence coefficients of the voiceprint time domain features, voiceprint frequency domain features and semantic segment features, and obtain the weighted comprehensive confidence. Based on the relationship between the weighted comprehensive confidence and the preset confidence screening threshold, decide whether to use the multimodal voiceprint features as the modeling data set. The weighted comprehensive confidence is denoted as , and the confidence screening threshold is denoted as ; Will ≥ The multimodal voiceprint feature screening is used as the modeling data set, and the < Multimodal voiceprint features.

[0009] Furthermore, a hybrid deep learning model is constructed, and a modeling data set is input to perform voiceprint modeling, thereby generating a voiceprint feature library. The process of setting a target voiceprint template of the voiceprint feature library includes: Select GRU recurrent neural network as the model architecture, supplement the GRU recurrent neural network architecture with LSTM, build model gating, and then build a preliminary hybrid deep learning model; Construct the loss function corresponding to the preliminary hybrid deep learning model, train the preliminary hybrid deep learning model through different loss functions, and establish the final hybrid deep learning model after the model evaluation indicators meet expectations; Input the modeling data set into the hybrid deep learning model. After the hybrid deep learning model analyzes the modeling data set, it constructs the data hierarchy and data index corresponding to the animal voiceprint, performs voiceprint modeling according to the data hierarchy and data index, and then constructs a voiceprint feature library for storing animal voiceprints. For the voiceprints of known species, the mean of all voiceprint embedding vectors corresponding to the voiceprints of known species is extracted as the target voiceprint template of the known species. The template confidence of the target voiceprint template corresponding to the known species is counted, and the confidence threshold is preset. If the template confidence is lower than the confidence threshold, the target voiceprint template of the corresponding known species is reconstructed, otherwise, no operation is performed.

[0010] Furthermore, the process of recording the audio voiceprint of the animal to be identified into the voiceprint feature library, calculating the voiceprint similarity, and marking the target voiceprint segment that meets the target voiceprint template includes: Select the animal audio voiceprint to be recognized and input it into the voiceprint feature library. Calculate the similarity with each target voiceprint template associated with the voiceprint feature library for the animal audio voiceprint to be recognized in turn, obtain the template vector corresponding to each target voiceprint template, and obtain the voiceprint vector corresponding to the animal audio voiceprint to be recognized. According to the template vector and the voiceprint vector, obtain the cosine similarity between the currently recognized animal audio voiceprint and each target voiceprint template, and use it as the voiceprint similarity. For each target voiceprint template, mark the part of the animal audio voiceprint with a voiceprint similarity greater than the preset similarity threshold as the target voiceprint segment that conforms to the corresponding target voiceprint template, and do not process the part of the animal audio voiceprint with a voiceprint similarity less than or equal to the similarity threshold.

[0011] Further, the process of judging whether there is a fuzzy segment area in the target voiceprint segment and deciding whether to perform context semantic completion according to the judgment result and outputting the recognition result of the final complete voiceprint information includes: Set fuzzy judgment indicators; The fuzzy judgment indicators include time domain fuzzy indicators, frequency domain fuzzy indicators, and semantic fuzzy indicators; If the target voiceprint segment does not meet at least any one of the time domain fuzzy indicator, frequency domain fuzzy indicator, or semantic fuzzy indicator, it is judged that there is a fuzzy segment area in the target voiceprint segment, and it is decided to perform context semantic completion on the target voiceprint segment. Otherwise, it is judged that there is no fuzzy segment area in the target voiceprint segment, and no context semantic completion is performed; When the context semantic completion of the fuzzy segment area corresponding to the target voiceprint segment is completed, output the recognition result of the final complete voiceprint information of the target voiceprint segment.

[0012] Further, an intelligent deep learning system for animal audio voiceprint recognition, the system includes: Animal audio acquisition and preprocessing module, set distributed acquisition terminals, collect animal audio at multiple monitoring nodes in the target area, process the animal audio into standardized audio information packets, filter out non-target sound sources in the standardized audio information packets, and extract target sound sources; Voiceprint feature extraction and confidence evaluation module, perform multi-modal feature extraction on the target sound source to obtain corresponding multi-modal voiceprint features, perform confidence evaluation on the multi-modal voiceprint features, and obtain all multi-modal voiceprint features that meet the confidence screening threshold as the modeling data set; Voiceprint modeling and target voiceprint matching module, build a hybrid deep learning model, input the modeling data set for voiceprint modeling, and then generate a voiceprint feature library, set the target voiceprint template of the voiceprint feature library, input the animal audio voiceprint to be recognized into the voiceprint feature library, calculate the voiceprint similarity, and mark the target voiceprint segment that conforms to the target voiceprint template; The voiceprint blur judgment and repair module determines whether there is a blurred segment area in the target voiceprint segment, decides whether to perform context semantic completion according to the judgment result, and outputs the recognition result of the final complete voiceprint information.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: Through multiple monitoring nodes arranged by the distributed acquisition terminal, large-scale and real-time acquisition of animal audio in the target area is realized, improving the efficiency and coverage of data acquisition, processing of standardized audio information packets and extraction of target sound sources, effectively filtering out non-target sound sources, reducing noise interference in subsequent processing, and improving the accuracy of voiceprint recognition; By adopting multi-modal feature extraction technology, it can comprehensively capture features of multiple modalities in the time domain, frequency domain and semantic aspects of animal voiceprints, enriching the dimension of voiceprint information and improving the representation ability of voiceprint features; The confidence evaluation method ensures the quality of the modeling data set, and only voiceprint features that meet the confidence screening threshold will be used for modeling, thereby improving the generalization ability and recognition accuracy of the voiceprint model; The constructed hybrid deep learning model can more effectively learn and represent complex voiceprint features, and the generated voiceprint feature library has a high recognition rate and robustness. Description of the Drawings

[0014] Figure 1 It is a flow chart of the intelligent deep learning method for animal audio voiceprint recognition of the present invention. Detailed Embodiments

[0015] As Figure 1 shown, an intelligent deep learning method for animal audio voiceprint recognition includes the following steps: Step S1: Set up a distributed acquisition terminal, collect animal audio at multiple monitoring nodes in the target area, process the animal audio into standardized audio information packets, filter out non-target sound sources in the standardized audio information packets, and extract target sound sources; Step S2: Perform multi-modal feature extraction on the target sound source to obtain corresponding multi-modal voiceprint features, perform confidence evaluation on the multi-modal voiceprint features, and obtain all multi-modal voiceprint features that meet the confidence screening threshold as the modeling data set; Step S3: Construct a hybrid deep learning model, input the modeling data set for voiceprint modeling, then generate a voiceprint feature library, set the target voiceprint template of the voiceprint feature library, input the animal audio voiceprint to be recognized into the voiceprint feature library, calculate the voiceprint similarity, and mark the target voiceprint segment that conforms to the target voiceprint template; Step S4: Determine whether there is a blurred segment area in the target voiceprint segment, decide whether to perform context semantic completion according to the judgment result, and output the recognition result of the final complete voiceprint information.

[0016] It should be further noted that in the specific implementation process, the process of setting up a distributed acquisition terminal to collect animal audio at multiple monitoring nodes in the target area and processing the animal audio into a standardized audio information packet includes: Set up a distributed acquisition terminal, which consists of several terminal recording devices. Select multiple monitoring nodes in the target area, and deploy a terminal recording device at each monitoring node to collect animal audio through the terminal recording device at each monitoring node; Integrate the node number at each monitoring node, the device number of the terminal recording device at the monitoring node, and the position coordinates in the target area where the monitoring node is located as the identity identification information of the animal audio at the monitoring node, and record the identity identification information as ID; Then: ID = <id1, id2, P(x, y)>; Among them, id1 represents the node number of the monitoring node, id2 represents the device number of the terminal recording device at the monitoring node, P(x, y) represents the position coordinates of the monitoring node in the target area, x represents the longitude coordinate in the position coordinates, and y represents the latitude coordinate in the position coordinates; Set up a master control terminal. Each terminal recording device is communicatively connected to the master control terminal. Each terminal recording device uploads its respective real-time working data into the master control terminal. After the master control terminal analyzes the real-time working data, it locates the terminal recording device with abnormal operation; Send the position coordinates of the terminal recording device in the target area to the maintenance personnel, and the maintenance personnel perform equipment maintenance on the terminal recording device; Integrate the animal audio at all monitoring nodes into a preset blank information packet, and then convert the blank information packet into an audio information packet corresponding to the stored animal audio. Perform audio noise reduction and format standardization on the audio information packet, and then convert the audio information packet into a standardized audio information packet.

[0017] It should be further noted that in the specific implementation process, the process of filtering out non-target sound sources in the standardized audio information packet and extracting target sound sources includes: Construct a target sound source feature parameter library, which is used to store the core sound source parameters corresponding to the animal voiceprints to be recognized. Among them, the animal voiceprints to be recognized include bird song voiceprints, mammalian roar voiceprints, insect song voiceprints, fish information sound voiceprints, and amphibian call voiceprints; The core sound source parameters of the animal voiceprint include fundamental frequency range, harmonic structure, time-domain energy distribution, and duration. The core sound source parameters are used to characterize the relevant feature information of the voiceprints corresponding to different types of animals; Construct a non-target sound source feature template library, which is used to record the sound source feature templates corresponding to different types of background noises. The background noises are interference voiceprints that do not need to be recognized, and the types of background noises include environmental noises and human activity noises; Among them, environmental noises include wind sounds, water flow sounds, thunder sounds, and rain sounds, etc.; Human activity noises include mechanical operation sounds, human voices, and musical instrument sounds; Obtain the sound source spectra corresponding to the target sound source and the non-target sound source respectively, perform probability modeling on the target sound source and the non-target sound source based on the Gaussian mixture model, and generate the spectral analysis models for the target sound source and the non-target sound source respectively; Bind the spectral analysis model of the target sound source with the target sound source feature parameter library; Bind the spectral analysis model of the non-target sound source with the non-target sound source feature template library; Perform multi-channel blind source separation on the standardized audio information packet, apply the independent component analysis algorithm to decompose the standardized audio information packet into several independent sound sources, calculate the matching degree between each independent sound source and the target sound source feature parameter library through the spectral analysis model of the target sound source, and record the matching degree as S-match. Set a matching degree threshold and record the matching degree threshold as T-match; When S-match ≥ T-match, determine the corresponding independent sound source as the target sound source, and extract the current target sound source in the standardized audio information packet; When S-match < T-match, determine the corresponding independent sound source as the non-target sound source, and filter out the current non-target sound source in the standardized audio information packet.

[0018] It should be further noted that in the specific implementation process, the process of performing multi-modal feature extraction on the target sound source to obtain the corresponding multi-modal voiceprint features and performing confidence evaluation on the multi-modal voiceprint features to obtain all multi-modal voiceprint features that meet the confidence screening threshold as the modeling data set includes: The multi-modal feature extraction performed on the target sound source includes time-domain feature extraction, frequency-domain feature extraction, and semantic feature extraction; Obtain the time-domain waveform information corresponding to the target sound source through time-domain feature extraction, and simultaneously construct the time-domain waveform diagram corresponding to the target sound source according to the time-domain waveform information. The time-domain waveform diagram is used to record the voiceprint time-domain features of the target sound source; Obtain the frequency-domain spectral line information corresponding to the target sound source through frequency-domain feature extraction, and simultaneously construct the frequency-domain spectral line diagram corresponding to the target sound source according to the frequency-domain spectral line information. The frequency-domain spectral line diagram is used to record the voiceprint frequency-domain features of the target sound source; Integrate the time-domain waveform diagram and the frequency-domain spectral line diagram to construct the feature spectral diagram of the target sound source; The horizontal axis of the characteristic spectrogram represents frequency, and the vertical axis represents amplitude (or energy); Through the spectrogram, the distribution of different frequency components in the target sound source can be visually seen. In animal audio recognition, the calls of different species of animals often have different frequency distribution characteristics. Therefore, the characteristic spectrogram is an important basis for distinguishing the calls of different animals; The spectrogram is also used to show the change of frequency components in the target sound source over time. For animal calls, their spectral characteristics may change as the call continues, such as the frequency rising, falling, or remaining unchanged, etc.; these changes are of great significance for identifying the species and behaviors of animals.

[0019] The content of time-domain feature extraction is as follows: The target sound source is segmented into several time-domain frames, and all the time-domain frames that are entirely in the high-frequency band are marked, and the time-domain frames in the high-frequency band are pre-emphasized. The purpose of pre-emphasis is to compensate for the attenuation of the high-frequency components of the target sound source during transmission; Each time-domain frame is windowed by applying a window function. The types of window functions include Hamming window, rectangular window, Hanning window, and Blackman window. The frame signal corresponding to each time-domain frame is extracted, and the window function is multiplied by each frame signal. Among them, the window length of the window function is equal to the signal length of the frame signal being multiplied; The window value of the window function at the middle position of the frame signal is set to the maximum value; An attenuation relationship formula of the window function from the maximum value at the middle position of the frame signal to the minimum value at the frame end point is constructed; By windowing with the window function, the truncation effect generated by the frame signal corresponding to the time-domain frame at the frame edge can be reduced, and the occurrence of spectral leakage can be prevented; Feature calculation is performed on each time-domain frame, and then the amplitude feature, energy feature, duration feature, sound start point, waveform feature, harmonic structure feature, and envelope feature corresponding to each time-domain frame are obtained. These are integrated as the voiceprint time-domain features corresponding to the target sound source, and the time-domain waveform information of the change of the voiceprint time-domain features over time is constructed; The content of frequency-domain feature extraction is as follows: The target sound source is segmented and converted into several frequency-domain signals according to the frequency band distribution. The frequency-domain signals are pre-emphasized and windowed in the same way as the time-domain frames are processed, and then the frequency-domain signals are converted into several frequency-domain frames. Fast Fourier transform is performed on each frequency-domain frame to obtain the amplitude and phase information of the frequency components corresponding to each frequency-domain frame; The spectral centroid, spectral bandwidth, spectral flatness, spectral roll-off, spectral flux, and frequency-domain peak corresponding to each frequency-domain frame are extracted, and after integration, they are used as the voiceprint frequency-domain features of the target sound source, and the frequency-domain spectral line information of the change of the voiceprint frequency-domain features over frequency is constructed; Obtain the semantic segment features corresponding to the target sound source through semantic feature extraction; The content of semantic feature extraction is as follows: Based on a pre-trained deep learning model, extract high-level semantic embedding vectors from the target sound source. The high-level semantic embedding vectors are used to represent the behavioral intentions of animal vocalizations, and the behavioral intentions include courtship or alertness; Combine an acoustic event detection algorithm to label the semantic labels of the audio segments corresponding to the animal vocalizations. The semantic labels are used to record the acoustic types of the audio segments, and the acoustic types include long calls and short calls; Integrate the high-level semantic embedding vectors and acoustic types obtained corresponding to the animal vocalizations as semantic segment features; Summarize the vocalization time-domain features, vocalization frequency-domain features, and semantic segment features of the target sound source as the multi-modal vocalization features of the target sound source; The content of the confidence evaluation of the multi-modal vocalization features is as follows: Define the confidence evaluation indicators corresponding to the vocalization time-domain features, vocalization frequency-domain features, and semantic segment features included in the multi-modal vocalization features, and generate the corresponding confidence coefficients according to the confidence evaluation indicators; Fuse the confidence coefficients of the vocalization time-domain features, vocalization frequency-domain features, and semantic segment features respectively through weighting to obtain a weighted comprehensive confidence, and determine whether to use the multi-modal vocalization features as a modeling data set according to the size relationship between the weighted comprehensive confidence and a preset confidence screening threshold; The confidence evaluation indicators of the vocalization time-domain features include signal-to-noise ratio, time-domain energy variance, and zero-crossing rate; The confidence coefficient corresponding to the vocalization time-domain features is denoted as , The calculation formula of ; Among them, SNR represents the value of the signal-to-noise ratio, is the weight coefficient for calculating the confidence coefficient of the vocalization time-domain features, which is adjusted according to the actual data situation; The confidence evaluation indicators of the vocalization frequency-domain features include spectral clarity, spectral smoothness, and harmonic structure integrity; The confidence coefficient corresponding to the vocalization frequency-domain features is denoted as , The calculation formula of ; Among them, is the weight coefficient for calculating the confidence coefficient of the vocalization frequency-domain features, and its weight coefficient is optimized according to experiments; The confidence evaluation indicators of the semantic segment features include the label consistency rate between the semantic label and the acoustic event label and the model confidence probability of the pre-trained deep learning model; The confidence coefficient of the semantic segment feature is denoted as , , and its calculation formula is as follows: ; Among them, and are the weight coefficients for calculating the confidence coefficient of the semantic segment feature. When the semantic label completely matches the acoustic event, the label consistency rate is 1. When the semantic label is completely irrelevant to the acoustic event, the label consistency rate is 0. The value range of the model confidence probability is (0, 1); The content of weighted fusion of the confidence coefficient is as follows: First, normalize the confidence coefficients of the voiceprint time-domain feature, voiceprint frequency-domain feature, and semantic segment feature to the interval [0, 1], and then assign the value weights of the voiceprint time-domain feature, voiceprint frequency-domain feature, and semantic segment feature respectively, so as to obtain the weighted comprehensive confidence, and denote the weighted comprehensive confidence as , , and its calculation is as follows: ; , and are the value weights assigned to the voiceprint time-domain feature, voiceprint frequency-domain feature, and semantic segment feature respectively. The value weights can be assigned according to the importance of different modal features. For example: = 0.5, = 0.3 and = 0.2; Among them, ; Preset the confidence screening threshold and denote the confidence screening threshold as ; Use the multi-modal voiceprint feature screening with ≥ as the modeling data set, and eliminate the multi-modal voiceprint features with < .

[0020] It should be further noted that in the specific implementation process, a hybrid deep learning model is constructed, and the modeling data set is input for voiceprint modeling, and then a voiceprint feature library is generated. The process of setting the target voiceprint template of the voiceprint feature library includes: Select the GRU recurrent neural network as the model architecture, supplement the architecture of the GRU recurrent neural network through LSTM, construct the model gates corresponding to the model architecture in the initial state, and the types of the model gates include reset gates, update gates, input gates, and forget gates, so as to construct a preliminary hybrid deep learning model; It should be noted that the functions of the reset gate, update gate, input gate, and forget gate are as follows: Reset gate: Controls the degree to which the state information of the previous moment should be forgotten and determines how much past information needs to be forgotten at the current moment; Update gate: Determines the influence degree of the state information of the previous moment on the current moment and controls how much memory of the previous moment should be retained in the current hidden state; Input gate: It determines how much of the current input is saved into the new cell state; Forget gate: It determines how much of the cell state of the previous moment is retained at the current moment.

[0021] Construct the loss function corresponding to the preliminary hybrid deep learning model. The loss function includes a contrastive loss function, a triplet loss function, and a cross-entropy loss function. Train the preliminary hybrid deep learning model through different loss functions until the model evaluation metrics meet the expectations, and then establish the final hybrid deep learning model; Input the modeling dataset into the hybrid deep learning model. After the hybrid deep learning model analyzes the modeling dataset, construct the data hierarchy and data index corresponding to the animal vocal fingerprint, and perform vocal fingerprint modeling according to the data hierarchy and data index, and then construct a vocal fingerprint feature library for storing animal vocal fingerprints; Among them, the data hierarchy includes a species layer, an individual layer, and a behavior layer; The species layer classifies the major categories of animals according to animal taxonomy (such as Aves, Mammalia), and the individual layer targets a certain tracked individual (such as a whale wearing a tracking tag), which is used to store its vocal fingerprint features and the timestamp when the vocal fingerprint features are obtained. The behavior layer is used to associate the vocal fingerprint features with behavior labels (such as courtship, alert); The data index is used as the retrieval identifier of a certain animal vocal fingerprint, and supports using the nearest neighbor algorithm to establish the feature vector index of the animal vocal fingerprint for fast retrieval of the animal vocal fingerprint; For the known species vocal fingerprint, extract the mean value of all vocal fingerprint embedding vectors corresponding to the known species vocal fingerprint as the target vocal fingerprint template of the known species, and statistically calculate the template confidence of the target vocal fingerprint template corresponding to the known species. Preset the confidence threshold. If the template confidence is lower than the confidence threshold, reconstruct the target vocal fingerprint template of the corresponding known species, otherwise, do nothing.

[0022] It should be further noted that in the specific implementation process, the process of inputting the animal audio vocal fingerprint to be recognized into the vocal fingerprint feature library, calculating the vocal fingerprint similarity, and marking the target vocal fingerprint segment that conforms to the target vocal fingerprint template includes: Select the animal audio vocal fingerprint to be recognized and input it into the vocal fingerprint feature library. Calculate the similarity with each target vocal fingerprint template associated with the vocal fingerprint feature library in turn for the animal audio vocal fingerprint to be recognized, obtain the template vector corresponding to each target vocal fingerprint template, and obtain the vocal fingerprint vector corresponding to the animal audio vocal fingerprint to be recognized; Furthermore, based on the template vector and the voiceprint vector, the cosine similarity between the currently recognized animal audio voiceprint and each target voiceprint template is obtained, and the cosine similarity is denoted as , and the calculation formula of the cosine similarity is as follows: ; where is the template vector, is the voiceprint vector; The cosine similarity is used as the voiceprint similarity between the animal audio voiceprint and each target voiceprint template. For each target voiceprint template, the part of the animal audio voiceprint with a voiceprint similarity greater than the preset similarity threshold is marked as the target voiceprint segment that conforms to the corresponding target voiceprint template, and no processing is performed on the part of the animal audio voiceprint with a voiceprint similarity less than or equal to the similarity threshold.

[0023] It should be further noted that in the specific implementation process, the process of determining whether there is a fuzzy segment area in the target voiceprint segment, deciding whether to perform context semantic completion according to the judgment result, and outputting the recognition result of the final complete voiceprint information includes: Set fuzzy judgment indicators; The fuzzy judgment indicators include time-domain fuzzy indicators, frequency-domain fuzzy indicators, and semantic fuzzy indicators; If the target voiceprint segment does not meet at least any one of the time-domain fuzzy indicator, frequency-domain fuzzy indicator, or semantic fuzzy indicator, it is determined that there is a fuzzy segment area in the target voiceprint segment, and it is decided to perform context semantic completion on the target voiceprint segment. Otherwise, it is determined that there is no fuzzy segment area in the target voiceprint segment, and no context semantic completion is performed; Perform fuzzy grading on the fuzzy segment area of the target voiceprint segment; When there is only single-modal fuzziness in the fuzzy segment area of the target voiceprint segment, and the proportion of the clear segment in the context of the target voiceprint segment is greater than or equal to 70%, it is marked that the fuzzy area is in mild fuzziness, and the automatic completion mechanism is triggered to perform context semantic completion on the fuzzy segment area of the target voiceprint segment; When the fuzzy segment area of the target voiceprint segment is multi-modal fuzzy, or the proportion of the clear segment in the context of the target voiceprint segment is less than 50%, it is marked that the fuzzy area is in severe fuzziness, and the manual annotation mechanism is triggered to perform context semantic completion on the fuzzy segment area of the target voiceprint segment; Explanation of single-modal fuzziness: The target voiceprint segment does not meet any one of the time-domain fuzzy indicator or frequency-domain fuzzy indicator; Explanation of multi-modal fuzziness: The target voiceprint segment does not meet at least any two of the time-domain fuzzy indicator, frequency-domain fuzzy indicator, and semantic fuzzy indicator; Explanation of multi-modal fuzziness: After completing the context semantic completion of the fuzzy segment area corresponding to the target voiceprint segment, output the recognition result of the final complete voiceprint information of the target voiceprint segment, that is, the voice content of the specific behavior of the specific species individual corresponding to the target voiceprint segment.

[0024] The present invention also provides an intelligent deep learning system for animal audio voiceprint recognition, which system includes: An animal audio acquisition and preprocessing module, which sets distributed acquisition terminals, acquires animal audio at multiple monitoring nodes in the target area, processes the animal audio into standardized audio information packets, filters out non-target sound sources in the standardized audio information packets, and extracts target sound sources; A voiceprint feature extraction and confidence evaluation module, which performs multi-modal feature extraction on the target sound source to obtain corresponding multi-modal voiceprint features, performs confidence evaluation on the multi-modal voiceprint features, and obtains all multi-modal voiceprint features that meet the confidence screening threshold as the modeling data set; A voiceprint modeling and target voiceprint matching module, which constructs a hybrid deep learning model, inputs the modeling data set for voiceprint modeling, then generates a voiceprint feature library, sets the target voiceprint template of the voiceprint feature library, inputs the animal audio voiceprint to be recognized into the voiceprint feature library, calculates the voiceprint similarity, and marks the target voiceprint segment that conforms to the target voiceprint template; A voiceprint fuzzy judgment and repair module, which judges whether there is a fuzzy segment area in the target voiceprint segment, decides whether to perform context semantic completion according to the judgment result, and outputs the recognition result of the final complete voiceprint information.

[0025] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. An intelligent deep learning method for animal audio voiceprint recognition, characterized in that: The following steps are involved: Step S1: Setting a distributed acquisition terminal to collect animal audio at multiple monitoring nodes in a target area, and processing the animal audio into a standardized audio information package, filtering out non-target sound sources in the standardized audio information package, and extracting the target sound source; Step S2: extract multimodal features of the target sound source to obtain corresponding multimodal voiceprint features, perform confidence assessment on the multimodal voiceprint features, and obtain all multimodal voiceprint features that meet the confidence screening threshold as a modeling data set; Step S3: construct a hybrid deep learning model, and input the modeling data set to perform voiceprint modeling, and then generate a voiceprint feature library, set the target voiceprint template of the voiceprint feature library, enter the animal audio voiceprint to be identified into the voiceprint feature library, calculate the voiceprint similarity, and mark the target voiceprint segment that meets the target voiceprint template; Step S4: Determine whether there is a fuzzy segment area in the target voiceprint segment, decide whether to perform context semantic completion based on the determination result, and output the recognition result of the final complete voiceprint information.

2. The intelligent deep learning method for animal audio voiceprint recognition according to claim 1, characterized in that: The process of setting up distributed acquisition terminals, collecting animal audio at multiple monitoring nodes in the target area, and processing the animal audio into standardized audio information packets includes: Set up a distributed collection terminal, which consists of several terminal recording devices. Select multiple monitoring nodes in the target area, and deploy a terminal recording device at each monitoring node to collect animal audio through the terminal recording device at each monitoring node. Integrate the animal audio at all monitoring nodes into a preset blank information package, convert the blank information package into an audio information package corresponding to the animal audio, perform audio noise reduction and format standardization on the audio information package, and then convert it into a standardized audio information package.

3. The intelligent deep learning method for animal audio voiceprint recognition according to claim 2, characterized in that: The process of filtering out non-target sound sources in the standardized audio information package and extracting the target sound source includes: Construct a target sound source feature parameter library and a non-target sound source feature template library; Acquire the sound source spectra of the target sound source and the non-target sound source, perform probability modeling for the target sound source and the non-target sound source based on the Gaussian mixture model, and generate their respective spectrum analysis models; Perform multi-channel blind source separation on the standardized audio information package, apply the independent component analysis algorithm to decompose the standardized audio information package into several independent sound sources, calculate the matching degree between each independent sound source and the target sound source feature parameter library through the spectrum analysis model of the target sound source, and record it as S-match, set the matching degree threshold, and record it as T-match; When S-match≥T-match, the corresponding independent sound source is determined to be the target sound source, and the current target sound source in the standardized audio information package is extracted; When S-match<T-match, the corresponding independent sound source is determined to be a non-target sound source, and the current non-target sound source in the standardized audio information package is filtered out.

4. The intelligent deep learning method for animal audio voiceprint recognition according to claim 3, characterized in that: The process of extracting multimodal features of the target sound source and obtaining the corresponding multimodal voiceprint features includes: The multimodal feature extraction of the target sound source includes time domain feature extraction, frequency domain feature extraction and semantic feature extraction; The time domain waveform information corresponding to the target sound source is obtained by extracting the time domain features, and a time domain waveform graph corresponding to the target sound source is constructed synchronously based on the time domain waveform information. The time domain waveform graph is used to record the time domain features of the voiceprint of the target sound source. The frequency domain spectral line information corresponding to the target sound source is obtained by frequency domain feature extraction, and the frequency domain spectral line graph corresponding to the target sound source is simultaneously constructed based on the frequency domain spectral line information. The frequency domain spectral line graph is used to record the frequency domain characteristics of the voiceprint of the target sound source; The semantic segment features corresponding to the target sound source are obtained by extracting semantic features; The voiceprint time domain features, voiceprint frequency domain features and semantic segment features of the target sound source are summarized as the multimodal voiceprint features of the target sound source.

5. The intelligent deep learning method for animal audio voiceprint recognition according to claim 4, characterized in that: The process of performing confidence evaluation on multimodal voiceprint features and obtaining all multimodal voiceprint features that meet the confidence screening threshold as a modeling data set includes: Define the confidence evaluation indexes of the voiceprint time domain features, voiceprint frequency domain features and semantic segment features, generate the corresponding confidence coefficients according to the confidence evaluation indexes, perform weighted fusion on the confidence coefficients of the voiceprint time domain features, voiceprint frequency domain features and semantic segment features, and obtain the weighted comprehensive confidence. Based on the relationship between the weighted comprehensive confidence and the preset confidence screening threshold, decide whether to use the multimodal voiceprint features as the modeling data set. The weighted comprehensive confidence is denoted as , and the confidence screening threshold is denoted as ; Will ≥ The multimodal voiceprint feature screening is used as the modeling data set, and the < Multimodal voiceprint features.

6. The intelligent deep learning method for animal audio voiceprint recognition according to claim 5, characterized in that: Construct a hybrid deep learning model, input the modeling data set for voiceprint modeling, and then generate a voiceprint feature library. The process of setting the target voiceprint template of the voiceprint feature library includes: Select GRU recurrent neural network as the model architecture, supplement the GRU recurrent neural network architecture with LSTM, build model gating, and then build a preliminary hybrid deep learning model; Construct the loss function corresponding to the preliminary hybrid deep learning model, train the preliminary hybrid deep learning model through different loss functions, and establish the final hybrid deep learning model after the model evaluation indicators meet expectations; Input the modeling data set into the hybrid deep learning model. After the hybrid deep learning model analyzes the modeling data set, it constructs the data hierarchy and data index corresponding to the animal voiceprint, performs voiceprint modeling according to the data hierarchy and data index, and then constructs a voiceprint feature library for storing animal voiceprints. For the voiceprints of known species, the mean of all voiceprint embedding vectors corresponding to the voiceprints of known species is extracted as the target voiceprint template of the known species. The template confidence of the target voiceprint template corresponding to the known species is counted, and the confidence threshold is preset. If the template confidence is lower than the confidence threshold, the target voiceprint template of the corresponding known species is reconstructed, otherwise, no operation is performed.

7. The intelligent deep learning method for animal audio voiceprint recognition according to claim 6, characterized in that: The process of recording the audio voiceprint of the animal to be identified into the voiceprint feature library, calculating the voiceprint similarity, and marking the target voiceprint segment that meets the target voiceprint template includes: Select the audio voiceprint of the animal to be identified and input it into the voiceprint feature library, calculate the similarity of the audio voiceprint of the animal to be identified and each target voiceprint template associated with the voiceprint feature library in turn, obtain the template vector corresponding to each target voiceprint template, and obtain the voiceprint vector corresponding to the audio voiceprint of the animal to be identified; According to the template vector and the voiceprint vector, the cosine similarity between the currently identified animal audio voiceprint and each target voiceprint template is obtained and used as the voiceprint similarity. For each target voiceprint template, some animal audio voiceprints whose voiceprint similarity is greater than the preset similarity threshold are marked as target voiceprint fragments that meet the corresponding target voiceprint template, and no processing is performed on some animal audio voiceprints whose voiceprint similarity is less than or equal to the similarity threshold.

8. The intelligent deep learning method for animal audio voiceprint recognition according to claim 7, characterized in that: The process of judging whether there is a fuzzy segment area in the target voiceprint segment, deciding whether to perform contextual semantic completion based on the judgment result, and outputting the final recognition result of the complete voiceprint information includes: Set fuzzy judgment indicators; Fuzzy judgment indicators include time domain fuzzy indicators, frequency domain fuzzy indicators and semantic fuzzy indicators; If the target voiceprint segment does not meet at least any one of the time domain fuzziness index, the frequency domain fuzziness index or the semantic fuzziness index, it is determined that the target voiceprint segment has a fuzzy segment area, and it is decided to perform contextual semantic completion on the target voiceprint segment; otherwise, it is determined that the target voiceprint segment does not have a fuzzy segment area, and no contextual semantic completion is performed; After completing the contextual semantics completion of the fuzzy segment area corresponding to the target voiceprint segment, the recognition result of the final complete voiceprint information of the target voiceprint segment is output.

9. An intelligent deep learning system for animal audio voiceprint recognition, used to implement the intelligent deep learning method according to any one of claims 1 to 8, characterized in that: The system includes: The animal audio collection and preprocessing module sets up a distributed collection terminal to collect animal audio at multiple monitoring nodes in the target area, and processes the animal audio into a standardized audio information package, filters out non-target sound sources in the standardized audio information package, and extracts the target sound source; The voiceprint feature extraction and confidence assessment module extracts multimodal features from the target sound source, obtains the corresponding multimodal voiceprint features, performs confidence assessment on the multimodal voiceprint features, and obtains all multimodal voiceprint features that meet the confidence screening threshold as the modeling data set; The voiceprint modeling and target voiceprint matching module builds a hybrid deep learning model, inputs the modeling data set for voiceprint modeling, generates a voiceprint feature library, sets the target voiceprint template of the voiceprint feature library, enters the audio voiceprint of the animal to be identified into the voiceprint feature library, calculates the voiceprint similarity, and marks the target voiceprint segment that meets the target voiceprint template; The voiceprint fuzzy judgment and repair module determines whether there is a fuzzy segment area in the target voiceprint segment, decides whether to perform contextual semantic completion based on the judgment result, and outputs the recognition result of the final complete voiceprint information.

Citation Information

Cited By

  • Building energy-saving performance AI dynamic evaluation method

    CN121212833A