A Bird Song Species Recognition Method Based on Dual-Channel Speech Enhancement

The bird singing signal is reduced through dual-channel voice enhancement technology, combined with feature extraction and classifier, and the problem of low accuracy of bird singing recognition under low signal-to-noise ratio is solved, achieving efficient bird singing species recognition.

CN115116461BActive Publication Date: 2025-07-29NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210861205.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-07-29
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

The prior art has low accuracy in bird singing signals under low signal-to-noise ratio, resulting in large identification errors.

Method used

Using a dual-channel voice enhancement method, the improved generalized side lobe eliminater and post-filter module are used to reduce the noise of bird singing signals. Combined with feature extraction, codebook construction and dimensionality reduction technology, a training database is established and classified and identified.

Benefits of technology

By enhancing the signal-to-noise ratio, the accuracy of bird singing species recognition is significantly improved and the implementation process is simplified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116461B_ABST
    Figure CN115116461B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying bird species based on dual-channel speech enhancement. The method first performs preliminary noise reduction on the collected dual-channel bird song signals through an improved generalized sidelobe canceller with a leakage suppression and signal recovery module added, and then further reduces the noise through dual-channel post-filtering to obtain a purer desired bird song signal; the enhanced bird song signals are divided into a training set and a test set. After the training bird songs are preprocessed, feature extracted, codebook constructed, encoded, and dimension-reduced, a feature database of the training bird songs is established; for the bird songs to be identified, after preprocessing, feature extraction, encoding, dimension-reducing transformation, and classification, the identified bird species can be obtained. The present invention has strong practicability, is economical and convenient, can make full use of the information of the collected dual-channel bird song signals, and can improve the signal-to-noise ratio of the bird song signals through beamforming algorithms for signal enhancement processing, thereby improving the accuracy of bird species identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of species monitoring and voice signal recognition, and particularly relates to a bird species recognition method based on dual-channel voice enhancement. Background Art

[0002] The diversity of bird species is an important indicator of species diversity and ecological environment protection. Acoustic monitoring can monitor the activities of birds without disturbing their activities. Bioacoustic research can be a useful tool for investigating and monitoring birds, and can also evaluate the impact of human activities on bird species. In this context, monitoring bird species by applying sound recognition technology has always been a widely concerned field in biological research.

[0003] In recent years, with the research of scholars at home and abroad, many bird species recognition methods based on bird calls have emerged. Common methods include: 1) Bird call recognition methods based on template matching; the dynamic time warping algorithm is one of the most representative algorithms, which can obtain a relatively high recognition effect, but due to the problem of excessive computational complexity, the application of this method is affected; 2) Bird call recognition methods based on feature extraction; by extracting features from bird call signals, the classification of bird species can be realized. Common features include Mel frequency cepstral coefficients (MFCC), linear predictive coding (LPC) coefficients, LPC reflection coefficients, etc. After classification by classifiers such as SVM and KNN, the recognized bird species can be obtained; 3) Bird call recognition methods based on deep learning: Recently, with the rapid development of deep learning technology, bird call recognition methods based on deep learning have received more and more attention. At present, common deep learning models include convolutional neural network models such as VGG and Resnet. Due to the memory characteristics of recurrent neural networks, models such as LSTM and GRU have received more and more attention.

[0004] However, research has shown that the above methods are all greatly affected by the signal-to-noise ratio of the sample set. When the signal-to-noise ratio is relatively high, many methods can show good classification performance, but as the signal-to-noise ratio decreases, the recognition rate also decreases to varying degrees. Since various noises may be included in the collection process of bird calls, the signal-to-noise ratio of bird call segments is relatively low.

[0005] Currently, common bird call sound recorders include Song meter, voice recorders or smartphones, etc. The signals collected by these devices are dual-channel signals. In previous research, the collected dual-channel signals were usually directly converted into single-channel signals for subsequent recognition. However, since dual-channel signals contain more useful information than single-channel signals, the present invention uses dual-channel voice enhancement technology to enhance the collected bird call signals.

[0006] Therefore, a major problem with the existing technology is that when the signal-to-noise ratio of the bird song signal is low, the recognition accuracy is low, resulting in a large recognition error. Summary of the Invention

[0007] The object of the present invention is to provide a method for identifying bird species based on dual-channel speech enhancement to address the problems existing in the above-mentioned existing technology, enhance the signal-to-noise ratio of the collected bird song signal, and thus improve the recognition accuracy.

[0008] The technical solution for achieving the object of the present invention is: a method for identifying bird species based on dual-channel speech enhancement, the method comprising the following steps:

[0009] Step 1, uniformly sample the collected dual-channel bird song signal;

[0010] Step 2, filter and enhance the bird song signal obtained in Step 1 to form a sample set;

[0011] Step 3, divide the sample set into a training set and a test set;

[0012] Step 4, preprocess, extract features, construct a codebook, encode, and reduce the dimension of the bird song signal in the training set to obtain a feature data set of the training bird song;

[0013] Step 5, combine the feature data set of the training bird song obtained in Step 4, preprocess, extract features, perform VLAD encoding, reduce the dimension of the bird song signal to be recognized, and then classify it through a classifier to obtain the species of the bird song to be recognized.

[0014] Further, the filtering and enhancement processing of the bird song signal obtained in Step 1 in Step 2 specifically includes:

[0015] Step 2-1, preliminarily denoise the bird song signal using an improved generalized sidelobe canceller;

[0016] The generalized sidelobe canceller includes: a fixed beamformer for transmitting signals in the desired direction; a blocking matrix for generating a reference noise signal; an adaptive noise canceller for eliminating noise; a leakage suppression module for applying spectral gain to the output of the blocking matrix to generate a modified reference noise; a signal recovery module for adding a quantitative microphone signal to the output of the fixed beamformer;

[0017] Step 2-2, perform post-filtering processing through a post-filtering module, specifically: use the transient beam reference ratio TBRR to test the hypothesis to determine whether the input contains the desired speech signal, and then obtain the enhanced single-channel bird song signal after noise spectrum estimation and spectral enhancement estimation.

[0018] Further, Step 4 specifically includes the following steps:

[0019] Step 4-1, preprocess the bird song signals in the training set, specifically:

[0020] Step 4-1-1, first divide the input signal into windows with a window length of 2 s, and overlap 1.75 s between adjacent windows to obtain a number of windows named texture windows;

[0021] Step 4-1-2, calculate the energy P of each texture window and convert the unit to dB, and set the threshold P TH to filter out the silent windows without bird songs in the texture windows, P TH The formula for is:

[0022] P TH = P max - 20

[0023] where P max is the energy of the texture window with the maximum power in the input bird song signal;

[0024] Calculate the weight of each texture window as w:

[0025] w = max{P - P TH , 0}

[0026] Step 4-1-3, for the texture windows with weight w > 0, according to 512 sampling points per frame and 256 sample points overlapping between adjacent frames, find the frame f with the maximum energy max , take 127 frames before and after f max , if it is less than 127 frames, perform zero-padding to obtain a segment with a length of N R = 65536, denoted as the recognition window;

[0027] Step 4-1-4, perform normalization processing on the recognition window x[n], and the formula is:

[0028]

[0029] where μ is the mean of each recognition window and σ is the standard deviation;

[0030] Step 4-2, perform discrete wavelet feature extraction on the preprocessed signal, specifically including the following steps:

[0031] Step 4-2-1, perform discrete wavelet decomposition on the signal to obtain a low-frequency sub-band A1 and a high-frequency sub-band D1, and further decompose the low-frequency sub-band to obtain a low-frequency sub-band A2 and a high-frequency sub-band D2. After L times of decomposition in this way, L high-frequency sub-bands D l and a low-frequency sub-band A L , l = 1, 2,... L;

[0032] Step 4-2-2: Remove the high-frequency subband D1 that exceeds the Nyquist sampling rate, and for the L-1 high-frequency subbands D l Extract the coefficients at the same moment to form an instantaneous acoustic unit, and then perform max pooling on the coefficients of each instantaneous acoustic unit to obtain a compact representation CU t , and after performing the mean subtraction operation, obtain the feature descriptor f t ; Assume that the length of D L is N L , then t = 1, 2,..., N L , f t is given by the formula:

[0033] f t = [f t [1], f t [2],..., f t [L-1]] T

[0034] In the formula, f t [1], f t [2],..., f t [L-1] respectively represent the features extracted from the L-1 high-frequency subbands;

[0035] Step 4-3: Perform k-means clustering on the feature descriptor obtained in Step 4-2 to obtain a codebook C of size K. The formula for the codebook is:

[0036] C = {c1, c2,..., c K}

[0037] In the formula, c k is the k-th codeword, k = 1, 2,..., K;

[0038] Step 4-4: Assign all the feature descriptors f t to the τ closest codewords c k . For each codeword c k , obtain a residual vector r k :

[0039]

[0040] Then concatenate the k residual vectors into a K(L-1)-dimensional feature vector v that represents the long-term features:

[0041]

[0042] Step 4-5: Perform dimensionality reduction on the feature vector obtained in Step 4-4 to establish a feature dataset for training bird songs. The steps are as follows:

[0043] Step 4-5-1: Obtain the covariance matrix ∑ of all training bird song features, then calculate the eigenvalues of the covariance matrix and sort them in descending order; take the first d-dimensional features that account for α of the total sum of all eigenvalues. The value of d is: PCA

[0044]

[0045] The eigenvectors corresponding to the first d-dimensional eigenvalues are the PCA transformation matrix A PCA :

[0046] A PCA = [a1, a2,... a d

[0047] In the formula, a j is the eigenvector corresponding to the j-th eigenvalue λ j ;

[0048] Then, the eigenvector after PCA dimensionality reduction of the eigenvector obtained in Step 4-4 is:

[0049]

[0050] Step 4-5-2: Further reduce the dimension of the vector after PCA dimensionality reduction through LDA to obtain an eigenvector y with dimension (S - 1):

[0051]

[0052] where S is the number of bird song species, and the LDA transformation formula is:

[0053]

[0054] In the formula, S B is the between-class scatter matrix, and S W is the within-class scatter matrix.

[0055] Furthermore, in Step 5, combining the feature dataset of the training bird songs obtained in Step 4, preprocess, extract features, perform VLAD encoding, reduce the dimension on the bird song signal to be recognized, and then classify it through a classifier to obtain the species of the bird song to be recognized, including the following steps:

[0056] Step 5-1: Preprocess the bird song signal to be recognized, and the operation process is the same as that in Step 4-1;

[0057] Step 5-2: Extract discrete wavelet features from the signal processed in Step 5-1, and the operation process is the same as that in Step 4-2;

[0058] ​​Step 5-3: Encode the feature descriptors after feature extraction. The reference codebook is the one constructed for bird calls in step 4-2 training, and the operation process is the same as that in step 4-4;

[0059] Step 5-4: Perform dimensionality reduction on the encoded high-dimensional vectors. First, perform PCA transformation, and then perform LDA transformation to obtain (S-1)-dimensional low-dimensional features;

[0060] Step 5-5: For a segment of the fragment to be recognized, first divide it into W recognition windows through preprocessing, and then obtain W low-dimensional features y after the above-mentioned feature extraction, encoding, and dimensionality reduction. i Specifically, it is expressed as:

[0061] y i =[y i [1], y i [2],..., y i [S-1]] T

[0062] Then calculate the shortest distance from y i to each type of bird call in the training bird call feature dataset to obtain d i (s), i = 1, 2,... W;

[0063] Finally, calculate the shortest weighted distance of the W recognition windows, and the category with the shortest distance is the recognized bird species s id :

[0064]

[0065] where w i is the weight corresponding to the Wth recognition window.

[0066] Compared with the prior art, the significant advantages of the present invention are:

[0067] 1) Utilize the characteristic that the bird call signal collected by a common sound recorder is a dual-channel signal. Compared with a single-channel signal, the dual-channel signal contains more useful information. Through the dual-channel voice enhancement technology, noise can be effectively removed, thereby improving the signal-to-noise ratio of the signal.

[0068] 2) Divide the enhanced signal into a training set and a test set. After training, a feature database is established. The classification results of the test set show that the recognition accuracy of the enhanced signal is significantly improved.

[0069] 3) The process of the present invention is convenient and easy to implement.

[0070] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0071] Figure 1This is the flowchart of the bird call recognition method based on dual-channel speech enhancement of the present invention.

[0072] Figure 2 This is the block diagram of the improved generalized sidelobe canceller in the present invention.

[0073] Figure 3 This is the schematic diagram of the dual-channel post-filtering module.

[0074] Figure 4 This is the block diagram of the bird call recognition system in the present invention. Specific embodiments

[0075] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0076] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0077] In one embodiment, in combination with Figure 1 , the bird call species recognition method based on dual-channel speech enhancement proposed by the present invention includes the following steps:

[0078] Step 1: After performing an improved beamformer and post-filtering on the dual-channel noisy bird call signal collected, an enhanced bird call signal is obtained. In combination with Figure 2 , Step 1 is specifically as follows:

[0079] Step 1-1: In combination with the improved generalized sidelobe canceller, after the dual-channel signal is processed by a fixed beamformer and a blocking matrix, and then through a leakage suppression module and a signal recovery module, a preliminarily enhanced speech signal is obtained through an adaptive noise canceller;

[0080] Step 1-2: Then, through the post-filtering module, further noise reduction is performed to obtain the enhanced bird call signal. The formula used is:

[0081] z i (t) = x(t) + d is (t) + d it(t), i = 1, 2

[0082] Among them, x(t) is the target signal, d is (t) and d it (t) respectively correspond to the interference signals of the i-th microphone. The signal is subjected to short-time Fourier transform (STFT) to obtain the frequency-domain representation of the signal:

[0083] Z(k, l) = AX(k, l) + D s (k, l) + D t (k, l)

[0084] Among them, A = [1, 1] T , k represents the frequency bin number, and l represents the frame number.

[0085] The signal is multiplied by a fixed beamformer to obtain a reference signal in the desired direction. The formula of the fixed beamformer is:

[0086]

[0087] Among them, Δ k is the uncertainty of the arrival of the target signal.

[0088] It is multiplied by a blocking matrix to obtain a reference noise signal, and then after being processed by an adaptive noise canceller, an enhanced signal is obtained. The blocking matrix B(k) and the adaptive noise canceller H(k) are expressed as:

[0089]

[0090]

[0091] Among them, Γ s (k, l) is the PSD matrix of the input noise signal relative to the spatial correlation function.

[0092] Although the generalized sidelobe canceller shows good performance, a certain amount of steering vectors is inevitable, which causes the speech signal to leak through the blocking matrix and be distorted in the output of the fixed beamformer. Therefore, a leakage suppression module is added to suppress the leakage in the output of the blocking matrix, and a signal recovery module is added to recover the desired signal in the fixed beamformer.

[0093] The improved fixed beamformer and blocking matrix are expressed as:

[0094]

[0095] G u (k) is in a form similar to the square root Wiener filter:

[0096]

[0097] Among them, α < 1 represents a tuning parameter for attenuating the desired signal components in the output of the blocking matrix.

[0098]

[0099] Among them, X selected (k) is the selected voice component for recovering the leakage, and a gain function similar to the square root Wiener filter is applied to suppress the leakage of the voice component in the output of the blocking matrix into the output of the blocking matrix; the desired signal with attenuation is recovered by adding some components of the microphone signals back to the output of the fixed beamformer.

[0100]

[0101] Among them, β > 1 is an adjustment parameter.

[0102] The output signal Y(k) is obtained and expressed as:

[0103]

[0104] Combined with Figure 3 , the dual-channel post-filtering module detects the desired source component at the output of the beamformer, indicating whether the dominant signal is noise or the desired signal at this time, and generates an estimate of the probability of the absence of the prior signal; based on the Gaussian statistical model and the decision-directed estimator of the prior signal-to-noise ratio under signal presence uncertainty, an estimator of the signal presence probability is derived. Among them, the formula for the local non-stationarity of the beamformer and the reference signal is:

[0105]

[0106]

[0107] Among them, S is the smoothing factor, and M is the estimate of the background pseudo-stationary noise PSD.

[0108] The signal presence probability estimator is introduced as noise into the components of the PSD estimator. Finally, spectral enhancement of the beamformer output is achieved by applying the optimal modified log-spectral amplitude (OM-LSA) gain function. This gain minimizes the mean square error of the log-spectrum under signal presence uncertainty. Finally, the enhanced clean bird song signal is obtained:

[0109]

[0110] Among them, G(k, l) is the OM-LSA gain function.

[0111] Step 2: For the enhanced bird song signals obtained in Step 1, different species are randomly divided into a training set and a test set according to the ratio of 2:1 by stratified sampling. The training set is used to establish a feature database during the training phase, and the test set is used to test the classification performance of the bird song recognition system.

[0112] Step 3: Preprocess, extract features, construct a codebook, encode, and reduce the dimension of the bird song signals in the training set to obtain a feature dataset of the training bird songs, combined with Figure 4 , specifically including:

[0113] Step 3-1: Preprocess the data in the training set, including cutting the signal into 2s texture windows, and then calculating the weight of each texture window; for the texture windows with weights greater than 0, further find the energy-aggregated part of each texture window, named the recognition window; then perform a normalization operation on each recognition window to obtain a normalized vector x nor .

[0114] Step 3-2: After L-level decomposition using the Daubechies 9 / 7 filter with DWT, obtain L high-frequency subbands D l (l = 1, 2, … L) and a low-frequency subband A L . For all high-frequency subbands, group the DWT coefficients sharing common time instants together to form instantaneous acoustic units. Among them, the number of coefficients in subband D l is twice that of subband D l+1 . After max pooling, obtain a compact time-frequency feature f t .

[0115] f t = [f t [1], f t [2],..., f t [L - 1]] T

[0116] Step 3-3: Perform K-means clustering on the features extracted from all training bird songs to obtain a codebook C of size K:

[0117] C = {c1, c2,..., c K}

[0118] Step 3-4: For a series of feature descriptors f t (t = 1, 2, …, N L ) extracted from a texture window, assign each feature descriptor to the τ nearest codewords in the codebook. For each audio word c k , form a single residual vector r k :

[0119]

[0120] Aggregate K vectors into a single vector v:

[0121]

[0122] After L2 normalization, the normalized vector v = v / ||v||2 is obtained.

[0123] Steps 3 - 5: First, project the N feature vectors obtained after VLAD encoding of the training set into a low - dimensional space through PCA transformation, so that the resulting within - class matrix becomes non - singular, and then perform linear discriminant analysis (LDA) to obtain low - dimensional vectors. L

[0124] A PCA = [a1, a2,... a d

[0125] The vector x after dimensionality reduction is obtained through PCA transformation:

[0126]

[0127]

[0128] where S B is the between - class scatter matrix and S W is the within - class scatter matrix. After LDA transformation, an (S - 1) - dimensional vector is obtained, where S is the number of species.

[0129]

[0130] The test samples are pre - processed, feature - extracted, VLAD - encoded, dimension - reduced as described above, and then classified by a classifier to obtain the species of the bird song to be recognized.

[0131] Step 4: In the recognition stage, first divide the input bird song into W overlapping recognition windows, and the features of each recognition window extracted are represented as y i = [y i [1], y i [2],..., y i [S - 1]] T , (i = 1, 2,…, W); Represent all the features extracted from the training set as For each recognition window T i in the input audio segment, find the shortest Euclidean distance to each class in the s classes:

[0132]

[0133] ​​For the input bird call segment, the bird call species s to be recognized is determined by finding the class with the minimum weighted distance. id :

[0134]

[0135] where w i is the weight corresponding to each recognition window T i .

[0136] In one embodiment, a bird call species recognition system based on dual-channel speech enhancement is provided. The system includes:

[0137] A first module for uniformly sampling the acquired dual-channel bird call signals;

[0138] A second module for filtering and enhancing the bird call signals obtained by the sampling module to form a sample set;

[0139] A third module for dividing the sample set into a training set and a test set;

[0140] A fourth module for preprocessing, feature extraction, codebook construction, encoding, and dimensionality reduction of the bird call signals in the training set to obtain a feature data set of the training bird calls;

[0141] A fifth module, in combination with the feature data set of the training bird calls obtained by the fourth module, performs preprocessing, feature extraction, VLAD encoding, and dimensionality reduction on the bird call signals to be recognized, and then obtains the species of the bird call to be recognized after classification by a classifier.

[0142] For the specific limitations of the bird call species recognition system based on dual-channel speech enhancement, reference can be made to the limitations of the bird call species recognition method based on dual-channel speech enhancement in the above text, which will not be elaborated here. Each module in the above bird call species recognition system based on dual-channel speech enhancement can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0143] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0144] Step 1, uniformly sampling the acquired dual-channel bird call signals;

[0145] Step 2, filtering and enhancing the bird call signals obtained in Step 1 to form a sample set;

[0146] Step 3: Divide the sample set into a training set and a test set;

[0147] Step 4: Preprocess, extract features, construct a codebook, encode, and reduce the dimension of the bird song signals in the training set to obtain a feature data set of the training bird songs;

[0148] Step 5: Combine the feature data set of the training bird songs obtained in Step 4, preprocess, extract features, perform VLAD encoding, reduce the dimension on the bird song signals to be recognized, and then classify them through a classifier to obtain the species of the bird song signals to be recognized.

[0149] For the specific limitations of each step, reference can be made to the limitations on the bird species recognition method based on dual-channel speech enhancement in the above text, which will not be elaborated here.

[0150] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0151] Step 1: Uniformly sample the collected dual-channel bird song signals;

[0152] Step 2: Filter and enhance the bird song signals obtained in Step 1 to form a sample set;

[0153] Step 3: Divide the sample set into a training set and a test set;

[0154] Step 4: Preprocess, extract features, construct a codebook, encode, and reduce the dimension of the bird song signals in the training set to obtain a feature data set of the training bird songs;

[0155] Step 5: Combine the feature data set of the training bird songs obtained in Step 4, preprocess, extract features, perform VLAD encoding, reduce the dimension on the bird song signals to be recognized, and then classify them through a classifier to obtain the species of the bird song signals to be recognized.

[0156] For the specific limitations of each step, reference can be made to the limitations on the bird species recognition method based on dual-channel speech enhancement in the above text, which will not be elaborated here.

[0157] The present invention has strong practicability, is economical and convenient, can make full use of the information of the collected dual-channel bird song signals, and can improve the signal-to-noise ratio of the bird song signals by enhancing the signals through a beamforming algorithm, thereby improving the accuracy of bird species recognition, which is of great significance for monitoring the bird species diversity in the ecosystem and protecting rare species.

[0158] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A bird species recognition method based on dual-channel voice enhancement, characterized in that, The method includes the following steps: Step 1: Uniformly sample the collected dual-channel bird song signals; Step 2: Filter and enhance the bird song signals obtained in Step 1 to form a sample set; Step 3: Divide the sample set into a training set and a test set; Step 4: Preprocess, extract features, construct a codebook, encode, and reduce the dimension of the bird song signals in the training set to obtain a feature data set of the training bird songs; Step 5: Combine the feature data set of the training bird songs obtained in Step 4, preprocess, extract features, perform VLAD encoding, reduce the dimension of the bird song signals to be recognized, and then classify them through a classifier to obtain the species of the bird songs to be recognized; Step 4 specifically includes the following steps: Step 4-1: Preprocess the bird song signals in the training set, specifically: Step 4-1-1: First, divide the input signal into windows with a window length of 2 s, and there is an overlap of 1.75 s between adjacent windows to obtain several windows named texture windows; Step 4-1-2, calculate the energy P of each texture window, convert the unit to dB, and set the threshold P TH to filter out the silent windows without bird songs in the texture window, P TH The formula for P TH = P max - 20 Among them, P max is the energy of the texture window with the largest power in the input bird song signal; Calculate the weight of each texture window as w: w = max{P - P TH , 0} Step 4-1-3: For the texture window with weight w > 0, calculate the frame f with the maximum energy according to 512 sampling points per frame and 256 overlapping sample points between adjacent frames. max , take f max and the 127 frames before and after it. If there are less than 127 frames, pad with zeros to obtain a segment of length N R = 65536, denoted as the recognition window. Step 4-1-4: Normalize the recognition window x[n], and the formula is: where μ is the mean of each recognition window and σ is the standard deviation; Step 4-2: Perform discrete wavelet feature extraction on the signals obtained after preprocessing, specifically including the following steps: Step 4-2-1: Perform discrete wavelet decomposition on the signal to obtain a low-frequency sub-band A1 and a high-frequency sub-band D1. Further decompose the low-frequency sub-band to obtain a low-frequency sub-band A2 and a high-frequency sub-band D2. After L times of such decomposition, L high-frequency sub-bands D l and a low-frequency sub-band A L , where l = 1, 2, …, L; Step 4-2-2: Remove the high-frequency subband D1 that exceeds the Nyquist sampling rate, and for the L-1 high-frequency subbands D l Extract the coefficients at the same moment to form an instantaneous acoustic unit, and then perform max pooling on the coefficients of each instantaneous acoustic unit to obtain a compact representation CU t , and after subtracting the mean value operation, obtain the feature descriptor f t ; Assume that the length of D L is N L , then t = 1, 2,..., N L , and the formula for f t is: f t = [f t [1], f t [2],..., f t [L - 1]] T where f t [1], f t [2],..., f t [L - 1] respectively represent the features extracted from L - 1 high-frequency subbands; Step 4-3: Perform k-means clustering on the feature descriptors obtained in Step 4-2 to obtain a codebook C of size K, and the codebook formula is: C = {c1, c2,..., c K} where c k is the k-th codeword, k = 1, 2, ..., K; Step 4-4, assign all the feature descriptors f t , to the τ codewords c k that are closest in distance, and for each codeword c k , obtain a residual vector r k : Then splice the k residual vectors into a feature vector v of K(L - 1) dimensions expressing long-term features: Step 4-5: Perform dimensionality reduction processing on the feature vectors obtained in Step 4-4 to establish a feature data set of the training bird songs, and the steps are as follows: Step 4-5-1: Obtain the covariance matrix ∑ of all training bird song features, then calculate the eigenvalues of the covariance matrix and sort them in descending order; take the first d-dimensional features that account for α of the total sum of all eigenvalues. The value of d is determined by the following formula: PCA ​ The eigenvector corresponding to the first d eigenvalues is the PCA transformation matrix A PCA : A PCA = [a1, a2,... a d ​ where a j is the eigenvector corresponding to the j-th eigenvalue λ j ; Then the feature vector after PCA dimensionality reduction of the feature vectors obtained in Step 4-4 is: Step 4-5-2: Further reduce the dimension of the vector after PCA dimensionality reduction through LDA to obtain a feature vector y of (S - 1) dimensions; where S is the number of bird song species, and the LDA conversion formula is: where S B is the between-class scatter matrix, and S W is the within-class scatter matrix; In Step 5, combining the feature data set of the training bird songs obtained in Step 4, preprocessing, extracting features, performing VLAD encoding, reducing the dimension of the bird song signals to be recognized, and then classifying them through a classifier to obtain the species of the bird songs to be recognized includes the following steps: Step 5-1: Preprocess the bird song signals to be recognized, and the operation process is the same as that in Step 4-1; Step 5-2: Perform discrete wavelet feature extraction on the signals processed in Step 5-1, and the operation process is the same as that in Step 4-2; Step 5-3: Encode the feature descriptors after feature extraction, and the reference codebook is the codebook constructed for the training bird songs in Step 4-2, and the operation process is the same as that in Step 4-4; Step 5-4: Perform dimensionality reduction processing on the encoded high-dimensional vectors, first perform PCA conversion, and then perform LDA conversion to obtain low-dimensional features of (S - 1) dimensions; Step 5-5: For a segment to be recognized, first, it is preprocessed into W recognition windows, and then, after the above-mentioned feature extraction, encoding, and dimensionality reduction, W low-dimensional features y are obtained, which are specifically represented as: i , specifically represented as: y i = [y i [1], y i [2],..., y i [S - 1]] T Calculate y again i The shortest distance to each bird call in the training bird call feature dataset, obtaining d i (s), i = 1, 2, … W; Finally, calculate the shortest weighted distance of the W recognition windows, and the category with the shortest distance is the identified bird species s id : Among them, w i is the weight corresponding to the Wth recognition window.

2. The method for identifying bird species based on dual-channel voice enhancement according to claim 1, wherein In Step 1, uniform sampling is performed, the sampling rate is 44100 Hz, and the sampling precision is 16 bit.

3. The method for bird species recognition based on dual-channel voice enhancement according to claim 2, characterized in that, In Step 2, filtering and enhancing the bird song signals obtained in Step 1 specifically includes: Step 2-1: Use an improved generalized sidelobe canceller to perform preliminary noise reduction on the bird song signals; The generalized sidelobe canceller includes: a fixed beamformer for transmitting signals in the desired direction; a blocking matrix for generating a reference noise signal; an adaptive noise canceller for eliminating noise; a leakage suppression module for applying a spectral gain to the output of the blocking matrix to generate a modified reference noise; a signal recovery module for adding a quantified microphone signal to the output of the fixed beamformer; Step 2-2: Perform post-filtering processing through a post-filtering module. Specifically, use the transient beam reference ratio (TBRR) to test the hypothesis to determine whether the input contains the required speech signal, and then obtain the enhanced single-channel bird song signal after noise spectrum estimation and spectral enhancement estimation.

4. A bird species recognition system based on dual-channel speech enhancement for the method according to any one of claims 1 to 3, characterized in that, The system includes: A first module for uniformly sampling the acquired two-channel bird song signals; A second module for filtering and enhancing the bird song signals obtained by the sampling module to form a sample set; A third module for dividing the sample set into a training set and a test set; A fourth module for preprocessing, feature extraction, codebook construction, encoding, and dimensionality reduction of the bird song signals in the training set to obtain a feature data set of the training bird songs; A fifth module for preprocessing, feature extraction, VLAD encoding, and dimensionality reduction of the bird song signals to be recognized in combination with the feature data set of the training bird songs obtained by the fourth module, and then obtaining the species of the bird songs to be recognized after classification by a classifier.

5. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Birdsong recognition method and system based on spatial orientation, computer equipment and medium

    CN113314127A