Black-leaf monkey chirp monitoring method based on machine learning, medium and system
By constructing a monitoring network topology, performing signal separation and multidimensional feature extraction, and combining lightweight neural networks and graph neural networks, the problem of low recognition in black leaf monkey population monitoring was solved, achieving efficient and accurate population dynamic monitoring.
Patent Information
- Application Number
- CN202511031214.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies have low identification rates in monitoring black leaf monkey populations, affecting the accuracy of monitoring results. Traditional methods are time-consuming, labor-intensive, and greatly affected by environmental and human factors.
A machine learning-based approach was used to construct a monitoring network topology, optimize the audio acquisition path, separate the call signals, extract multi-dimensional features, and build a fusion classification model, including lightweight neural networks and graph neural networks, to identify call types and assess population density.
This method enables efficient and accurate monitoring of the black leaf monkey population, suppresses environmental noise interference, accurately depicts population dynamics, and improves the reliability and practicality of monitoring.
Smart Images

Figure CN120932675A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of sound monitoring technology, and specifically relates to a method, medium, and system for monitoring the sound of black leaf monkeys based on machine learning. Background Technology
[0002] As a rare and endangered primate, the black-leaf monkey's population distribution and dynamic changes have always been a key focus in wildlife conservation. Traditional methods for surveying black-leaf monkey populations primarily rely on manual field observations, which have several limitations. First, black-leaf monkeys inhabit dense tropical forests, making field surveys arduous and unsafe. Second, surveys require significant human and material resources, limiting both time and spatial coverage. Third, survey results are easily influenced by observer experience and weather conditions, making it difficult to objectively reflect the true state of the population. Therefore, there is an urgent need to develop a reliable and efficient black-leaf monkey population monitoring technology to accurately grasp population dynamics and provide a scientific basis for formulating effective conservation strategies.
[0003] In recent years, with the continuous advancement of acoustic monitoring technology, sound recognition-based wildlife monitoring methods have been widely applied. Existing research has shown that the vocal signals of different species possess significant individual identification characteristics and population features. By analyzing animal vocal signals, species can be accurately identified and population sizes inferred. For black leaf monkeys, their rich and diverse vocal behaviors offer unique advantages in reflecting individual status and community structure. Therefore, establishing an automated monitoring technology based on black leaf monkey vocalizations is undoubtedly a feasible and effective solution.
[0004] Existing wildlife call monitoring technologies often classify and analyze the density of wild animals through simple spectrum analysis, lacking targeted analysis based on the macroscopic characteristics of the black leaf monkey's calls, resulting in low recognition and affecting the accuracy of monitoring results. Summary of the Invention
[0005] In view of this, the present invention provides a method, medium and system for monitoring the calls of black leaf monkeys based on machine learning, which can solve the problem that the low recognition rate of the existing technology affects the accuracy of the monitoring results.
[0006] This invention is implemented as follows:
[0007] The first aspect of the present invention provides a method for monitoring the calls of black leaf monkeys based on machine learning, comprising the following steps:
[0008] S01. Construct a topology diagram of the black leaf monkey habitat monitoring network, wherein vertices in the monitoring network topology diagram represent audio acquisition devices, and edges in the monitoring network topology diagram represent communication links between the audio acquisition devices.
[0009] S02. Determine the audio acquisition path based on the geographical location information of the audio acquisition devices and the environmental noise level in the monitoring network topology diagram;
[0010] S03. Acquire the original audio signal using the audio acquisition device, and decompose the original audio signal to obtain the individual sound signal component and the superimposed sound signal component.
[0011] S04. Establish the temporal relationship matrix between the individual sound signal component and the superimposed sound signal component, and calculate the sound signal stability coefficient matrix.
[0012] S05. Based on the time sequence relationship matrix and the sound signal stability coefficient matrix, the superimposed sound signal components are separated to obtain independent sound signals;
[0013] S06. Perform signal state sharpening processing on the individual sound signal component and the independent sound signal to form a sound signal sample;
[0014] S07. The sound signal sample is subjected to feature extraction and evaluation. Mel frequency cepstral coefficient features, fundamental frequency features, and energy features are extracted. The black leaf monkey sound feature evaluation equation set is used for evaluation. The Mel frequency cepstral coefficient features, the fundamental frequency features, the energy features and the evaluation results of the black leaf monkey sound feature evaluation equation set are combined to form a feature vector.
[0015] S08. Establish a song classification model, which includes a first machine learning model, a second machine learning model, and a fusion model. The first machine learning model is used to process the Mel frequency cepstral coefficient features, the fundamental frequency features, and the energy features. The second machine learning model is used to process the evaluation results of the evaluation equation set of black leaf monkey song features. The fusion model is used to fuse the output results of the first machine learning model and the output results of the second machine learning model.
[0016] S09. The feature vector is processed using the first machine learning model to obtain the probability of sound type recognition;
[0017] S10. Calculate the contribution value of the sound signal in the feature vector to the sound type identification using the second machine learning model;
[0018] S11. Optimize and train the second machine learning model according to the loss function between the probability of sound type recognition and the contribution value of sound type recognition;
[0019] S12. Input the call types identified at different monitoring points and the distribution of the call type identification contribution values into the pre-trained black leaf monkey population density assessment model to obtain the black leaf monkey population density distribution matrix and output it.
[0020] Based on the above technical solution, the black leaf monkey call monitoring method based on machine learning of the present invention can be further improved as follows:
[0021] The evaluation equation set for the vocal characteristics of black leaf monkeys includes equations for the smoothness of frequency changes, the length of vocal duration, the melodiousness of vocalizations, and the characteristics of diurnal distribution.
[0022] Furthermore, the frequency change smoothness equation is used to quantify the smoothness of frequency changes. It takes frequency difference data between adjacent time windows as input and outputs a frequency smoothness index. The main parameters of the frequency change smoothness equation include a frequency difference threshold coefficient, a time window length coefficient, a frequency gradient weight coefficient, and a smoothness benchmark coefficient. The frequency difference threshold coefficient and the time window length coefficient are derived from the fundamental frequency feature, while the frequency gradient weight coefficient and the smoothness benchmark coefficient are derived from the energy feature.
[0023] Furthermore, the sound duration equation is used to quantify the sound duration characteristics. It takes sound duration data and energy envelope data as input and outputs a sound duration index. The main parameters of the sound duration equation include a standard duration coefficient, an energy attenuation coefficient, a time scaling coefficient, and an energy fluctuation weighting coefficient. The standard duration coefficient and the time scaling coefficient are derived from the fundamental frequency characteristics, and the energy attenuation coefficient and the energy fluctuation weighting coefficient are derived from the energy characteristics.
[0024] Furthermore, the sonic melodicity equation is used to quantify the intonation and rhythm characteristics of a sound. It takes harmonic structure data and spectral variation data as input and outputs a melodicity index. The main parameters of the sonic melodicity equation include harmonic richness coefficient, spectral fluctuation coefficient, and timbre variation coefficient. The harmonic richness coefficient is derived from the Mel frequency cepstral coefficient characteristics, and the spectral fluctuation coefficient and the timbre variation coefficient are derived from the fundamental frequency characteristics.
[0025] Furthermore, the daytime distribution characteristic equation is used to describe the changes in sound characteristics at different times within 24 hours. It takes 24-hour sound sample data as input and outputs time-varying characteristic indicators. The main parameters of the daytime distribution characteristic equation include diurnal cycle coefficient, activity peak coefficient, time period weight coefficient, and seasonal adjustment coefficient. The diurnal cycle coefficient and the activity peak coefficient are derived from the fundamental frequency characteristics, and the time period weight coefficient and the seasonal adjustment coefficient are derived from the energy characteristics.
[0026] Furthermore, the audio acquisition path includes the following steps: constructing an acquisition path graph based on the monitoring network topology graph, and using a breadth-first search algorithm to mark the passage weights between the audio acquisition devices; prioritizing according to the ambient noise level, setting monitoring areas with ambient noise levels below 30 dB as high priority, monitoring areas with ambient noise levels above 30 dB but below 50 dB as medium priority, and monitoring areas with ambient noise levels above 50 dB as low priority; performing state transition calculations in the acquisition path graph, traversing all state nodes, obtaining the optimal state transition sequence, and forming the audio acquisition path.
[0027] Furthermore, the first machine learning model, the second machine learning model, the fusion model, and the black leaf monkey population density assessment model all employ lightweight neural networks.
[0028] The audio acquisition path includes the following steps: constructing an acquisition path graph based on the monitoring network topology, and marking the passage weights between the audio acquisition devices using a breadth-first search algorithm; prioritizing monitoring areas based on the ambient noise level, setting monitoring areas with ambient noise levels below 30 dB as high priority, monitoring areas with ambient noise levels above 30 dB but below 50 dB as medium priority, and monitoring areas with ambient noise levels above 50 dB as low priority; performing state transition calculations in the acquisition path graph, traversing all state nodes, obtaining the optimal state transition sequence, and forming the audio acquisition path.
[0029] The first machine learning model uses a lightweight convolutional neural network, specifically MobileNetV3, whose structure includes one 3×3 convolutional layer, 12 inverted residual bottleneck blocks, one 1×1 convolutional layer, one average pooling layer, and one fully connected layer.
[0030] The second machine learning model uses a lightweight graph neural network, specifically GraphSAGE, whose structure includes three message passing layers, two feature transformation layers, and one fully connected layer.
[0031] The fusion model uses a lightweight attention network, specifically SENet, whose structure includes one global average pooling layer, two fully connected layers, one feature channel weight calculation module, and one feature recalibration module.
[0032] The black leaf monkey population density assessment model uses a lightweight recurrent neural network, specifically a GRU, whose structure includes one embedding layer, two GRU layers, one temporal attention layer, and one fully connected layer.
[0033] The steps for establishing the training dataset for the first machine learning model include: collecting black leaf monkey vocalization data and labeling the type of each vocalization data segment; performing noise reduction preprocessing on the vocalization data and extracting Mel frequency cepstral coefficient features, fundamental frequency features, and energy features; and dividing the extracted features into training set, validation set, and test set in an 8:1:1 ratio.
[0034] The steps for establishing the training dataset for the second machine learning model include: constructing a sound feature relationship graph, with each sound sample as a node in the graph; calculating the similarity between sound samples and constructing an adjacency matrix; and dividing the data into a training set, a validation set, and a test set in an 8:1:1 ratio.
[0035] The steps for establishing the training dataset for the fusion model include: collecting the output results of the first machine learning model and the second machine learning model; constructing a feature fusion matrix and labeling the true labels; and dividing the data into a training set, a validation set, and a test set in an 8:1:1 ratio.
[0036] The steps for establishing the training dataset for the black leaf monkey population density assessment model include: collecting field survey data on black leaf monkey population density in different regions; combining the call recognition results with geographical location information to construct spatiotemporal sequence data; and dividing the data into training set, validation set, and test set in an 8:1:1 ratio.
[0037] The training steps of the first machine learning model include: initializing the network parameters of the first machine learning model; optimizing the network parameters using the stochastic gradient descent algorithm with a learning rate of 0.001; calculating the loss value using the cross-entropy loss function; and repeating the training for 200 rounds, reducing the learning rate once every 50 rounds.
[0038] The training steps of the second machine learning model include: initializing the network parameters of the second machine learning model; optimizing the network parameters using the Adam optimization algorithm, with the learning rate set to 0.0005; calculating the loss value using the mean squared error loss function; and repeating the training for 150 rounds, reducing the learning rate once every 30 rounds.
[0039] The training steps of the fusion model include: initializing the network parameters of the fusion model; optimizing the network parameters using the Adam optimization algorithm with a learning rate of 0.0002; calculating the loss value using the weighted cross-entropy loss function; and repeating the training for 100 rounds, reducing the learning rate once every 25 rounds.
[0040] The training steps of the black leaf monkey population density assessment model include: initializing the network parameters of the black leaf monkey population density assessment model; optimizing the network parameters using the RMSprop optimization algorithm, with the learning rate set to 0.001; calculating the loss value using the mean square logarithmic error loss function; and repeating the training for 180 rounds, reducing the learning rate once every 45 rounds.
[0041] A second aspect of the present invention provides a computer-readable storage medium storing program instructions that, when executed in a computer, perform the aforementioned machine learning-based method for monitoring the calls of black leaf monkeys.
[0042] A third aspect of the present invention provides a black leaf monkey call monitoring system based on machine learning, wherein the system includes the aforementioned computer-readable storage medium, and the system is any one of a computer, a server, or a microcontroller, wherein the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0043] Compared with existing technologies, the beneficial effects of the black leaf monkey call monitoring method, medium, and system provided by this invention are as follows: This invention proposes a black leaf monkey call monitoring method based on machine learning, which can overcome the above-mentioned problems in existing technologies and achieve efficient and accurate monitoring of black leaf monkey populations. This method features innovative designs in constructing a reasonable monitoring network topology, optimizing audio acquisition paths, separating call signals, extracting multi-dimensional features, and constructing a fusion classification model. First, a monitoring network topology is constructed using a geographic information system, and the acquisition path is optimized according to the environmental noise level to ensure the acquisition of high signal-to-noise ratio raw audio signals. Then, independent component analysis and other methods are used to separate independent individual black leaf monkey calls from complex superimposed signals, significantly reducing noise interference. Next, multi-dimensional features including frequency, energy, and harmonics are extracted, and the call features are quantitatively analyzed using a black leaf monkey call feature evaluation equation system to obtain reliable feature vectors. Finally, a multi-model classification framework integrating deep learning and graph neural networks is constructed, which can not only accurately identify call types but also quantitatively evaluate the contribution of each feature to the classification results, providing a basis for subsequent population density assessment. Through the above innovative design, this method can effectively suppress the influence of complex environmental noise, overcome the difficulties caused by individual differences, accurately depict the dynamic changes of the population, greatly improve the reliability and practicality of black leaf monkey population monitoring, and solve the problem that the existing technology has low recognition and affects the accuracy of monitoring results. Attached Figure Description
[0044] Figure 1 A flowchart of the method provided by the present invention. Detailed Implementation
[0045] like Figure 1 The diagram shown is a flowchart of a machine learning-based method for monitoring the calls of black leaf monkeys provided by this invention. The method includes the following steps:
[0046] S01. Construct a topology diagram of the black leaf monkey habitat monitoring network, where vertices in the monitoring network topology diagram represent audio acquisition devices, and edges in the monitoring network topology diagram represent communication links between audio acquisition devices.
[0047] S02. Determine the audio acquisition path based on the geographical location information of the audio acquisition devices and the environmental noise level in the monitoring network topology diagram;
[0048] S03. Use an audio acquisition device to acquire the original audio signal, decompose and process the original audio signal to obtain the individual sound signal component and the superimposed sound signal component.
[0049] S04. Establish the temporal relationship matrix between the individual sound signal components and the superimposed sound signal components, and calculate the sound signal stability coefficient matrix.
[0050] S05. Based on the time sequence matrix and the sound signal stability coefficient matrix, the superimposed sound signal components are separated to obtain independent sound signals;
[0051] S06. The individual sound signal components and independent sound signals are subjected to signal state sharpening processing to form sound signal samples.
[0052] S07. Extract and evaluate features from the vocal signal samples. Extract Mel frequency cepstral coefficient features, fundamental frequency features, and energy features. Use the evaluation equation set of black leaf monkey vocal features for evaluation. Combine the Mel frequency cepstral coefficient features, fundamental frequency features, energy features, and evaluation results of the evaluation equation set of black leaf monkey vocal features to form a feature vector.
[0053] S08. Establish a song classification model, which includes a first machine learning model, a second machine learning model, and a fusion model. The first machine learning model is used to process the Mel frequency cepstral coefficient features, fundamental frequency features, and energy features. The second machine learning model is used to process the evaluation results of the evaluation equation system for the song features of black leaf monkeys. The fusion model is used to fuse the output results of the first machine learning model and the output results of the second machine learning model.
[0054] S09. The feature vector is processed using the first machine learning model to obtain the probability of sound type recognition.
[0055] S10. Calculate the contribution value of the sound signal in the feature vector to the sound type identification using the second machine learning model;
[0056] S11. Optimize and train the second machine learning model based on the loss function between the probability of sound type recognition and the contribution value of sound type recognition.
[0057] S12. Input the sound types and sound type identification contribution values identified at different monitoring points into the pre-trained black leaf monkey population density assessment model to obtain the black leaf monkey population density distribution matrix and output it.
[0058] The specific implementation methods of the above steps are described in detail below:
[0059] The specific implementation of step S01 involves constructing a monitoring network topology map based on a habitat geographic information system. First, it's necessary to acquire geographic information data about the black leaf monkey habitat, including topography, vegetation cover, and water system distribution, and determine the deployment locations of audio acquisition equipment. High-resolution remote sensing imagery is acquired using the Google Earth engine platform, combined with field survey data, to establish a three-dimensional geographic information model of the habitat. Then, based on the communication range and signal strength of the audio acquisition equipment, a minimum spanning tree algorithm is used to construct the network topology map. A communication range threshold of 800 meters and a signal strength threshold of -60 dB / mW are recommended. Finally, a graph coloring algorithm is used to allocate communication channels in the network topology map to avoid signal interference between adjacent nodes. The minimum frequency interval between communication channels should be no less than 20 MHz. The main purpose of this step is to establish a reasonable monitoring network structure to ensure the effective acquisition and transmission of audio data.
[0060] The specific implementation of step S02 involves determining the optimal audio acquisition path based on the monitoring network topology. First, an environmental acoustic modeling method is used to construct a habitat noise distribution map, considering the impact of topography, vegetation, and meteorology on sound propagation. A sound wave propagation attenuation model is used to calculate the environmental noise level at different locations, considering factors such as geometric divergence loss, atmospheric absorption loss, and ground effect loss. Then, an ant colony optimization algorithm is used to design the audio acquisition path to minimize accumulated noise interference along the path. During path optimization, the environmental noise level is used as heuristic information, and a noise threshold of 45 dB is recommended. Finally, the feasibility of the optimized acquisition path is verified to ensure that it meets hardware constraints such as device power consumption and storage capacity. The main purpose of this step is to design a scientifically sound audio acquisition scheme while considering the impact of environmental noise.
[0061] The specific implementation of step S03 involves decomposing the original audio signal. First, the acquired original audio signal undergoes preprocessing, including noise reduction and denoising. An adaptive noise cancellation algorithm using Wiener filtering is employed, with a signal-to-noise ratio improvement threshold of no less than 10 dB. Then, an empirical mode decomposition algorithm is used to decompose the preprocessed signal into several intrinsic mode functions (IMFs). Based on frequency characteristics, these IMFs are categorized into single-unit sound signal components and superimposed sound signal components. The frequency range of single-unit sound signal components is relatively concentrated, while the superimposed sound signal components exhibit significant frequency aliasing. Finally, Hilbert transform is used to perform envelope analysis on the decomposed signal components, extracting the instantaneous frequency and instantaneous amplitude features. The main purpose of this step is to decompose the complex original audio signal into basic components that facilitate subsequent processing.
[0062] The specific implementation of step S04 involves constructing a signal feature relationship matrix. First, the temporal correlation between individual sound signal components and superimposed sound signal components is calculated using the time-series relationship matrix formula. A sliding time window method is used to calculate the cross-correlation function between the signals, with a recommended window length of 256 milliseconds and an overlap rate of 50%. Then, the stability characteristics of each signal component are calculated using the sound signal stability coefficient matrix formula, considering stability indices in both frequency and amplitude dimensions. Finally, the time-series relationship matrix and the stability coefficient matrix are combined, and a weighted fusion method is used to obtain a comprehensive feature matrix. The weighting coefficients are determined using a particle swarm optimization algorithm. The main purpose of this step is to establish the correlation between signal components, providing a basis for subsequent signal separation.
[0063] The specific implementation of step S05 is based on blind source separation of the superimposed call signals using the independent component analysis algorithm. First, principal component analysis is used to pre-whiten the signals, reducing the correlation between them. Then, by combining information from the time-series relation matrix and the stability coefficient matrix, an objective function is constructed, and a fast independent component analysis algorithm is used for signal separation, with an iterative convergence threshold set to 0.0001. Finally, the separated independent call signals undergo post-processing, including dereverberation and phase correction, and Kalman filtering is used for signal enhancement. The main purpose of this step is to separate the independent black leaf monkey call signals from the superimposed signals.
[0064] The specific implementation of step S06 involves sharpening the sound signal. First, wavelet transform is used to decompose the signal into multiple scales, selecting the Dobessi 5 wavelet as the basis function and setting the decomposition level to 4. Then, nonlinear mapping is applied to the coefficients at each scale, and a soft thresholding function is used for coefficient shrinkage. The threshold is selected using the maximum likelihood estimation method. Finally, the processed wavelet coefficients are reconstructed to obtain the sharpened sound signal, and spectral subtraction is used to eliminate residual noise. The main purpose of this step is to improve the clarity and recognizability of the sound signal.
[0065] The specific implementation of step S07 involves feature extraction and evaluation of the call signal. First, Mel-frequency cepstral analysis is used to extract the acoustic features of the signal, dividing it into 32 Mel-frequency bands and extracting 13th-order Mel-frequency cepstral coefficients. Then, autocorrelation analysis is used to extract the fundamental frequency features of the signal, and fast Fourier transform is used to calculate the signal's energy characteristics. Next, the extracted features are input into a set of evaluation equations for the black-leaf monkey's call characteristics, including four equations: frequency variation smoothness, call duration, call melodicity, and diurnal distribution characteristics. Finally, all features are standardized and combined into a feature vector, and principal component analysis is used for feature dimensionality reduction. The main purpose of this step is to extract and evaluate indicators that can effectively characterize the call features of the black-leaf monkey.
[0066] The specific implementation of step S08 involves constructing a multi-model fusion sound classification framework. First, a deep neural network is used to construct the first machine learning model, with a network structure consisting of four convolutional layers and two fully connected layers, using a modified linear unit function as the activation function. Then, a random forest algorithm is used to construct the second machine learning model, with 100 decision trees and a maximum depth of 10. Finally, a fusion model based on stacked ensemble learning is designed, using logistic regression as the meta-learner to fuse the outputs of the two base models. The main objective of this step is to establish a classification model framework that can effectively utilize different feature information.
[0067] The specific implementation of step S09 involves using a first machine learning model for sound type recognition. First, the feature vector is input into a pre-trained deep neural network, and the recognition probability of each type is calculated through forward propagation. Then, a soft maximum function is used to normalize the network output, resulting in a standardized probability distribution. Finally, the cross-entropy loss function is used to evaluate the model's recognition performance, with a loss threshold set to 0.1. The main purpose of this step is to achieve preliminary sound type recognition based on acoustic features.
[0068] The specific implementation of step S10 involves using a second machine learning model to analyze feature contribution. First, the feature vectors are input into a random forest model, and the Gini importance index for each feature in call type recognition is calculated. Then, a permutation importance analysis method is used to verify the feature contribution, with 100 permutations performed. Finally, based on the feature importance ranking results, the normalized contribution value of each feature is calculated. The main purpose of this step is to quantitatively evaluate the importance of different features in call recognition.
[0069] The specific implementation of step S11 involves optimizing the second machine learning model. First, a joint loss function is constructed based on the probability of sound type recognition and feature contribution values, including a classification loss term and a feature importance loss term. Then, the model parameters are optimized using stochastic gradient descent, with the initial learning rate set to 0.001, and cosine annealing is used to adjust the learning rate. Finally, an early stopping strategy is employed to prevent overfitting; training is stopped if the loss on the validation set does not improve for five consecutive epochs. The main purpose of this step is to improve the model's generalization ability and feature utilization efficiency.
[0070] The specific implementation of step S12 involves assessing the population density distribution of black leaf monkeys. First, the call types and their contribution values identified at each monitoring point are input into a pre-trained density assessment model, constructed using a Gaussian process regression algorithm. Then, considering spatial autocorrelation, Kriging interpolation is used to spatially interpolate the population density, with a search radius set to 1000 meters. Finally, a population density distribution matrix is generated, and an adaptive threshold segmentation method is used to classify the density into five levels. The main purpose of this step is to achieve spatial distribution assessment of the black leaf monkey population density.
[0071] The calculation process or mathematical model involved in this invention will be described in detail below.
[0072] 1. Timing relationship matrix and acoustic signal stability coefficient matrix (S04):
[0073] The time series relation matrix is specifically represented as follows:
[0074]
[0075] In the formula, R ij The temporal relationship strength between the i-th signal and the j-th signal; s i (k) represents the amplitude of the i-th signal at time point k; s j (k+τ) represents the amplitude of the j-th signal at time point k+τ; N is the number of sampling points; τ is the time delay; α ij This is a correction factor, ranging from 0.1 to 0.3.
[0076] The stability coefficient matrix of the sound signal is specifically represented as follows:
[0077]
[0078] In the formula, S ij f represents the stability coefficients of the i-th and j-th signals; i ,f j The fundamental frequencies of the i-th and j-th signals are respectively; a i ,a j σ represents the amplitudes of the i-th and j-th signals, respectively;f ,σ a γ represents the standard deviation of frequency and amplitude; β1 and β2 are weighting coefficients, both ranging from 0 to 1; ij This is a correction term, with a range of 0.05 to 0.15.
[0079] 2. The equation for the degree of frequency variation easing is expressed as follows:
[0080]
[0081] In the formula, F s The smoothness of frequency change; Δf i f represents the frequency difference in the i-th time window; t This is the frequency difference threshold coefficient; E represents the energy gradient for the i-th time window. t λ1 and λ2 are the energy gradient thresholds; λ1 and λ2 are the weighting coefficients, both ranging from 0 to 1; M is the number of time windows; δ f This is the error term, with a range of 0.1 to 0.2.
[0082] 3. The equation for the length of the sound is expressed as follows:
[0083]
[0084] In the formula, L d T represents the length of the sound; T represents the actual duration. s E is the standard duration factor. f E is the final energy value; E0 is the initial energy value; E i The energy value at the i-th time point; E represents the average energy value. v ω1, ω2, ω3 are weighting coefficients, all ranging from 0 to 1; η1, η2 are power coefficients, ranging from 1 to 3; K is the number of sampling points; δ l This is the error term, ranging from 0.05 to 0.15.
[0085] 4. The equation for the melodiousness of a bird's song is expressed as follows:
[0086]
[0087] In the formula, M t A is an index of the melodiousness of the song; i A1 is the amplitude of the i-th harmonic component; A1 is the fundamental amplitude; H is the number of harmonics; ΔS j S represents the change in the j-th spectral segment; t δ is the threshold for spectral variation; P is the number of spectral segments; C(t) is the function of timbre variation over time; φ1, φ2, φ3 are weighting coefficients, all ranging from 0 to 1; mThis is the error term, with a range of 0.1 to 0.2.
[0088] 5. The specific expression of the diurnal distribution characteristic equation is as follows:
[0089]
[0090] In the formula, D t This is a time-varying characteristic index; t is 24-hour time; t p This is the peak time for the event; For phase difference; σ t δ is the standard deviation of the time distribution; W(t) is the time period weighting function; S(m) is the monthly seasonal adjustment function; θ1, θ2, θ3 are weighting coefficients, all ranging from 0 to 1; δ d This is the error term, with a range of 0.05 to 0.15.
[0091] The principles and significance of these equations are explained below:
[0092] 1. The time series relationship matrix adopts the form of normalized cross-correlation function, which takes into account the time delay characteristics between signals and can effectively describe the degree of correlation between different signals; the exponential function is used to characterize the similarity decay law of signal features.
[0093] 2. The frequency change smoothness equation uses an exponential form to describe the smoothness of frequency changes, taking into account both frequency difference and energy gradient dimensions. The exponential function can well express the nonlinear decay characteristics of the feature.
[0094] 3. The length equation of the sound describes the duration characteristics using a power function and the energy change characteristics using an exponential function, thus comprehensively considering the contributions of both duration and energy dimensions.
[0095] 4. The equation for the melodiousness of a sound includes three parts: harmonic structure, frequency spectrum variation, and timbre variation. These three characteristics are described using ratios, exponents, and integrals, respectively, to comprehensively reflect the intonation and rhythm of the sound.
[0096] 5. The diurnal distribution characteristic equation uses a sine function to describe the diurnal periodic variation, a Gaussian function to describe the peak characteristics of activity, and a product term to describe the seasonal influence, thus comprehensively reflecting the temporal distribution pattern.
[0097] In summary, this monitoring method, through a series of steps including monitoring network construction, signal processing, feature extraction, and model recognition, achieves automatic identification of black leaf monkey calls and population density assessment. This method fully considers the acoustic and spatiotemporal distribution characteristics of black leaf monkey calls and employs various machine learning algorithms for model construction and optimization, demonstrating strong practicality and reliability.
[0098] A second aspect of the present invention provides a computer-readable storage medium storing program instructions that, when executed in a computer, perform the aforementioned machine learning-based method for monitoring the calls of black leaf monkeys.
[0099] A third aspect of the present invention provides a black leaf monkey call monitoring system based on machine learning, wherein the system includes the aforementioned computer-readable storage medium, and the system is any one of a computer, a server, or a microcontroller, wherein the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0100] Specifically, the principle of this invention is to utilize advanced signal processing and pattern recognition technologies to extract highly correlated population features from complex acoustic signals and establish an accurate population monitoring model. The specific principle is as follows:
[0101] First, by constructing a monitoring network topology and optimizing the audio acquisition path, high-quality raw audio signals are ensured. In the topology, each vertex represents an audio acquisition device, and edges represent communication links between devices. The minimum spanning tree algorithm is used to construct the topology, and the weight of each edge is evaluated based on the environmental noise level, thereby designing an acquisition path that minimizes noise interference.
[0102] Secondly, the collected raw audio signals were decomposed and feature extracted. Considering the relatively melodious and prolonged nature of the black leaf monkey's calls, this invention first uses an empirical mode decomposition algorithm to decompose the signal into individual call components and superimposed call components, significantly reducing background noise interference. Then, for the individual call components and independent call components, multi-dimensional features such as frequency, energy, and harmonic structure were extracted, and these features were quantitatively analyzed using a set of evaluation equations for black leaf monkey call characteristics to obtain reliable feature vectors. These evaluation equations fully consider the time-frequency characteristics and ecological behavioral features of the black leaf monkey's calls, and can better characterize the physiological state and community structure of the population.
[0103] Finally, based on the aforementioned feature vectors, a multi-model fusion framework for song classification was constructed. The first machine learning model uses a lightweight convolutional neural network to process acoustic features, while the second model uses a graph neural network to process the evaluation results of black leaf monkey song features. The outputs of the two models are then fused through an attention mechanism to obtain the final recognition result. This fusion model not only fully utilizes information from features of different dimensions but also quantitatively assesses the importance of each feature to the recognition result, thus providing a reliable basis for subsequent population density estimation.
[0104] The following is a specific embodiment 1 of the present invention. The detailed implementation of each step in this embodiment 1 is described below: Step S01 is implemented by constructing a monitoring network topology map based on a habitat geographic information system. First, geographic information data of the black leaf monkey habitat needs to be acquired, including topography, vegetation cover, water system distribution, etc., and the deployment locations of audio acquisition equipment need to be determined. High-resolution remote sensing images are acquired using the Google Earth engine platform, combined with field survey data, to establish a three-dimensional geographic information model of the habitat. Then, based on the communication range and signal strength of the audio acquisition equipment, a minimum spanning tree algorithm is used to construct the network topology map. The recommended communication range threshold is set to 800 meters, and the recommended signal strength threshold is set to -60 dB / mW. Finally, a graph coloring algorithm is used to allocate communication channels in the network topology map to avoid signal interference between adjacent nodes. The minimum frequency interval between communication channels should not be less than 20 MHz. The main purpose of this step is to establish a reasonable monitoring network structure to ensure the effective acquisition and transmission of audio data.
[0105] The specific implementation of step S02 involves determining the optimal audio acquisition path based on the monitoring network topology. First, an environmental acoustic modeling method is used to construct a habitat noise distribution map, considering the impact of topography, vegetation, and meteorology on sound propagation. A sound wave propagation attenuation model is used to calculate the environmental noise level at different locations, considering factors such as geometric divergence loss, atmospheric absorption loss, and ground effect loss. Then, an ant colony optimization algorithm is used to design the audio acquisition path to minimize accumulated noise interference along the path. During path optimization, the environmental noise level is used as heuristic information, and a noise threshold of 45 dB is recommended. Finally, the feasibility of the optimized acquisition path is verified to ensure that it meets hardware constraints such as device power consumption and storage capacity. The main purpose of this step is to design a scientifically sound audio acquisition scheme while considering the impact of environmental noise.
[0106] The specific implementation of step S03 involves decomposing the original audio signal. First, the acquired original audio signal undergoes preprocessing, including noise reduction and denoising. An adaptive noise cancellation algorithm using Wiener filtering is employed, with a signal-to-noise ratio improvement threshold of at least 10 dB. Then, an empirical mode decomposition algorithm is used to decompose the preprocessed signal into several intrinsic mode functions (IMFs). Based on frequency characteristics, these IMFs are categorized into single-unit sound signal components and superimposed sound signal components. The frequency range of single-unit sound signal components is relatively concentrated, while the superimposed sound signal components exhibit significant frequency aliasing. Finally, Hilbert transform is used to perform envelope analysis on the decomposed signal components, extracting the instantaneous frequency and amplitude characteristics of the signal. The main purpose of this step is to decompose the complex original audio signal into basic components that facilitate subsequent processing.
[0107] The specific implementation of step S04 is to construct a signal feature relationship matrix. First, based on the formula for the time series relationship matrix... The temporal correlation between the individual sound signal component and the superimposed sound signal components is calculated using a sliding time window method to determine the cross-correlation function between the signals. A window length of 256 milliseconds and an overlap rate of 50% are recommended. Then, the sound signal stability coefficient matrix formula is applied. The stability characteristics of each signal component are calculated, considering stability indices in both frequency and amplitude dimensions. Finally, the time-series relationship matrix and the stability coefficient matrix are combined using a weighted fusion method to obtain a comprehensive feature matrix, with the weighting coefficients determined by a particle swarm optimization algorithm. The main purpose of this step is to establish the correlation between signal components, providing a basis for subsequent signal separation.
[0108] The specific implementation of step S05 is to perform blind source separation of the superimposed call signals based on the independent component analysis algorithm. First, principal component analysis is used to pre-whiten the signals, reducing the correlation between them. Then, by combining information from the time-series relation matrix and the stability coefficient matrix, an objective function is constructed, and a fast independent component analysis algorithm is used for signal separation, with the iterative convergence threshold set to 0.0001. Finally, the separated independent call signals are post-processed, including dereverberation and phase correction, and a Kalman filter algorithm is used for signal enhancement. The main purpose of this step is to separate the independent black leaf monkey call signals from the superimposed signals.
[0109] The specific implementation of step S06 involves sharpening the acoustic signal. First, wavelet transform is used to decompose the signal into multiple scales, selecting the Dobessi 5 wavelet as the basis function and setting the decomposition level to four. Then, nonlinear mapping is applied to the coefficients at each scale, and a soft thresholding function is used for coefficient shrinkage. The threshold is selected using the maximum likelihood estimation method. Finally, the processed wavelet coefficients are reconstructed to obtain the sharpened acoustic signal, and spectral subtraction is used to eliminate residual noise. The main purpose of this step is to improve the clarity and recognizability of the acoustic signal.
[0110] The specific implementation of step S07 involves extracting and evaluating the features of the call signal. First, Mel-frequency cepstral analysis is used to extract the acoustic features of the signal, dividing the signal into 32 Mel-frequency bands and extracting the 13th-order Mel-frequency cepstral coefficients. Then, autocorrelation analysis is used to extract the fundamental frequency features of the signal, and fast Fourier transform is used to calculate the energy features of the signal. Next, the extracted features are input into the evaluation equation set for the black-leaf monkey call features, including the frequency change smoothness equation. Equation of the length of the sound Equation of Melodiousness and the characteristic equation of daytime distribution Four equations were used. Finally, all features were standardized and combined into an eigenvector, and principal component analysis was employed for dimensionality reduction. The main purpose of this step was to extract and evaluate indicators that effectively characterize the vocal features of the black leaf monkey.
[0111] The specific implementation of step S08 involves constructing a multi-model fusion-based sound classification framework. First, a deep neural network is used to construct the first machine learning model, with a network structure consisting of four convolutional layers and two fully connected layers, using a modified linear unit function as the activation function. Then, a random forest algorithm is used to construct the second machine learning model, with 100 decision trees and a maximum depth of 10. Finally, a fusion model based on stacked ensemble learning is designed, using logistic regression as the meta-learner to fuse the outputs of the two base models. The main objective of this step is to establish a classification model framework that can effectively utilize different feature information.
[0112] The specific implementation of step S09 involves using a first machine learning model to identify the type of sound. First, the feature vector is input into a pre-trained deep neural network, and the recognition probability of each type is calculated through forward propagation. Then, a soft maximum function is used to normalize the network output, resulting in a standardized probability distribution. Finally, the cross-entropy loss function is used to evaluate the model's recognition performance, with a loss threshold set to 0.1. The main purpose of this step is to achieve preliminary identification of the sound type based on acoustic features.
[0113] The specific implementation of step S10 involves using a second machine learning model to analyze feature contribution. First, the feature vectors are input into a random forest model, and the Gini importance index for each feature in call type recognition is calculated. Then, a permutation importance analysis method is used to verify the feature contribution, with 100 permutations performed. Finally, based on the feature importance ranking results, the normalized contribution value of each feature is calculated. The main purpose of this step is to quantitatively evaluate the importance of different features in call recognition.
[0114] Step S11 involves optimizing the second machine learning model. First, a joint loss function is constructed based on the sound type recognition probability and feature contribution values, including a classification loss term and a feature importance loss term. Then, the model parameters are optimized using stochastic gradient descent, with the initial learning rate set to 0.001, and cosine annealing is used to adjust the learning rate. Finally, an early stopping strategy is employed to prevent overfitting; training is stopped if the loss on the validation set does not improve for five consecutive epochs. The main purpose of this step is to improve the model's generalization ability and feature utilization efficiency.
[0115] Step S12 involves assessing the population density distribution of black leaf monkeys. First, the call types and their contribution values identified at each monitoring point are input into a pre-trained density assessment model, constructed using a Gaussian process regression algorithm. Then, considering spatial autocorrelation, Kriging interpolation is used to spatially interpolate the population density, with a search radius set to 1000 meters. Finally, a population density distribution matrix is generated, and an adaptive threshold segmentation method is used to classify the density into five levels. The main purpose of this step is to assess the spatial distribution of the black leaf monkey population density.
[0116] To better understand and implement this invention, Example 2 of a specific application scenario is provided below: A management unit of a protected area decided to adopt the machine learning-based black leaf monkey call monitoring method proposed in this invention to conduct long-term dynamic monitoring of the black leaf monkey population within the protected area. This protected area is located in a tropical mountain forest region with complex and diverse terrain and dense vegetation, making it one of the important habitats for black leaf monkeys.
[0117] First, relevant technical personnel conducted a comprehensive survey and modeling of the geographical environment within the protected area. Using high-resolution remote sensing image data, combined with information from field investigations, a three-dimensional geographic information model of the protected area was constructed on a GIS platform, including key elements such as topography, vegetation cover, and water system distribution.
[0118] Then, based on the activity area and habitat characteristics of the black-leaf macaque, 20 audio acquisition devices were strategically deployed within the reserve. These devices all employ low-power, high-sensitivity sound transmission components and are equipped with data buffers and wireless communication modules. Through testing the communication coverage and signal strength of the devices, technicians constructed a monitoring network topology using the minimum spanning tree algorithm, as shown in Table 1. In the topology, each vertex represents an audio acquisition device, and edges represent communication links between devices. To avoid interference between adjacent nodes, technicians also used a graph coloring algorithm to rationally allocate communication channels, ensuring a frequency interval of no less than 20MHz between adjacent nodes.
[0119] Table 1 Monitoring Network Topology Diagram
[0120] node Neighbor nodes Link weight (dBm) 1 2,5 -58,-62 2 1,3,6 -55,-59,-61 3 2,4,7 -57,-60,-63 4 3,8 -59,-64 5 1,6,9 -60,-57,-62 6 2,5,7,10 -56,-58,-59,-63 7 3,6,8,11 -62,-57,-60,-64 8 4,7,12 -61,-58,-63 9 5,10,13 -59,-61,-65 10 6,9,11,14 -57,-60,-62,-64 11 7,10,12,15 -59,-58,-61,-63 12 8,11,16 -60,-59,-62 13 9,14,17 -63,-61,-64 14 10,13,15,18 -59,-60,-62,-65 15 11,14,16,19 -58,-61,-60,-63 16 12,15,20 -59,-58,-62 17 13,18 -61,-63 18 14,17,19 -59,-60,-62 19 15,18,20 -57,-60,-59 20 16,19 -58,-57
[0121] To obtain high-quality audio data, technicians also conducted modeling and analysis of the environmental noise distribution within the protected area. Noise levels within the protected area vary significantly, mainly influenced by factors such as topography and vegetation cover. Based on the noise distribution map, technicians designed audio acquisition paths using an ant colony algorithm to minimize cumulative noise interference along the paths. After multiple iterations and optimizations, the final acquisition paths are shown in Table 2, where the noise level in priority 1 area is below 30 dB, priority 2 area is 30-50 dB, and priority 3 area is above 50 dB.
[0122] Table 2 Audio Acquisition Path
[0123]
[0124]
[0125] Based on the determined acquisition path, technicians deployed audio acquisition equipment within the protected area and used wireless transmission technology to transmit the raw audio signals to the monitoring center for further processing. At the monitoring center, technicians first preprocessed the raw audio signals, employing Wiener filtering for adaptive noise reduction, improving the signal-to-noise ratio to over 10 dB. Then, they used empirical mode decomposition (EMD) to decompose the preprocessed signal into individual sound components and superimposed sound components. Analysis of the frequency characteristics of these two types of components revealed that the frequency range of the individual sound components was relatively concentrated, while the superimposed components exhibited significant frequency aliasing.
[0126] Next, the technicians constructed a corresponding feature relationship matrix based on the temporal relationship between the individual sound components and the superimposed sound components. First, the cross-correlation function of the two signal components was calculated using the sliding time window method, with a window length of 256 ms and an overlap rate of 50%, resulting in the temporal relationship matrix R:
[0127]
[0128] Secondly, regarding the frequency and amplitude stability of each signal component, the technicians calculated the sound signal stability coefficient matrix S:
[0129]
[0130] Finally, the technicians used a weighted fusion method to combine R and S to obtain the comprehensive feature relation matrix F:
[0131]
[0132] With this feature relation matrix, technicians used a fast independent component analysis algorithm to perform blind source separation on the superimposed call components, successfully separating the independent black leaf monkey call signals. For the separated call signals, technicians then used wavelet transform for situational sharpening to improve signal clarity and recognizability.
[0133] Next, technicians extracted multi-dimensional features such as the Mel frequency cepstral coefficients, fundamental frequency, and energy, and conducted a detailed analysis using a specially designed set of evaluation equations for the black leaf monkey's vocal characteristics. These evaluation equations include indicators such as the smoothness of frequency changes, the length of the vocalization, the melodicity of the vocalization, and diurnal distribution characteristics, which can comprehensively characterize the physiological and behavioral features of the black leaf monkey's vocalizations. Taking the smoothness of frequency changes as an example, its calculation formula is as follows:
[0134]
[0135] Where, Δf i This represents the frequency difference in the i-th time window. This represents the energy gradient at the i-th time window. Through a comprehensive evaluation of various characteristic indicators, technicians obtained a highly reliable feature vector.
[0136] Based on the aforementioned feature vectors, researchers constructed a multi-model fusion framework for classifying the calls of black leaf monkeys. The first machine learning model employs the lightweight convolutional neural network MobileNetV3, which efficiently processes acoustic features; the second model uses a graph neural network GraphSAGE, which utilizes the correlation information contained in the feature evaluation results; the outputs of the two base models are then fused through an SENet fusion model to obtain the final recognition result. During training, a weighted cross-entropy loss function and a learning rate decay strategy were implemented, resulting in a classification accuracy of over 95% on the test set.
[0137] With accurate call identification results, technicians further utilized this information to assess the population density distribution of black-leaf monkeys within the reserve. First, the call types and their contribution values identified at different monitoring points were input into a pre-trained GRU recurrent neural network model, constructed using a Gaussian process regression algorithm. Then, considering the spatial autocorrelation of population density, technicians used Kriging interpolation to spatially interpolate the density data, ultimately generating a black-leaf monkey population density distribution matrix within the reserve.
[0138] It should be noted that the variables involved in this invention are explained in detail in Table 3 below.
[0139] Table 3. Variable Explanation Table
[0140]
[0141] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for monitoring the calls of black leaf monkeys based on machine learning, characterized in that, include: A topology map of the black leaf monkey habitat monitoring network was constructed, where vertices represent audio acquisition devices and edges represent communication links between audio acquisition devices. Based on the geographical location information of the audio acquisition devices and the environmental noise level in the monitoring network topology map, the audio acquisition path was determined. The original audio signal is acquired using an audio acquisition device, and the original audio signal is decomposed and processed to obtain the individual sound signal component and the superimposed sound signal component. Establish the temporal relationship matrix between the individual sound signal components and the superimposed sound signal components, and calculate the sound signal stability coefficient matrix. Based on the temporal relationship matrix and the call signal stability coefficient matrix, the superimposed call signal components are separated to obtain independent call signals. The individual call signal components and independent call signals are then subjected to signal state sharpening processing to form call signal samples. Feature extraction and evaluation are performed on the call signal samples, extracting Mel frequency cepstral coefficient features, fundamental frequency features, energy features, and combining the evaluation results of the black-leaf monkey call feature evaluation equation set to form feature vectors. A call classification model is established, including a first machine learning model, a second machine learning model, and a fusion model. The first machine learning model is used to process the feature vectors to obtain the call type recognition probability. The second machine learning model is used to calculate the contribution value of the call signal in the feature vector to the call type recognition. Based on the loss function between the call type recognition probability and the call type recognition contribution value, the second machine learning model is optimized and trained. The call types identified at different monitoring points and the distribution of call type recognition contribution values are input into a pre-trained black-leaf monkey population density assessment model to obtain and output the black-leaf monkey population density distribution matrix.
2. The method for monitoring the calls of black leaf monkeys based on machine learning according to claim 1, characterized in that, The evaluation equation set for the vocal characteristics of black leaf monkeys includes equations for the smoothness of frequency changes, the length of vocal duration, the melodiousness of vocalizations, and the characteristics of diurnal distribution.
3. The method for monitoring the calls of black leaf monkeys based on machine learning according to claim 2, characterized in that, The frequency change smoothness equation is used to quantify the smoothness of frequency changes. It takes frequency difference data between adjacent time windows as input and outputs a frequency smoothness index. The main parameters of the frequency change smoothness equation include frequency difference threshold coefficient, time window length coefficient, frequency gradient weight coefficient, and smoothness benchmark coefficient. The frequency difference threshold coefficient and the time window length coefficient are derived from the fundamental frequency feature, and the frequency gradient weight coefficient and the smoothness benchmark coefficient are derived from the energy feature.
4. The method for monitoring the calls of black leaf monkeys based on machine learning according to claim 3, characterized in that, The sound duration equation is used to quantify the duration characteristics of the sound. It takes sound duration data and energy envelope data as input and outputs a sound duration index. The main parameters of the sound duration equation include a standard duration coefficient, an energy attenuation coefficient, a time scaling coefficient, and an energy fluctuation weighting coefficient. The standard duration coefficient and the time scaling coefficient are derived from the fundamental frequency characteristics, and the energy attenuation coefficient and the energy fluctuation weighting coefficient are derived from the energy characteristics.
5. The method for monitoring the calls of black leaf monkeys based on machine learning according to claim 4, characterized in that, The sonic melodicity equation is used to quantify the intonation and cadence characteristics of a sound. It takes harmonic structure data and spectral variation data as input and outputs a melodicity index. The main parameters of the sonic melodicity equation include harmonic richness coefficient, spectral fluctuation coefficient, and timbre variation coefficient. The harmonic richness coefficient is derived from the Mel frequency cepstral coefficient characteristics, and the spectral fluctuation coefficient and the timbre variation coefficient are derived from the fundamental frequency characteristics.
6. The method for monitoring the calls of black leaf monkeys based on machine learning according to claim 5, characterized in that, The daytime distribution characteristic equation is used to describe the changes in sound characteristics at different times within 24 hours. It takes 24-hour sound sample data as input and outputs time-varying characteristic indicators. The main parameters of the daytime distribution characteristic equation include diurnal cycle coefficient, activity peak coefficient, time period weight coefficient, and seasonal adjustment coefficient. The diurnal cycle coefficient and the activity peak coefficient are derived from the fundamental frequency characteristics, and the time period weight coefficient and the seasonal adjustment coefficient are derived from the energy characteristics.
7. The method for monitoring the calls of black leaf monkeys based on machine learning according to claim 6, characterized in that, The audio acquisition path includes the following steps: constructing an acquisition path graph based on the monitoring network topology, and marking the passage weights between the audio acquisition devices using a breadth-first search algorithm; prioritizing monitoring areas based on the ambient noise level, setting monitoring areas with ambient noise levels below 30 dB as high priority, monitoring areas with ambient noise levels above 30 dB but below 50 dB as medium priority, and monitoring areas with ambient noise levels above 50 dB as low priority; performing state transition calculations in the acquisition path graph, traversing all state nodes, obtaining the optimal state transition sequence, and forming the audio acquisition path.
8. A machine learning-based method for monitoring the calls of black leaf monkeys according to claim 7, characterized in that, The first machine learning model is used to process the Mel frequency cepstral coefficient features, fundamental frequency features, and energy features. The second machine learning model is used to process the evaluation results of the evaluation equation set of black leaf monkey vocal features. The fusion model is used to fuse the output results of the first machine learning model and the output results of the second machine learning model. The first machine learning model, the second machine learning model, the fusion model, and the black leaf monkey population density assessment model all use lightweight neural networks.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed in a computer, are used to perform a machine learning-based method for monitoring the calls of black leaf monkeys as described in any one of claims 1-8.
10. A machine learning-based system for monitoring the calls of black leaf monkeys, characterized in that, The system includes the computer-readable storage medium of claim 9, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.