Natural sound recognition device of online monitor

By using an online monitoring instrument with a natural sound recognition device, sound signals from two monitoring points are collected simultaneously. The dominant period of the cicada swarm is decomposed and determined. A Butterworth filter is used to separate the cicada swarm signal, which solves the problem of the cicada swarm signal masking the natural sound and achieves efficient natural sound recognition.

CN121122292AInactive Publication Date: 2025-12-12ABELOO BUILDING MATERIALS BEIJING
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511300067.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional methods cannot effectively identify the problem of cicada swarm signals masking other natural sounds, which limits the application of natural sound recognition technology in scenarios such as ecological monitoring.

Method used

An online monitoring instrument with a natural sound recognition device was used to simultaneously collect sound signals from two monitoring points. The signals were decomposed to obtain the frequency band energy distribution and cross-point coherence. The core narrowband was determined by local maxima relation. The dominant period of the cicada swarm was confirmed by energy relation determination and periodic resonance determination. The cicada swarm signal was separated by a Butterworth bandpass filter, while other natural sound source signals were preserved.

Benefits of technology

It accurately identifies cicada swarm signals, eliminates cicada swarm interference, and ensures the integrity of other natural sound signals, thereby improving the specificity and accuracy of natural sound recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122292A_ABST
    Figure CN121122292A_ABST
Patent Text Reader

Abstract

The invention discloses a natural sound recognition device for an online monitor, and relates to the technical field of natural sound recognition, and the device comprises an acquisition module which is used for synchronously acquiring sound signals of double monitoring points; the decomposition module is used for decomposing the sound signal to obtain energy distribution and cross-point coherence of each frequency band; the core narrowband determination module is used for determining a cross-point coherent core narrowband; the energy relation judgment module is used for confirming whether inter-band slope inversion exists in the core narrow band or not; the periodic resonance judgment module is used for determining whether periodic resonance exists or not; the cicada group identification module is used for identifying a cicada group dominant time period; and the sound correction module is used for performing frequency spectrum separation and filtering on the sound signals in the dominant period of the cicada group, and keeping other natural sound source signals to obtain a corrected identification result. According to the invention, the dominant time period of the cicada group can be accurately identified, and the concealment of other natural sounds by cicada group signals is eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural voice recognition technology, and in particular to a natural voice recognition device for an online monitoring instrument. Background Technology

[0002] Natural sounds are important information carriers of the ecological environment, encompassing the characteristics of various sound sources such as birdsong, flowing water, and wind. Their accurate identification and analysis play an irreplaceable role in fields such as ecological monitoring, biodiversity assessment, and environmental quality evaluation. In actual natural environments, multiple sound sources often coexist, causing sound signals to superimpose in both time and frequency dimensions, posing challenges to accurate identification. Among these, cicada swarms are one of the most prominent sources of interference during specific periods, such as summer. Their signal characteristics make them particularly effective at masking other natural sounds. Cicada swarms rely on high-frequency vibrating organs for vocalization, and their signals are mostly concentrated in a narrow frequency band. The superposition effect when cicadas gather in groups makes the energy of this frequency band much higher than other natural sounds, creating energy suppression. Furthermore, the distinct periodicity of cicada calls continuously occupies the acoustic environment over time, further compressing the identifiable space of other natural sounds.

[0003] Traditional methods often rely on energy analysis within a single frequency band. However, the strong narrowband characteristics of cicada swarm signals mean that the energy in that band is dominated by cicada chirping, completely masking weak signals from other natural sounds and making them impossible to detect. Furthermore, traditional filtering techniques often employ fixed-band bandpass or bandstop processing. If the core frequency band of the cicada swarm signal is not accurately located, it may either inadvertently delete valid natural sounds in adjacent frequency bands or fail to completely remove cicada interference, leaving residual signals that still affect recognition. Therefore, in environments with active cicada swarms, traditional methods have consistently failed to address the problem of cicada signals masking other natural sounds, severely limiting the application of natural sound recognition technology in scenarios such as ecological monitoring. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies where cicada swarm signals mask other natural sounds, and to propose an online monitoring instrument for natural sound recognition.

[0005] To address the problems existing in the prior art, the present invention adopts the following technical solution:

[0006] An online monitoring instrument natural voice recognition device includes:

[0007] The acquisition module is used to simultaneously acquire acoustic signals from two monitoring points;

[0008] The decomposition module is used to decompose the acoustic signal to obtain the energy distribution of each frequency band and cross-point coherence;

[0009] The core narrowband determination module is used to determine the cross-point coherent core narrowband through local maxima relations;

[0010] The energy relationship determination module is used to determine whether there is an inter-band slope reversal in the core narrowband based on the relationship between the energy of the core narrowband and adjacent frequency bands.

[0011] The periodic resonance determination module is used to extract the envelope on the core narrowband and confirm the existence of periodic resonance by the consistency between the envelope autocorrelation and the cross-point envelope crosscorrelation.

[0012] The cicada swarm identification module is used to identify the corresponding time period as the cicada swarm's dominant period when cross-point coherence, inter-band slope inversion, and periodic resonance are all confirmed.

[0013] The sound correction module is used to perform spectral separation and filtering on the sound signal during the period when the cicada swarm dominates, while retaining other natural sound source signals to obtain the corrected recognition result.

[0014] Preferably, the simultaneous acquisition of acoustic signals from two monitoring points includes:

[0015] Acquire the acoustic signal at monitoring point A and the acoustic signal at monitoring point B at the same sampling rate;

[0016] By calculating the cross-correlation function of the signals from the two monitoring points during the observation period, the maximum delay is determined. Based on the maximum delay, the acoustic signal of monitoring point B is time-shifted and compensated to obtain the time-aligned acoustic signal of monitoring point B.

[0017] Preferably, the acoustic signal is decomposed to obtain the energy distribution of each frequency band and cross-point coherence, including:

[0018] Short-time Fourier transforms are performed on the acoustic signal at monitoring point A and the time-aligned acoustic signal at monitoring point B to obtain the spectral signals of monitoring point A and monitoring point B, respectively.

[0019] The frequency band energy distribution of the spectrum signal at monitoring point A and the spectrum signal at monitoring point B are calculated respectively to obtain the frequency band energy distribution of monitoring point A and the frequency band energy distribution of monitoring point B.

[0020] Cross-point coherence calculations are performed based on the spectral signals from monitoring points A and B to obtain the cross-point coherence coefficient.

[0021] Preferably, the core narrowband of cross-point coherence is determined through local maxima relations, including:

[0022] The cross-point coherence is discretized into coherent sequences according to frequency;

[0023] The data in the coherent sequence are evaluated to identify local maxima, and these local maxima are then merged into the candidate set.

[0024] The frequency of the maximum value in the candidate set is taken as the core center frequency. If there are ties for the maximum value, the one with the lower center frequency is selected as the core center frequency.

[0025] Starting from the core center frequency, the nearest local minimum points are located in the low-frequency and high-frequency directions respectively as boundaries. If there is no local minimum point at the endpoint, the frequency endpoint is taken as the boundary to determine the core narrowband domain.

[0026] Preferably, based on the relationship between the energy of the core narrowband and adjacent frequency bands, it is confirmed whether there is an inter-band slope inversion in the core narrowband, including:

[0027] Based on the energy of the core narrowband and adjacent frequency bands, the slope of the energy difference is obtained by calculating the logarithmic slope.

[0028] The slope is determined, and if the slope is less than zero, it is determined that there is an inter-band slope reversal.

[0029] Preferably, the envelope is extracted on the core narrowband, and the consistency between the envelope autocorrelation and the cross-point envelope crosscorrelation is used to confirm the existence of periodic resonance, including:

[0030] The envelope signal of the core narrowband frequency is extracted by Hilbert transform to obtain the envelope signals of monitoring point A and monitoring point B.

[0031] Calculate the autocorrelation function and cross-correlation function of the envelope signals of monitoring point A and monitoring point B, and obtain the local maxima of the autocorrelation and cross-correlation.

[0032] The local maxima of autocorrelation and cross-correlation are determined. If the autocorrelation and cross-correlation functions have maxima at the same time delay point, then periodic resonance is determined to exist.

[0033] Preferably, when cross-point coherence, inter-band slope inversion, and periodic resonance are all confirmed, the corresponding time period is identified as the cicada swarm-dominated period, including:

[0034] If the monitoring period meets the condition that the cross-point coherence is greater than the preset coherence threshold and there is inter-band slope reversal and periodic resonance, then this period is regarded as the period when the cicada swarm is dominant.

[0035] Preferably, the acoustic signal during the dominant period of the cicada swarm is subjected to spectral separation and filtering, while retaining other natural sound source signals, to obtain a corrected recognition result, including:

[0036] Based on the core narrowband, the transfer function of the Butterworth bandpass filter is constructed, where the transfer function of the Butterworth bandpass filter is:

[0037]

[0038] In the formula, H(f) is the transfer function of the filter, which represents the signal attenuation factor at frequency f, f is the frequency, f0 is the center frequency of the filter, B is the bandwidth of the filter, and N is the filter order.

[0039] The acoustic signal during the dominant cicada swarm period is input into a Butterworth bandpass filter to extract the cicada swarm signal component. The cicada swarm signal component is then subtracted from the acoustic signal during the dominant cicada swarm period to obtain the filtered natural acoustic signal.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] 1. In this invention, synchronous acquisition ensures that the time axis of the acoustic signals at the two monitoring points is strictly aligned, eliminating cross-point analysis errors caused by acquisition delay. The acoustic signal is decomposed into a spectral signal, frequency band energy, and cross-point coherence, solving the limitation of time-domain analysis in distinguishing different frequency components and accurately capturing the narrowband characteristics of cicada swarm signals.

[0042] 2. In this invention, by filtering local maxima and locating boundaries, the core narrowband with the highest cross-point coherence is locked from the broadband, eliminating interference from low-coherence frequency bands. The analysis range is focused on the frequency range most likely to contain cicada swarm signals, which can adapt to the frequency differences of signals from different cicada species. The energy logarithmic slope is used to determine the characteristic that the energy of the core narrowband is significantly higher than that of adjacent frequency bands, accurately distinguishing between strong narrowband cicada swarm signals and broadband natural sounds. Envelope autocorrelation is used to capture the periodicity of cicada swarm signals and distinguish between periodic signals and random noise, solving the problem that simple frequency domain analysis is difficult to identify irregular noise. The consistency requirement of cross-point envelope cross-correlation and autocorrelation ensures that the identified periodic signals are real, large-scale sound sources.

[0043] 3. In this invention, a three-dimensional verification mechanism is formed in the spatial, frequency, and temporal domains through cross-point coherence, inter-band slope reversal, and periodic resonance, which greatly improves the specificity of cicada swarm identification, filters out the dominant period of the cicada swarm, provides a clear processing range for the subsequent sound correction module, and accurately extracts and separates the cicada swarm signal through a Butterworth filter matched with the core narrowband, maintaining the integrity of the natural sound signal. Attached Figure Description

[0044] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0045] Figure 1 This is a functional block diagram of an online monitoring instrument natural sound recognition device provided in an embodiment of the present invention. Detailed Implementation

[0046] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0047] Example: This example provides a natural sound recognition device for an online monitoring instrument. See [link to example]. Figure 1 Specifically, including:

[0048] The acquisition module is used to simultaneously acquire acoustic signals from two monitoring points;

[0049] In an embodiment of the present invention, the synchronous acquisition of acoustic signals from two monitoring points includes:

[0050] Acquire the acoustic signal at monitoring point A and the acoustic signal at monitoring point B at the same sampling rate;

[0051] Specifically, two acoustic sensors deployed within the same canopy zone are used to collect acoustic signals from the environment. The two sensors are of the same model to ensure consistent sound response characteristics and are set to the same sampling rate. The signals are stored in discrete time series format. The sensors must start and stop collecting data simultaneously to avoid initial offset caused by start-up time differences.

[0052] By calculating the cross-correlation function of the signals from the two monitoring points during the observation period, the maximum delay is determined. Based on the maximum delay, the acoustic signal of monitoring point B is time-shifted and compensated to obtain the time-aligned acoustic signal of monitoring point B.

[0053] Specifically, due to differences in sensor deployment locations and signal transmission delays, the two original signals may have a time offset, which needs to be eliminated through processing. A cross-correlation function can be used to measure the similarity of the two signals at different time offsets; the calculation formula is as follows:

[0054]

[0055] In the formula, τ is the time offset, and R AB (τ) represents the cross-correlation function value of the acoustic signals at monitoring point A and monitoring point B, N is the total number of sound samples, m is the sample index, and x A (m) represents the amplitude of the acoustic signal at monitoring point A at the m-th sample point, x B (m+τ) represents the amplitude of the acoustic signal at monitoring point B at the (m+τ)th sample point;

[0056] By iterating through all possible time offsets, the offset that maximizes the cross-correlation function is found and taken as the maximum delay. The acoustic signal at monitoring point B is then shifted according to this maximum delay. If the maximum delay is greater than zero, the acoustic signal at monitoring point B is shifted forward by the maximum delay amount (i.e., the first maximum delay amount of signal samples are deleted, and subsequent samples are shifted forward sequentially). If the maximum delay is less than zero, the acoustic signal at monitoring point B is shifted backward by the maximum delay amount (i.e., the maximum delay amount of zero-value samples are added before the signal), and the positions of subsequent samples remain unchanged. After alignment processing, the consistency of the two signals in the time dimension can be ensured, providing a reliable synchronous data source for subsequent steps such as frequency band analysis and coherence calculation.

[0057] The decomposition module is used to decompose the acoustic signal to obtain the energy distribution of each frequency band and cross-point coherence;

[0058] In embodiments of the present invention, the acoustic signal is decomposed to obtain the energy distribution of each frequency band and cross-point coherence, including:

[0059] Short-time Fourier transforms are performed on the acoustic signal at monitoring point A and the time-aligned acoustic signal at monitoring point B to obtain the spectral signals of monitoring point A and monitoring point B, respectively.

[0060] Specifically, a short-time Fourier transform can be used to map the time-domain signal to a two-dimensional plane of time and frequency, enabling the analysis of the frequency components of the signal at different times. A Hanning window is used to weight intra-frame samples to suppress spectral leakage caused by signal abrupt changes. The window length is set to 512 to balance frequency and time resolution, and the overlap rate is set to 50% to avoid information loss between time frames. The frame shift is half the length of the Hanning window. The acoustic signal from monitoring point A and the time-aligned acoustic signal from monitoring point B are divided into frames by sliding the window. After applying a Hanning window to each frame, a Fourier transform is performed. The specific formula is as follows:

[0061]

[0062] In the formula, X(f,n) is a complex-form spectral signal representing the signal amplitude intensity at frequency f of the monitoring point in the nth frame, k is the intra-frame sample index, L is the window length, w(k) is the discrete value of the Hanning window, j is the imaginary unit, f is the frequency, and f s Where n is the sampling rate, n is the time frame index, x is the acoustic signal at the monitoring point, and x(n×L / 2+k) is the amplitude of the time-domain signal at a specific sample point.

[0063] The frequency band energy distribution of the spectrum signal at monitoring point A and the spectrum signal at monitoring point B are calculated respectively to obtain the frequency band energy distribution of monitoring point A and the frequency band energy distribution of monitoring point B.

[0064] Specifically, the spectral signal is divided into frequency bands based on a 1 / 3 octave band filter bank. The bandwidth of the 1 / 3 octave band division increases proportionally with the center frequency, which can capture the narrow frequency band energy concentrated in the cicada swarm signal while also taking into account the energy distribution of wideband natural sound. The energy value of each frequency band is calculated to quantify the signal strength in different frequency ranges. The formula for calculating the frequency band energy distribution is as follows:

[0065]

[0066] In the formula, E(f) b (n) represents the monitoring point at the center frequency f in the nth frame. b The total energy within the frequency band, X(f,n) is a complex spectral signal representing the signal amplitude intensity at frequency f of the monitoring point in the nth frame, where f is the signal strength. b Here, n is the center frequency of the b-th band in the 1 / 3 octave filter bank, n is the time frame index, and f is the frequency. b,high For the upper limit of the b-th frequency band, f = f b,low Δf represents the lower limit of the b-th frequency band, and Δf represents the frequency resolution.

[0067] Cross-point coherence calculation is performed based on the spectral signals from monitoring points A and B to obtain the cross-point coherence coefficient;

[0068] Specifically, large-scale sound sources such as cicada swarms exhibit high signal coherence at two monitoring points, while local sound sources, such as a single bird call or noise, have lower coherence. This can be effectively distinguished using the coherence coefficient. Furthermore, compared to energy analysis at a single monitoring point, cross-point coherence introduces spatial dimension features, reducing misjudgments caused by single-point interference and providing a reliable spatial consistency basis for determining the dominant period of cicada swarms. The cross-point coherence coefficient is calculated by analyzing the power spectral density relationship between the two signals, measuring the signal similarity of the same frequency band at the two monitoring points. The formula for calculating the cross-point coherence coefficient is as follows:

[0069]

[0070] In the formula, C AB (f b (n) represents the center frequency f of monitoring point A and monitoring point B in the nth frame. b The cross-point coherence coefficient within the frequency band, f b Here, n is the center frequency of the b-th band in the 1 / 3 octave filter bank, and f is the time frame index. b,high For the upper limit of the b-th frequency band, f = f b,low S is the lower limit of the b-th frequency band. AA (f,n) represents the self-power spectral density at monitoring point A, S BB (f,n) represents the self-power spectral density at monitoring point B, Δf is the frequency resolution, and S AB(f,n) represents the cross-power spectral density between monitoring points A and B.

[0071] The core narrowband determination module is used to determine the cross-point coherent core narrowband through local maxima relations;

[0072] In embodiments of the present invention, determining the core narrowband of cross-point coherence through local maxima relations includes:

[0073] The cross-point coherence is discretized into coherent sequences according to frequency;

[0074] The data in the coherent sequence are evaluated to identify local maxima, and these local maxima are then merged into the candidate set.

[0075] Specifically, the cross-point coherence coefficients of continuous frequency dimensions are discretized according to frequency band indices. The center frequency is replaced by the frequency band index to construct a discrete sequence. The sequence is arranged in ascending order of frequency band index to maintain the monotonicity of the frequency, forming a coherent sequence with the frequency band number as the horizontal axis, which facilitates the determination and calculation of local extrema. For any frequency band index in the coherent sequence, if the cross-point coherence coefficient of the index is greater than the cross-point coherence coefficient of its adjacent index, the index is determined to be a local maximum. All frequency band indices determined to be local maxima and their corresponding coherence coefficient values ​​are stored in the candidate set. By retaining all local maxima, the misjudgment of a single peak caused by noise interference is avoided, and multiple candidates are provided for subsequent screening.

[0076] The frequency of the maximum value in the candidate set is taken as the core center frequency. If there are ties for the maximum value, the one with the lower center frequency is selected as the core center frequency.

[0077] Starting from the core center frequency, the nearest local minimum points are located in the low-frequency and high-frequency directions respectively as boundaries. If there is no local minimum point at the endpoint, the frequency endpoint is taken as the boundary to determine the core narrowband band.

[0078] Specifically, the frequency band index with the largest coherence coefficient is searched in the candidate set. If multiple frequency band indices have equal coherence coefficients and are all maximum values, the frequency band index with the smallest coherence coefficient is selected as the core center frequency. The frequency band corresponding to the maximum value is the region with the highest coherence, corresponding to the dominant frequency band of large-scale homogeneous sound sources such as cicada swarms, which can ensure the representativeness of the core narrowband. When there are ties for the maximum value, the low frequency should be selected, which conforms to the characteristic that cicada swarm signals are mostly concentrated in the mid-to-low frequency band in nature, which can reduce the interference of high-frequency noise bands. Starting from the frequency band index corresponding to the selected core center frequency, the frequency band index is traversed to the left to find the nearest local minimum. The nearest local minimum must satisfy that it is less than the cross-point coherence coefficient of its adjacent index. If there is still no local minimum after traversing to the left endpoint, the left endpoint is taken as the left boundary. Starting from the frequency band index corresponding to the selected core center frequency, the frequency band index is traversed to the right to find the nearest local minimum. If there is still no local minimum after traversing to the right endpoint, the right endpoint is taken as the right boundary. The left boundary, the right boundary, and the area in between are taken as the core narrowband domain.

[0079] The energy relationship determination module is used to determine whether there is an inter-band slope reversal in the core narrowband based on the relationship between the energy of the core narrowband and adjacent frequency bands.

[0080] In embodiments of the present invention, determining whether there is an inter-band slope inversion in the core narrowband based on the relationship between the energy of the core narrowband and adjacent frequency bands includes:

[0081] Based on the energy of the core narrowband and adjacent frequency bands, the slope of the energy difference is obtained by calculating the logarithmic slope.

[0082] The slope is determined; if the slope is less than zero, it is determined that there is an inter-band slope reversal.

[0083] Specifically, the energy of the frequency band corresponding to the core center frequency is selected as the core narrowband energy. The adjacent indices of the core center frequency are used as adjacent frequency bands. The energy of the left and right adjacent frequency bands is obtained respectively. The slope of the energy change is calculated. Logarithmic transformation is used to eliminate the influence of absolute energy differences, converting the absolute energy differences into the slope value of the relative change trend. This quantifies the energy distribution pattern of the core narrowband and adjacent frequency bands. The specific formula is as follows:

[0084]

[0085] In the formula, k(n) is the logarithmic slope of the energy, E c E represents the energy value of the core narrowband. L E represents the energy value of the adjacent frequency band to the left. R represents the energy value of the adjacent frequency band on the right, and n is the time frame index;

[0086] Since the slope value reflects the symmetry of energy distribution, when the core narrowband is a strong narrowband signal, the energy presents a pattern of high in the middle and low on both sides. At this time, the energy of the frequency band corresponding to the core center frequency is greater than the energy of the adjacent frequency band on the left and the adjacent frequency band on the right, and the corresponding slope value is less than zero. Therefore, when the slope value is less than zero, it is determined that there is an inter-band slope reversal in the core narrowband; otherwise, it is determined that there is no inter-band slope reversal.

[0087] The periodic resonance determination module is used to extract the envelope on the core narrowband and confirm the existence of periodic resonance by the consistency between the envelope autocorrelation and the cross-point envelope crosscorrelation.

[0088] In an embodiment of the present invention, the envelope is extracted on the core narrowband, and the consistency between the envelope autocorrelation and the cross-point envelope crosscorrelation is used to confirm whether periodic resonance exists, including:

[0089] The envelope signal of the core narrowband frequency is extracted by Hilbert transform to obtain the envelope signals of monitoring point A and monitoring point B.

[0090] Specifically, the original time-domain signal is bandpass filtered to retain only the frequency components within the core narrowband, resulting in the core narrowband time-domain signals for monitoring point A and monitoring point B. A Hilbert transform is then performed on the core narrowband time-domain signals to obtain the analytic signals for monitoring point A and B. The magnitude of the analytic signal is the envelope signal. The periodicity of the cicada swarm signal is mainly reflected in the periodic fluctuations in amplitude; the envelope signal can directly capture this change, avoiding interference from high-frequency carrier waves in the periodic analysis.

[0091] Calculate the autocorrelation function and cross-correlation function of the envelope signals of monitoring point A and monitoring point B, and obtain the local maxima of the autocorrelation and cross-correlation.

[0092] The local maxima of autocorrelation and cross-correlation are determined. If the autocorrelation and cross-correlation functions have maxima at the same time delay point, then periodic resonance is determined to exist.

[0093] Specifically, the autocorrelation formula is used to calculate the correlation between the envelope signals of monitoring points A and B and themselves at different time delays.

[0094]

[0095] In the formula, R a (τ;n) represents the envelope autocorrelation function value of the nth frame at a delay of τ, where τ is the time offset, t0 is the window start time, T is the window length, a(t;n) is the envelope value at time t within the nth frame, and n is the time frame index;

[0096] The cross-correlation formula is:

[0097]

[0098] In the formula, Let be the cross-correlation function value of the envelope signals of monitoring points A and B at a delay of τ in the nth frame, where τ is the time offset, t0 is the window start time, T is the window length, n is the time frame index, and a A (t; n) represents the envelope value of detection point A at time t within the nth frame, a B '(t+τ;n) is the envelope value of detection point B at time t+τ within the nth frame;

[0099] By peak detection, the time offsets corresponding to the autocorrelation peaks and cross-correlation peaks of monitoring points A and B are determined. If there is a common time offset that is simultaneously the maximum value of the autocorrelation and cross-correlation of monitoring points A and B, then periodic resonance is determined to exist; otherwise, periodic resonance is determined not to exist.

[0100] The cicada swarm identification module is used to identify the corresponding time period as the cicada swarm's dominant period when cross-point coherence, inter-band slope inversion, and periodic resonance are all confirmed.

[0101] In embodiments of the present invention, when cross-point coherence, inter-band slope inversion, and periodic resonance are all confirmed, the corresponding time period is identified as the cicada swarm-dominated period, including:

[0102] If the monitoring period meets the condition that the cross-point coherence is greater than the preset coherence threshold and there is inter-band slope reversal and periodic resonance, then the period is regarded as the dominant period of the cicada swarm.

[0103] Specifically, for each time frame, the coherence coefficient of its core narrowband frequency band is selected, and the results of interband slope inversion and periodic resonance are associated with the time frame to form a frame-level feature determination list. If the coherence coefficient is greater than the coherence threshold and interband slope inversion and periodic resonance exist, the time frame is marked as a candidate frame for cicada swarm dominance. The coherence threshold is obtained by collecting cicada chirping and background noise samples in the target area and performing statistical analysis on the coherence values ​​of the cicada swarm activity period. If multiple consecutive time frames are candidate frames for cicada swarm dominance, their corresponding time periods are merged into a continuous cicada swarm dominance time period.

[0104] The sound correction module is used to perform spectrum separation and filtering on the sound signal during the period when the cicada swarm dominates, while retaining other natural sound source signals to obtain the corrected recognition result.

[0105] In an embodiment of the present invention, the acoustic signal during the dominant period of the cicada swarm is subjected to spectral separation and filtering, while retaining other natural sound source signals, to obtain a corrected recognition result, including:

[0106] Based on the core narrowband, construct the transfer function of the Butterworth bandpass filter;

[0107] The acoustic signal during the dominant cicada swarm period is input into a Butterworth bandpass filter to extract the cicada swarm signal component. The cicada swarm signal component is then subtracted from the acoustic signal during the dominant cicada swarm period to obtain the filtered other natural acoustic signals.

[0108] Specifically, the lower and upper limits of the core narrowband are calculated using geometric mean values ​​to obtain the center frequency of the filter, ensuring symmetrical coverage of the core narrowband. The bandwidth of the filter is obtained by subtracting the lower limit frequency from the upper limit frequency of the core narrowband, reflecting the frequency span of the core narrowband. The transfer function of the Butterworth bandpass filter is constructed using the center frequency and bandwidth. The transfer function formula is as follows:

[0109]

[0110] In the formula, H(f) is the transfer function of the filter, which represents the signal attenuation factor at frequency f, f is the frequency, f0 is the center frequency of the filter, B is the bandwidth of the filter, and N is the filter order.

[0111] The acoustic signal during the dominant cicada swarm period is input into a Butterworth bandpass filter. A Fourier transform is performed on the acoustic signal to obtain its spectrum. The spectrum is multiplied by the filter transfer function to obtain the spectrum of the cicada swarm signal component. An inverse Fourier transform is performed on the spectrum of the cicada swarm signal component to obtain the cicada swarm signal component in the time domain. The cicada swarm signal component is subtracted from the original acoustic signal to obtain the filtered other natural acoustic signals.

[0112] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An on-line monitor natural sound recognition device, characterized by, include: The acquisition module is used to simultaneously acquire acoustic signals from two monitoring points; The decomposition module is used to decompose the acoustic signal to obtain the energy distribution of each frequency band and cross-point coherence; The core narrowband determination module is used to determine the cross-point coherent core narrowband through local maxima relations; The energy relationship determination module is used to determine whether there is an inter-band slope reversal in the core narrowband based on the relationship between the energy of the core narrowband and adjacent frequency bands. The periodic resonance determination module is used to extract the envelope on the core narrowband and confirm the existence of periodic resonance by the consistency between the envelope autocorrelation and the cross-point envelope crosscorrelation. The cicada swarm identification module is used to identify the corresponding time period as the cicada swarm's dominant period when cross-point coherence, inter-band slope inversion, and periodic resonance are all confirmed. The sound correction module is used to perform spectral separation and filtering on the sound signal during the period when the cicada swarm dominates, while retaining other natural sound source signals to obtain the corrected recognition result.

2. The natural sound recognition device of an on-line monitor according to claim 1, wherein Simultaneous acquisition of acoustic signals from two monitoring points, including: Acquire the acoustic signal at monitoring point A and the acoustic signal at monitoring point B at the same sampling rate; By calculating the cross-correlation function of the signals from the two monitoring points during the observation period, the maximum delay is determined. Based on the maximum delay, the acoustic signal of monitoring point B is time-shifted and compensated to obtain the time-aligned acoustic signal of monitoring point B.

3. The online monitoring instrument natural voice recognition device according to claim 1, characterized in that, The acoustic signal is decomposed to obtain the energy distribution of each frequency band and cross-point coherence, including: Short-time Fourier transforms are performed on the acoustic signal at monitoring point A and the time-aligned acoustic signal at monitoring point B to obtain the spectral signals of monitoring point A and monitoring point B, respectively. The frequency band energy distribution of the spectrum signal at monitoring point A and the spectrum signal at monitoring point B are calculated respectively to obtain the frequency band energy distribution of monitoring point A and the frequency band energy distribution of monitoring point B. Cross-point coherence calculations are performed based on the spectral signals from monitoring points A and B to obtain the cross-point coherence coefficient.

4. The online monitoring instrument natural voice recognition device according to claim 1, characterized in that, The core narrowband of cross-point coherence is determined by local maxima relations, including: The cross-point coherence is discretized into coherent sequences according to frequency; The data in the coherent sequence are evaluated to identify local maxima, and these local maxima are then merged into the candidate set. The frequency of the maximum value in the candidate set is taken as the core center frequency. If there are ties for the maximum value, the one with the lower center frequency is selected as the core center frequency. Starting from the core center frequency, the nearest local minimum points are located in the low-frequency and high-frequency directions respectively as boundaries. If there is no local minimum point at the endpoint, the frequency endpoint is taken as the boundary to determine the core narrowband domain.

5. The online monitoring instrument natural voice recognition device according to claim 1, characterized in that, Based on the relationship between the energy of the core narrowband and adjacent frequency bands, it is determined whether there is an inter-band slope inversion in the core narrowband, including: Based on the energy of the core narrowband and adjacent frequency bands, the slope of the energy difference is obtained by calculating the logarithmic slope. The slope is determined, and if the slope is less than zero, it is determined that there is an inter-band slope reversal.

6. The online monitoring instrument natural voice recognition device according to claim 1, characterized in that, The envelope is extracted from the core narrowband, and the consistency between envelope autocorrelation and cross-point envelope cross-correlation is used to confirm the existence of periodic resonance, including: The envelope signal of the core narrowband frequency is extracted by Hilbert transform to obtain the envelope signals of monitoring point A and monitoring point B. Calculate the autocorrelation function and cross-correlation function of the envelope signals of monitoring point A and monitoring point B, and obtain the local maxima of the autocorrelation and cross-correlation. The local maxima of autocorrelation and cross-correlation are determined. If the autocorrelation and cross-correlation functions have maxima at the same time delay point, then periodic resonance is determined to exist.

7. The online monitoring instrument natural voice recognition device according to claim 1, characterized in that, With cross-point coherence, inter-band slope inversion, and periodic resonance all confirmed, the corresponding time period is identified as the cicada swarm-dominated period, including: If the monitoring period meets the condition that the cross-point coherence is greater than the preset coherence threshold and there is inter-band slope reversal and periodic resonance, then this period is regarded as the period when the cicada swarm is dominant.

8. The online monitoring instrument natural sound recognition device according to claim 1, characterized in that, The acoustic signal during the dominant cicada swarm period is subjected to spectral separation and filtering, while other natural sound source signals are preserved, resulting in a corrected identification result, including: Based on the core narrowband, the transfer function of the Butterworth bandpass filter is constructed, where the transfer function of the Butterworth bandpass filter is: In the formula, H(f) is the transfer function of the filter, which represents the signal attenuation factor at frequency f, f is the frequency, f0 is the center frequency of the filter, B is the bandwidth of the filter, and N is the filter order. The acoustic signal during the dominant cicada swarm period is input into a Butterworth bandpass filter to extract the cicada swarm signal component. The cicada swarm signal component is then subtracted from the acoustic signal during the dominant cicada swarm period to obtain the filtered natural acoustic signal.

Citation Information

Patent Citations

  • System and method for wind detection and suppression

    CN103348686A

  • Method and installation for processing a sequence of signals for polyphonic note recognition

    CN107210029A

  • Sound source orientation system

    CN118275972A

  • Acoustic automatic recognition system and method for birds in wetland environment

    CN120220702A

  • Determining an upperband signal from a narrowband signal

    US20110099004A1