A method, apparatus, device and medium for monitoring voice information of a target area
The method addresses the challenge of large noise data volumes in audio-based ecological monitoring by filtering noise and using sound pressure corrections and Fourier transforms to enhance data precision and reduce transmission complexity.
Patent Information
- Application Number
- CN202210224155.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-07
AI Technical Summary
When the prior art passes sound mode in ecological environment monitoring, the noise sound data is large, and the data transmission and analysis are difficult to effectively screen out useful sound data.
The sound information of the target area is obtained through the sound acquisition device, the sound pressure level correction value is calculated, and the noise screening process is performed, including the screening of background noise, rain noise and wind noise. The similarity comparison is used using pre-stored template data to determine useful sound data and generate monitoring data.
It improves the accuracy of noise screening, reduces the difficulty of data transmission and analysis, ensures that the monitoring data is useful sound data, and improves convenience.
Smart Images

Figure CN114758674B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing, and in particular to a method, device, equipment and medium for monitoring sound information of a target area. Background Art
[0002] In recent years, ecological environment monitoring based on acoustic information has attracted more and more attention. Environmental sounds contain a lot of information and can reflect the dynamic changes of the current environment to a certain extent. For example, the sounds that need attention are: 1) such as the sounds made by animals. Through the analysis of sound signals, we can know the activity level of animals in the area and further understand the species, abundance, and abundance; 2) gunshots. Through the gunshots of poachers, we can know when and where poaching is taking place. However, this way of monitoring the ecological environment through sound can theoretically be done 24 hours a day, but this uninterrupted method will generate a large amount of sound data containing noise. When storing or analyzing a large amount of sound data, it is difficult to transmit data or analyze the sound data of interest. Summary of the invention
[0003] In view of this, in order to solve at least one of the above technical problems, an object of the present invention is to provide a convenient method, device, equipment and medium for monitoring sound information of a target area.
[0004] The technical solution adopted in the embodiment of the present invention is:
[0005] A method for monitoring sound information in a target area, comprising:
[0006] Acquire first sound information of the target area through a sound collection device, and collect second sound information generated by a monitoring speaker at the position of the sound collection device;
[0007] Calculating a sound pressure level correction value according to the second sound information;
[0008] Performing noise screening processing according to the first sound information and the sound pressure level correction value to determine target sound data after noise screening; the noise screened by the noise screening processing includes at least one of background noise, rain noise and wind noise;
[0009] A similarity comparison process is performed based on the target sound data and the pre-stored template data to determine useful sound data, and monitoring data is generated based on the useful sound data.
[0010] Further, the first sound information includes a plurality of first data segments arranged in sequence, each of the first data segments has a plurality of data points, and the noise screening process is performed according to the first sound information and the sound pressure level correction value to determine the target sound data after the noise screening, including:
[0011] Determine a target data segment from the first data segment, and calculate the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value;
[0012] When the serial number corresponding to the target data segment is equal to the preset serial number threshold, determine the preset serial number threshold number of the smallest average sound pressure levels from the average sound pressure levels, and calculate the average sound pressure value of the preset serial number threshold number of the smallest average sound pressure levels;
[0013] Calculate the background noise value according to the smallest average sound pressure level and the average sound pressure value, and generate a background noise sound pressure level threshold according to the average sound pressure value;
[0014] When the average sound pressure level is less than or equal to the background noise sound pressure level threshold, use the target data segment as a candidate background noise segment and add it to the first set, determine a new target data segment, and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value;
[0015] When the average sound pressure level is greater than the background noise sound pressure level threshold and the number of candidate background noise segments in the first set is less than the count value threshold, add the candidate background noise segments in the first set to the second set and empty the first set, determine a new target data segment, and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value until all the first data segments are used as target data segments; the second set includes at least one first target data segment;
[0016] Perform at least one of rain noise screening and wind noise screening according to the first target data segment to obtain target sound data.
[0017] Further, the performing at least one of rain noise screening and wind noise screening according to the first target data segment to obtain target sound data includes:
[0018] Calculate the power spectrum of the first target data segment, and determine the first power value ranked first, the second power value ranked second, the first frequency corresponding to the first power value, and the second frequency corresponding to the second power value after arranging the power spectrum from largest to smallest;
[0019] Obtain the sampling rate of the first sound information, and calculate the weighted frequency according to the first power value, the second power value, the first frequency, the second frequency, the sampling rate, and the number of data points;
[0020] When the weighted frequency is greater than or equal to the rain noise threshold, determine the first target data segment as the second target data segment;
[0021] Or,
[0022] When the weighted frequency is less than the rain noise threshold, determine a first frequency band centered on the weighted frequency; the frequency band has an upper frequency and a lower frequency;
[0023] Calculate a first average power value of the power spectrum in the first frequency band, and calculate the variance of the power spectrum according to the first average power value;
[0024] Determine a second frequency band according to the upper frequency and a preset multiple of the upper frequency, and calculate a second average power value of the power spectrum in the second frequency band;
[0025] When a first ratio of the first average power value to the second average power value is less than a first preset threshold, and a second ratio of the first average power value to the variance is less than a second preset threshold, determine the first target data segment as the second target data segment;
[0026] Perform wind noise screening according to the second target data segment to obtain target sound data.
[0027] Further, the performing wind noise screening according to the second target data segment to obtain target sound data includes:
[0028] When the weighted frequency of the second target data segment is greater than or equal to the wind noise threshold, use the second target data segment as the target sound data;
[0029] Or,
[0030] When the weighted frequency of the second target data segment is less than the wind noise threshold, calculate the modulus value of the second target data segment, and calculate the variance of the second target data segment according to the maximum modulus value;
[0031] When the variance of the second target data segment is greater than or equal to the variance threshold, use the second target data segment as the target sound data.
[0032] Further, the calculating the sound pressure level correction value according to the second sound information includes:
[0033] Obtain the recording length of the second sound information; the second data segment of the second sound information;
[0034] Calculate the sum value of the squares of the second data segment;
[0035] Determine a correction parameter according to the ratio of the sum value to the recording length;
[0036] Determine the sound pressure level correction value according to the difference between the preset value and the correction parameter.
[0037] Further, the generation steps of the pre-stored template data include:
[0038] Obtain the concerned sound data;
[0039] Classify the concerned sound data according to the preset number of frequency points to obtain a number of audio segments;
[0040] Perform the first Fourier transform processing on the audio segments and calculate the first modulus value of the first Fourier transform processing result;
[0041] Calculate the average value of the first modulus value according to the preset number of frequency points, and normalize the maximum average value to obtain the pre-stored template data.
[0042] Further, the generation of the monitoring data according to the useful sound data includes:
[0043] Calculate the sound information according to the useful sound data, and perform the second Fourier transform processing on the useful sound data; the sound information includes at least one of an acoustic index, an acoustic diversity index, an acoustic uniformity index, an acoustic richness index, a bioacoustic index, an acoustic entropy index, and a normalized difference soundscape index;
[0044] Generate monitoring data according to the sound information, the second Fourier transform processing result, and the average sound pressure level corresponding to the useful sound data; the monitoring data is used to be uploaded to a monitoring platform for display.
[0045] An embodiment of the present invention further provides a sound information monitoring device for a target area, including:
[0046] An acquisition module, configured to acquire first sound information of a target area through a sound acquisition device, and acquire second sound information generated at the position of the sound acquisition device by a monitoring speaker;
[0047] A calculation module, configured to calculate a sound pressure level correction value according to the second sound information;
[0048] A screening module, configured to perform noise screening processing according to the first sound information and the sound pressure level correction value to determine target sound data after noise screening; the noise screened by the noise screening processing includes at least one of background noise, rain noise, and wind noise;
[0049] A generation module, configured to perform a similarity comparison process according to the target sound data and the pre-stored template data to determine useful sound data, and generate monitoring data according to the useful sound data.
[0050] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method.
[0051] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set, or an instruction set is stored, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the method.
[0052] The beneficial effects of the present invention are as follows: The first sound information of the target area is obtained through the sound collection device, and the second sound information generated at the position of the sound collection device by listening to the speaker is collected. The sound pressure level correction value is calculated according to the second sound information, and noise screening processing is performed according to the first sound information and the sound pressure level correction value to determine the target sound data after noise screening, which is beneficial to improving the accuracy of noise screening processing; similarity comparison processing is performed according to the target sound data and the pre-stored template data to determine the useful sound data, and monitoring data is generated according to the useful sound data, so that the final monitoring data used for monitoring is the useful sound data concerned after screening out noise. When the sound collection device transmits data and when analyzing the monitoring data, the amount of data transmitted and the difficulty of analysis are reduced, and the convenience is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic flow chart of the steps of the method for monitoring the sound information of the target area of the present invention;
[0054] Figure 2 It is a schematic diagram of the system structure of a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0056] The terms "first", "second", "third", "fourth", etc. in the description, claims and drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0057] Reference to "embodiment" herein means that a particular feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0058] As Figure 1 shown, an embodiment of the present invention provides a method for monitoring sound information of a target area, including steps S100 - S400:
[0059] S100. Obtain first sound information of the target area through a sound collection device, and collect second sound information generated at the position of the sound collection device by a monitoring speaker.
[0060] As Figure 2As shown in the figure, information can be collected through multiple sound collection devices. Each sound collection device has an omnidirectional pointing characteristic and includes a microphone, an amplifier, an analog-to-digital converter A / D, a microprocessor, a communication module, and an SD memory card. The microprocessor communicates with the monitoring platform through the communication module (such as wifi, Bluetooth, radio frequency, etc.). In the embodiment of the present invention, a microphone is used for sound collection, with a sensitivity ≥ 50mV / Pa, a frequency range of 100Hz - 16000Hz, a sound pressure level range of 30 - 115dB, a signal-to-noise ratio ≥ 60dB, and a harmonic distortion ≤ 10%; the amplifier amplifies the sound signal received by the microphone. When the sound pressure level is between 30 - 95dB, the amplification factor is fixed at ns times. It should be noted that when the sound pressure level is greater than 95dB, the circuit has a compression limiting function with a ratio of nv:1, and the values of ns and nv are determined by experiments. The output voltage range of the amplifier is the same as the input voltage range of the A / D. For example: when the sound pressure level is between 30 - 95, such as 95dB, the output voltage of the microphone is 30mV. Assuming the amplification factor ns = 10, the output of the amplifier is 30mV × 10 = 300mV; if the sound pressure level is 100dB, the output voltage of the microphone is 50mV. Assuming the amplification factor ns = 10 and nv = 2, the output of the amplifier is calculated in two parts. The output corresponding to 95dB of the microphone is 30mV, and the amplification factor is fixed at ns, so the corresponding output of the amplifier is 30mV × ns = 300mV. The remaining 50mV - 30mV = 20mV is caused by exceeding 95dB. This part of the output is not amplified by ns times but is compression limited, 20mV × ns / nv = 20 × 10 / 2 = 100mV. Thus, the total output is 300mV + 100mV = 400mV. In addition, the A / D converts the analog sound signal output by the amplifier into a digital signal, with a sampling rate fs ≥ 32000 and a sampling bit length of at least 16 bits. The signal read into the microprocessor from the output of the A / D is recorded as the first sound information x(n).
[0061] In the embodiment of the present invention, a monitoring speaker plays a sine signal of 1000Hz, so that the sound pressure level generated at the microphone of the sound collection device is 94dB, and thus the signal output by the A / D is read into the microprocessor, that is, the second sound information, recorded as x r (n), and the recording length is taken as N r = fs × 10.
[0062] S200. Calculate the sound pressure level correction value according to the second sound information.
[0063] Optionally, step S200 includes steps S210 - S240:
[0064] S210. Obtain the recording length of the second sound information.
[0065] It should be noted that the second sound information includes a second data segment, and the recording length of the second data segment is Nr = fs × 10.
[0066] S220. Calculate the sum value of the squares of the second data segment.
[0067] S230. Determine the correction parameter according to the ratio of the sum value to the recording length.
[0068] Specifically, the correction parameter L pr The calculation formula is: L pr = 20 × lg(pr), where:
[0069]
[0070] S240. Determine the sound pressure level correction value according to the difference between the preset value and the correction parameter.
[0071] Specifically, the calculation formula of the sound pressure level correction value ΔL is:
[0072] ΔL = A - L pr , where L pr = 20 × lg(pr)
[0073] where A is the preset value, including but not limited to 94.
[0074] S300. Perform noise screening processing according to the first sound information and the sound pressure level correction value, and determine the target sound data after noise screening.
[0075] It should be noted that the noises screened by the noise screening processing include background noise, rain noise, and wind noise. In other embodiments, one or more of them may be included, and the order of screening noises can be adjusted as needed. In the embodiments of the present invention, the order of screening noises is taken as an example of background noise, rain noise, and wind noise, and no specific limitation is made. It can be understood that the target sound data is the sound information remaining after screening background noise, rain noise, and wind noise from the first sound information.
[0076] Optionally, step S300 includes steps S311 - S316:
[0077] S311. Determine a target data segment from the first data segment, and calculate the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value.
[0078] It should be noted that the first sound information x(n) includes several j' first data segments arranged in sequence, and each first data segment has multiple data blocks x i (n), and each data block x i(n) has multiple data points. i is the serial number of the data block. In the microprocessor, every N data points in the received first sound signal are used as a data block x i (n), a first data segment has multiple data blocks x i (n), where n = 0, 1, 2,..., N - 1, i = 1, 2,..., M, there are a total of M data blocks for extracting the first sound information. It should be noted that the duration of the first sound information is set as needed, for example, between 1 second and 10 seconds. For example: Suppose the signal sampling rate is 32k (32000), and 100s of data is collected. If every 10s is used as a first data segment, then there are 10 first data segments in total, that is, j' is 10. The serial number of the first data segment is represented by j, and the value range of j is 1, 2, 3,..., 10. If every 0.01s of the signal is used as a data block, then each 10s first data segment is divided into 1000 data blocks. The serial number of the data block is represented by i, and the value range of i is 1, 2, 3,..., 1000; each data block is 0.01s, the number of data points is 32000 * 0.01 = 320, and the serial number of the data points is represented by n, and the value range of n is 0, 1, 2,..., 319.
[0079] In the embodiment of the present invention, a target data segment is determined from the first data segments. For example, the j-th first data segment is determined as the target data segment from j' first data segments, where j = 1, 2... j'. Then, according to the number of data points of the target data segment and the sound pressure level correction value, the average sound pressure level L p (j) is calculated, where x in the following formula (1) i (n) refers to the data block under the j-th target data segment. Specifically:
[0080] L p (j) = 10lg E + ΔL, where
[0081] S312. When the serial number corresponding to the target data segment is equal to the preset serial number threshold, determine the preset serial number threshold number of the smallest average sound pressure levels from the average sound pressure levels, and calculate the average sound pressure value of the preset serial number threshold number of the smallest average sound pressure levels.
[0082] It should be noted that the preset serial number threshold NL can be set as needed. Optionally, when j = NL, determine the preset serial number threshold number NL of the smallest average sound pressure levels MIN_L p from the average sound pressure levels. NL is set as needed (including but not limited to 600). For example, NL is 2, that is, determine the smallest average sound pressure level and the second smallest average sound pressure level, and calculate the average sound pressure value MEAN_L p of the NL smallest average sound pressure levels MIN_L p .
[0083] S313. Calculate the background noise value based on the minimum average sound pressure level and the average sound pressure value, and generate a background noise sound pressure level threshold based on the average sound pressure value.
[0084] It should be noted that generating the background noise sound pressure level threshold means generating the background noise sound pressure level threshold when there is no background noise sound pressure level threshold, or updating the initial value of the background noise sound pressure level threshold to obtain a new background noise sound pressure level threshold when the background noise sound pressure level threshold has an initial value. In the embodiments of the present invention, the update of the initial value of the background noise sound pressure level threshold is taken as an example.
[0085] Optionally, calculate the background noise value BG_L p = αMIN_L p +(1 - α)MEAN_L p , where α is obtained by debugging according to the actual scenario, and the default value includes but is not limited to 0.8. Among them, assuming that the default value of the background noise sound pressure level threshold is TH_'L p , then the background noise sound pressure level threshold TH_L p = βTH_'L p +(1 - β)BG_L p , where β is obtained by debugging according to the actual scenario, and the default value includes but is not limited to 0.7.
[0086] S314. When the average sound pressure level is less than or equal to the background noise sound pressure level threshold, take the target data segment as a candidate background noise segment and add it to the first set, determine a new target data segment, and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value.
[0087] Optionally, when the average sound pressure level L p (j) ≤ TH_L p , at this time, the target data segment is a candidate background noise segment and is added to the first set. At this time, update the count value of the candidate background noise segment counter CandidateBNCnt = CandidateBNCnt (the initial value is 0)+1. This count value represents the number of candidate background noise segments in the first set. Then, determine a new target data segment (for example, j + 1) from the first data segment according to the serial number, and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value in step S311.
[0088] S315. When the average sound pressure level is greater than the background noise sound pressure level threshold and the number of candidate background noise segments in the first set is less than the count value threshold, add the candidate background noise segments in the first set to the second set, clear the first set, determine a new target data segment, and return the step of calculating the average sound pressure level of the target data segment based on the number of data points of the target data segment and the sound pressure level correction value, until all the first data segments are used as target data segments; the second set includes at least one first target data segment.
[0089] Optionally, when the average sound pressure level L p (j)>TH_L p , check the value of CandidateBNCnt, that is, the number of candidate background noise segments in the first set. When CandidateBNCnt≥count value threshold TH_BNN, the first CandidateCnt segments of signals are background noise, that is, all candidate background noise segments in the first set are background noise. These background noise segments need to be discarded and not processed subsequently, and set CandidateBNCnt = 0, that is, clear the first set, where the count value threshold TH_BNN is determined by experiments, and the default value includes but is not limited to 3.
[0090] It can be understood that when CandidateBNCnt is less than the count value threshold TH_BNN, the first CandidateCnt segments of signals are not background noise, that is, all candidate background noise segments in the first set are not background noise. These background noise segments need to be retained. Therefore, add the candidate background noise segments in the first set to the second set and clear the first set, that is, set CandidateBNCnt = 0. Then, determine a new target data segment from the first data segments according to the serial number (for example, j + B (B = 1, 2, 3, 4...)), and return the step of calculating the average sound pressure level of the target data segment based on the number of data points of the target data segment and the sound pressure level correction value in step S311, until all the first data segments are used as target data segments. At this time, the final second set is obtained. At this time, the candidate background noise segments in the second set are not background noise, denoted as the first target data segments, and the second set may include at least one first target data segment.
[0091] S316. Perform at least one of rain noise screening and wind noise screening based on the first target data segment to obtain the target sound data.
[0092] It should be noted that the embodiments of the present invention include rain noise screening and wind noise screening, and one of them may be included in other embodiments, which is not specifically limited.
[0093] Optionally, the steps of rain noise screening include step S321 or S322, and include step S323. Specifically, step S321 includes steps S3211 - S3213, and S322 includes steps S3221 - S3224:
[0094] S3211. Calculate the power spectrum of the first target data segment, and determine the first power value ranked first, the second power value ranked second, the first frequency corresponding to the first power value, and the second frequency corresponding to the second power value after arranging the power spectrum from largest to smallest.
[0095] In the embodiment of the present invention, taking the jth first data segment as the first target data segment as an example for illustration, calculate the power spectrum of the first target data segment k = 0, 1, 2,..., N / 2 or N - 1, X i (k) is the result obtained by performing a discrete Fourier transform DFT on x i (n) in formula (1), and determine the first power value P(k m1 )(i.e., the maximum value) ranked first, the second power value P(k m2 )(i.e., the second - largest value) ranked second, the first frequency k m1 and the second frequency k m2 .
[0096] S3212. Obtain the sampling rate of the first sound information, and calculate the weighted frequency according to the first power value, the second power value, the first frequency, the second frequency, the sampling rate, and the number of data points.
[0097] Specifically, the calculation formula for the weighted frequency f Pm is:
[0098]
[0099] where fs is the sampling rate and N is the number of data points.
[0100] S3213. When the weighted frequency is greater than or equal to the rain noise threshold, determine the first target data segment as the second target data segment.
[0101] Optionally, when the weighted frequency f Pm ≥TH_F rain (rain noise threshold), then the first target data segment is not rain noise, denoted as the second target data segment, and further wind noise screening is required. It should be noted that the rain noise threshold can be set as needed, for example, the default value is 1500.
[0102] S3221. When the weighted frequency is less than the rain noise threshold, determine the first frequency band centered on the weighted frequency; the frequency band has an upper frequency and a lower frequency.
[0103] When the weighted frequency f Pm <TH_F rain (rain noise threshold), take a first frequency band centered on the weighted frequency f Pm The lower limit frequency of the first frequency band is f L =f Pm -Δf, and the upper limit frequency is f H =f Pm +Δf, where Δf can be adjusted by experiments, for example, the default value is 300.
[0104] S3222. Calculate the first average power value of the power spectrum in the first frequency band, and calculate the variance of the power spectrum according to the first average power value.
[0105] S3223. Determine a second frequency band according to the upper limit frequency and a preset multiple of the upper limit frequency, and calculate the second average power value of the power spectrum in the second frequency band.
[0106] Optionally, convert f L , f H to the corresponding frequency index values. The formula is: The symbol refers to rounding down. Then, take the average value of the power spectrum P(k). For k L ≤k≤k H , the average value is denoted as meanP m (the first average power value), and calculate the variance of the power spectrum according to the first average power value, denoted as σ; for the second frequency band: k H <k≤3k H (taking 3 times as an example for the preset multiple, not limited to 3 times), the average value is denoted as meanP s (the second average power value). Take the variance of P(k), denoted as σ, where k L ≤k≤k H .
[0107] S3224. When the first ratio of the first average power value to the second average power value is less than the first preset threshold, and the second ratio of the first average power value to the variance is less than the second preset threshold, determine the first target data segment as the second target data segment.
[0108] When the first ratio of the first average power value to the second average power value is less than the first preset threshold TH MEAN and the second ratio of the first average power value to the variance is less than the second preset threshold TH_STD:
[0109]
[0110]
[0111] If the current data segment is not rain noise, the first target data segment is determined as the second target data segment, and further wind noise screening is required; otherwise, it is rain noise, and a new target data segment (for example, j + B (B = 1, 2, 3, 4...)) is determined from the first data segment according to the serial number, and the step of calculating the average sound pressure level of the target data segment based on the number of data points and the sound pressure level correction value of the target data segment in step S311 is returned. Optionally, TH_MEAN and TH_STD can be adjusted by experiments, for example, the default values are 2 and 3 respectively.
[0112] S323. Perform wind noise screening based on the second target data segment to obtain the target sound data.
[0113] Optionally, step S323 includes step S3231, or includes S3232 - S3233:
[0114] S3231. When the weighted frequency of the second target data segment is greater than or equal to the wind noise threshold, the second target data segment is used as the target sound data.
[0115] Optionally, when the weighted frequency f Pm ≥TH_E wind (wind noise threshold), then the second target data segment is not wind noise, and the second target data segment is used as the target sound data. It should be noted that the target sound data finally obtained in the screening iteration process can include several second target data segments; TH_F wind can be adjusted by experiments, for example, the default value is 750.
[0116] S3232. When the weighted frequency of the second target data segment is less than the wind noise threshold, calculate the modulus value of the second target data segment, and calculate the variance of the second target data segment based on the maximum modulus value.
[0117] S3233. When the variance of the second target data segment is greater than or equal to the variance threshold, the second target data segment is used as the target sound data.
[0118] Optionally, when the weighted frequency f Pm <TH_F wind (wind noise threshold), then the second target data segment is wind noise, calculate the modulus value of the second target data segment, and find the maximum modulus value and store it in the array am(i); where i = 1, 2,..., M. Take the variance of am(i), denoted as δ am (variance of the second target data segment). If δ am ≥TH_δ am(Variance threshold), if so, the current second target data segment is not wind noise, and the second target data segment is used as the target sound data; otherwise, it is wind noise, and a new target data segment is determined from the first data segment according to the serial number (for example, j + B (B = 1, 2, 3, 4... M - 1)), and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value in step S311. Optionally, TH_δ am Can be adjusted by experiment, for example, the default value is 0.08.
[0119] S400. Perform a similarity comparison process on the target sound data and the pre-stored template data to determine the useful sound data, and generate monitoring data according to the useful sound data.
[0120] Optionally, the generation steps of the pre-stored template data include S411 - S414:
[0121] S411. Obtain the sound data of interest.
[0122] S412. Classify the sound data of interest according to the preset number of frequency points to obtain several audio segments.
[0123] S413. Perform the first Fourier transform process on the audio segments and calculate the first modulus value of the first Fourier transform process result.
[0124] S414. Calculate the average value of the first modulus value according to the preset number of frequency points, and normalize the maximum average value to obtain the pre-stored template data.
[0125] Optionally, the sound data of interest can be set according to requirements. For example, the sound data files of rare animal sounds, poaching gunshots, blasting sounds, etc. Record the data of the sound data file as y(n), and each N points (the default value of N is 512) of y(n) is an audio segment. Therefore, y(n) has several audio segments. Perform the first Fourier transform process (including but not limited to discrete Fourier transform) on the audio segments, and calculate the first modulus value of the first Fourier transform process result. Then calculate the average value of the first modulus value according to the preset number of frequency points N, and normalize the maximum average value to obtain the pre-stored template data Y avg (k), k = 0, 1, 2,... N / 2 represents the frequency points. It should be noted that the pre-stored template data is stored in the SD memory card.
[0126] Optionally, the steps of generating monitoring data according to the useful sound data include S421 - S422:
[0127] S421. Calculate the sound information according to the useful sound data and perform the second Fourier transform process on the useful sound data.
[0128] Optionally, the useful sound data is considered as the sound data with the sound of interest. The useful sound data is stored in the SD memory card in the wav file format for calling when needed. Optionally, the sound information includes but is not limited to the Acoustic Index ACI, the Acoustic Complexity Index (ACI), the Acoustic Diversity Index (ADI), the Acoustic Evenness Index (AEI), the Acoustic Richness Index (AR), the Bioacoustic Index (BI), the Acoustic Entropy Index (H), the Normalized Difference Soundscape Index (NDSI); the useful sound data is subjected to a second Fourier transform process, including but not limited to the discrete Fourier transform, to obtain the second Fourier transform process result X l (k), where l is the serial number of the second target data segment retained in the useful sound data.
[0129] S422. Generate monitoring data according to the sound information, the second Fourier transform process result, and the average sound pressure level corresponding to the useful sound data.
[0130] Optionally, the monitoring data is used to be uploaded to the monitoring platform for display. Specifically, the sound information, the modulus value of the second Fourier transform process result, and the average sound pressure level corresponding to the useful sound data are used as the monitoring data. During the non-useful sound signal time period, that is, background noise, rain noise, and wind noise, that is, during the idle period of the microprocessor, the monitoring data is uploaded to the monitoring platform through the communication module for display.
[0131] Optionally, the monitoring platform can process and display the monitoring data. For example: the monitoring platform draws and displays a spatio-temporal distribution map according to the monitoring data of multiple sound acquisition devices; in addition, for a sound that may be of special concern, the modulus value of the second Fourier transform process result is further identified and confirmed by the method of neural network, and the discrimination result is output and displayed. In addition, the spatio-temporal distribution map and the discrimination result can be displayed or given an early warning through the monitoring platform or sent to the terminal.
[0132] Optionally, the step of performing a similarity comparison process according to the target sound data and the pre-stored template data is specifically: calculating the modulus value result k = 0, 1, 2,..., N / 2, where X i (k) is the result obtained by Fourier transform calculation of the target sound data, and find Xabs The maximum value of (k), max[X abs (k)], normalizes X abs (k), that is respectively compares the similarity with the pre-stored template data Y avg (k) stored in the SD memory card in advance: calculates A d (k) = X avg (k) - Y avg (k), and then calculates the sum of squares of A d (k) and the variance Optionally, if and σ Ad <TH_σAD, the current target sound data may be a sound that needs special attention, that is, useful sound data. Optionally, each sound that needs special attention has a corresponding third threshold TH_AD and a fourth threshold TH_σAD, and the specific values are determined through experiments.
[0133] The method for monitoring the sound information of the target area in the embodiment of the present invention can identify invalid sound signals in the 24-hour recording every day, such as background noise, rain noise, and wind noise, and determine the useful sound data of concern for transmission monitoring, greatly reducing the amount of monitoring data, greatly reducing the power consumption and bandwidth of remote wireless transmission, and being able to give an early warning when a sound signal that needs special attention is collected, with good monitoring effects.
[0134] The embodiment of the present invention also provides a device for monitoring the sound information of a target area, including:
[0135] An acquisition module, configured to acquire the first sound information of the target area through a sound acquisition device, and acquire the second sound information generated at the position of the sound acquisition device through a monitoring speaker;
[0136] A calculation module, configured to calculate a sound pressure level correction value according to the second sound information;
[0137] A screening module, configured to perform noise screening processing according to the first sound information and the sound pressure level correction value to determine the target sound data after noise screening; the noise screened by the noise screening processing includes at least one of background noise, rain noise, and wind noise;
[0138] A generation module, configured to perform similarity comparison processing according to the target sound data and the pre-stored template data to determine useful sound data, and generate monitoring data according to the useful sound data.
[0139] The content in the above method embodiment is applicable to the device embodiment of the present invention. The functions specifically implemented by the device embodiment of the present invention are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method embodiment.
[0140] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method for monitoring sound information in the target area in the foregoing embodiment. The electronic device in the embodiment of the present invention includes, but is not limited to, any intelligent terminal such as a mobile phone, a tablet computer, a computer, and an in-vehicle computer.
[0141] The content in the foregoing method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those in the foregoing method embodiments, and the beneficial effects achieved are also the same as those in the foregoing method embodiments.
[0142] An embodiment of the present invention further provides a computer-readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method for monitoring sound information in the target area in the foregoing embodiment.
[0143] An embodiment of the present invention further provides a computer program product or a computer program, which includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for monitoring sound information in the target area in the foregoing embodiment.
[0144] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to the clearly listed those steps or units, but may include other steps or units that are not clearly listed or are inherent to these process, method, product or device.
[0145] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (piece) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one (piece) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0146] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0147] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0148] The above, the above embodiments are only used to illustrate the technical solution of this application and are not intended to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A method for monitoring voice information of a target area, characterized in that, Including: Obtaining first sound information of a target area through a sound acquisition device, and collecting second sound information generated at the position of the sound acquisition device by a monitoring speaker; Calculating a sound pressure level correction value according to the second sound information; Performing noise screening processing according to the first sound information and the sound pressure level correction value to determine target sound data after noise screening; The noise screened by the noise screening processing includes at least one of background noise, rain noise, and wind noise; Performing a similarity comparison process according to the target sound data and pre-stored template data to determine useful sound data, and generating monitoring data according to the useful sound data; The first sound information includes a plurality of first data segments arranged in sequence, and each of the first data segments has a plurality of data points. The performing noise screening processing according to the first sound information and the sound pressure level correction value to determine target sound data after noise screening includes: Determining a target data segment from the first data segments, and calculating an average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value; When the serial number corresponding to the target data segment is equal to a preset serial number threshold, determining a preset serial number threshold number of the smallest average sound pressure levels from the average sound pressure levels, and calculating an average sound pressure value of the preset serial number threshold number of the smallest average sound pressure levels; Calculating a background noise value according to the smallest average sound pressure level and the average sound pressure value, and generating a background noise sound pressure level threshold according to the average sound pressure value; When the average sound pressure level is less than or equal to the background noise sound pressure level threshold, taking the target data segment as a candidate background noise segment and adding it to a first set, determining a new target data segment, and returning to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value; When the average sound pressure level is greater than the background noise sound pressure level threshold and the number of candidate background noise segments in the first set is less than a count threshold, adding the candidate background noise segments in the first set to a second set and clearing the first set, determining a new target data segment, and returning to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value until all the first data segments are used as target data segments; the second set includes at least one first target data segment; Performing at least one of rain noise screening and wind noise screening according to the first target data segment to obtain target sound data.
2. The method for monitoring the sound information of the target area according to claim 1, wherein: The performing at least one of rain noise screening and wind noise screening according to the first target data segment to obtain target sound data includes: Calculating a power spectrum of the first target data segment, and determining a first power value ranked first, a second power value ranked second, a first frequency corresponding to the first power value, and a second frequency corresponding to the second power value after arranging the power spectrum from largest to smallest; Obtain the sampling rate of the first sound information, and calculate a weighted frequency according to the first power value, the second power value, the first frequency, the second frequency, the sampling rate, and the number of data points. When the weighted frequency is greater than or equal to the rain noise threshold, determine the first target data segment as the second target data segment. Or, When the weighted frequency is less than the rain noise threshold, determine a first frequency band centered on the weighted frequency; the frequency band has an upper limit frequency and a lower limit frequency. Calculate a first average power value of the power spectrum in the first frequency band, and calculate the variance of the power spectrum according to the first average power value. Determine a second frequency band according to the upper limit frequency and a preset multiple of the upper limit frequency, and calculate a second average power value of the power spectrum in the second frequency band. When a first ratio of the first average power value to the second average power value is less than a first preset threshold, and a second ratio of the first average power value to the variance is less than a second preset threshold, determine the first target data segment as the second target data segment. Perform wind noise screening according to the second target data segment to obtain target sound data.
3. The method for monitoring the sound information of the target area according to claim 2, wherein: The performing wind noise screening according to the second target data segment to obtain target sound data includes: When the weighted frequency of the second target data segment is greater than or equal to the wind noise threshold, use the second target data segment as the target sound data. Or, When the weighted frequency of the second target data segment is less than the wind noise threshold, calculate the modulus value of the second target data segment, and calculate the variance of the second target data segment according to the maximum modulus value. When the variance of the second target data segment is greater than or equal to the variance threshold, use the second target data segment as the target sound data.
4. The method for monitoring the sound information of the target area according to claim 1, characterized in that: The calculating the sound pressure level correction value according to the second sound information includes: Obtain the recording length of the second sound information; the second sound information includes a second data segment. Calculate the sum value of the squares of the second data segment. Determine a correction parameter according to the ratio of the sum value to the recording length. Determine the sound pressure level correction value according to the difference between a preset value and the correction parameter.
5. The method for monitoring the sound information of the target area according to claim 1, characterized in that: The generating step of the pre-stored template data includes: Obtain the concerned sound data. Classify the concerned sound data according to a preset number of frequency points to obtain a plurality of audio segments. Perform a first Fourier transform process on the audio segments, and calculate a first modulus value of the first Fourier transform process result. Calculate the average value of the first modulus value according to the preset number of frequency points, and perform normalization processing on the maximum average value to obtain the pre-stored template data.
6. The method for monitoring the sound information of the target area according to any one of claims 2-5, characterized in that: The generating the monitoring data according to the useful sound data includes: Calculate sound information according to the useful sound data, and perform a second Fourier transform process on the useful sound data; the sound information includes at least one of an acoustic index, an acoustic diversity index, an acoustic uniformity index, an acoustic richness index, a bioacoustic index, an acoustic entropy index, and a normalized difference soundscape index. Generate monitoring data based on the sound information, the second Fourier transform processing result, and the average sound pressure level corresponding to the useful sound data; the monitoring data is used to be uploaded to a monitoring platform for display.
7. An apparatus for monitoring sound information of a target area, characterized in that, It includes: An acquisition module, configured to acquire first sound information of a target area through a sound acquisition device, and acquire second sound information generated at the position of the sound acquisition device by a monitoring speaker; A calculation module, configured to calculate a sound pressure level correction value according to the second sound information; A screening module, configured to perform noise screening processing according to the first sound information and the sound pressure level correction value to determine target sound data after noise screening; The noise screened by the noise screening processing includes at least one of background noise, rain noise, and wind noise; A generation module, configured to perform similarity comparison processing according to the target sound data and pre-stored template data to determine useful sound data, and generate monitoring data according to the useful sound data; Wherein, the first sound information includes a plurality of first data segments arranged in sequence, each of the first data segments has a plurality of data points, and the screening module is specifically configured to: Determine a target data segment from the first data segments, and calculate the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value; When the sequence number corresponding to the target data segment is equal to a preset sequence number threshold, determine the preset sequence number threshold number of the smallest average sound pressure levels from the average sound pressure levels, and calculate the average sound pressure value of the preset sequence number threshold number of the smallest average sound pressure levels; Calculate the background noise value according to the smallest average sound pressure level and the average sound pressure value, and generate a background noise sound pressure level threshold according to the average sound pressure value; When the average sound pressure level is less than or equal to the background noise sound pressure level threshold, use the target data segment as a candidate background noise segment and add it to the first set, determine a new target data segment, and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value; When the average sound pressure level is greater than the background noise sound pressure level threshold and the number of candidate background noise segments in the first set is less than a count threshold, add the candidate background noise segments in the first set to the second set and clear the first set, determine a new target data segment, and return to the step of calculating the average sound pressure level of the target data segment according to the number of data points of the target data segment and the sound pressure level correction value until all the first data segments are used as target data segments; the second set includes at least one first target data segment; Perform at least one of rain noise screening and wind noise screening according to the first target data segment to obtain target sound data.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method and system for monitoring abnormal sound of pigs
CN110189756A
Sound pressure level intelligent control method based on sound reinforcement system
CN111328008A