An Adaptive Background Noise Detection Method, System and Medium
By performing fast Fourier conversion and steady-state and dynamic statistical analysis on the sound signals in the background noise detection interval, background noise is identified, and the problem of low detection accuracy in the prior art is solved, and high-accurate background noise detection is achieved.
Patent Information
- Application Number
- CN202111512446.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing background noise detection methods cannot be adjusted in time according to changes in environmental noise, resulting in low detection accuracy.
By obtaining the sound signal in the background noise detection interval, performing fast Fourier conversion and steady-state and dynamic statistical analysis, we can determine whether the spectral amplitude and variation number of the sound signal are stable, and we can identify background noise.
Timely adjustments are achieved according to changes in environmental noise, and the accuracy of background noise detection is improved.
Smart Images

Figure CN114220446B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of noise detection, and particularly relates to an adaptive background noise detection method, system and medium. Background Art
[0002] In systems such as sound recording, communication, and detection, background noise is often accompanied, which affects the sound quality. One is to cause auditory interference, and the other is to reduce the detection accuracy. Therefore, the elimination or suppression of background noise has become a very important issue.
[0003] At present, there are endless sound noise reduction methods, but the noise reduction methods based on a single microphone are very limited. Usually, it is based on the method of spectral subtraction, and this method requires an effective estimation of background noise to achieve its maximum effect. Otherwise, it may have the opposite effect and generate extra noise. The voice activity detection (VAD) in the prior art is usually judged based on the sound energy level. However, due to the change of noise in the environment, this detection method cannot make timely adjustments according to the change of noise, so the detection accuracy is not ideal.
[0004] Therefore, it is particularly important to provide a background noise detection method, system and medium that can make timely adjustments according to the change of the detected environmental noise, so as to achieve high detection accuracy. Summary of the Invention
[0005] In order to solve the technical problem that the voice activity detection method in the prior art cannot make timely adjustments according to the change of the detected environmental noise, resulting in low detection accuracy, the present invention proposes an adaptive background noise detection method, system and medium.
[0006] According to the first aspect of the present application, an adaptive background noise detection method is proposed, including the following steps:
[0007] S1. Obtain the sound signal within the background noise detection interval;
[0008] S2. Perform a fast Fourier transform on each frame of the sound signal to estimate the spectral amplitude of the sound signal;
[0009] S3. Perform a steady-state statistical analysis on the spectral amplitude of the sound signal to calculate the change of the spectral amplitude of the sound signal;
[0010] S4. Perform a dynamic statistical analysis on the spectral amplitude of the sound signal to calculate the change of the spectral amplitude variance of the sound signal; and
[0011] S5. Determine whether the sound signal is background noise according to the change of the spectral amplitude of the sound signal and the change of the variance of the spectral amplitude of the sound signal.
[0012] By collecting the sound signal within the background noise detection interval and performing steady-state statistical analysis and dynamic statistical analysis on the sound signal, according to the analysis results, if it is judged that both the spectral amplitude and the variance of the spectral amplitude of the sound signal remain in a stable state, it indicates that the sound signal is background noise, and record the parameters of the sound signal as the basis for eliminating background noise. This method can make timely adjustments according to the change of the detected environmental noise, effectively identify background noise, and has high detection accuracy.
[0013] Preferably, the step S3 specifically includes:
[0014] S31. Calculate the average spectral amplitude of the sound signal within 1 second and continuously record it for 5 seconds;
[0015] S32. Calculate the first average standard deviation between the average spectral amplitudes of the corresponding 5 sound signals within 5 seconds;
[0016] S33. Determine whether the first average standard deviation is less than the first threshold. If so, it indicates that the spectral amplitude of the sound signal remains in a stable state.
[0017] By comparing the second average standard deviation between the average spectral amplitudes of 5 sound signals within 5 seconds with the first threshold, it can reflect whether the spectral amplitude of the sound signal remains in a stable state during this time, thus serving as one of the bases for determining whether the sound signal is background noise.
[0018] Preferably, the step S4 specifically includes:
[0019] S41. Calculate the variance of the average spectral amplitude of the sound signal within 1 second and continuously record it for 5 seconds. The specific calculation formula is:
[0020]
[0021] Where V x,j (k) represents the variance of the average spectral amplitude of the sound signal within the jth second, X i (k) represents the spectral amplitude of the ith frame signal, X a,j (k) represents the average spectral amplitude of the sound signal within the jth second, and N represents the number of frames in the 1-second sound signal;
[0022] S42. Calculate the second average standard deviation between the variances of the average spectral amplitudes of the corresponding 5 sound signals within 5 seconds;
[0023] S43. Determine whether the second average standard deviation is less than the second threshold. If so, it indicates that the spectral amplitude variance of the sound signal remains in a stable state.
[0024] By comparing the second average standard deviation between the average spectral amplitude variances of 5 sound signals within 5 seconds with the second threshold, it can be reflected whether the spectral amplitude variance of the sound signal remains in a stable state during this time, thus serving as one of the bases for determining whether the sound signal is background noise.
[0025] Preferably, the step S1 specifically includes:
[0026] S11. Collect a sound signal and perform preprocessing on the sound signal.
[0027] S12. Set the initial detection state of the sound signal and perform preprocessing on the sound signal.
[0028] S13. Estimate the sound energy of the sound signal and determine whether the sound energy of the sound signal is less than the third threshold. If so, the sound signal belongs to the background noise detection interval, and step S2 is executed. If not, return to the above step S12.
[0029] By estimating the sound energy of the processed sound signal and comparing it with the third threshold, if the sound energy of the sound signal is less than the third threshold, it indicates that the sound signal is within the background noise detection interval.
[0030] Further preferably, the step S11 specifically includes:
[0031] S111. Collect a sound signal and convert the sound signal into a voltage signal.
[0032] S112. Perform amplification processing on the voltage signal.
[0033] S113. Perform filtering processing on the amplified voltage signal to adjust the spectral response of the voltage signal.
[0034] S114. Convert the voltage signal into a digital signal.
[0035] After the preprocessing step, the collected sound signal is converted into a digital signal recognizable in the subsequent detection steps.
[0036] Further preferably, setting the initial detection state of the sound signal in the step S12 specifically includes:
[0037] S121. Set the sampling time, calculate the average value of the initial parameters of the sound signal within the sampling time as the initial parameters of the sound signal, and dynamically update the initial parameters of the sound signal.
[0038] Further preferably, the preprocessing of the sound signal in step S12 specifically includes:
[0039] S122. Extract each frame of the signal from the sound signal and perform spectral equalization processing on the sound signal.
[0040] Extracting each frame of the signal from the sound signal can reduce the distortion on the spectrum. In addition to compensating for the distortion during the sound collection process, spectral equalization processing can also be used to emphasize a certain frequency band and increase or decrease the weight of background noise detection in this frequency band.
[0041] Further preferably, the estimation of the sound energy of the sound signal in step S13 specifically includes: estimating the sound energy of each frame of the signal in the sound signal, and the specific estimation formula is:
[0042]
[0043] where E i represents the sound energy of the i-th frame of the signal, x(n) represents the frame signal corresponding to the i-th frame, and k represents the total number of frames of the sound signal.
[0044] Further preferably, the setting standard of the third threshold in step S13 is: taking 4 times the average sound energy of the sound signal as the third threshold, where the minimum sound energy of each frame of the signal in the sound signal within the first 5 seconds is taken as the average sound energy of the sound signal.
[0045] By comparing the sound energy of the sound signal with the third threshold, it is possible to identify which sound signals belong to the interval of background noise detection.
[0046] Preferably, step S5 specifically includes:
[0047] S51. Judge whether the spectral amplitude of the sound signal remains in a stable state. If so, execute step S52; if not, return to step S121;
[0048] S52. Judge whether the variance of the spectral amplitude of the sound signal remains in a stable state. If so, the sound signal is background noise, execute step S53; if not, return to step S121;
[0049] S53. Record and update the parameters of the sound signal into the background noise data, and return to step S121.
[0050] Through the above steps, when both the spectral amplitude of the sound signal is maintained in a stable state and the variance of the spectral amplitude of the sound signal is maintained in a stable state, it is determined that the sound signal is background noise, and the parameters of the sound signal are recorded and updated.
[0051] Further preferably, after the step S4 and before the step S5, the following steps are further included:
[0052] S4a. Store the statistical data of the stable statistical analysis and the dynamic statistical analysis;
[0053] S5a. Determine whether the storage time of the statistical data exceeds a predetermined time. If so, execute step S5. If not, return to step S122.
[0054] According to the second aspect of the present application, an adaptive background noise detection system is proposed, including:
[0055] A sound acquisition device configured to acquire a sound signal and perform preprocessing on the sound signal;
[0056] A processor operation unit configured to set an initial detection state of the sound signal, perform preprocessing on the sound signal, estimate the sound energy of the sound signal, and determine whether the sound energy of the sound signal is less than a third threshold. If not, reset the initial detection state of the sound signal and perform preprocessing on the sound signal. If so, estimate the spectral amplitude of the sound signal, perform stable statistical analysis and dynamic statistical analysis on the sound signal, and determine whether the sound signal is background noise according to the analysis results of the stable statistical analysis and the dynamic statistical analysis. If so, record and update the parameters of the sound signal into the background noise data, reset the initial detection state of the sound signal, and perform preprocessing on the sound signal. If not, directly reset the initial detection state of the sound signal and perform preprocessing on the sound signal;
[0057] A memory unit configured to store programs or tables required in the processor operation unit and temporarily store data during the operation process;
[0058] A statistical record storage unit configured to store the background noise data.
[0059] Preferably, the sound acquisition device specifically includes:
[0060] A microphone sound collection device configured to collect the sound signal and convert the sound signal into a voltage signal;
[0061] An amplifier configured to amplify the voltage signal;
[0062] A filter configured to filter the amplified voltage signal.
[0063] An analog-to-digital converter configured to convert the filtered voltage signal into a digital signal.
[0064] According to a third aspect of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, it implements the adaptive background noise detection method described in the first aspect of the present application.
[0065] The present application provides an adaptive background noise detection method, system and medium. The method includes collecting a sound signal through a sound collection device and performing preprocessing, setting an initial detection state and preprocessing on the sound signal, estimating the sound energy of the processed sound signal, and comparing it with a third threshold. If the sound energy of the sound signal is less than the third threshold, it indicates that the sound signal is within the background noise detection interval. Then, steady-state statistical analysis and dynamic statistical analysis are continued on the sound signal. According to the analysis results, if both the spectral amplitude and the spectral amplitude variance of the sound signal remain in a stable state, the sound signal is determined to be background noise, and the parameters of the sound signal are recorded and updated into the background noise data as the basis for eliminating the background noise. This method can make timely adjustments according to the changes in the detected environmental noise, effectively identify the background noise, and has a high detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, serve to explain the principles of the invention. Other embodiments and many of the intended advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale relative to each other. Like reference numerals refer to corresponding like parts.
[0067] Figure 1 is a flowchart of an adaptive background noise detection method according to an embodiment of the present invention;
[0068] Figure 2 is a flowchart of background noise detection according to a specific embodiment of the present invention;
[0069] Figure 3 is a flowchart of preprocessing according to a specific embodiment of the present invention;
[0070] Figure 4 is a flowchart of steady-state statistical analysis according to a specific embodiment of the present invention;
[0071] Figure 5It is a flowchart of dynamic statistical analysis according to a specific embodiment of the present invention;
[0072] Figure 6 It is a flowchart for determining whether a sound signal is background noise according to a specific embodiment of the present invention;
[0073] Figure 7 It is a system block diagram of an adaptive background noise detection system according to an embodiment of the present invention;
[0074] Figure 8 It is an architecture diagram of a sound collection device according to a specific embodiment of the present invention.
[0075] Explanation of reference numerals: 1, sound collection device; 2, processor operation unit; 3, memory unit; 4, statistical record storage unit; 11, microphone sound collection device; 12, amplifier; 13, filter; 14, analog-to-digital converter. Detailed implementation manners
[0076] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without some of these specific details. The following description of the embodiments is only provided to provide a better understanding of the present invention by showing examples of the present invention.
[0077] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.
[0078] According to the first aspect of the present application, an adaptive background noise detection method is proposed. Figure 1 A flowchart of the adaptive background noise detection method according to an embodiment of the present invention is shown. As Figure 1 shown, the adaptive background noise detection method includes the following steps:
[0079] S1. Obtain the sound signal within the background noise detection interval.
[0080] In a specific embodiment, a sound acquisition device is used to collect the sound signal in the environment, and this sound signal is within the background noise detection interval. Figure 2 The flowchart of background noise detection according to a specific embodiment of the present invention is shown. As Figure 2 shown, the specific detection process is as follows:
[0081] S11. Collect the sound signal and preprocess the sound signal.
[0082] Figure 3 The flowchart of preprocessing according to a specific embodiment of the present invention is shown. As Figure 3 shown, in a specific embodiment, the preprocessing specifically includes:
[0083] S111. Collect the sound signal and convert the sound signal into a voltage signal.
[0084] In this embodiment, the process of converting the sound signal into a voltage (analog) signal is also completed on the sound acquisition device.
[0085] S112. Amplify the voltage signal.
[0086] Since the signal intensity of the sound signal is weak after being converted into a voltage signal, it is necessary to amplify it.
[0087] S113. Filter the amplified voltage signal to adjust the spectral response of the voltage signal.
[0088] In the filtering process, the spectral response of the voltage signal can be adjusted to perform sound enhancement and equalization processing.
[0089] S114. Convert the voltage signal into a digital signal.
[0090] Converting the analog digital signal into a digital signal prepares for subsequent signal processing.
[0091] Continue to refer to Figure 1 and Figure 2 , after S11:
[0092] S12. Set the initial detection state of the sound signal and preprocess the sound signal.
[0093] In a specific embodiment, setting the initial detection state of the sound signal is specifically manifested as:
[0094] S121. Set the sampling time, and statistically calculate the average value of the initial parameters of the sound signal within the sampling time as the initial parameters of the sound signal, and dynamically update the initial parameters of the sound signal.
[0095] By statistically calculating each parameter of the sound signal within a short period of time, such as spectral amplitude and sound energy, and taking the average value of each parameter as the initial parameters of the sound signal, the subsequent initial parameters are dynamically updated as the statistical data changes.
[0096] The preprocessing of the sound signal is specifically manifested as:
[0097] S122. Extract each frame of the signal from the sound signal and perform spectral equalization processing on the sound signal.
[0098] In the detection process, each frame of the sound signal will be detected once. The method of extracting each frame of the signal from the sound signal includes but is not limited to capturing using a Hanning Window, thereby reducing the distortion on the signal spectrum. The spectral equalization processing is implemented through a digital filter. The spectral equalization processing can not only compensate for the distortion when the sound acquisition device acquires the sound signal, but also be used to intensify a certain frequency band, thereby increasing or decreasing the weight of noise detection in this frequency band and improving the detection efficiency.
[0099] Continue to refer to Figure 2 , after step S12:
[0100] S13. Estimate the sound energy of the sound signal, and determine whether the sound energy of the sound signal is less than the third threshold. If so, the sound signal belongs to the background noise detection interval, and step S2 is executed; if not, return to the above step S12.
[0101] In a specific embodiment, the estimation of the sound energy of the sound signal is specifically manifested as: estimating the sound energy of each frame of the signal in the sound signal, and the specific estimation formula is:
[0102]
[0103] where, E i represents the sound energy of the i-th frame number, x(n) represents the frame signal corresponding to the i-th frame, and k represents the total number of frames of the sound signal.
[0104] In the following text, it is uniformly assumed that the sampling frequency of the sound acquisition device for the sound signal is 16000Hz, the number of frames of the sound signal is 512, the duration of each frame of the signal is 32 milliseconds, and there are approximately 32 frames in 1 second of the sound signal. Then, the estimated sound energy Ei of the i-th frame signal is:
[0105]
[0106] In a specific embodiment, the setting criterion for the third threshold is as follows:
[0107] Assume that the initial average sound energy of the sound signal is E o , and take the minimum sound energy in each frame of the sound signal (160-frame signal) within the first 5 seconds as the value of E o , that is:
[0108]
[0109] Assume that the first threshold is E th , and take 4 times the average sound energy of the sound signal as the first threshold, that is:
[0110] E th = 4E0
[0111] When it is judged that the sound energy of the sound signal is greater than the third threshold, it indicates that the sound signal is not within the background noise detection interval, and then return to step S121 to reset the initial detection state of the sound signal; when it is judged that the sound energy of the sound frame is within the first threshold, it indicates that the sound signal is within the background noise detection interval, and enter the next step S2.
[0112] Continue to refer to Figure 1 , after step S1:
[0113] S2. Perform a fast Fourier transform on each frame signal in the sound signal to estimate the spectral amplitude of the sound signal.
[0114] Performing a fast Fourier transform on each frame signal x(n) in the sound signal gives:
[0115] X i (k) = |FFT{x(n)}|, k = 1, 2, 3,..., 512
[0116] where X i (k) represents the spectral amplitude of the i-th frame signal. According to the spectral amplitude of each frame signal, the spectral amplitude of the entire sound signal can be obtained.
[0117] Continue to refer to Figure 1 , after step S2:
[0118] S3. Perform a steady-state statistical analysis on the spectral amplitude of the sound signal to calculate the change situation of the spectral amplitude of the sound signal.
[0119] Figure 4 shows a flowchart of the steady-state statistical analysis according to a specific embodiment of the present invention. As Figure 4 shown, in a specific embodiment, the steady-state statistical analysis specifically includes:
[0120] S31. Calculate the average spectral amplitude of the sound signal within 1 second and continuously record it for 5 seconds.
[0121] The calculation formula for the average spectral amplitude of the sound signal within 1 second (32-frame signal) is:
[0122]
[0123] where X a,j (k) represents the average spectral amplitude of the sound signal within the j-th second.
[0124] S32. Calculate the first average standard deviation between the average spectral amplitudes of 5 corresponding sound signals within 5 seconds.
[0125] Define the first average standard deviation as SD X , and its specific calculation formula is:
[0126]
[0127] where
[0128]
[0129] S33. Determine whether the first average standard deviation is less than the first threshold. If so, it indicates that the spectral amplitude of the sound signal remains in a stable state.
[0130] In a specific embodiment, the first threshold is set to:
[0131]
[0132] It should be noted that γ is an empirical value parameter. Since there are various noises such as rain sounds and fan sounds in the background noise, the value of γ is dynamically changing, and it is usually between 0.5 and 2. When the background noise is more stable, the value of γ is lower. In this embodiment, γ is specifically taken as 1.
[0133] Therefore, when it is determined that
[0134]
[0135] it indicates that the spectral amplitude of the sound signal statistically remains in a stable state.
[0136] Continue to refer to Figure 1 , after step S3:
[0137] S4. Conduct dynamic statistical analysis on the spectral amplitude of the sound signal and calculate the change in the variance of the spectral amplitude of the sound signal.
[0138] Figure 5Shows a flowchart of dynamic statistical analysis according to a specific embodiment of the present invention, as Figure 5 shown. In a specific embodiment, the dynamic statistical analysis specifically includes:
[0139] S41. Calculate the average spectral amplitude variance of the sound signal within 1 second and continuously record it for 5 seconds. The specific calculation formula is:
[0140]
[0141] where V x,j (k) represents the average spectral amplitude variance of the sound signal within the j-th second, X i (k) represents the spectral amplitude of the i-th frame signal, X a,j (k) represents the average spectral amplitude of the sound signal within the j-th second, and N represents the number of frames in the 1-second sound signal.
[0142] Therefore, in this embodiment, the average spectral amplitude variance of the sound signal within 1 second is:
[0143]
[0144] S42. Calculate the second average standard deviation between the average spectral amplitude variances of the corresponding 5 sound signals within 5 seconds.
[0145] Define the second average standard deviation as SD V , and its specific calculation formula is:
[0146]
[0147] where
[0148]
[0149] S43. Determine whether the second average standard deviation is less than the second threshold. If so, it indicates that the spectral amplitude variance of the sound signal remains in a stable state.
[0150] In a specific embodiment, the second threshold is set to:
[0151]
[0152] It should be noted that β is also an empirical value parameter, and its value-taking method is the same as that of γ above, which will not be elaborated here.
[0153] Therefore, when it is judged that
[0154]
[0155] it indicates that the spectral amplitude variance of the sound signal conforms to maintaining a stable state statistically.
[0156] Continue to refer to Figure 1 , after step S4:
[0157] S5. Determine whether the sound signal is background noise according to the change of the spectral amplitude of the sound signal and the change of the variance of the spectral amplitude of the sound signal.
[0158] Figure 6 The figure shows a flowchart for determining whether a sound signal is background noise according to a specific embodiment of the present invention. As Figure 6 shown, in a specific embodiment, step S5 specifically includes:
[0159] S51. Determine whether the spectral amplitude of the sound signal remains in a stable state. If so, execute step S52; if not, return to step S121.
[0160] S52. Determine whether the variance of the spectral amplitude of the sound signal remains in a stable state. If so, the sound signal is background noise, and execute step S53; if not, return to step S121.
[0161] S53. Record and update the parameters of the sound signal into the background noise data, and return to step S121.
[0162] Through the dual judgment of steady-state statistical analysis and dynamic statistical analysis, when both the spectral amplitude of the sound signal remains in a stable state and the variance of the spectral amplitude of the sound signal remains in a stable state, it is determined that the sound signal is background noise, and the parameters of the sound signal are recorded as the basis for background noise cancellation.
[0163] So far, a round of background noise detection is completed, and a new round of detection is entered again.
[0164] Continue to refer to Figure 2 , in a preferred embodiment, after step S43 and before step S51, there is also included:
[0165] S4a. Store the statistical data of steady-state statistical analysis and dynamic statistical analysis.
[0166] S5a. Determine whether the storage time of the statistical data exceeds a predetermined time. If so, execute step S5; if not, return to step S122.
[0167] Only when the statistical data accumulates and exceeds the predetermined time will it enter the next step to determine whether the sound signal is background noise. Otherwise, the steady-state statistical analysis and dynamic statistical analysis of the sound signal will continue until the statistical data time reaches the predetermined time. In this embodiment, the predetermined time is set to 5 seconds.
[0168] According to a second aspect of the present application, an adaptive background noise detection system is proposed, and this detection system is built based on the above detection method. Figure 7 The system block diagram of the adaptive background noise detection system according to an embodiment of the present invention is shown, as Figure 7 shown, this system includes:
[0169] A sound acquisition device 1, configured to acquire a sound signal and perform preprocessing on the sound signal.
[0170] A processor operation unit 2, configured to set an initial detection state of the sound signal, perform preprocessing on the sound signal, estimate the sound energy of the sound signal, and determine whether the sound energy of the sound signal is less than a third threshold. If not, reset the initial detection state of the sound signal and perform preprocessing on the sound signal. If so, estimate the spectral amplitude of the sound signal and perform steady-state statistical analysis and dynamic statistical analysis on the sound signal; according to the analysis results of the steady-state statistical analysis and the dynamic statistical analysis, determine whether the sound signal is background noise. If so, record and update the parameters of the sound signal into the background noise data, and reset the initial detection state of the sound signal and perform preprocessing on the sound signal. If not, directly reset the initial detection state of the sound signal and perform preprocessing on the sound signal.
[0171] A memory unit 3, configured to store the programs or tables required in the processor operation unit 2 and temporarily store the data during the operation process;
[0172] A statistical record storage unit 4, configured to store background noise data.
[0173] Figure 8 The architecture diagram of the sound acquisition device in a specific embodiment of the present invention is shown, as Figure 8 shown, the sound acquisition device 1 specifically includes:
[0174] A microphone sound collection device 11, configured to collect a sound signal and convert the sound signal into a voltage signal;
[0175] An amplifier 12, configured to amplify the voltage signal, and moreover, the amplifier can be preset with a variety of different sensitivities according to the usage requirements, so as to quickly adjust the voltage signal to an appropriate size;
[0176] A filter 13, configured to perform filtering processing on the amplified voltage signal;
[0177] An analog-to-digital converter 14, configured to convert the filtered voltage signal into a digital signal.
[0178] According to a third aspect of the present application, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the adaptive background noise detection method described above.
[0179] The present invention provides an adaptive background noise detection method, system, and medium. The method includes collecting a sound signal through a sound acquisition device and performing preprocessing, setting an initial detection state and preprocessing on the sound signal, estimating the sound energy of the processed sound signal, and comparing it with a third threshold. If the sound energy of the sound signal is less than the third threshold, it indicates that the sound signal is within the background noise detection interval. Then, steady-state statistical analysis and dynamic statistical analysis are continued on the sound signal. According to the analysis results, if both the spectral amplitude and spectral amplitude variance of the sound signal remain in a stable state, the sound signal is determined to be background noise, and the parameters of the sound signal are recorded and updated into the background noise data as the basis for eliminating background noise. This method can make timely adjustments according to the changes in the detected environmental noise, effectively identify background noise, and has a high detection accuracy.
[0180] In the embodiments of the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device / system / method embodiments described above are merely illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0181] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0183] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0184] Obviously, those skilled in the art can make various modifications and changes to the embodiments of the present invention without departing from the spirit and scope of the present invention. In this way, if these modifications and changes are within the scope of the claims of the present invention and their equivalent forms, the present invention also aims to cover these modifications and changes. The word "comprising" does not exclude the presence of other elements or steps not listed in the claims. The simple fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. An adaptive background noise detection method, characterized in that, Including the following steps: S1: Obtain the sound signal within the background noise detection interval; S2: Perform a fast Fourier transform on each frame of the sound signal to estimate the spectral amplitude of the sound signal; S3: Conduct a steady-state statistical analysis on the spectral amplitude of the sound signal to calculate the change in the spectral amplitude of the sound signal; specifically including: S31: Calculate the average spectral amplitude of the sound signal within 1 second and continuously record it for 5 seconds; S32: Calculate the first average standard deviation between the average spectral amplitudes corresponding to 5 of the sound signals within 5 seconds; S33: Determine whether the first average standard deviation is less than the first threshold. If so, it indicates that the spectral amplitude of the sound signal remains in a stable state; S4: Conduct a dynamic statistical analysis on the spectral amplitude of the sound signal to calculate the change in the spectral amplitude variance of the sound signal; specifically including: S41: Calculate the average spectral amplitude variance of the sound signal within 1 second and continuously record it for 5 seconds. The specific calculation formula is as follows: Among them, Vx,j(k) represents the average spectral amplitude variance of the sound signal within the j-th second, Xi(k) represents the spectral amplitude of the i-th frame signal, Xa,j(k) represents the average spectral amplitude of the sound signal within the j-th second, and N represents the number of frames in the 1-second sound signal; S42: Calculate the second average standard deviation between the average spectral amplitude variances corresponding to 5 of the sound signals within 5 seconds; S43: Determine whether the second average standard deviation is less than the second threshold. If so, it indicates that the spectral amplitude variance of the sound signal remains in a stable state; S5: Determine whether the sound signal is background noise based on the change in the spectral amplitude of the sound signal and the change in the spectral amplitude variance of the sound signal.
2. The method according to claim 1, wherein The specific steps of S1 include: S11: Collect the sound signal and preprocess the sound signal; S12: Set the initial detection state of the sound signal and perform preprocessing on the sound signal; S13: Estimate the sound energy of the sound signal and determine whether the sound energy of the sound signal is less than the third threshold. If so, the sound signal is within the background noise detection interval, and step S2 is executed. If not, return to the above step S12.
3. The method according to claim 2, wherein The specific steps of S11 include: S111: Collect the sound signal and convert the sound signal into a voltage signal; S112: Amplify the voltage signal; S113: Filter the amplified voltage signal to adjust the spectral response of the voltage signal; S114: Convert the voltage signal into a digital signal.
4. The method according to claim 2, wherein The specific setting of the initial detection state of the sound signal in step S12 includes: S121: Set the sampling time, calculate the average value of the initial parameters of the sound signal within the sampling time as the initial parameters of the sound signal, and dynamically update the initial parameters of the sound signal.
5. The method according to claim 4, wherein The specific preprocessing of the sound signal in step S12 includes: S122: Extract each frame of the sound signal and perform spectral equalization on the sound signal.
6. The method according to claim 5, wherein The specific estimation of the sound energy of the sound signal in step S13 includes: Estimate the sound energy of each frame of the sound signal. The specific estimation formula is: Among them, E i represents the sound energy of the i-th frame signal, x(n) represents the frame signal corresponding to the i-th frame, and k represents the total number of frames of the sound signal.
7. The method according to claim 6, wherein The setting criterion of the third threshold in the step S13 is: taking 4 times the average sound energy of the sound signal as the third threshold, wherein the minimum sound energy in each frame of the sound signal within the first 5 seconds is taken as the average sound energy of the sound signal.
8. The method according to claim 7, wherein The step S5 specifically includes: S51: Judging whether the spectral amplitude of the sound signal maintains a stable state. If so, execute step S52; if not, return to step S121. S52: Judging whether the variance of the spectral amplitude of the sound signal maintains a stable state. If so, the sound signal is background noise, and execute step S53; if not, return to step S121. S53: Record and update the parameters of the sound signal into the background noise data, and return to step S121.
9. The method according to claim 5, characterized in that Before the step S5 and after the step S4, it further includes: S4a: Storing the statistical data of the steady-state statistical analysis and the dynamic statistical analysis. S5a: Judging whether the storage time of the statistical data exceeds a predetermined time. If so, execute step S5; if not, return to step S122.
10. An adaptive background noise detection system, characterized in that, Used to implement the method as described in any one of claims 1-9, including: A sound collection device configured to collect a sound signal and perform preprocessing on the sound signal. A processor arithmetic unit configured to set an initial detection state of the sound signal, perform preprocessing on the sound signal, estimate the sound energy of the sound signal, and judge whether the sound energy of the sound signal is less than the third threshold. If not, reset the initial detection state of the sound signal and perform preprocessing on the sound signal. If so, estimate the spectral amplitude of the sound signal, perform steady-state statistical analysis and dynamic statistical analysis on the sound signal, and judge whether the sound signal is background noise according to the analysis results of the steady-state statistical analysis and the dynamic statistical analysis. If so, record and update the parameters of the sound signal into the background noise data, reset the initial detection state of the sound signal, and perform preprocessing on the sound signal. If not, directly reset the initial detection state of the sound signal and perform preprocessing on the sound signal. A memory unit configured to store the programs or tables required in the processor arithmetic unit and temporarily store the data during the operation process. A statistical record storage unit configured to store the background noise data.
11. The system according to claim 10, wherein The sound collection device specifically includes: A microphone sound collection device configured to collect the sound signal and convert the sound signal into a voltage signal. An amplifier configured to amplify the voltage signal. A filter configured to perform filtering processing on the amplified voltage signal. An analog-to-digital converter configured to convert the filtered voltage signal into a digital signal.
12. A computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Broadband background noise and voice separation detection system and method
CN106504760A