A method, device and medium for detecting sound energy

By performing sound frame extraction and spectrum processing on the sound signal, calculating the crest factor and performing energy compensation, the problems of low sensitivity of sound energy detection and unstable measurement values ​​in the prior art are solved, and more accurate energy change estimation and abnormal judgment are achieved.

CN115019813BActive Publication Date: 2025-06-17ジャン州立達信光電子科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210610682.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-06-17
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The existing sound energy detection methods are not sensitive and have unstable measurement values, making it difficult to accurately estimate instantaneous energy changes in real time.

Method used

By collecting sound data, converting it into sound signals, and performing sound frame extraction and spectrum equalization processing, calculating the spectrum amplitude and energy of the sound frame signal, calculating the crest factor, and compensating energy distortion based on the comparison of the crest factor and the preset threshold value, and finally determining whether the sound energy is abnormal.

Benefits of technology

It improves the sensitivity of sound energy detection and the stability of measurement values, can more accurately estimate instantaneous energy changes, and effectively determine whether the sound energy is abnormal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019813B_ABST
    Figure CN115019813B_ABST
Patent Text Reader

Abstract

The present application proposes a method for detecting sound energy, including: S1, collecting sound data and converting it into a sound signal, performing frame extraction and spectrum equalization processing on the sound signal to generate a plurality of frame signals; S2, calculating the spectrum amplitude of the frame signals and obtaining the sound energy through weighted estimation; S3, calculating the root mean square of the waveform data of the frame signals to obtain an energy parameter, and the ratio of the maximum spectrum amplitude of the frame signals to the energy parameter is the crest factor; S4, comparing the crest factor with a preset crest threshold to determine whether the crest factor exceeds the crest threshold. If not, adding a preset compensation value to the sound energy for energy distortion compensation; S5, comparing the compensated sound energy with a preset energy threshold, and judging whether the sound energy is abnormal according to the comparison result. The sound energy detection method of the present application has a relatively high detection sensitivity and relatively stable measurement values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of sound detection, and particularly to a method, device and medium for sound energy detection. Background Art

[0002] Existing sound energy detection methods generally use sound pressure meters available on the market. They usually estimate the sound pressure value by averaging a short period of time and adding three weighting methods of A, B, and C for auditory effects to obtain measurement values such as dBA, dBB, and dBC applicable to different volume ranges. However, these sound pressure meters often cannot accurately estimate in real time due to too short sounds, are relatively insensitive to instantaneous energy changes, and have large deviations in each numerical measurement of the same sound, and their measurement accuracy is worrying.

[0003] In view of this, it is particularly important to provide a sound energy detection method, device and medium with higher detection sensitivity and relatively stable measurement values. Summary of the Invention

[0004] In order to solve the technical problems of low sensitivity and unstable measurement values in the existing sound energy detection methods, this application proposes a method, device and medium for sound energy detection.

[0005] According to the first aspect of this application, a method for sound energy detection is proposed, including:

[0006] S1. Collect sound data and convert it into a sound signal, perform frame extraction and spectrum equalization processing on the sound signal to generate a plurality of frame signals with a certain length;

[0007] S2. Calculate the spectral amplitude of the frame signal in different frequency bands, and obtain the sound energy of the frame signal through weighted estimation;

[0008] S3. Calculate the root mean square of the waveform data of the frame signal to obtain the energy parameter of the frame signal, and the ratio of the maximum spectral amplitude of the frame signal to the energy parameter is the crest factor of the frame signal;

[0009] S4. Compare the crest factor of the frame signal with a preset crest threshold to determine whether the crest factor of the frame signal exceeds the crest threshold. If not, add a preset compensation value to the sound energy for energy distortion compensation; and

[0010] S5. Compare the compensated sound energy with a preset energy threshold, and determine whether the sound energy is abnormal according to the comparison result.

[0011] Preferably, the step S3 specifically includes: equally dividing the frame signal into a plurality of sub-frame signals for calculating the spectral amplitude and energy parameters, and the ratio of the maximum spectral amplitude of the sub-frame signal to the corresponding energy parameter is the crest factor of the sub-frame signal.

[0012] Preferably, the comparison of the crest factor with a preset crest threshold in the step S4 is specifically manifested as: comparing the minimum crest factor among a plurality of the sub-frame signals with the crest threshold.

[0013] Preferably, the step S4 specifically includes: determining whether the maximum spectral amplitude of a plurality of the sub-frame signals reaches the maximum saturation value, correspondingly setting different levels of the crest threshold and the compensation value according to the number of the sub-frame signals equal to or reaching the maximum saturation value, comparing the minimum crest factor among a plurality of the sub-frame signals with the corresponding-level crest threshold, determining whether the minimum crest factor of the sub-frame signal exceeds the corresponding-level crest threshold, and if not, adding the corresponding-level compensation value to the sound energy for energy distortion compensation.

[0014] Preferably, the frame signal is equally divided into 4 sub-frame signals, the crest threshold and the compensation value are correspondingly set with 4 levels, and the distortion compensation calculation of the sound energy specifically includes:

[0015] When Amax cnt ≥1 and CF min ≤T1, E = E A *G1;

[0016] When Amax cnt ≥2 and CF min ≤T2, E = E A *G2;

[0017] When Amax cnt ≥3 and CF min ≤T3, E = E A *G3;

[0018] When Amax cnt = 4 and CF min ≤T4, E = E A *G4;

[0019] Wherein, Amax cnt represents the number of the maximum spectral amplitudes of 4 sub-frame signals reaching the maximum saturation value, CF min represents the minimum crest factor among 4 sub-frame signals, T1, T2, T3, and T4 respectively represent the crest thresholds of 4 levels, G1, G2, G3, and G4 respectively represent the compensation values of 4 levels, and EA represents the initial sound energy, and E represents the sound energy after compensation;

[0020] The specific values of the peak threshold and the compensation value are as follows:

[0021] T1 = 1.32, G1 = 1.585;

[0022] T2 = 1.15, G1 = 2.512;

[0023] T3 = 1.065, G3 = 3.98;

[0024] T4 = 1.035, G4 = 7.94.

[0025] Preferably, the step S5 specifically includes: presetting different levels of the energy threshold, comparing the compensated sound energy with different levels of the energy threshold, and judging whether the sound energy is abnormal according to the comparison result.

[0026] Preferably, the step S2 specifically includes: performing a fast Fourier transform on the frame signal to calculate the spectral amplitude, and estimating the sound energy of the frame signal through an A-weighting value.

[0027] Preferably, in the step S1, the frame extraction specifically includes: extracting the sound signal into a plurality of frame signals by using a Hamming window; the spectral equalization specifically includes: compensating for the distortion of the sound data during the acquisition process by using a digital filter, so that the sound signal has a nearly equal spectral amplitude in each frequency band.

[0028] According to the second aspect of the present application, a sound energy detection device is proposed, including:

[0029] A preprocessing unit configured to collect sound data and convert it into a sound signal, perform frame extraction and spectral equalization processing on the sound signal, and generate a plurality of frame signals with a certain length;

[0030] A sound energy estimation unit configured to calculate the spectral amplitude of the frame signal in different frequency bands and estimate the sound energy of the frame signal through weighted estimation;

[0031] A crest factor calculation unit configured to calculate the root mean square of the waveform data of the frame signal to obtain the energy parameter of the frame signal, and the ratio of the maximum spectral amplitude of the frame signal to the energy parameter is the crest factor of the frame signal;

[0032] A sound energy compensation unit, configured to compare the crest factor of the frame signal with a preset crest threshold, determine whether the crest factor of the frame signal exceeds the crest threshold, and if not, add a preset compensation value to the sound energy for energy distortion compensation;

[0033] A sound energy abnormality determination unit, configured to compare the compensated sound energy with a preset energy threshold, and determine whether the sound energy is abnormal according to the comparison result.

[0034] According to the third aspect of the present application, a computer-readable storage medium is provided, which stores a computer program, and the computer program, when executed by a processor, implements the sound energy detection method as described in the first aspect of the present application.

[0035] The present application provides a sound energy detection method, device and medium, which equally divide a sound signal into multiple frame signals, calculate the crest factor of the frame signal by estimating the spectral amplitude and sound energy of the frame signal, and perform different degrees of compensation on the sound energy of the frame signal according to the saturation of the spectral amplitude of the frame signal and the magnitude of the crest factor, improving the problem of unstable measurement values of sound energy. And by setting different levels of energy thresholds to compare with the compensated sound energy, it is determined whether the detected sound energy is abnormal, solving the problem of poor detection sensitivity of sound energy. Description of the Drawings

[0036] The drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present application. Other embodiments and many of the expected advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale relative to each other. The same reference numerals refer to corresponding like parts.

[0037] Figure 1 is a flowchart of the sound energy detection method according to an embodiment of the present application;

[0038] Figure 2 is a block diagram of the architecture of a sound collection device according to a specific embodiment of the present application;

[0039] Figure 3 is a flowchart of the sound energy detection method according to a specific embodiment of the present application;

[0040] Figure 4 is a system block diagram of the sound energy detection device according to an embodiment of the present application.

[0041] Explanation of the accompanying drawings: 10, sound collection device; 101, microphone receiving device; 102, amplifier; 103, filter; 104, analog-to-digital converter; 20, pre-processing unit; 30, sound energy estimation unit; 40, crest factor calculation unit; 50, sound energy compensation unit; 60, sound energy anomaly judgment unit. DETAILED DESCRIPTION

[0042] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present application and are not configured to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0043] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0044] According to a first aspect of the present application, a sound energy detection method is proposed. Figure 1 FIG. 4 shows a flow chart of a sound energy detection method according to an embodiment of the present application. Figure 1 As shown, the detection method comprises the following steps:

[0045] S1. Collect sound data and convert it into sound signals, perform sound frame extraction and spectrum equalization on the sound signals, and generate several sound frame signals with a certain length.

[0046] In a specific embodiment, the sound data is collected by a sound collection device and converted into a sound signal. Figure 2 The structure diagram of the sound collection device according to a specific embodiment of the present application is shown as follows: Figure 2As shown, the sound acquisition device 10 includes a microphone sound collection device 101, an amplifier 102, a filter 103, and an analog-to-digital converter 104. Among them, the microphone sound collection device 101 is used to collect sound data and convert it into a voltage signal; the amplifier 102 amplifies the voltage signal, and the amplifier 102 can be preset with a variety of different sensitivities according to usage requirements, so as to quickly adjust the voltage signal to an appropriate size; the filter 103 filters the amplified voltage signal to eliminate some noise; the analog-to-digital converter 104 converts the filtered voltage signal into a digital signal, that is, a sound signal.

[0047] In a specific embodiment, the frame extraction process specifically includes: using a Hamming Window to equally divide the sound signal into a number of frame signals.

[0048] In a specific embodiment, the spectrum equalization process specifically includes: using a digital filter to compensate for the distortion of the sound data during the acquisition process, such as enhancing the high-frequency part of the sound data, so that the sound signal has a spectrum amplitude of approximately the same size in each frequency band.

[0049] Continue to refer to Figure 1 , after step S1,

[0050] S2. Calculate the spectrum amplitude of the frame signal in different frequency bands, and obtain the sound energy of the frame signal through weighted estimation.

[0051] Figure 3 Shows a flowchart of a sound energy detection method according to a specific embodiment of the present application. As Figure 3 shown, step S2 specifically includes:

[0052] S21. Perform a fast Fourier transform on the frame signal to calculate the spectrum amplitude.

[0053] Assume that the sampling frequency of the sound acquisition device for the sound signal is 16000Hz, the sound signal is equally divided into 512 frame signals, and the duration of each frame signal is 32ms.

[0054] In a specific embodiment, the formula for calculating the spectrum amplitude of the frame signal is as follows:

[0055] X(k) = FFT{x(n)}, n = 0, 1, 2,..., N - 1; k = 0, 1, 2,..., K - 1

[0056] Among them, x(n) is the nth frame signal, X(k) is the spectrum amplitude of the nth frame signal, and N and K are both 512.

[0057] S22. Obtain the sound energy of the frame signal through A-weighted level estimation.

[0058] In a specific embodiment, taking the A-weighted level value as an example, the spectral amplitude of the frame signal is weighted and calculated to obtain the sound energy of the frame signal. The specific calculation formula is as follows:

[0059]

[0060] Where W A is the weighted value for estimating the A-level energy.

[0061] Continuing to refer to Figure 1 and Figure 3 after step S2,

[0062] S3. Calculate the root mean square of the waveform data of the frame signal to obtain the energy parameter of the frame signal. The ratio of the maximum spectral amplitude of the frame signal to the energy parameter is the crest factor of the frame signal.

[0063] Combined with reference to Figure 3 step S3 is divided into the following steps:

[0064] S31. Calculate the maximum spectral amplitude of the frame signal;

[0065] S32. Calculate the energy parameter of the frame signal and calculate the crest factor of the frame signal according to the maximum spectral amplitude and the energy parameter.

[0066] In a specific embodiment, each frame signal is further equally divided into several sub-frame signals, and then the maximum spectral amplitude, energy parameter, and crest factor of each sub-frame signal are calculated. In this embodiment, taking 4 sub-frame signals as an example, let the i maximum spectral amplitude of the i-th sub-frame signal be Amax i , then:

[0067] Amax i = max(x(n + 128×(1 - 1))), n = 0, 1, 2,... 128, i = 1, 2, 3, 4

[0068] Let the i energy parameter (i.e., the root mean square of the waveform data) of the i-th sub-frame signal be RMS i , then:

[0069]

[0070] Let the i crest factor of the i-th sub-frame signal be CF i , then:

[0071]

[0072] Continuing to refer toFigure 1 and Figure 3 , after step S3,

[0073] S4. Compare the crest factor of the frame signal with a preset crest threshold to determine whether the crest factor of the frame signal exceeds the crest threshold. If not, add a preset compensation value to the sound energy for energy distortion compensation.

[0074] Taking a sine wave as an example, for a non-distorted sine wave, its crest factor is approximately equal to 1.414. When the sound energy increases and exceeds the maximum saturation value, the sine wave waveform will be truncated. If it is continuously amplified, the distortion will become more serious until it approaches a square wave, and the crest factor of the square wave is equal to 1. Therefore, the crest factor can be used to estimate the degree of waveform distortion and compensate the sound energy estimation accordingly.

[0075] In a specific embodiment, determine whether the maximum spectral amplitude Amax of a number of sub-frame signals i reaches the maximum saturation value. According to the number Amax of sub-frame signals that reach the maximum saturation value cnt , set different levels of crest thresholds T i , i = 1, 2, 3, 4 and compensation values G i , i = 1, 2, 3, 4. Take the minimum crest factor CF of a number of sub-frame signals min and compare it with the corresponding level of crest threshold T i to determine whether the minimum crest factor CF of the sub-frame signal min exceeds the corresponding level of crest threshold T i . If not, add the corresponding level of compensation value G A to the sound energy E of the frame signal i for energy distortion compensation to obtain the compensated sound energy E. The specific process is as follows:

[0076] Since in this embodiment, a frame signal is equally divided into 4 sub-frame signals, first judge one by one whether the maximum spectral amplitude Amax of the 4 sub-frame signals i is greater than or equal to the maximum saturation value, Amax cnt = 1, 2, 3, 4 (the case where the number is 0 is not subject to compensation calculation). Then:

[0077] When Amax cnt ≥ 1 and CF min ≤ T1, let E = E A * G1;

[0078] When Amax cnt ≥ 2 and CF min ≤ T2, let E = E A * G2;

[0079] When amax cnt ≥ 3 and CF min ≤ T3, let E = E A *G3;

[0080] When amax cnt = 4 and CF min ≤ T4, let E = E A *G4;

[0081] Among them, T1, T2, T3, T4 and G1, G2, G3, G4 are empirical values obtained by the inventor through multiple experiments, and their value sizes can be adjusted according to the actual situation. Taking the specific detected sound sample as an example, the following is the sound energy calculation compensation table, as shown in Table 1 below:

[0082]

[0083] Table 1

[0084] It can be seen from Table 1 that when the minimum crest factor CF min is smaller, the signal distortion is greater, and the estimated sound pressure (sound energy) error is also greater. And, based on the experimental data in Table 1, the crest threshold T i and the compensation value G i are inversely deduced. In this embodiment, T1 = 1.32, G1 = 1.585 (equivalent to a compensation of 2 dB); T2 = 1.15, G1 = 2.512 (equivalent to a compensation of 4 dB); T3 = 1.065, G3 = 3.98 (equivalent to a compensation of 6 dB); T4 = 1.035, G4 = 7.94 (equivalent to a compensation of 9 dB).

[0085] Continue to refer to Figure 1 and Figure 3 , after step S4,

[0086] S5. Compare the compensated sound energy with a preset energy threshold, and judge whether the sound energy is abnormal according to the comparison result.

[0087] In a specific embodiment, the energy threshold is also preset to several different levels. Assuming that the sound energy to be detected is above 80 dB, and the corresponding sound energy is E t , then according to the logarithmic operation of the sound pressure level value SPL, we can get: E t = 10 (80 / 10) = 100000000. Set the energy thresholds to E t , 1 / 2E t , 1 / 4E t these three levels, and the sound energy E of the E cnt frame signalsA are respectively compared with the energy thresholds E of three levels t , 1 / 2E t , 1 / 4E t , and judge and calculate:

[0088] 1. When E A ≥ E t , perform cumulative calculation on this frame signal and record it in Count_E;

[0089] 2. When E t > E A ≥ 1 / 2E t , perform cumulative calculation on this frame signal and record it in Count2_E;

[0090] 3. When 1 / 2E t > E A ≥ 1 / 4E t , perform cumulative calculation on this frame signal and record it in Count4_E;

[0091] 4. When E A < 1 / 4E t , perform cumulative calculation on this frame signal and record it in Count0;

[0092] And when there are two consecutive frame signals with sound energy E A < 1 / 4E t , then when comparing the next frame signal:

[0093] When E A < E t , continue to perform cumulative calculation on this frame signal and record it in Count0;

[0094] When E A ≥ E t , perform cumulative calculation on this frame signal and record it in Count_E, and at the same time clear Count0.

[0095] After all E cnt frame signals are judged, when judging:

[0096] Count_E + Count2_E ≥ E cntl , and Count_E + Count2_E + Count4_E ≥ E cnt2 , then judge that the sound energy of this sound signal is abnormal and send a notification reminder.

[0097] Among them, E cnt , E cntl , E cnt2is a preset empirical value related to the sensitivity of sound detection. The following is a detection of a sound sample with a sound energy of 80 dB, and the sound data measured under 3 different sensitivities E cnt 、E cnt1 、E cnt2 is shown in Table 2 below:

[0098] Sensitivity <![CDATA[E cnt > <![CDATA[E cnt1 > <![CDATA[E cnt2 > Minimum sound length (ms) Instantaneous average sound pressure (dB) 1-second average sound pressure (dB) 1 5 2 3 96 77.66 67.38 2 16 8 12 384 76.61 72.35 3 32 16 24 768 76.41 75.16

[0099] Table 2

[0100] As can be seen from Table 2, the smaller the values of E cnt 、E cnt1 、E cnt2 , the shorter the time required for the instantaneous average sound pressure (sound energy) detected, and the greater the average sound pressure (sound energy) within 1 second, indicating the higher the sensitivity of sound detection; vice versa. Taking the first sensitivity as an example, when the number of frame signals E cnt to be detected is 5, as long as E cnt1 reaches 2 and E cnt2 reaches 3, it is determined that the sound energy of the sound signal is abnormal and a notification reminder is sent.

[0101] According to the second aspect of the present application, a sound energy detection device is proposed, and this detection device is built based on the detection method of the first aspect of the present application. Figure 4 shows a system block diagram of the sound energy detection device according to an embodiment of the present application, as Figure 4 shown, this detection device includes:

[0102] A preprocessing unit 20, configured to collect sound data and convert it into a sound signal, perform frame extraction and spectrum equalization processing on the sound signal, and generate a plurality of frame signals with a certain length;

[0103] A sound energy estimation unit 30, configured to calculate the spectral amplitude of the frame signal in different frequency bands and obtain the sound energy of the frame signal through weighted estimation;

[0104] A crest factor calculation unit 40, configured to calculate the root mean square of the waveform data of the frame signal to obtain the energy parameter of the frame signal, and the ratio of the maximum spectral amplitude of the frame signal to the energy parameter is the crest factor of the frame signal;

[0105] A sound energy compensation unit 50, configured to compare the crest factor of the frame signal with a preset crest threshold, determine whether the crest factor of the frame signal exceeds the crest threshold, and if not, add a preset compensation value to the sound energy for energy distortion compensation;

[0106] The sound energy abnormality determination unit 60 is configured to compare the compensated sound energy with a preset energy threshold, and determine whether the sound energy is abnormal according to the comparison result.

[0107] According to the third aspect of the present application, a computer-readable storage medium is proposed, which stores a computer program. When the computer program is executed by a processor, the sound energy detection method in the first aspect of the present application is implemented.

[0108] The present application proposes a sound energy detection method, device and medium. By equally dividing a sound signal into multiple frame signals, and then further equally dividing the frame signals into several sub-frame signals, calculating the minimum crest factor of the sub-frame signals, and compensating the sound energy of the original frame signals to different degrees according to the saturation of the spectral amplitude of the sub-frame signals and the magnitude of the minimum crest factor, the problem of unstable sound energy measurement values is improved. And by setting different levels of energy thresholds to compare with the compensated sound energy, it is determined whether the detected sound energy is abnormal, and the problem of poor detection sensitivity of sound energy is solved.

[0109] In the embodiments of the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the above-described device / system / method embodiments are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.

[0110] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0111] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0112] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0113] Obviously, those skilled in the art can make various modifications and changes to the embodiments of the present application without departing from the spirit and scope of the present application. In this way, if these modifications and changes are within the scope of the claims of the present application and their equivalent forms, the present application also aims to cover these modifications and changes. The word "comprising" does not exclude the presence of other elements or steps not listed in the claims. The simple fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. A method for detecting sound energy, characterized in that, Including: S1. Collect voice data and convert it into a voice signal, perform frame extraction and spectrum equalization processing on the voice signal, and generate a number of frame signals with a certain length; S2. Calculate the spectral amplitude of the frame signal in different frequency bands, and obtain the voice energy of the frame signal through weighted estimation; S3. Calculate the root mean square of the waveform data of the frame signal to obtain the energy parameter of the frame signal, and the ratio of the maximum spectral amplitude of the frame signal to the energy parameter is the crest factor of the frame signal; specifically including: equally dividing the frame signal into a number of sub-frame signals for spectral amplitude and energy parameter calculation, and the ratio of the maximum spectral amplitude of the sub-frame signal to the corresponding energy parameter is the crest factor of the sub-frame signal; S4. Compare the crest factor of the frame signal with a preset crest threshold, and judge whether the crest factor of the frame signal exceeds the crest threshold. If not, add a preset compensation value to the voice energy for energy distortion compensation; among them, comparing the crest factor with the preset crest threshold is specifically manifested as: comparing the minimum crest factor among a number of the sub-frame signals with the crest threshold; and S5. Compare the compensated voice energy with a preset energy threshold, and judge whether the voice energy is abnormal according to the comparison result.

2. The method for detecting sound energy according to claim 1, characterized in that, The step S4 specifically includes: judging whether the maximum spectral amplitude of a number of the sub-frame signals reaches the maximum saturation value, correspondingly setting different levels of the crest threshold and the compensation value according to the number of the sub-frame signals equal to or reaching the maximum saturation value, comparing the minimum crest factor among a number of the sub-frame signals with the crest threshold of the corresponding level, and judging whether the minimum crest factor of the sub-frame signal exceeds the crest threshold of the corresponding level. If not, add the compensation value of the corresponding level to the voice energy for energy distortion compensation.

3. The method for detecting sound energy according to claim 2, characterized in that, The frame signal is equally divided into 4 sub-frame signals, and the crest threshold and the compensation value are correspondingly set with 4 levels. The distortion compensation calculation of the voice energy specifically includes: When Amax cnt ≥ 1 and CF min ≤ T1, E = E A * G1; When Amax cnt ≥ 2 and CF min ≤ T2, E = E A *G2; When Amax cnt ≥ 3 and CF min ≤ T3, then E = E A * G3; When Amax cnt = 4 and CF min ≤ T4, E = E A * G4; Among them, Amax cnt represents the number of the maximum spectral amplitudes of the four sub - vowel frame signals reaching the maximum saturation value, CF min represents the minimum crest factor among the four sub - vowel frame signals, T1, T2, T3, and T4 respectively represent the crest thresholds of four levels, G1, G2, G3, and G4 respectively represent the compensation values of four levels, E A represents the initial sound energy, and E represents the sound energy after compensation; The specific values of the crest threshold and the compensation value are: T1 = 1.32, G1 = 1.585; T2 = 1.15, G1 = 2.512; T3 = 1.065, G3 = 3.98; T4 = 1.035, G4 = 7.

94.

4. The method for detecting sound energy according to claim 1, characterized in that, The step S5 specifically includes: presetting different levels of the energy threshold, comparing the compensated voice energy with different levels of the energy threshold, and judging whether the voice energy is abnormal according to the comparison result.

5. The method for detecting sound energy according to claim 1, characterized in that, The step S2 specifically includes: performing a fast Fourier transform on the frame signal to calculate the spectral amplitude, and obtaining the voice energy of the frame signal through A-weighted level estimation.

6. The method for detecting sound energy according to claim 1, characterized in that, In the step S1, the frame extraction specifically includes: Using a Hamming window to extract the voice signal into a number of the frame signals; The spectrum equalization specifically includes: A digital filter is used to compensate for the distortion of the sound data during the acquisition process, so that the sound signal has a spectral amplitude of nearly the same size in each frequency band.

7. A device for detecting sound energy, characterized in that, It includes: A preprocessing unit configured to collect sound data and convert it into a sound signal, perform frame extraction and spectral equalization processing on the sound signal, and generate a plurality of frame signals with a certain length; A sound energy estimation unit configured to calculate the spectral amplitude of the frame signal in different frequency bands and obtain the sound energy of the frame signal through weighted estimation; A crest factor calculation unit configured to calculate the root mean square of the waveform data of the frame signal to obtain the energy parameter of the frame signal, and the ratio of the maximum spectral amplitude of the frame signal to the energy parameter is the crest factor of the frame signal; specifically includes: equally dividing the frame signal into a plurality of sub-frame signals for spectral amplitude and energy parameter calculation, and the ratio of the maximum spectral amplitude of the sub-frame signal to the corresponding energy parameter is the crest factor of the sub-frame signal; A sound energy compensation unit configured to compare the crest factor of the frame signal with a preset crest threshold to determine whether the crest factor of the frame signal exceeds the crest threshold. If not, add a preset compensation value to the sound energy for energy distortion compensation; wherein, comparing the crest factor with the preset crest threshold is specifically manifested as: comparing the minimum crest factor among a plurality of the sub-frame signals with the crest threshold; A sound energy abnormality determination unit configured to compare the compensated sound energy with a preset energy threshold and determine whether the sound energy is abnormal according to the comparison result.

8. A computer-readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Adaptive background noise detection method and system and medium

    CN114220446A

  • Apparatus and method for determining abnormal sound and program

    JP2004333200A