Speech detection method and apparatus, storage medium, and electronic device
By analyzing the time-domain energy and time-domain harmonic content of the voice frame signal and combining it with frequency point updates, the howling frame signal can be accurately identified, solving the problem of low howling detection accuracy and improving call quality.
Patent Information
- Application Number
- CN202210404900.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-04-18
AI Technical Summary
In current voice network calls, the accuracy of howling detection is low, and existing methods have problems with missed detections or false detections, which affect call quality.
By acquiring the voice frame signal to be detected, the suspected howling frame signal is initially judged by the time domain energy and time domain harmonic content. The target reference howling frequency point is updated by combining the candidate howling frequency point and the initial reference howling frequency point, and the howling frame signal is finally determined.
It improves the accuracy of howling detection, effectively alleviates the technical problem of low howling detection accuracy, and ensures call quality.
Smart Images

Figure CN114694676B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, and in particular to a voice detection method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the development of communication technology, network calls through mobile terminals have become one of the indispensable communication methods for people, and thus the quality of network calls is becoming increasingly important.
[0003] In the current voice network call process, the signal of a near-end receiver is transmitted back to a near-end microphone through an acoustic path, and then returned from an acoustic path of a far-end through a network environment to form feedback (i.e., howling), which seriously affects the quality of network calls. In order to avoid the adverse effects of howling on the quality of network calls, it is necessary to detect and suppress howling in a timely manner. Currently, methods such as peak-to-average power ratio, peak-to-harmonic power ratio, and inter-frame amplitude spectrum slope deviation are used for howling detection. However, the above methods generally have the problems of missed detection or false detection when detecting howling, resulting in low accuracy of howling detection. SUMMARY
[0004] The present application provides a voice detection method, device, storage medium and electronic device to alleviate the technical problem of low accuracy of current howling detection.
[0005] To solve the above technical problem, the present application provides the following technical solutions:
[0006] The present application provides a voice detection method, comprising:
[0007] obtaining a to-be-detected voice frame signal;
[0008] When the time-domain energy of the to-be-detected voice frame signal and the time-domain harmonic content satisfy a suspected howling condition, determining that the to-be-detected voice frame signal is a suspected howling frame signal;
[0009] According to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which an initial reference howling frequency point is located, updating the initial reference howling frequency point to obtain a target reference howling frequency point;
[0010] When the target reference howling frequency point satisfies a howling condition, determining that the suspected howling frame signal is a howling frame signal.
[0011] The step of obtaining a to-be-detected voice frame signal comprises:
[0012] obtaining a to-be-detected voice signal;
[0013] performing frame processing and windowing processing on the to-be-detected voice signal to obtain at least one to-be-detected voice frame signal.
[0014] The step of determining that the to-be-detected voice frame signal is a suspected howling frame signal when the time domain energy of the to-be-detected voice frame signal and the time domain harmonic content meet a suspected howling condition comprises:
[0015] The time domain energy of the to-be-detected voice frame signal is calculated according to the signal length of the to-be-detected voice frame signal, and the time domain harmonic content of the to-be-detected voice frame signal is determined according to the normalized autocorrelation function of the to-be-detected voice frame signal.
[0016] When the time domain energy is greater than or equal to a time domain energy threshold value and the time domain harmonic content is greater than or equal to a time domain harmonic content threshold value, it is determined that the time domain energy and the time domain harmonic content meet the suspected howling condition, and it is determined that the to-be-detected voice frame signal is the suspected howling frame signal.
[0017] The step of updating the initial reference howling frequency point to obtain a target reference howling frequency point according to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located further comprises:
[0018] An amplitude spectrum of the suspected howling frame signal is obtained.
[0019] The amplitude spectrum is smoothed, and an extreme value in the smoothed amplitude spectrum is determined.
[0020] When the extreme value is greater than or equal to a candidate howling threshold value, the extreme value is determined to be the candidate howling frequency point; wherein the candidate howling threshold value comprises a candidate howling frequency point absolute value and an amplitude spectrum total energy relative value.
[0021] The step of updating the initial reference howling frequency point to obtain a target reference howling frequency point according to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located comprises:
[0022] When the frequency point interval in which the candidate howling frequency point is located is the same as the frequency point interval in which the initial reference howling frequency point is located, the candidate howling frequency point and the initial reference howling frequency point are merged to obtain the target reference howling frequency point; wherein the target reference howling frequency point comprises the candidate howling frequency point and the initial reference howling frequency point.
[0023] The howling condition comprises an initial howling condition and a target howling condition, and the step of determining that the suspected howling frame signal is a howling frame signal when the target reference howling frequency point meets a howling condition comprises:
[0024] when the duration of the target reference howling frequency point is greater than or equal to a preset time, it is determined that the target reference howling frequency point satisfies the initial howling condition;
[0025] when the target reference howling frequency point satisfying the initial howling condition satisfies the target howling condition, the suspected howling frame signal is determined as the howling frame signal.
[0026] The step of determining the suspected howling frame signal as the howling frame signal when the target reference howling frequency point satisfying the initial howling condition satisfies the target howling condition comprises:
[0027] when the number of occurrences of the target reference howling frequency point satisfying the initial howling condition continuously increases in a preset period, it is determined that the target reference howling frequency point satisfying the initial howling condition satisfies the target howling condition, and the suspected howling frame signal is determined as the howling frame signal.
[0028] The embodiment of the present application further provides a voice detection device, comprising:
[0029] a voice frame signal acquisition module, configured to acquire a to-be-detected voice frame signal;
[0030] a suspected howling frame signal determination module, configured to determine the to-be-detected voice frame signal as a suspected howling frame signal when the time domain energy and the time domain harmonic content of the to-be-detected voice frame signal satisfy a suspected howling condition;
[0031] a target reference howling frequency point acquisition module, configured to update an initial reference howling frequency point according to a frequency point interval in which a candidate howling frequency point in the suspected howling frame signal is located and a frequency point interval in which the initial reference howling frequency point is located, to obtain a target reference howling frequency point;
[0032] a determination module, configured to determine the suspected howling frame signal as a howling frame signal when the target reference howling frequency point satisfies a howling condition.
[0033] The embodiment of the present application further provides a computer readable storage medium, wherein a plurality of instructions are stored in the computer readable storage medium, and the instructions are suitable for being loaded by a processor to execute steps in the voice detection method.
[0034] The embodiment of the present application further provides an electronic device, comprising a processor and a memory, wherein the processor is electrically connected with the memory, the memory is used for storing instructions and data, and the processor is used for executing steps in the voice detection method.
[0035] Beneficial effects: The embodiments of this application provide a speech detection method, device, storage medium and electronic device. The method performs initial howling detection on the speech frame signal to be detected based on time-domain energy and time-domain harmonic content to determine the suspected howling frame signal. Then, it performs secondary howling detection on the suspected howling frame signal based on the target reference howling frequency point to determine the howling frame signal. This effectively improves the accuracy of howling detection and alleviates the technical problem of low accuracy of current howling detection. Attached Figure Description
[0036] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.
[0037] Figure 1 This is a flowchart illustrating the speech detection method provided in the embodiments of this application.
[0038] Figure 2a This is a histogram of the initial reference howling frequency point and the number of times the frequency point occurs, provided in the embodiments of this application.
[0039] Figure 2b This is a histogram of the target reference howling frequency point and the number of times the frequency point occurs, provided in the embodiments of this application.
[0040] Figure 3 This is another target reference howling frequency point - frequency point occurrence frequency histogram provided in the embodiments of this application.
[0041] Figure 4 This is a schematic diagram of the structure of the voice detection device provided in the embodiments of this application.
[0042] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.
[0043] Figure 6 This is another structural schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0045] This application provides a voice detection method, apparatus, storage medium, and electronic device.
[0046] like Figure 1 As shown, Figure 1is a flowchart of a voice detection method provided by an embodiment of the present application, and the specific flow can be as follows:
[0047] S101. Obtain a to-be-detected voice frame signal.
[0048] The to-be-detected voice frame signal is a frame signal that needs to be subjected to howling detection. Specifically, in a current voice network call process, the signal of a near-end microphone is returned to a near-end microphone through an acoustic path, and then returned from an acoustic path of a far end through a network environment, thereby forming feedback (i.e., howling), which is usually manifested as repeated multiple times of the same voice content, and seriously affects the quality of network call. In order to avoid the adverse effects of howling on the quality of network call, it is necessary to detect howling in time and suppress it.
[0049] Optionally, the step S101 specifically includes:
[0050] Obtain a to-be-detected voice signal.
[0051] Perform frame processing and windowing processing on the to-be-detected voice signal to obtain at least one to-be-detected voice frame signal.
[0052] The to-be-detected voice signal includes a voice signal extracted in a network call process. Specifically, since the to-be-detected voice signal is unstable as a whole, in order to ensure the reliability of subsequent howling detection, it is necessary to ensure the stability of the detected to-be-detected voice signal as much as possible, and therefore it is necessary to perform frame processing on the to-be-detected voice signal to divide it into multiple to-be-detected voice frame signals. Since the beginning and end of each to-be-detected voice frame signal will have discontinuity, the more to-be-detected voice frame signals are divided, the greater the error with the original to-be-detected voice signal. In order to avoid the impact of the error on subsequent howling detection, each to-be-detected voice frame signal is subjected to windowing processing to make each to-be-detected voice frame signal continuous, and each to-be-detected voice frame signal exhibits the characteristics of a periodic function.
[0053] Optionally, in the present embodiment, a Hanning window is used as a window function for windowing processing, and the obtaining process of the to-be-detected voice frame signal can be embodied by s l (n)=s(n)·win(n) where l is the number of the to-be-detected voice frame signal, n is the number of sampling points, s l (n) is the to-be-detected voice frame signal, s(n) is the to-be-detected voice signal, and win(n) is the window function.
[0054] S102. When the time domain energy of the to-be-detected voice frame signal and the time domain harmonic content satisfy a suspected howling condition, determine that the to-be-detected voice frame signal is a suspected howling frame signal.
[0055] The time domain energy is used to represent the energy size of the to-be-detected voice frame signal in the time domain, the time domain harmonic content is used to represent the harmonic content of the to-be-detected voice frame signal in the time domain, and the suspected howling condition is a basis for preliminarily judging whether the to-be-detected voice frame signal is likely to be a howling frame signal (i.e., a suspected howling frame signal).
[0056] Optionally, in the embodiment, the time domain energy of the to-be-detected voice frame signal is first calculated according to the signal length of the to-be-detected voice frame signal, and the time domain harmonic content of the to-be-detected voice frame signal is determined according to the normalized autocorrelation function of the to-be-detected voice frame signal. When the time domain energy is greater than or equal to the time domain energy threshold, and the time domain harmonic content is greater than or equal to the time domain harmonic content threshold, it is determined that the time domain energy and the time domain harmonic content satisfy the suspected howling condition, and the to-be-detected voice frame signal is determined to be a suspected howling frame signal.
[0057] Specifically, wherein N is the signal length of the to-be-detected voice frame signal; wherein the normalized autocorrelation function of the to-be-detected voice frame signal is:
[0058]
[0059] wherein m is the lag amount of the normalized autocorrelation function, and M is the maximum lag amount of the normalized autocorrelation function. is the signal length of the to-be-detected voice frame signal, F s is the sampling frequency, preferably, F s = 16 kHz, L is the total number of to-be-detected voice frame signals, and M0 is the first zero-crossing point. l
[0060] For example, the time domain energy threshold of the to-be-detected voice frame signal A is 80, the time domain harmonic content threshold is 0.8, the time domain energy is 100, and the time domain harmonic content is 0.9, which are calculated by the above formula. Since the time domain energy is greater than the time domain energy threshold, and the time domain harmonic content is greater than the time domain harmonic content threshold, it is determined that the time domain energy and the time domain harmonic content satisfy the suspected howling condition, and the to-be-detected voice frame signal A is determined to be a suspected howling frame signal.
[0061] S103. According to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located, and the frequency point interval in which the initial reference howling frequency point is located, the initial reference howling frequency point is updated to obtain a target reference howling frequency point.
[0062] The candidate howling frequency point is an extreme value in the amplitude spectrum of the suspected howling frame signal, the initial reference howling frequency point is used for monitoring the number of occurrences of the candidate howling frequency point, and the target reference howling frequency point is a basis for judging whether the suspected howling frame signal is a howling frame signal (i.e., a basis for analyzing the monitoring result of the number of occurrences of the candidate howling frequency point).
[0063] Specifically, the voice frame signal to be detected has been preliminarily screened in the above steps to obtain a suspected howling frame signal that may be a howling frame signal. In order to ensure the accuracy of howling detection, the suspected howling frame is further detected in this embodiment to further screen out a real howling frame signal.
[0064] In actual application, the frequencies are divided into 890 MHz, 890.2 MHz, 890.4 MHz, 890.6 MHz, 890.8 MHz, 891 MHz,..., 915 MHz, a total of 125 wireless frequency bands according to a frequency interval of 200 kHz, and each frequency band is numbered from 1, 2, 3, 4,..., 125. These numbers of the frequency bands are frequency points.
[0065] Further, before the step S103, the method further comprises:
[0066] obtaining an amplitude spectrum of the suspected howling frame signal;
[0067] performing smoothing processing on the amplitude spectrum, and determining an extreme value in the amplitude spectrum after the smoothing processing;
[0068] when the extreme value is greater than or equal to a candidate howling threshold value, determining that the extreme value is a candidate howling frequency point; wherein the candidate howling threshold value comprises a candidate howling frequency point absolute value and an amplitude spectrum total energy relative value.
[0069] In the amplitude spectrum, the curve often has slight fluctuations, which can easily adversely affect the subsequent howling detection process. Therefore, in this embodiment, the amplitude spectrum of the suspected howling frame signal is first smoothed to make the curve in the amplitude spectrum smoother and more stable.
[0070] Specifically, first, the Fourier transform of s l (n) is performed, and then the absolute value thereof is obtained to obtain the amplitude spectrum of the suspected howling frame signal, i.e., the amplitude spectrum of the suspected howling frame signal = S l (w) = abs(FFT(s l (n))), and then the amplitude spectrum is smoothed, i.e., the smoothed greater than or equal to the candidate howling frequency point absolute value and the amplitude spectrum total energy relative value, it is determined that the extreme value is a candidate howling frequency point. Optionally, an upper limit of the number of extreme values can be set according to actual requirements.
[0071] For example, the absolute value of the candidate howling frequency point is 82, the relative value of the total energy of the amplitude spectrum is 98, the upper limit of the number of extreme values is 4, and the determined extreme values are 80, 99, 103 and 115. Since 95 is less than the relative value of the total energy of the amplitude spectrum, 99, 103 and 115 are all greater than the absolute value of the candidate howling frequency point and the relative value of the total energy of the amplitude spectrum, 99, 103 and 115 are determined as the candidate howling frequency points.
[0072] Further, in one embodiment, when the candidate howling frequency point is in the same frequency point interval as the initial reference howling frequency point (which can be set according to actual needs), the candidate howling frequency point and the initial reference howling frequency point are merged to obtain a target reference howling frequency point (including the candidate howling frequency point and the initial reference howling frequency point).
[0073] Optionally, the initial reference howling frequency points are displayed in a histogram, and the histogram can represent the number of occurrences of each initial reference howling frequency point. When the candidate howling frequency point is in the same frequency point interval as the initial reference howling frequency point, the candidate howling frequency point and the initial reference howling frequency point are merged to obtain a target reference howling frequency point, and the number of occurrences of the frequency points in the overlapping and newly added parts is accumulated, the number of occurrences of the frequency points in the non-overlapping part is reduced, and the number of occurrences of each target reference howling frequency point is displayed in the histogram.
[0074] For example, the frequency point intervals [f1, f2, f3, f4], [f5, f6, f7, f8]……[f122, f123, f124, f125] are divided in advance, as shown in Figures 2a-2b The initial reference howling frequency points include [f1, f2, f3], and the number of occurrences corresponding to t1 is [2, 2, 4]. The candidate howling frequency points include [f2, f3, f4], and since the initial reference howling frequency points and the candidate howling frequency points are in the same frequency point interval, the two are merged to obtain the target reference howling frequency point [f1, f2, f3, f4]. Since f2 and f3 are frequency points in the overlapping part and f4 is a newly added frequency point, the number of occurrences of f2, f3 and f4 is increased by 1, and f1 is a frequency point in the non-overlapping part, so the number of occurrences of f1 is reduced by 1. That is, the number of occurrences corresponding to the target reference howling frequency point [f1, f2, f3, f4] at t2 is [1, 3, 5, 1].
[0075] In another embodiment, when the candidate howling frequency point is different from the frequency point interval in which the initial reference howling frequency point is located (i.e., the candidate howling frequency point and the initial reference howling frequency point are both different and not close), another set of initial reference howling frequency points is added to be compared with the candidate howling frequency point until the candidate howling frequency point and the initial reference howling frequency point are located in the same frequency point interval, the candidate howling frequency point and the initial reference howling frequency point are combined, and the target reference howling frequency point is obtained.
[0076] For example, the frequency point intervals [f1, f2, f3, f4], [f5, f6, f7, f8],..., [f122, f123, f124, f125] are divided in advance, the initial reference howling frequency point A includes [f1, f2, f3], the candidate howling frequency point includes [f9, f10, f11], and since the initial reference howling frequency point and the candidate howling frequency point are located in different frequency point intervals, a set of initial reference howling frequency points B [f10, f11, f12] is added. At this time, the candidate howling frequency point and the initial reference howling frequency point B are located in the same frequency point interval, so they are combined to obtain the target reference howling frequency point [f9, f10, f11, f12]. Among them, since f10 and f11 are frequency points of the superimposed part, f9 is a newly added frequency point, so the occurrence times of f9, f10 and f11 are increased by 1 respectively, and f12 is a frequency point of the non-superimposed part, so the occurrence time of f12 is reduced by 1.
[0077] S104. When the target reference howling frequency point satisfies the howling condition, the suspected howling frame signal is determined as a howling frame signal.
[0078] The howling condition is a basis for judging whether the suspected howling frame signal is a howling frame signal.
[0079] Specifically, in the embodiment, the howling condition includes an initial howling condition and a target howling condition. When the duration of the target reference howling frequency point is greater than or equal to a preset time, it is determined that the target reference howling frequency point satisfies the initial howling condition. When the target reference howling frequency point satisfying the initial howling condition satisfies the target howling condition, it is determined that the suspected howling frame signal is a howling frame signal. Alternatively, when the occurrence times of the target reference howling frequency point satisfying the initial howling condition continuously increase in a preset period, it is determined that the target reference howling frequency point satisfying the initial howling condition satisfies the target howling condition.
[0080] For example, as shown in FIG. 2, the suspected howling frame signal is determined as a howling frame signal when the target reference howling frequency point satisfies the howling condition. Figure 3As shown, the preset time is set as 3, the occurrence times of the target reference howling frequency points [f1, f2, f3, f4] at t1 are [2, 2, 4, 0], so the duration of f1, f2 and f3 is 1 (i.e. f1, f2 and f3 exist at t1), and the duration of f4 is 0; the occurrence times at t2 are [1, 3, 5, 1], so the duration of f1, f2 and f3 is 2 (i.e. f1, f2 and f3 exist from t1 to t2), and the duration of f4 is 1 (i.e. f4 exists from t2); the occurrence times at t3 are [0, 4, 6, 2], so the duration of f1 is 2 (i.e. f1 exists from t2 to t3), the duration of f2 and f3 is 3 (i.e. f2 and f3 exist from t1 to t3), and the duration of f4 is 2 (i.e. f4 exists from t2 to t3), and since the duration of f2 and f3 is equal to 3, it is determined that f2 and f3 satisfy the initial howling condition.
[0081] Further, it is assumed that the preset period is 3 (i.e. t1-t3), since the occurrence times of f2 in the period t1-t3 continuously increase (from 2 to 3, and then from 3 to 4), and the occurrence times of f3 in the period t2-t4 also continuously increase (from 4 to 5, and then from 5 to 6), it is determined that f2 and f3 satisfy the target howling condition, and the suspected howling frame signal is determined as a howling frame signal.
[0082] Optionally, after detecting the howling frame signal, a howling warning can be outputted, so as to subsequently and timely suppress the howling frame signal, thereby improving the communication quality.
[0083] As known from the above, the voice detection method provided in the present application first acquires a to-be-detected voice frame signal, and when the time domain energy and the time domain harmonic content of the to-be-detected voice frame signal satisfy a suspected howling condition, the to-be-detected voice frame signal is determined as a suspected howling frame signal, then the initial reference howling frequency point is updated according to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located, to obtain a target reference howling frequency point, and finally when the target reference howling frequency point satisfies a howling condition, the suspected howling frame signal is determined as a howling frame signal. The to-be-detected voice frame signal is initially detected based on the time domain energy and the time domain harmonic content to determine the suspected howling frame signal, and then the suspected howling frame signal is detected again based on the target reference howling frequency point to determine the howling frame signal, thereby effectively improving the accuracy of howling detection, and thus relieving the technical problem that the current howling detection accuracy is low.
[0084] According to the method described in the above embodiment, the present embodiment will be further described from the perspective of a voice detection device.
[0085] Please refer to Figure 4 ,Figure 4 The voice detection device provided by the embodiment of the present application is specifically described, which can comprise a voice frame signal acquisition module 10, a suspected howling frame signal determination module 20, a target reference howling frequency point acquisition module 30 and a determination module 40, wherein:
[0086] (1) The voice frame signal acquisition module 10
[0087] The voice frame signal acquisition module 10 is used for acquiring a voice frame signal to be detected.
[0088] The voice frame signal acquisition module 10 is specifically used for:
[0089] acquiring a voice signal to be detected;
[0090] performing frame processing and window processing on the voice signal to be detected to obtain at least one voice frame signal to be detected.
[0091] (2) The suspected howling frame signal determination module 20
[0092] The suspected howling frame signal determination module 20 is used for determining that the voice frame signal to be detected is a suspected howling frame signal when the time domain energy of the voice frame signal to be detected and the time domain harmonic content meet a suspected howling condition.
[0093] The suspected howling frame signal determination module 20 is specifically used for:
[0094] calculating the time domain energy of the voice frame signal to be detected according to the signal length of the voice frame signal to be detected, and determining the time domain harmonic content of the voice frame signal to be detected according to the normalized autocorrelation function of the voice frame signal to be detected;
[0095] when the time domain energy is greater than or equal to a time domain energy threshold value and the time domain harmonic content is greater than or equal to a time domain harmonic content threshold value, it is determined that the time domain energy and the time domain harmonic content meet the suspected howling condition, and it is determined that the voice frame signal to be detected is a suspected howling frame signal.
[0096] (3) The target reference howling frequency point acquisition module 30
[0097] The target reference howling frequency point acquisition module 30 is used for updating an initial reference howling frequency point to obtain a target reference howling frequency point according to the frequency point interval in which a candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located.
[0098] The target reference howling frequency point acquisition module 30 is specifically used for:
[0099] When the candidate howling frequency point is in the same frequency point interval as the initial reference howling frequency point, the candidate howling frequency point and the initial reference howling frequency point are merged to obtain a target reference howling frequency point; wherein, the target reference howling frequency point includes the candidate howling frequency point and the initial reference howling frequency point.
[0100] (4) determining module 40
[0101] The determining module 40 is configured to determine that the suspected howling frame signal is a howling frame signal when the target reference howling frequency point satisfies a howling condition.
[0102] The howling condition includes an initial howling condition and a target howling condition, and the determining module 40 is specifically configured to:
[0103] When the duration of the target reference howling frequency point is greater than or equal to a preset time, it is determined that the target reference howling frequency point satisfies the initial howling condition.
[0104] When the target reference howling frequency point that satisfies the initial howling condition satisfies the target howling condition, it is determined that the suspected howling frame signal is a howling frame signal.
[0105] Specifically, the determining module 40 is further configured to:
[0106] When the number of occurrences of the target reference howling frequency point that satisfies the initial howling condition continuously increases within a preset period, it is determined that the target reference howling frequency point that satisfies the initial howling condition satisfies the howling condition and satisfies the target howling condition, and it is determined that the suspected howling frame signal is a howling frame signal.
[0107] In specific implementation, each of the above modules can be implemented as an independent entity, or can be combined as the same or several entities, and the specific implementation of each of the above modules can be referred to the method embodiments above, which will not be repeated here.
[0108] From the above, the voice detection device provided by the application first acquires the to-be-detected voice frame signal through the voice frame signal acquisition module 10, determines that the to-be-detected voice frame signal is a suspected howling frame signal when the time domain energy and the time domain harmonic content of the to-be-detected voice frame signal meet the suspected howling condition through the suspected howling frame signal determination module 20, then updates the initial reference howling frequency point according to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located through the target reference howling frequency point acquisition module 30 to obtain the target reference howling frequency point, and finally determines that the suspected howling frame signal is a howling frame signal when the target reference howling frequency point meets the howling condition through the determination module 40. The to-be-detected voice frame signal is first detected for howling based on the time domain energy and the time domain harmonic content to determine the suspected howling frame signal, and then the suspected howling frame signal is detected for howling again based on the target reference howling frequency point to determine the howling frame signal, so that the accuracy of howling detection is effectively improved, thereby relieving the technical problem that the current howling detection accuracy is low.
[0109] Correspondingly, the embodiment of the application also provides a voice detection system, which comprises any voice detection device provided by the embodiment of the application, and the voice detection device can be integrated in an electronic device.
[0110] The to-be-detected voice frame signal is acquired; the to-be-detected voice frame signal is determined to be a suspected howling frame signal when the time domain energy and the time domain harmonic content of the to-be-detected voice frame signal meet the suspected howling condition; the initial reference howling frequency point is updated according to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located to obtain the target reference howling frequency point; and the suspected howling frame signal is determined to be a howling frame signal when the target reference howling frequency point meets the howling condition.
[0111] The specific implementation of each device can be referred to the foregoing embodiments, which will not be described here.
[0112] Since the voice detection system can comprise any voice detection device provided by the embodiment of the application, the beneficial effects that can be achieved by any voice detection device provided by the embodiment of the application can be achieved, which will be described in detail in the foregoing embodiments, which will not be described here.
[0113] In addition, the embodiment of the application also provides an electronic device. Figure 5 As shown in the figure, the electronic device 500 comprises a processor 501 and a memory 502. The processor 501 is electrically connected with the memory 502.
[0114] The processor 501 is the control center of the electronic device 500, connects each part of the whole electronic device by using various interfaces and lines, executes various functions of the electronic device and processes data by running or loading the application programs stored in the memory 502 and calling the data stored in the memory 502, and thus monitors the whole electronic device.
[0115] In the embodiment, the processor 501 in the electronic device 500 loads the instructions corresponding to the process of one or more application programs into the memory 502 and runs the application programs stored in the memory 502 by the processor 501 to realize various functions according to the following steps:
[0116] Obtain a to-be-detected voice frame signal;
[0117] When the time domain energy of the to-be-detected voice frame signal and the time domain harmonic content meet a suspected howling condition, determine that the to-be-detected voice frame signal is a suspected howling frame signal;
[0118] According to the frequency point interval in which the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval in which the initial reference howling frequency point is located, update the initial reference howling frequency point to obtain a target reference howling frequency point;
[0119] When the target reference howling frequency point meets a howling condition, determine that the suspected howling frame signal is a howling frame signal.
[0120] Figure 6 A specific structure block diagram of an electronic device provided by the embodiment of the present application is shown, and the electronic device can be used to implement the voice detection method provided in the above embodiments.
[0121] The RF circuit 610 is used to receive and send electromagnetic waves, and to convert electromagnetic waves and electrical signals to each other, so as to communicate with a communication network or other devices. The RF circuit 610 can include various existing circuit elements for performing these functions, for example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and the like. The RF circuit 610 can communicate with various networks, such as the Internet, an intranet, a wireless network, or communicate with other devices through a wireless network. The wireless network can include a cellular telephone network, a wireless local or metropolitan area network. The wireless network can use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short message service, and any other suitable communication protocol, even those not yet developed as of the date of the application.
[0122] The memory 620 can be used to store software programs and modules, and the processor 680 can execute various functions and data processing by running the software programs and modules stored in the memory 620, that is, to implement the function of storing 5G capability information. The memory 620 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 620 can further include a memory remotely disposed relative to the processor 680, which can be connected to the electronic device 600 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0123] The input unit 630 can be configured to receive input of numbers or characters, and generate key, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, the input unit 630 can include a touch-sensitive surface 631 and other input devices 632. The touch-sensitive surface 631, also known as a touch display or touchpad, can collect touch operations (such as operations of a user using a finger, a stylus, or any suitable object or accessory near the touch-sensitive surface 631) on or near the touch-sensitive surface 631 and drive corresponding connections according to pre-set programs. Optionally, the touch-sensitive surface 631 can include two parts of touch detection devices and touch controllers. Among them, the touch detection devices detect the touch position of the user and detect the signals generated by the touch operation, and transmit the signals to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch coordinates, and sends it to the processor 680, and can receive commands from the processor 680 and execute them. In addition, the touch-sensitive surface 631 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 631, the input unit 630 can also include other input devices 632. Specifically, the other input devices 632 can include one or more of a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc.
[0124] The display unit 640 can be configured to display information input by a user or information provided to a user and various graphical user interfaces of the electronic device 600, which can be composed of graphics, text, icons, video, and any combination thereof. The display unit 640 can include a display panel 641, which can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface 631 can cover the display panel 641, and when the touch-sensitive surface 631 detects a touch operation on or near it, it transmits to the processor 680 to determine the type of touch event, and then the processor 680 provides corresponding visual output on the display panel 641 according to the type of touch event. Although in the above description, the touch-sensitive surface 631 and the display panel 641 are implemented as two independent components to realize input and output functions, in some embodiments, the touch-sensitive surface 631 and the display panel 641 can be integrated to realize input and output functions. Figure 6
[0125] The electronic device 600 can also include at least one sensor 650, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor can include an ambient light sensor that can adjust the brightness of the display panel 641 according to the brightness of ambient light, and a proximity sensor that can turn off the display panel 641 and / or the backlight when the electronic device 600 is moved to the ear. As one of the motion sensors, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes), and when at rest, can detect the magnitude and direction of gravity, and can be used for identifying the posture of the mobile phone (such as switching between landscape and portrait, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), and the like. As for other sensors that the electronic device 600 can also be configured, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, and the like, will not be described here.
[0126] The audio circuit 660, the speaker 661, and the microphone 662 can provide an audio interface between the user and the electronic device 600. The audio circuit 660 can convert received audio data into an electrical signal, transmit the electrical signal to the speaker 661, and convert the electrical signal into a sound signal output by the speaker 661. On the other hand, the microphone 662 converts a sound signal collected into an electrical signal, and the audio circuit 660 receives the electrical signal, converts the electrical signal into audio data, and outputs the audio data to the processor 680 for processing. The audio data is processed by the processor 680 and transmitted to another terminal via the RF circuit 610, or output to the memory 620 for further processing. The audio circuit 660 can also include a jack for providing communication between an external earphone and the electronic device 600.
[0127] The electronic device 600 can help the user to send and receive emails, browse web pages, and access streaming media, etc. through the transmission module 670 (such as a Wi-Fi module), which provides the user with wireless broadband Internet access. Although Figure 6 The transmission module 670 is shown, but it can be understood that it does not belong to the necessary components of the electronic device 600, and can be omitted as needed without changing the essence of the application.
[0128] The processor 680 is the control center of the electronic device 600, which connects all parts of the mobile phone through various interfaces and lines, executes various functions of the electronic device 600 and processes data by running or executing software programs and / or modules stored in the memory 620, and calling data stored in the memory 620. Optionally, the processor 680 can include one or more processing cores; in some embodiments, the processor 680 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 680.
[0129] The electronic device 600 also includes a power supply 690, such as a battery, for powering various components of the electronic device 600. In some embodiments, the power supply 690 can be or include any of a variety of battery types or configurations. In some embodiments, the power supply 690 can be or include a renewable energy source. In some embodiments, the power supply 690 can be or include a power management system to manage charging, discharging, and power consumption of the power supply 690. In some embodiments, the power supply 690 can be or include one or more DC or AC power sources, recharging systems, power failure detection circuitry, power conversion or inverter circuitry, power status indicators, and any other components associated with the supply and consumption of power in the electronic device 600.
[0130] Although not shown, the electronic device 600 can also include a camera (e.g., a front-facing camera, a rear-facing camera), a Bluetooth module, and the like, which are not described in detail herein. In particular embodiments, the display unit of the electronic device is a touch screen display, and the electronic device further includes a memory, and one or more programs stored in the memory and configured for execution by the one or more processors, the one or more programs including instructions for:
[0131] obtaining a to-be-detected speech frame signal;
[0132] determining that the to-be-detected speech frame signal is a suspected howling frame signal when a time domain energy of the to-be-detected speech frame signal and a time domain harmonic content satisfy a suspected howling condition;
[0133] updating an initial reference howling frequency point to obtain a target reference howling frequency point according to a frequency point interval in which a candidate howling frequency point in the suspected howling frame signal is located and a frequency point interval in which the initial reference howling frequency point is located;
[0134] determining that the suspected howling frame signal is a howling frame signal when the target reference howling frequency point satisfies a howling condition.
[0135] In specific implementations, the above various modules can be implemented as independent entities, or can be combined as the same or a plurality of entities, and the specific implementation of the above various modules can be referred to the method embodiments described above, which will not be described herein.
[0136] Those skilled in the art can understand that all or part of the steps of the various methods in the above embodiments can be completed by instructions, or by related hardware controlled by the instructions, and the instructions can be stored in a computer readable storage medium and loaded and executed by a processor. Therefore, the embodiments of the present application provide a storage medium, which stores a plurality of instructions, and the instructions can be loaded by a processor to execute the steps in any voice detection method provided by the embodiments of the present application.
[0137] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or the like.
[0138] Due to the instructions stored in the storage medium, the steps of any of the voice detection methods provided by the embodiments of the present application can be executed, thus achieving the beneficial effects of any of the voice detection methods provided by the embodiments of the present application. Details are described in the foregoing embodiments, which will not be repeated here.
[0139] The specific implementation of each operation can refer to the foregoing embodiments, which will not be repeated here.
[0140] To sum up, although the present application has disclosed the above preferred embodiments, the above preferred embodiments are not used to limit the present application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application is defined by the scope of the claims.
Claims
1. A voice detection method characterized by, The method comprises the steps of: obtaining a to-be-detected voice frame signal; when the time domain energy of the to-be-detected voice frame signal and the time domain harmonic content meet a suspected howling condition, determining that the to-be-detected voice frame signal is a suspected howling frame signal, wherein the time domain harmonic content is used to represent the harmonic content of the to-be-detected voice frame signal in the time domain, and the time domain harmonic content is determined according to the maximum value after the first zero-crossing point of the normalized autocorrelation function of the to-be-detected voice frame signal; according to the comparison between the frequency point interval where the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval where the initial reference howling frequency point is located, merging the candidate howling frequency point and the initial reference howling frequency point located in the same frequency point interval, accumulating the occurrence frequency of the frequency points located in the same frequency point interval, accumulating the occurrence frequency of the frequency points in the newly added part of the frequency point interval where the candidate howling frequency point is located relative to the frequency point interval where the initial reference howling frequency point is located, and reducing the occurrence frequency of the frequency points in the reduced part of the frequency point interval where the candidate howling frequency point is located relative to the frequency point interval where the initial reference howling frequency point is located, to update the initial reference howling frequency point and obtain a target reference howling frequency point; when the target reference howling frequency point meets a howling condition, determining that the suspected howling frame signal is a howling frame signal; wherein, before the step of updating the initial reference howling frequency point according to the frequency point interval where the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval where the initial reference howling frequency point is located to obtain the target reference howling frequency point, the method further comprises the steps of: obtaining the amplitude spectrum of the suspected howling frame signal; performing smoothing processing on the amplitude spectrum and determining the extreme value in the amplitude spectrum after the smoothing processing; when the extreme value is greater than or equal to a candidate howling threshold, determining that the extreme value is the candidate howling frequency point; wherein the candidate howling threshold comprises a candidate howling frequency point absolute value and an amplitude spectrum total energy relative value.
2. The voice detection method of claim 1, wherein, The step of obtaining the to-be-detected voice frame signal comprises the steps of: obtaining a to-be-detected voice signal; performing frame processing and windowing processing on the to-be-detected voice signal to obtain at least one to-be-detected voice frame signal.
3. The voice detection method of claim 2, wherein, The step of determining that the to-be-detected voice frame signal is a suspected howling frame signal when the time domain energy of the to-be-detected voice frame signal and the time domain harmonic content meet a suspected howling condition comprises the steps of: calculating the time domain energy of the to-be-detected voice frame signal according to the signal length of the to-be-detected voice frame signal, and determining the time domain harmonic content of the to-be-detected voice frame signal according to the normalized autocorrelation function of the to-be-detected voice frame signal; when the time domain energy is greater than or equal to a time domain energy threshold and the time domain harmonic content is greater than or equal to a time domain harmonic content threshold, determining that the time domain energy and the time domain harmonic content meet the suspected howling condition, and determining that the to-be-detected voice frame signal is the suspected howling frame signal.
4. The voice detection method of claim 1, wherein, The step of updating the initial reference howling frequency point according to the frequency point interval where the candidate howling frequency point in the suspected howling frame signal is located and the frequency point interval where the initial reference howling frequency point is located to obtain the target reference howling frequency point comprises the steps of: When the candidate howling frequency point is the same as the frequency point interval where the initial reference howling frequency point is located, the candidate howling frequency point and the initial reference howling frequency point are merged to obtain the target reference howling frequency point; wherein the target reference howling frequency point includes the candidate howling frequency point and the initial reference howling frequency point.
5. The voice detection method of claim 4, wherein, The howling condition includes an initial howling condition and a target howling condition, and the step of determining the suspected howling frame signal as the howling frame signal when the target reference howling frequency point meets the howling condition includes: When the duration of the target reference howling frequency point is greater than or equal to a preset time, it is determined that the target reference howling frequency point meets the initial howling condition; When the target reference howling frequency point that meets the initial howling condition meets the target howling condition, the suspected howling frame signal is determined as the howling frame signal.
6. The voice detection method of claim 5, wherein, The step of determining the suspected howling frame signal as the howling frame signal when the target reference howling frequency point that meets the initial howling condition meets the target howling condition includes: When the number of occurrences of the target reference howling frequency point that meets the initial howling condition continuously increases within a preset period, it is determined that the target reference howling frequency point that meets the initial howling condition meets the target howling condition, and the suspected howling frame signal is determined as the howling frame signal.
7. A voice detection apparatus characterized by comprising: includes: A voice frame signal acquisition module is configured to acquire a to-be-detected voice frame signal. A suspected howling frame signal determination module is configured to determine the to-be-detected voice frame signal as a suspected howling frame signal when the time domain energy of the to-be-detected voice frame signal and the time domain harmonic content meet a suspected howling condition, wherein the time domain harmonic content is used to represent the harmonic content of the to-be-detected voice frame signal in the time domain, and the time domain harmonic content is determined according to the maximum value after the first zero-crossing point of the normalized autocorrelation function of the to-be-detected voice frame signal. A target reference howling frequency point acquisition module is configured to merge the candidate howling frequency point and the initial reference howling frequency point in the same frequency point interval according to the comparison between the frequency point interval where the candidate howling frequency point is located and the frequency point interval where the initial reference howling frequency point is located, accumulate the number of occurrences of the frequency points in the same frequency point interval, accumulate the number of occurrences of the frequency points in the newly added part of the frequency point interval where the candidate howling frequency point is located relative to the frequency point interval where the initial reference howling frequency point is located, and reduce the number of occurrences of the frequency points in the reduced part of the frequency point interval where the candidate howling frequency point is located relative to the frequency point interval where the initial reference howling frequency point is located, to update the initial reference howling frequency point and obtain the target reference howling frequency point. A determination module is configured to determine the suspected howling frame signal as the howling frame signal when the target reference howling frequency point meets the howling condition. The target reference howling frequency point acquisition module is further configured to acquire an amplitude spectrum of the suspected howling frame signal, perform smoothing processing on the amplitude spectrum, and determine an extreme value in the amplitude spectrum after the smoothing processing; and when the extreme value is greater than or equal to a candidate howling threshold, determine that the extreme value is the candidate howling frequency point; wherein the candidate howling threshold includes a candidate howling frequency point absolute value and an amplitude spectrum total energy relative value.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a plurality of instructions adapted to be loaded by the processor to perform the steps in the speech detection method of any one of claims 1 to 6.
9. An electronic device, comprising: The device comprises a processor and a memory, the processor is electrically connected with the memory, the memory is used for storing instructions and data, and the processor is used for performing the steps in the speech detection method of any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic sound feedback detection and elimination method of real-time communication system
CN106373587A
Audio signal processing method and hearing aid
CN110677796A
Howling detection method, voice communication method and related device
CN113450812A