Squeal Detection Method, Detection Device, Terminal, and Storage Medium

By obtaining the correlation between the target and historical audio signal power spectrum of the audio playback device, using covariance correlation coefficient and frame processing technology, the problem of low accuracy of the existing howling detection method is solved, and efficient and accurate howling detection is achieved.

CN114520948BActive Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011295093.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-18
Publication Date
2025-07-22
Estimated Expiration
2040-11-18

AI Technical Summary

Technical Problem

The existing howling detection methods have low accuracy, and are prone to false detection and missed detection, making it difficult to accurately judge howling under different acoustic environments and equipment characteristics.

Method used

By obtaining the power spectrum correlation between the target audio signal sent by the audio playback device and the historical audio signal, the covariance correlation coefficient is used to determine whether the audio playback device is screaming, and the detection accuracy is improved by combining frame processing and smooth filtering technology.

Benefits of technology

It improves the accuracy and efficiency of howling detection, reduces misjudgment and missed detection, and quickly recognizes howling and suppresses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114520948B_ABST
    Figure CN114520948B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a howling detection method, a detection device, a terminal, and a storage medium. Among them, the howling detection method obtains a target audio signal sent by an audio playback device to an audio receiving device at a first time, obtains a historical audio signal sent by the audio playback device to the audio receiving device at a second time, determines a target power spectrum according to the target audio signal, determines a historical power spectrum according to the historical audio signal, and when a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold, it is determined that the audio playback device has howled. Thus, the similarity between the target audio signal and the historical audio signal can be judged through the first correlation degree, and further whether the audio playback device has howled can be judged, which can improve the accuracy of howling detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly to a howling detection method, a detection device, a terminal and a storage medium. Background Art

[0002] Howling is a phenomenon of acoustic feedback. It is caused by the microphone repeatedly picking up and amplifying the sound reproduced by the speaker to form positive feedback. When the volume exceeds a certain limit, howling will occur in a certain frequency band.

[0003] The existing howling detection is mainly realized by making an abnormality judgment on the time-frequency domain energy of the audio signal. However, due to the differences in the acoustic environment during a call, the acoustic characteristics of the device, and the vocal characteristics of different people, even if an audio signal conforms to the time-frequency domain energy characteristics of howling, it does not necessarily mean that there is howling, thus prone to false detection; and when making a howling judgment, it is also difficult to determine a precise judgment threshold. If a conservative (tend to not misjudge) judgment threshold is set, although it can reduce the damage to normal sounds, it is prone to missed detection. It can be seen that the accuracy of the existing howling detection method is relatively low and has great limitations. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.

[0005] Embodiments of the present invention provide a howling detection method, a detection device, a terminal and a storage medium, which can improve the accuracy of howling detection.

[0006] In a first aspect, an embodiment of the present invention provides a howling detection method, including:

[0007] Obtain a target audio signal sent by an audio playback device to an audio receiving device at a first time;

[0008] Obtain a historical audio signal sent by the audio playback device to the audio receiving device at a second time, where the second time is earlier than the first time;

[0009] Determine a target power spectrum according to the target audio signal, and determine a historical power spectrum according to the historical audio signal;

[0010] When a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold, it is determined that the audio playback device has howled.

[0011] In a second aspect, an embodiment of the present invention further provides a howling detection device, including:

[0012] A first audio acquisition unit, configured to acquire a target audio signal sent by an audio playback device to an audio receiving device at a first time;

[0013] A second audio acquisition unit, configured to acquire a historical audio signal sent by the audio playback device to the audio receiving device at a second time, where the second time is earlier than the first time;

[0014] A power spectrum acquisition unit, configured to determine a target power spectrum according to the target audio signal and determine a historical power spectrum according to the historical audio signal;

[0015] A howling determination unit, configured to determine that the audio playback device has howled when a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold.

[0016] In a third aspect, an embodiment of the present invention further provides a howling detection device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the howling detection method described in the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present invention further provides a terminal, including the howling detection device described in the first aspect or the howling detection device described in the second aspect.

[0018] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The storage medium stores a program, and when the program is executed by a processor, the howling detection method described in the first aspect is implemented.

[0019] The embodiments of the present invention at least include the following beneficial effects: By acquiring a target audio signal sent by an audio playback device to an audio receiving device at a first time, acquiring a historical audio signal sent by the audio playback device to the audio receiving device at a second time, determining a target power spectrum according to the target audio signal, determining a historical power spectrum according to the historical audio signal, and determining that the audio playback device has howled when a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold, the similarity between the target audio signal and the historical audio signal can be judged through the first correlation degree, and then whether the audio playback device has howled can be judged, which can improve the accuracy of howling detection. Moreover, by judging the similarity between the target audio signal and the historical audio signal through the first correlation degree between the target power spectrum and the historical power spectrum, it has the advantage of fast convergence speed, which is beneficial to improving the efficiency of howling detection.

[0020] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the specification, claims and drawings. Description of the Drawings

[0021] The drawings are used to provide a further understanding of the technical solution of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention, and do not constitute a limitation to the technical solution of the present invention.

[0022] Figure 1 It is a schematic diagram of the cause of howling provided by an embodiment of the present invention.

[0023] Figure 2 It is a schematic diagram of a network architecture provided by an embodiment of the present invention.

[0024] Figure 3 It is a flowchart of a howling detection method provided by an embodiment of the present invention.

[0025] Figure 4 It is a specific step flowchart of obtaining a target power spectrum according to a target audio signal and obtaining a historical power spectrum according to a historical audio signal provided by an embodiment of the present invention.

[0026] Figure 5 It is a specific step flowchart of determining that a howling occurs in an audio playback device when a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold value provided by an embodiment of the present invention.

[0027] Figure 6 It is a specific step flowchart of determining that a howling occurs in an audio playback device when the first correlation degree is greater than the threshold value provided by an embodiment of the present invention.

[0028] Figure 7 It is a specific step flowchart of determining the first correlation degree between the target power spectrum and the historical power spectrum according to a plurality of second correlation degrees provided by an embodiment of the present invention.

[0029] Figure 8 It is a specific step flowchart of an example of a howling detection method provided by an embodiment of the present invention.

[0030] Figure 9 It is a schematic structural diagram of a howling detection device provided by an embodiment of the present invention.

[0031] Figure 10 It is a schematic structural diagram of a power spectrum acquisition unit provided by an embodiment of the present invention.

[0032] Figure 11 It is a schematic structural diagram of a howling determination unit provided by an embodiment of the present invention.

[0033] Figure 12 It is a schematic structural diagram of a howling determination unit provided by an embodiment of the present invention.

[0034] Figure 13 It is a schematic structural diagram of a second correlation determination unit provided by an embodiment of the present invention.

[0035] Figure 14 It is a schematic structural diagram of a terminal provided by an embodiment of the present invention.

[0036] Figure 15 It is another schematic structural diagram of a howling detection device provided by an embodiment of the present invention. Detailed implementation manners

[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] It should be understood that in the description of the embodiments of the present invention, the meaning of "a plurality of (or multiple)" is more than two. Understandings such as "greater than", "less than", and "exceeding" do not include the present number, and understandings such as "above", "below", and "within" include the present number. If there is a description of "first", "second", etc., it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0039] Howling is a phenomenon of acoustic feedback. It is caused by the microphone (i.e., the microphone) repeatedly picking up and amplifying the sound reproduced by the speaker to form positive feedback. When the volume exceeds a certain limit, howling will occur in a certain frequency band. The auditory sensation of howling is very harsh and unbearable. In practical applications, for example, during a call, if howling occurs, it will seriously affect the call quality. As an example, the howling phenomenon during a call often occurs in a multi-person meeting (some participants are relatively close). After howling appears, due to the large interference of howling, the call process may be directly interrupted and unable to proceed. Therefore, the howling problem is an extremely serious call failure problem. And howling detection is a prerequisite for effectively suppressing howling. If howling detection misses a detection, the howling will continue and become louder and harsher; on the contrary, if howling detection has a false detection, the normal voice will be damaged, which is also not conducive to the call.

[0040] Refer to Figure 1, taking the terminal as an example for illustration, the audio signal played by the second terminal speaker enters the first terminal microphone after being transmitted through the air and is collected by the first terminal. The audio signal collected by the first terminal is transmitted back to the second terminal through the network, and after being processed by voice enhancement inside the second terminal, it is played by the second terminal speaker again and is collected by the first terminal microphone again, repeating this cycle. When the closed-loop gain of this closed-loop process is greater than 1, the audio signal will be amplified every time it circulates. After multiple circulations, the volume of the audio signal may be greater than the original signal and even reach a harsh volume level. Such periodically or quasi-periodically continuously loop-played audio signals are called howls.

[0041] Existing howl detection is mainly achieved by making abnormality judgments on the time-frequency domain energy of audio signals. Common howl detection can be divided into two methods: time-domain detection and frequency-domain detection.

[0042] Time-domain detection is to make judgments based on the sudden change of the signal energy of howls and the possible reciprocating periodicity of howl signals. When the signal energy suddenly becomes larger and exceeds the preset threshold or a periodic energy transient signal appears, it is judged as a howl. Due to the single judgment basis in the time domain, only based on the overall energy, with little information, it is easy to misjudge speech as a howl.

[0043] Frequency-domain detection mainly has the following two types:

[0044] One is the peak-to-average ratio judgment: that is, taking the ratio of the power peak value of the current frame to the average power of the current frame as the judgment basis, comparing the ratio of the power peak value of the current frame to the average power of the audio signal with the preset threshold. If it exceeds the preset threshold, it is judged that a howl has occurred.

[0045] The other is the multi-reference frequency point judgment: different from the peak-to-average ratio judgment, the multi-reference frequency point judgment does not use the average power value of the current frame, but divides each frame into multiple frequency bands in the frequency domain, and different frequency bands are set with different reference power values. Calculate the ratio of the power value of each frequency point in each frequency band to the reference power value corresponding to that frequency band. If the ratio exceeds the preset threshold, it is judged that a howl has occurred.

[0046] However, if the peak-to-average ratio judgment method is used to detect howls, due to the large power difference of audio signals at different frequency points, for example, the low-frequency power value is large while the high-frequency power value is small, or the power changes in different frequency bands fluctuate greatly, resulting in large fluctuations in the average power value of each frame and unstable peak values. For example, for some voiced sounds, more than 70% of the energy is below 1.5 kHz, while the energy above 1.5 kHz is very weak. In this way, the calculated average value will be pulled down by the energy above 1.5 kHz, resulting in a high peak value. In this way, voiced sounds are easily misjudged as howls. Therefore, the peak-to-average ratio judgment method will lead to a high probability of misjudging howl frequency points.

[0047] If the howling is detected by using the method of multi-reference frequency point judgment, although different frequency bands are divided to avoid the mutual influence caused by large differences in band power, the accuracy of howling judgment is improved compared with the method of peak-to-average ratio judgment. However, the method of multi-reference frequency point judgment ignores the objective fact that there are large differences in the power spectrum distributions of different speech categories. For example, speech can be divided into voiceless and voiced sounds, and there are obvious differences in the distribution ratios of the power spectra of voiceless and voiced sounds in the entire frequency domain space. The frequency domain energy of voiceless sounds is concentrated in the middle and high frequencies, while the low-frequency energy is relatively low. On the contrary, the proportion of low-frequency energy of voiced sounds is relatively high, and the middle and high-frequency energy is relatively low. If the method of multi-reference frequency point judgment does not make a refined distinction of speech categories, it will lead to voiceless sounds being easily misjudged as howling.

[0048] In addition, due to the differences in the acoustic environment during a call, the acoustic characteristics of devices, and the vocal characteristics of different voices, it is also very difficult to set the preset thresholds of the above-mentioned detection methods more accurately. Therefore, the accuracy of the existing howling detection methods needs to be improved.

[0049] To solve the above problems, the embodiments of the present invention provide a howling detection method, a detection device, a terminal, and a storage medium, which can improve the accuracy of howling detection.

[0050] Refer to Figure 2 , which is a schematic structural diagram of a network architecture provided by an embodiment of the present invention. The network architecture may include a server 2001 and a terminal cluster 2002. The terminal cluster 2002 may include multiple terminals, specifically terminals 2002a, 2002b, 2002c... 2002n. The howling detection method provided by the embodiments of the present invention may use any terminal in the terminal cluster 2002 as the execution subject, or use the server 2001 as the execution subject. Or, in other embodiments, the howling detection method provided by the embodiments of the present invention may also use an independent howling detection device as the execution subject. Among them, the terminal 2002a may be a desktop terminal or a mobile terminal, and the mobile terminal may specifically be at least one of a mobile phone, a tablet computer, and a laptop computer. The server 2001 may be an independent server providing voice call support for the terminal cluster 2002 or a cluster server composed of multiple servers. The howling detection method in this embodiment may also be applied to an application scenario including only the terminal cluster 2002 or including only the terminal cluster 2002 and the server 2001, such as a human-robot chat program or an artificial intelligence program set in the terminal cluster 2002 or the server 2001.

[0051] Refer to Figure 3 , which is a flowchart of a howling detection method provided by an embodiment of the present invention. The howling detection method may include but is not limited to the following steps 301, 302, 303, and 304.

[0052] Step 301: Obtain a target audio signal sent by a first-time audio playback device to an audio receiving device.

[0053] It can be understood that the audio playback device is used to play audio signals. Exemplarily, the audio playback device can be a terminal with a speaker or a separate speaker. Similarly, the audio receiving device is used to receive audio signals. The audio receiving device can be a terminal with a microphone or a separate microphone. Among them, the audio playback device and the audio receiving device can be the same device or different devices. For example, when the audio playback device and the audio receiving device are the same device, the audio playback device and the audio receiving device can be the speaker and the microphone on the same terminal; when the audio playback device and the audio receiving device are different devices, the audio playback device and the audio receiving device can be different terminals. Taking Figure 2 the network architecture shown as an example, the audio playback device can be the speaker and the microphone on terminal 2002a, or the audio playback device can be terminal 2002a, and the audio receiving device can be terminal 2002b. Of course, the audio playback device and the audio receiving device can also be Figure 2 other terminals in

[0054] Step 302: Obtain a historical audio signal sent by a second-time audio playback device to an audio receiving device.

[0055] Among them, the second time is earlier than the first time. Both the target audio signal and the historical audio signal are audio signals sent by the audio playback device. The target audio signal and the historical audio signal can be audio signals sent by the audio playback device at different times. The target audio signal can be the audio signal currently sent by the audio playback device, and the historical audio signal can be the audio signal previously sent by the audio playback device. Of course, the target audio signal can also be the audio signal previously sent by the audio playback device. In this case, the historical audio signal can be the audio signal sent even earlier than the target audio signal. It can be seen that the target audio signal can be the currently played audio signal or the recorded audio signal that has been played.

[0056] It can be understood that the above-mentioned first time and second time can refer to a certain moment or a certain time period, and the embodiments of the present invention do not make any limitations.

[0057] As an example, if the first time is the current time and the second time is earlier than the first time, the first time can be 8:00:00 and the second time can be 7:59:58; if the first time is a historical time, the second time is earlier than the first time, and the current time is 8:00:00, then the first time can be 7:59:58 and the second time can be 7:59:56. Of course, the above is only an exemplary illustration, and the embodiments of the present invention do not limit the first time and the second time.

[0058] Step 303: Determine the target power spectrum according to the target audio signal, and determine the historical power spectrum according to the historical audio signal.

[0059] Among them, the power spectrum represents the variation of the audio signal power with frequency, that is, the distribution of the signal power in the frequency domain. Determining the target power spectrum according to the target audio signal and determining the historical power spectrum according to the historical audio signal can be achieved by means such as Fourier transform.

[0060] Since the transmission times of the target audio signal and the historical audio signal are different, in order to facilitate obtaining the historical power spectrum of the historical audio signal and realizing the subsequent comparison between the target power spectrum and the historical power spectrum, therefore, in step 303, after determining the historical power spectrum according to the historical audio signal, the historical power spectrum of the historical audio signal is cached. At this time, when comparing the target power spectrum and the historical power spectrum, the historical power spectrum can be obtained from the cache area. Exemplarily, the data space size of the cache area can be the memory size occupied by the audio signal with a length of several seconds, such as the memory size occupied by the audio signal of 5 seconds. Of course, the data space size can be set according to the actual situation, and the embodiments of the present invention do not make any limitations.

[0061] Step 304: When the first correlation degree between the target power spectrum and the historical power spectrum is greater than the threshold, it is determined that the audio playback device has a howling sound.

[0062] It can be understood that the first correlation degree is used to characterize the similarity between the target power spectrum and the historical power spectrum. Since the howling sound has the characteristic of reciprocating transmission and playback, there is a certain similarity between the howling signals in each cycle. Therefore, by obtaining the first correlation degree between the target power spectrum and the historical power spectrum, the similarity between the target audio signal and the historical audio signal can be judged through this first correlation degree. When the first correlation degree between the target power spectrum and the historical power spectrum is relatively high, it can be further determined that the audio playback device has a howling sound.

[0063] It can be understood that the above threshold can be set according to the actual situation, and the embodiments of the present invention do not make any limitations.

[0064] In one embodiment, the above first correlation degree may adopt a covariance correlation coefficient, which can better reflect the similarity between the power spectra of the target audio signal and the historical audio signal. Among them, the first power absolute value of the target power spectrum and the second power absolute value of the historical power spectrum may be obtained, the first covariance between the first power absolute value and the second power absolute value may be obtained, the first standard deviation of the first power absolute value and the second standard deviation of the second power absolute value may be obtained, the first product between the first standard deviation and the second standard deviation may be obtained, and the first quotient value between the first covariance and the first product may be used as the first correlation degree.

[0065] In one embodiment, the calculation formula of the first correlation degree includes: Where ρ(X) is the first correlation degree, X1 is the target power spectrum, X2 is the historical power spectrum, where Conv(X1, X2) represents the covariance of X1 and X2, Conv(X1, X2)=E(X1X2)-E(X1)E(X2), where E(X1) represents the mean of X1, and similarly, E(X2) represents the mean of X2; D(X1)=(X1-E(X1)) 2 , and similarly, D(X2)=(X2-E(X2)) 2 .

[0066] By obtaining the target audio signal sent by the audio playback device to the audio receiving device at the first time, obtaining the historical audio signal sent by the audio playback device to the audio receiving device at the second time, determining the target power spectrum according to the target audio signal, determining the historical power spectrum according to the historical audio signal, when the first correlation degree between the target power spectrum and the historical power spectrum is greater than the threshold, it is determined that the audio playback device has a howling sound. Thus, the similarity between the target audio signal and the historical audio signal can be judged through the first correlation degree, and further whether the audio playback device has a howling sound can be judged, which can improve the accuracy of howling detection. Moreover, by judging the similarity between the target audio signal and the historical audio signal through the first correlation degree between the target power spectrum and the historical power spectrum, it has the advantage of fast convergence speed, which is beneficial to improving the efficiency of howling detection.

[0067] It can be understood that based on Figure 2 the network architecture shown, any one of the audio playback device, the audio receiving device, and the server can execute the above howling detection method. Therefore, obtaining the target audio signal and the historical audio signal can be obtained locally or from other devices, and the obtaining method can be transmission methods such as network and Bluetooth. The embodiments of the present invention do not make any limitations.

[0068] Exemplarily, based on Figure 2In the network architecture shown, the audio playback device can be the terminal 2002a, the audio receiving device can be the terminal 2002b, the target audio signal and the historical audio signal can be the audio signals sent by the terminal 2002a. If the entity executing the howling detection method is the terminal 2002a, the terminal 2002a can obtain the audio signal it sends as the target audio signal, and the terminal 2002a can obtain the audio signal cached by itself as the historical audio signal; if the entity executing the howling detection method is the terminal 2002b, the terminal 2002b can obtain the audio signal sent by the terminal 2002a received by itself as the target audio signal, and the terminal 2002b can obtain the audio signal sent by the terminal 2002a cached by itself as the historical audio signal; if the entity executing the howling detection method is the server 2001, the server can obtain the target audio signal and the historical audio signal through the network.

[0069] It should be noted that the howling detection method in the embodiments of the present application can be applied to the application scenarios of voice communication, and can also be applied to the human-computer interaction scenarios with speaker playback, such as intelligent devices such as intelligent robots with voice calls, smart speakers, and smart watches. The target audio signal and the historical audio signal can include but are not limited to audio signals such as user voice (including call voice), music, other background sounds, synthesized sounds, and prompt sounds.

[0070] In one embodiment, referring to Figure 4 , the above step 303 can further include step 401, step 402, step 403, step 404, and step 405.

[0071] Step 401: Perform frame splitting processing on the target audio signal and the historical audio signal.

[0072] Specifically, the window function can be used to perform frame addition and windowing on the target audio signal and the historical audio signal. In order to reduce spectral energy leakage, the window function can be used to achieve this. Exemplarily, one of the following window functions can be used: rectangular window, triangular window, Hanning window, Hamming window, etc. Among them, the rectangular window belongs to the zero-th power window of the time variable. The rectangular window is used the most, and it is customary that not adding a window means that the signal passes through the rectangular window. The advantage of this window is that the main lobe is relatively concentrated, and the disadvantage is that the side lobes are relatively high and there are negative side lobes, resulting in high-frequency interference and leakage in the transformation, and even negative spectrum phenomena. The triangular window is in the form of the first power of the power window. Compared with the rectangular window, the main lobe width is about twice that of the rectangular window, but the side lobes are small and there are no negative side lobes. The Hanning window can be regarded as the sum of the spectra of 3 rectangular time windows. The main lobe of the Hanning window is broadened and reduced, and the side lobes are significantly reduced, which is beneficial to reducing leakage. The Hamming window and the Hanning window are both cosine windows, only the weighting coefficients are different, and the weighting coefficient of the Hamming window can make the side lobes smaller.

[0073] As an example, the formula of the Hann window can be expressed as follows: where n belongs to [0, N - 1], and N is a positive integer.

[0074] It can be understood that the time length of each frame can be set according to the actual situation. For example, it can be set to 10 milliseconds or 20 milliseconds. The embodiments of the present invention do not make any limitations.

[0075] Step 402: Obtain the first sub-audio signal of the current frame in the target audio signal after frame division processing.

[0076] Wherein, the current frame can be one of the frames in the target audio signal after frame division processing. It can be understood that the current frame can also be regarded as the target audio signal only. In the actual detection process, each frame audio signal of the target audio signal after frame division processing can be obtained in sequence. When the first frame of the target audio signal is obtained, the current frame is the first frame, and the first frame of the target audio signal is the first sub-audio signal. Of course, in the actual detection process, the acquisition order of each frame audio signal in the target audio signal can also be set according to the actual situation.

[0077] In one embodiment, when the target audio signal and the historical audio signal are continuous, when the current frame is the second frame, the first frame can be correspondingly cached in the buffer area and become the historical audio signal, and so on.

[0078] Step 403: Obtain multiple second sub-audio signals before or after the current frame in the historical audio signal after frame division processing.

[0079] After the historical audio signal undergoes frame division processing, multiple frames of audio signals can be obtained. Multiple second sub-audio signals before or after the current frame in the historical audio signal after frame division processing are actually multiple offset frames of the current frame of the target audio signal towards the historical time. Among them, obtaining multiple second sub-audio signals before or after the current frame in the historical audio signal after frame division processing can be obtaining all the second sub-audio signals after the current frame in multiple frames of historical audio signals, or obtaining some of the second sub-audio signals after the current frame in the historical audio signal. The embodiments of the present invention do not make any limitations. Obtaining all the second sub-audio signals after the current frame in multiple frames of historical audio signals is beneficial to improving the accuracy of subsequent comparison.

[0080] It can be understood that according to the caching method of the power spectrum of the historical audio signal, the second sub-audio signal can be the historical audio signal before the current frame or the historical audio signal after the current frame. The embodiments of the present invention do not make any limitations.

[0081] Step 404: Obtain the first power spectrum of the first sub-audio signal and the second power spectrum of the second sub-audio signal.

[0082] Among them, obtaining the first power spectrum of the first sub-audio signal may be to perform a Fourier transform on the first sub-audio signal and calculate the corresponding absolute value of the power. The first sub-audio signal is the i-th frame of the target audio signal, and the calculation formula may include:

[0083] Among them, k is the total number of frequency points, k = 1, 2, 3,..., N.

[0084] Similarly, the calculation method of the second power spectrum X(i + j, k) of the second sub-audio signal is the same and will not be elaborated here.

[0085] As an example, the number of second sub-audio signals is nine, and correspondingly, the number of second correlation degrees is also nine. Obtaining multiple second correlation degrees between the first power spectrum and multiple second power spectra, that is, obtaining nine second correlation degrees between the first power spectrum and nine second power spectra respectively.

[0086] It should be added that the number of second sub-audio signals is not limited to the above nine, and it is only used as an exemplary explanation. The specific number can be determined according to the frame division processing in step 401.

[0087] In one embodiment, obtaining the first power spectrum of the first sub-audio signal and the multiple second power spectra of the multiple second sub-audio signals may be to obtain the first power spectrum of the first sub-audio signal and the multiple second power spectra of the multiple second sub-audio signals according to a preset frequency range, that is, the above k = N1 - N2, where N1 and N2 represent the preset frequency range. Since whistling usually does not occur in the full frequency band, by setting a preset frequency range and only using some frequency bands to judge the similarity, it is beneficial to improve the convergence speed of the whistling detection method and improve the efficiency of whistling detection.

[0088] Step 405, taking the first power spectrum as the target power spectrum of the target audio signal and taking the second power spectrum as the historical power spectrum of the historical audio signal.

[0089] Refer to Figure 5 , based on the above steps 401 to 405, the above step 304 may specifically include steps 501 to 503.

[0090] Step 501, respectively obtaining the second correlation degrees between the first power spectrum and the multiple second power spectra.

[0091] In one embodiment, the second covariance between the third absolute power values of each frequency point of the i-th frame in the first sub-audio signal and the fourth absolute power values of each frequency point of the j-th frame before or after the i-th frame in the second sub-audio signal can be obtained, the third standard deviation of the third absolute power values and the fourth standard deviation of the fourth absolute power values can be obtained, and the second product between the third standard deviation and the fourth standard deviation can be obtained. The second quotient between the second covariance and the second product is used as the second correlation degree, where both i and j are positive integers.

[0092] In one embodiment, the above second correlation degree can be obtained using the following calculation formula:

[0093] Where X(i, k) is the absolute power value of each frequency point of the i-th frame in the first sub-audio signal, X(i + j, k) is the absolute power value of each frequency point of the j-th frame before or after the i-th frame in the second sub-audio signal, k is the total number of frequency points, and i, j, and k are all positive integers.

[0094] Where Conv(X(i, k), X(i + j, k)) represents the covariance of X(i, k) and X(i + j, k), Conv(X(i, k), X(i + j, k)) = E(X(i, k)X(i + j, k)) - E(X(i, k))E(X(i + j, k)), where E(X(i, k)) represents the mean value of X(i, k), and similarly, E(X(i + j, k)) represents the mean value of X(i + j, k); D(X(i, k)) = (X(i, k) - E(X(i, k))) 2 , and similarly, D(X(i + j, k)) = (X(i + j, k) - E(X(i + j, k))) 2 .

[0095] The above ρ(j) is the covariance correlation coefficient, which can better reflect the similarity degree between the power spectra of the target audio signal and the historical audio signal.

[0096] By performing frame splitting on the target audio signal and the historical audio signal, multiple frames of target audio signals and multiple frames of historical audio signals are obtained. Then, the second correlation degrees of the first power spectrum and multiple second power spectra are respectively obtained. Since the frame splitting is performed on the target audio signal and the historical audio signal, the dimension of the audio signal is smaller. Compared with the solution of directly judging the overall similarity of the audio signal, the howling detection method provided by the embodiment of the present invention is beneficial to improving the accuracy when judging the similarity degree.

[0097] Step 502, determine the first correlation degree between the target power spectrum and the historical power spectrum according to multiple second correlation degrees.

[0098] Among them, determining the first correlation degree between the target power spectrum and the historical power spectrum according to multiple second correlation degrees may be using the second correlation degree corresponding to the third power spectrum as the first correlation degree, where the third power spectrum is the second power spectrum with the highest similarity to the first power spectrum. For example, the largest second correlation degree may be used as the first correlation degree. Using the second correlation degree corresponding to the second power spectrum with the highest similarity to the first power spectrum as the first correlation degree is beneficial to making the first correlation degree more reasonably reflect the similarity degree between the first power spectrum and the second power spectrum. Of course, in other embodiments, the similarity degree may not be directly judged by the magnitude of the second correlation degree. For example, some mathematical operations may be performed on the second correlation degree to convert it into another parameter, and then this parameter is used for judgment.

[0099] Step 503: When the first correlation degree is greater than the threshold, it is determined that the audio playback device has a howling sound.

[0100] Among them, when the first correlation degree is greater than the threshold and it is determined that the audio playback device has a howling sound, based on step 502, it may be that when the maximum value among multiple second correlation degrees is greater than the threshold, it is determined that the audio playback device has a howling sound. Therefore, that is, when the similarity between the first power spectrum of a certain frame of audio signal in the target audio signal and the second power spectrum of a certain frame of audio signal in the historical audio signal is relatively large, it can be determined that the audio playback device has a howling sound. It can be understood that the maximum value among the second correlation degrees represents the highest similarity degree between the first power spectrum and the second power spectrum.

[0101] It can be understood that in other embodiments, the threshold may also be compared every time the second correlation degree between the first power spectrum and a second power spectrum is obtained. Once the second correlation degree is greater than the threshold, it is determined that the audio playback device has a howling sound. The advantage of this method is that it is not necessary to obtain all the multiple second correlation degrees between the first power spectrum and multiple second power spectra. Therefore, it is beneficial to improve the convergence speed of the howling detection method and the efficiency of howling detection.

[0102] Refer to Figure 6 , in one embodiment, the above step 503 may further include step 601 and step 602.

[0103] Step 601: When the first correlation degree is greater than the threshold, obtain the time difference between the first sub-audio signal and the second sub-audio signal with the highest similarity degree according to the second correlation degree.

[0104] In the actual application process, even when the similarity between two audio signals is relatively high, there may still be a situation where both of these two audio signals are normal speech signals and no howling occurs at this time. Therefore, based on the above steps 501 to 503, in the embodiment of the present invention, after determining that the first correlation degree is greater than the threshold, the time difference between the first sub-audio signal and the second sub-audio signal with the highest similarity degree is further obtained according to the second correlation degree. Among them, the first sub-audio signal and the second sub-audio signal with the highest similarity degree can be the first sub-audio signal and the second sub-audio signal corresponding to the maximum value of the second correlation degree. Among them, the time difference between the first sub-audio signal and the second sub-audio signal can be regarded as the period when howling occurs.

[0105] In one embodiment, the time difference between the first sub-audio signal and the second sub-audio signal can be obtained through the frame distance between the first sub-audio signal and the second sub-audio signal. For example, when the frame distance between the first sub-audio signal and the second sub-audio signal is 5 frames, if the length of each frame is 10 milliseconds, the time difference between the first sub-audio signal and the second sub-audio signal is 50 milliseconds.

[0106] Step 602, when the time difference is within a preset period range, it is determined that the audio playback device has howled.

[0107] Since the period when howling occurs is generally very short, in order to exclude the situation of two normal speech signals with relatively high similarity, the time difference between the first sub-audio signal and the second sub-audio signal is compared with the preset period range. Only when the time difference is within the preset period range, it is determined that the audio playback device has howled, thereby improving the accuracy of howling detection and reducing the occurrence of misjudgment phenomena. As an example, the above period range can be within 2 seconds. It can be understood that the above period range can be set according to the actual situation, and the embodiment of the present invention does not make any limitations.

[0108] Refer to Figure 7 In one embodiment, the above step 502 may further include step 701 and step 702.

[0109] Step 701, perform a smoothing filtering process on multiple second correlation degrees.

[0110] After obtaining the second correlation degree, a smoothing filtering process can also be performed on the second correlation degree, so as to filter out the interference signals generated during the calculation of the second correlation degree and improve the accuracy of howling detection.

[0111] Among them, the calculation formula of the smoothing filtering process includes:

[0112] P(j) = (1 - α) * P(j) + α * ρ(j), where ρ(j) is the second correlation between the second sub-audio signal and the first sub-audio signal of the j-th frame before or after the current frame in the historical audio signal after frame processing, P(j) is the second correlation after smoothing filtering, j is a positive integer, and α is a positive number less than 1.

[0113] Step 702, determine the first correlation between the target power spectrum and the historical power spectrum according to multiple second correlations after smoothing filtering.

[0114] Among them, since the second correlation has been smoothed filtered, it is beneficial to improve the accuracy of the first correlation between the finally obtained target power spectrum and the historical power spectrum.

[0115] It should be added that the howling detection method in any of the above embodiments can be carried by the system program of the terminal or by the application program that needs to use the microphone and speaker in the terminal. The embodiments of the present application do not make specific limitations. When the howling detection method in any of the above embodiments is carried by the system program of the terminal, the user only needs to upgrade and update the system program of the terminal to use the howling detection method in any of the above embodiments, without having to perform factory debugging on the terminal, which is convenient for the user to update and use. When the howling detection method in any of the above embodiments is carried by the application program that needs to use the microphone and speaker in the terminal, only by installing the application program can the howling detection method in any of the above embodiments be used, and there is also no need to perform factory debugging on the terminal, which is convenient for the user to update and use. The application program can be an independent voice signal processing software or plug-in, or an algorithm program built into the voice call software, such as application programs for WeChat audio and video calls, live broadcasts, broadcasts, etc. The embodiments of the present application do not make limitations.

[0116] Based on Figure 2 the network architecture shown, where both the terminal 2002a and the terminal 2002b are installed with voice communication software capable of voice communication. The voice communication software can be chat, social software, or remote video conferencing software, etc. The howling detection method of the embodiments of the present invention will be described below taking the remote video conferencing scenario as an example.

[0117] As an example, in this application scenario, the distance between terminal 2002a and terminal 2002b is relatively close. Moreover, both terminal 2002a and terminal 2002b establish a voice communication connection through server 2001 to conduct a remote video conference with another terminal at the far end. Terminal 2002a and terminal 2002b respectively play the audio signals sent by the far-end terminal through their own speakers. Since the distance between terminal 2002a and terminal 2002b is relatively close, the audio signal played by the speaker of terminal 2002a will be picked up by the microphone of terminal 2002b. Terminal 2002b transmits the audio signal picked up by its microphone to the server through the network, and the server then transmits this audio signal to terminal 2001a. Terminal 2001a plays this audio signal again through its own speaker, and the microphone of terminal 2002b picks up this audio signal again... and so on. Once the volume of this audio signal reaches a certain limit, howling will occur, thus affecting the video call between terminal 2002a, terminal 2002b and the far-end terminal. To detect whether howling occurs in terminal 2002a, first, when terminal 2002a, terminal 2002b and the far-end terminal start a video call, terminal 2002a starts to collect the audio signal played by its own speaker and caches it in a buffer area preset in itself. Among them, the size of the buffer area can allow terminal 2002a to cache the audio signal for 5 seconds. Terminal 2002a continuously collects the target audio signal played by its own speaker, and calculates the correlation between the power spectrum of the current frame of the currently collected target audio signal and the power spectrum of each frame of the historical audio signal in the buffer area in turn, and takes the maximum value of these correlations and compares it with a threshold. Once the maximum value of these correlations reaches the threshold, the time difference between a certain frame of the historical audio signal with the maximum correlation and the current frame of the target audio signal is confirmed and obtained. Further, this time difference is compared with a preset cycle range. Once this time difference is within the preset cycle range, it is determined that howling occurs in terminal 2002a, and corresponding howling suppression actions need to be executed subsequently. The detected target audio signal can be cached as a historical audio signal. Through the above howling detection method, the similarity between the target audio signal and the historical audio signal is judged by the correlation between their power spectra, and then it is judged whether howling occurs in terminal 2002a. Compared with the existing howling detection methods, the accuracy of howling detection can be improved. Moreover, on the basis that the similarity reaches the threshold, the judgment of the howling cycle value is further introduced, which is conducive to further improving the accuracy of howling detection and reducing the probability of misjudgment. It should be noted that since the duration of each frame of the target audio signal or the historical audio is relatively short, the above howling detection method actually has a short processing time and a fast detection speed.

[0118] Refer to Figure 8, the howling detection method in the above example may specifically include the following steps 801, 802, 803, 804, 805, 806, and 807:

[0119] Step 801: Obtain the target audio signal played by the terminal speaker and the historical audio signal in the buffer;

[0120] Step 802: Perform frame splitting on the target audio signal and the historical audio signal to obtain the power spectra of the target audio signal and the historical audio signal;

[0121] Step 803: Obtain the correlation between the power spectrum of the current frame in the target audio signal and the power spectra of each frame in the historical audio signal;

[0122] Step 804: Determine whether the correlation between the power spectrum of the current frame in the target audio signal and the power spectrum of a certain frame in the historical audio signal is greater than the threshold. If so, jump to Step 805; otherwise, jump to Step 806;

[0123] Step 805: Determine whether the time difference between the current frame of the target audio signal and that frame in the historical audio signal is within the periodic range. If so, jump to Step 807; otherwise, jump to Step 806;

[0124] Step 806: Cache the current frame of the target audio signal in the buffer, continue to obtain another frame of the target audio signal as the current frame, and jump to Step 803;

[0125] Step 807: Confirm that howling occurs in the terminal.

[0126] Similarly, for the terminal 2002b, howling will also occur for similar reasons, and the terminal 2002b can also perform howling detection in the same way, which will not be elaborated here.

[0127] It can be understood that to detect whether howling occurs in the terminal 2002a, in addition to the terminal 2002a itself, it can also be detected by the terminal 2002b or the server 2001. For example, if detected by the terminal 2002b, the terminal 2002b can obtain the target audio signal played by the speaker of the terminal 2002a through wireless transmission methods such as Bluetooth and cache the historical audio signal; similarly, if detected by the server 2001, the server 2001 can obtain the target audio signal played by the speaker of the terminal 2002a through transmission methods such as the network and cache the historical audio signal.

[0128] As another example, based on Figure 2In the network architecture shown, there is also a situation where the audio signal played by the speaker of terminal 2002a is picked up by the microphone of terminal 2002a itself. Terminal 2002a transmits the audio signal picked up by its microphone to the server through the network. The server then transmits this audio signal to terminal 2001a, and terminal 2001a plays this audio signal again through its own speaker. The microphone of terminal 2002a picks up this audio signal again... and so on. Once the volume of this audio signal reaches a certain limit, howling will also occur, thus affecting the video call between terminal 2002a, terminal 2002b, and the remote terminal. Therefore, the howling detection method of the first example above can be used to determine whether terminal 2002a has howled. The specific detection principle is the same as that of the first example and will not be elaborated here.

[0129] Referring to Figure 9 , another embodiment of the present invention further provides a howling detection device 900, and the howling detection device 900 includes:

[0130] A first audio acquisition unit 901, configured to acquire a target audio signal sent by an audio playback device to an audio receiving device at a first time;

[0131] A first audio acquisition unit 902, configured to acquire a historical audio signal sent by the audio playback device to the audio receiving device at a second time, where the second time is earlier than the first time;

[0132] A power spectrum acquisition unit 903, configured to determine a target power spectrum according to the target audio signal and determine a historical power spectrum according to the historical audio signal;

[0133] A howling determination unit 904, configured to determine that the audio playback device has howled when a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold.

[0134] Among them, the audio playback device is used to play audio signals. Exemplarily, the audio playback device can be a terminal with a speaker or a separate speaker. Similarly, the audio receiving device is used to receive audio signals. The audio receiving device can be a terminal with a microphone or a separate microphone. Among them, the audio playback device and the audio receiving device can be either the same device or different devices. For example, when the audio playback device and the audio receiving device are the same device, the audio playback device and the audio receiving device can be the speaker and the microphone on the same terminal; when the audio playback device and the audio receiving device are different devices, the audio playback device and the audio receiving device can be different terminals. The target audio signal and the historical audio signal are both audio signals sent by the audio playback device. The target audio signal and the historical audio signal can be audio signals sent by the audio playback device at different times. The target audio signal can be the audio signal currently sent by the audio playback device, and the historical audio signal can be the audio signal previously sent by the audio playback device. Of course, the target audio signal can also be the audio signal previously sent by the audio playback device. In this case, the historical audio signal can be the audio signal sent even earlier than the target audio signal.

[0135] Among them, the power spectrum represents the variation of the audio signal power with frequency, that is, the distribution of the signal power in the frequency domain. The target power spectrum is determined based on the target audio signal, and the historical power spectrum is determined based on the historical audio signal, which can be achieved by means such as Fourier transform.

[0136] Since the sending times of the target audio signal and the historical audio signal are different, in order to facilitate obtaining the historical power spectrum of the historical audio signal and realizing the subsequent comparison between the target power spectrum and the historical power spectrum, therefore, after the power spectrum acquisition unit 903 determines the historical power spectrum based on the historical audio signal, it caches the historical power spectrum of the historical audio signal. At this time, when comparing the target power spectrum and the historical power spectrum, the historical power spectrum can be obtained from the buffer. Exemplarily, the data space size of the buffer can be the memory size occupied by the audio signal with a length of several seconds, such as the memory size occupied by the audio signal of 5 seconds. Of course, the data space size can be set according to the actual situation, and the embodiments of the present invention do not make any limitations.

[0137] The first correlation degree is used to characterize the similarity between the target power spectrum and the historical power spectrum. Since the howling has the characteristic of reciprocating transmission and playback, there is a certain similarity between the howling signals in each cycle. Therefore, the howling determination unit 904 determines the first correlation degree between the target power spectrum and the historical power spectrum. Through this first correlation degree, the similarity between the target audio signal and the historical audio signal can be judged. When the first correlation degree between the target power spectrum and the historical power spectrum is relatively high, it can be further determined that the audio playback device has howled.

[0138] The first audio acquisition unit 901 acquires a target audio signal sent by an audio playback device to an audio receiving device at a first time, and the first audio acquisition unit 902 acquires a historical audio signal sent by an audio playback device to the audio receiving device at a second time. The power spectrum acquisition unit 903 determines a target power spectrum based on the target audio signal, and the power spectrum acquisition unit 903 determines a historical power spectrum based on the historical audio signal. When the first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold value, the howling determination unit 904 determines that howling occurs in the audio playback device. Thus, the similarity between the target audio signal and the historical audio signal can be judged through the first correlation degree, and further, whether howling occurs in the audio playback device can be judged, which can improve the accuracy of howling detection. Moreover, judging the similarity between the target audio signal and the historical audio signal through the first correlation degree between the target power spectrum and the historical power spectrum has the advantage of fast convergence speed, which is beneficial to improving the efficiency of howling detection.

[0139] It should be noted that the howling detection device 900 in the embodiments of the present application can be applied to the application scenarios of voice communication, and can also be applied to the human-computer interaction scenarios with speaker playback, such as intelligent devices such as intelligent robots with voice calls, smart speakers, and smart watches. The target audio signal and the historical audio signal can include but are not limited to audio signals such as user voices (including call voices), music, other background sounds, synthetic sounds, and prompt sounds.

[0140] Refer to Figure 10 , in one embodiment, the power spectrum acquisition unit 903 further includes:

[0141] A framing unit 1001 for performing framing processing on the target audio signal and the historical audio signal;

[0142] A first sub-audio signal acquisition unit 1002 for acquiring a first sub-audio signal of the current frame in the target audio signal after framing processing;

[0143] A second sub-audio signal acquisition unit 1003 for acquiring a plurality of second sub-audio signals before or after the current frame in the historical audio signal after framing processing;

[0144] A power spectrum calculation unit 1004 for acquiring a first power spectrum of the first sub-audio signal and a second power spectrum of the second sub-audio signal;

[0145] A power spectrum determination unit 1005 for using the first power spectrum as the target power spectrum of the target audio signal and using the second power spectrum as the historical power spectrum of the historical audio signal.

[0146] Specifically, the first sub-audio signal acquisition unit 1002 can frame and window the target audio signal using a window function, and the second sub-audio signal acquisition unit 1003 can also frame and window the historical audio signal using a window function. To reduce spectral energy leakage, a window function can be used. Exemplarily, one of the following window functions can be used: rectangular window, triangular window, Hanning window, Hamming window, etc. Among them, the rectangular window belongs to the zero-power window of the time variable. The rectangular window is used the most, and conventionally, not windowing means that the signal passes through the rectangular window. The advantage of this window is that the main lobe is relatively concentrated, but the disadvantage is that the side lobes are relatively high and there are negative side lobes, resulting in high-frequency interference and leakage in the transformation, and even negative spectrum phenomena. The triangular window is in the form of the first power of the power window. Compared with the rectangular window, the main lobe width is approximately twice that of the rectangular window, but the side lobes are small and there are no negative side lobes. The Hanning window can be regarded as the sum of the spectra of 3 rectangular time windows. The main lobe of the Hanning window is broadened and reduced, and the side lobes are significantly reduced, which is beneficial to reducing leakage. The Hamming window and the Hanning window are both cosine windows, only the weighting coefficients are different, and the weighting coefficients of the Hamming window can make the side lobes smaller.

[0147] Among them, the current frame can be one of the frames in the target audio signal after frame processing. In the actual detection process, each frame of the audio signal of the target audio signal after frame processing can be sequentially acquired. When the first frame of the target audio signal is acquired, the current frame is the first frame, and the first frame of the target audio signal is the first sub-audio signal. Of course, in the actual detection process, the acquisition order of each frame of the audio signal in the target audio signal can also be set according to the actual situation.

[0148] After the historical audio signal is frame processed, multiple frames of audio signals can be obtained. The multiple second sub-audio signals before or after the current frame in the historical audio signal after frame processing are actually multiple offset frames of the current frame of the target audio signal in the historical time. Among them, obtaining multiple second sub-audio signals before or after the current frame in the historical audio signal after frame processing can be obtaining all the second sub-audio signals after the current frame in multiple frames of historical audio signals, or obtaining some of the second sub-audio signals after the current frame in the historical audio signal. The embodiments of the present invention do not make limitations. Obtaining all the second sub-audio signals after the current frame in multiple frames of historical audio signals is beneficial to improving the accuracy of subsequent comparison.

[0149] Among them, the power spectrum calculation unit 1004 obtaining the first power spectrum of the first sub-audio signal can be performing a Fourier transform on the first sub-audio signal and calculating the corresponding absolute value of the power. The acquisition method of the second sub-audio signal is the same and will not be elaborated here.

[0150] Refer to Figure 11 , in one embodiment, the howling determination unit 904 can further include:

[0151] The first correlation acquisition unit 1101 is configured to acquire the second correlation between the first power spectrum and multiple second power spectra respectively;

[0152] The second correlation determination unit 1102 is configured to determine the first correlation between the target power spectrum and the historical power spectrum according to the multiple second correlations;

[0153] The howling determination unit 1103 is configured to determine that the audio playback device has howled when the first correlation is greater than the threshold.

[0154] Since the target audio signal and the historical audio signal are framed by the framing unit to obtain multiple frames of target audio signals and multiple frames of historical audio signals, and then the first correlation acquisition unit 1101 acquires the second correlation between the first power spectrum and multiple second power spectra respectively, due to the framing processing of the target audio signal and the historical audio signal, the dimension of the audio signal is smaller. Compared with the solution of directly judging the overall similarity of the audio signal, the howling detection method provided by the embodiment of the present invention is beneficial to improving the accuracy when judging the similarity degree.

[0155] Wherein, the second correlation determination unit 1102 determines the first correlation between the target power spectrum and the historical power spectrum according to the multiple second correlations, which may be taking the second correlation corresponding to the second power spectrum with the highest similarity degree to the first power spectrum as the first correlation. For example, it may be taking the largest second correlation as the first correlation. Among them, taking the second correlation corresponding to the second power spectrum with the highest similarity degree to the first power spectrum as the first correlation is beneficial to making the first correlation more reasonably reflect the similarity degree between the first power spectrum and the second power spectrum. Of course, in other embodiments, it may not directly use the magnitude of the second correlation to judge the similarity degree. For example, some mathematical operations may be performed on the second correlation to convert it into another parameter, and then this parameter is used for judgment.

[0156] Wherein, when the first correlation is greater than the threshold, the howling determination unit 1103 determines that the audio playback device has howled, which may be determining that the audio playback device has howled when the maximum value among the multiple second correlations is greater than the threshold. Therefore, that is, when the similarity between the first power spectrum of a certain frame of the target audio signal and the second power spectrum of a certain frame of the historical audio signal is relatively large, the howling determination unit 1103 can determine that the audio playback device has howled.

[0157] It can be understood that in other embodiments, the howling judgment unit 1103 may also compare the threshold value every time it obtains the second correlation degree between the first power spectrum and a second power spectrum. Once the second correlation degree is greater than the threshold value, it is determined that the audio playback device has howled. The advantage of this method is that it is not necessary to obtain all the second correlation degrees between the first power spectrum and multiple second power spectra. Therefore, it is beneficial to improve the convergence speed of the howling detection method and improve the efficiency of howling detection.

[0158] Referring to Figure 12 , in one embodiment, the howling judgment unit 1103 may further include:

[0159] A time difference acquisition unit 1201, configured to obtain the time difference between the first sub-audio signal and the second sub-audio signal with the highest similarity degree according to the second correlation degree when the first correlation degree is greater than the threshold value;

[0160] A time difference judgment unit 1202, configured to determine that the audio playback device has howled when the time difference is within a preset period range.

[0161] In the actual application process, even when the similarity of two audio signals is relatively high, there may be a situation where both of these two audio signals are normal speech signals and there is no howling at this time. Therefore, after determining that the first correlation degree is greater than the threshold value, the time difference acquisition unit 1201 also obtains the time difference between the first sub-audio signal and the second sub-audio signal with the highest similarity degree according to the second correlation degree. Among them, the first sub-audio signal and the second sub-audio signal with the highest similarity degree may be the first sub-audio signal and the second sub-audio signal corresponding to the maximum value of the second correlation degree. Among them, the time difference between the first sub-audio signal and the second sub-audio signal can be regarded as the period of howling occurrence.

[0162] Since the period of howling occurrence is generally very short, in order to exclude the situation of two normal speech signals with relatively high similarity, the time difference judgment unit 1202 compares the time difference between the first sub-audio signal and the second sub-audio signal with a preset period range. When the time difference is within the preset period range, it is determined that the audio playback device has howled, thereby improving the accuracy of howling detection and reducing the occurrence of misjudgment phenomena.

[0163] Referring to Figure 13 , in one embodiment, the second correlation degree determination unit 1102 may further include:

[0164] A smoothing filter unit 1301, configured to perform smoothing filtering processing on multiple second correlation degrees;

[0165] The third correlation determination unit 1302 is configured to determine a first correlation between a target power spectrum and a historical power spectrum according to a plurality of second correlations after being smoothed and filtered.

[0166] After obtaining the second correlation through the first correlation acquisition unit, the second correlation can be further smoothed and filtered, so as to filter out the interference signals generated during the calculation of the second correlation, and improve the accuracy of howling detection. Since the second correlation has been smoothed and filtered by the smoothing filter unit 1301, it is beneficial to improve the accuracy of the first correlation between the target power spectrum and the historical power spectrum finally obtained by the third correlation determination unit 1302.

[0167] It should also be understood that the various embodiments provided in the embodiments of the present invention can be combined arbitrarily to achieve different technical effects.

[0168] Refer to Figure 14 , the embodiment of the present invention further provides a terminal 1400. The terminal includes the howling detection device 900 in the above embodiment. Therefore, the terminal 1400 can obtain the target audio signal sent by the audio playback device to the audio receiving device at the first time through the first audio acquisition unit 901, and the first audio acquisition unit 902 obtains the historical audio signal sent by the audio playback device to the audio receiving device at the second time. The power spectrum acquisition unit 903 determines the target power spectrum according to the target audio signal, and the power spectrum acquisition unit 903 determines the historical power spectrum according to the historical audio signal. When the first correlation between the target power spectrum and the historical power spectrum is greater than the threshold, the howling determination unit 904 determines that the audio playback device has howled. Thus, the similarity between the target audio signal and the historical audio signal can be judged through the first correlation, and then it can be judged whether the audio playback device has howled, which can improve the accuracy of howling detection. Moreover, by judging the similarity between the target audio signal and the historical audio signal through the first correlation between the target power spectrum and the historical power spectrum, it has the advantage of fast convergence speed, which is beneficial to improving the efficiency of howling detection.

[0169] Figure 15 Fig. shows the howling detection device 900 provided by the embodiment of the present invention. The howling detection device 900 includes: a memory 1501, a processor 1502, and a computer program stored on the memory 1501 and executable on the processor 1502. When the computer program runs, it is used to execute the above howling detection method.

[0170] The processor 1502 and the memory 1501 can be connected through a bus or other means.

[0171] The memory 1501, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the howling detection method described in the embodiments of the present invention. The processor 1502 realizes the above-mentioned howling detection method by running the non-transitory software programs and instructions stored in the memory 1501.

[0172] The memory 1501 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store the execution of the above-mentioned howling detection method. In addition, the memory 1501 may include a high-speed random access memory 1501, and may also include a non-transitory memory 1501, such as at least one storage device storage device, a flash memory device or other non-transitory solid-state storage devices. In some embodiments, the memory 1501 may optionally include a memory 1501 remotely provided with respect to the processor 1502, and these remote memories 1501 can be connected to the howling detection device 900 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0173] The non-transitory software programs and instructions required to implement the above-mentioned howling detection method are stored in the memory 1501. When executed by one or more processors 1502, the above-mentioned howling detection method is executed. For example, execute Figure 3 method steps 301 to 304 in Figure 4 method steps 401 to 405 in Figure 5 method steps 501 to 503 in Figure 6 method steps 601 to 602 in Figure 6 method steps 601 to 605 in Figure 7 method steps 701 to 702 in Figure 8 method steps 801 to 807 in

[0174] The embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for executing the above-mentioned howling detection method.

[0175] In one embodiment, the computer-readable storage medium stores computer-executable instructions, which are executed by one or more control processors, for example, executed by a processor 1502 in the above-mentioned howling detection device 900, so that the above-mentioned processor 1502 can execute the above-mentioned howling detection method. For example, execute Figure 3 method steps 301 to 304 in Figure 4 method steps 401 to 405 in Figure 5 method steps 501 to 503 in Figure 6Steps 601 to 602 in Figure 6 Steps 601 to 605 in Figure 7 Steps 701 to 702 in Figure 8 Steps 801 to 807 in

[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] Those of ordinary skill in the art can understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, storage device storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0178] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A howling detection method, comprising: Obtaining a target audio signal sent by an audio playback device to an audio receiving device at a first time; Obtaining a historical audio signal sent by the audio playback device to the audio receiving device at a second time, where the second time is earlier than the first time; Determining a target power spectrum according to the target audio signal, and determining a historical power spectrum according to the historical audio signal, where the target power spectrum and the historical power spectrum are the variation of the corresponding audio signal power with frequency, that is, the distribution of the signal power in the frequency domain; When a first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold, determining that the audio playback device has howled; The first correlation degree is a covariance correlation coefficient, and the first correlation degree is obtained through the following steps: Obtaining a first power absolute value of the target power spectrum and a second power absolute value of the historical power spectrum; Obtaining a first covariance between the first power absolute value and the second power absolute value; Obtaining a first standard deviation of the first power absolute value and a second standard deviation of the second power absolute value; Obtaining a first product between the first standard deviation and the second standard deviation; Taking a first quotient value between the first covariance and the first product as the first correlation degree.

2. The method according to claim 1, wherein After determining the historical power spectrum according to the historical audio signal, further comprising: Caching the historical power spectrum of the historical audio signal.

3. The method according to any one of claims 1 or 2, characterized in that, Determining the target power spectrum according to the target audio signal and determining the historical power spectrum according to the historical audio signal, including: Performing frame segmentation processing on the target audio signal and the historical audio signal; Obtaining a first sub-audio signal of a current frame in the target audio signal after frame segmentation processing; Obtaining a plurality of second sub-audio signals before or after the current frame in the historical audio signal after frame segmentation processing; Obtaining a first power spectrum of the first sub-audio signal and a second power spectrum of the second sub-audio signal; Taking the first power spectrum as the target power spectrum of the target audio signal and taking the second power spectrum as the historical power spectrum of the historical audio signal.

4. The method according to claim 3, wherein When the first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold, determining that the audio playback device has howled, including: Respectively obtaining second correlation degrees between the first power spectrum and a plurality of the second power spectra; Determining the first correlation degree between the target power spectrum and the historical power spectrum according to the plurality of second correlation degrees; When the first correlation degree is greater than a threshold, determining that the audio playback device has howled.

5. The method according to claim 4, wherein Determining the first correlation degree between the target power spectrum and the historical power spectrum according to the plurality of second correlation degrees, including: Taking the second correlation degree corresponding to a third power spectrum as the first correlation degree, where the third power spectrum is the second power spectrum with the highest similarity degree to the first power spectrum.

6. The method according to claim 4, wherein When the first correlation degree is greater than a threshold, determining that the audio playback device has howled, including: When the first correlation degree is greater than a threshold value, obtain the time difference between the first sub-audio signal and the second sub-audio signal with the highest similarity degree according to the second correlation degree; When the time difference is within a preset period range, determine that the audio playback device has a howling sound.

7. The method according to claim 4, wherein The second correlation degree is a covariance correlation coefficient, and the second correlation degree is obtained through the following steps: Obtain the absolute value of the third power of each frequency point in the i-th frame of the first sub-audio signal and the absolute value of the fourth power of each frequency point in the j-th frame before or after the i-th frame of the second sub-audio signal; Obtain the second covariance between the absolute value of the third power and the absolute value of the fourth power; Obtain the third standard deviation of the absolute value of the third power and the fourth standard deviation of the absolute value of the fourth power; Obtain the second product between the third standard deviation and the fourth standard deviation; Use the second quotient between the second covariance and the second product as the second correlation degree; Among them, , are all positive integers.

8. The method according to claim 3, characterized in that The obtaining the first power spectrum of the first sub-audio signal and the second power spectrum of the second sub-audio signal includes: Obtain the first power spectrum of the first sub-audio signal and the second power spectrum of the second sub-audio signal according to a preset frequency range.

9. The method according to claim 4, characterized in that, The determining the first correlation degree between the target power spectrum and the historical power spectrum according to a plurality of the second correlation degrees includes: Perform a smoothing filtering process on a plurality of the second correlation degrees, where the calculation formula of the smoothing filtering process includes: , wherein, is the second correlation degree between the second sub-audio signal of the j-th frame before or after the current frame and the first sub-audio signal in the historical audio signal after frame processing, is the second correlation degree after smoothing filtering processing, is a positive integer, is a positive number less than 1; Determine the first correlation degree between the target power spectrum and the historical power spectrum according to a plurality of the second correlation degrees after the smoothing filtering process.

10. The method according to claim 1, characterized in that, The target audio signal is the currently played audio signal or the recorded audio signal that has been played.

11. A howling detection device, characterized in that, Includes: A first audio acquisition unit, configured to acquire a target audio signal sent by an audio playback device to an audio receiving device at a first time; A second audio acquisition unit, configured to acquire a historical audio signal sent by the audio playback device to the audio receiving device at a second time, where the second time is earlier than the first time; A power spectrum acquisition unit, configured to determine a target power spectrum according to the target audio signal and determine a historical power spectrum according to the historical audio signal, where the target power spectrum and the historical power spectrum are the variation of the corresponding audio signal power with frequency, that is, the distribution of the signal power in the frequency domain; A howling determination unit, configured to determine that the audio playback device has a howling sound when the first correlation degree between the target power spectrum and the historical power spectrum is greater than a threshold value; the first correlation degree is a covariance correlation coefficient, and the first correlation degree is obtained through the following steps: obtain the absolute value of the first power of the target power spectrum and the absolute value of the second power of the historical power spectrum; obtain the first covariance between the absolute value of the first power and the absolute value of the second power; obtain the first standard deviation of the absolute value of the first power and the second standard deviation of the absolute value of the second power; obtain the first product between the first standard deviation and the second standard deviation; use the first quotient between the first covariance and the first product as the first correlation degree.

12. A howling detection device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the howling detection method described in any one of claims 1 to 10.

13. A terminal, characterized in that, It includes the howling detection device described in claim 11 or the howling detection device described in claim 12.

14. A computer-readable storage medium, characterized in that, The storage medium stores a program, and when the program is executed by a processor, it implements the howling detection method described in any one of claims 1-10.

Citation Information

Patent Citations

  • Method and device for removing headset scream

    CN103338419A