A method and apparatus for detecting howling in a communication system

By combining peak power spectrum stability measure, peak-to-average amplitude ratio stability measure and long-time frame amplitude squared coherence coefficient, the howling detection method solves the problem of high false detection rate or low detection rate of howling detection in communication systems, and achieves efficient howling suppression and system stability improvement.

CN116189700BActive Publication Date: 2026-04-14G NET INTEGRATED SERVICE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
G NET INTEGRATED SERVICE
Filing Date
2023-02-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing feedback detection technologies have high false detection rates or low detection rates in communication systems, leading to improper acoustic feedback control and affecting system stability and sound quality.

Method used

A howling detection method is adopted, which combines peak power spectrum stability measure and peak-to-average amplitude ratio stability measure with long-time frame amplitude squared coherence coefficient. Short-time spectrum is generated by short-time Fourier transform, howling frequency index is screened, and hierarchical decision rules are used to improve detection accuracy.

Benefits of technology

While reducing the probability of false positives, it improves the feedback detection rate, effectively suppresses feedback, and enhances system stability and sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189700B_ABST
    Figure CN116189700B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of communication system howling detection method and device, it is related to howling detection technique, the method is based on peak power spectrum stability measure detects whether there is first type howling frequency point index in the short time spectrum to be detected at t frame, if not, it is based on peak-average amplitude ratio stability measure and peak-harmonic power ratio detects whether there is second type howling frequency point index in the short time spectrum to be detected at t frame, first type howling frequency point index is the local peak frequency point index of peak power spectrum stability measure greater than preset peak power spectrum stability measure threshold, second type howling frequency point index is the local peak frequency point index corresponding to peak-average amplitude stability measure greater than preset peak-average amplitude stability measure threshold and peak-harmonic power ratio greater than preset peak-harmonic power ratio threshold, the scheme is combined with judgment using two kinds of howling frequency point index, effectively improve howling detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to howling detection technology, specifically to a method and apparatus for howling detection in a communication system. Background Technology

[0002] Both sound reinforcement systems and hands-free communication systems suffer from acoustic feedback problems, for example in... Figure 1 In the sound reinforcement system shown, when an audio signal is captured by a microphone, amplified, and played through a loudspeaker, the loudspeaker sound is typically fed back to the microphone via direct acoustic coupling or indirectly due to reverberation. This acoustic feedback path forms a closed-loop signal circuit. Figure 2 In the hands-free communication system shown, when the near-end user terminal A speaks, the speaker of the nearby user terminal B plays the voice signal received from user terminal A, which is also fed back to the microphone of user terminal A, forming a closed-loop signal loop 1 through the communication network. When the speaker of the far-end user terminal C plays the voice signal received from user terminal A, it is also fed back to the microphone of the nearby user terminal D, forming a closed-loop signal loop 2 through the communication network.

[0003] The existence of this closed signal loop prevents the system from operating reliably and stably, and causes a severely disturbing howling phenomenon. Therefore, it is necessary to first detect the howling condition promptly, and then send a detection indication signal to the system so that the system can perform relevant subsequent control processing.

[0004] A review of numerous domestic and international literatures reveals that current feedback detection technologies largely rely on extracting relevant time-domain and frequency-domain features from microphone-received signals. The principle is as follows: the time-domain (digital) signal x(n) received by the microphone is transformed into the short-time spectral signal X(k,t) to be detected at frame t using the Short-Time Fourier Transform (STFT) technique.

[0005] (1),

[0006] Where k = 0, 1, 2, …, N-1, t = 0, 1, 2, …, and n = 0, 1, 2, …, N-1 are the frequency index, signal frame index, and sample index, respectively; w(·) is a window function with N sample lengths, typically a Hamming, Hanning, or Blackman window function; x(n,t) is the nth sample in the t-th frame signal, i.e. Here, L is the number of samples for frame shift jumps.

[0007] For the short-time spectrum X(k,t) to be detected in frame t, the Peak Picking Algorithm (PPA) is applied to select P maximum peak frequency indices. As a set of alternative whistling frequency indexes ; For sets Each element in the algorithm calculates a corresponding feature parameter. If the value of this feature parameter exceeds a preset threshold, the element corresponding to that feature parameter is determined to be a howling frequency index. Toon van Waterschoot and Marc Moonen, in their paper "Comparative evaluation of howling detection criteria in notch-filter-based howling suppression" (J. Audio Eng. Soc., Vol.58, No. 11, November, 2010, pp. 923 - 940), provided a detailed review of the signal feature parameters used for howling detection. These feature parameters and the corresponding howling detection criteria are as follows:

[0008] 1) Peak-to-Threshold Power Ratio (PTPR): This feature is a frequency domain feature, defined as the index of candidate howling frequency points. The ratio of the spectral power at a given point to the fixed absolute power threshold P0, i.e.:

[0009] (2),

[0010] like If a howling sound occurs, then it is determined that a howling sound has occurred. Index of howling frequencies; here This is the set of frequency indexes for howling (the same applies below, without further explanation). This is the decision threshold for this feature parameter.

[0011] 2) Peak-to-Average Power Ratio (PAPR): This feature is a frequency domain characteristic, defined as the index of candidate howling frequency points. The ratio of the sample power spectrum at a given location to the average power of the microphone received signal, i.e.:

[0012] (3),

[0013] in, (4),

[0014] like If a howling sound occurs, then it is determined that a howling sound has occurred. Index of howling frequencies; here This is the decision threshold for this feature parameter.

[0015] 3) Peak-to-Harmonic Power Ratio (PHPR): This feature is a frequency domain characteristic, defined as the index of candidate howling frequency points. The ratio of the sample power spectrum at a given location to the power of its m-th harmonic component is:

[0016] (5),

[0017] Where m is usually chosen as , , , and Given that the spectral structure of a howl differs from that of a speech or audio signal with harmonic components, this characteristic is used to distinguish howls from signal components.

[0018] For the selected ,like If so, then a howling sound is determined to have occurred, and Index of howling frequencies; here This is the decision threshold for the feature parameter. This is the logical AND operator (the same applies below, without further explanation).

[0019] 4) Peak-to-Neighboring Power Ratio (PNPR): This feature is a frequency domain characteristic, defined as the index of candidate howling frequency points. The ratio of the sample power spectrum at a given frequency point to the power at its m-th nearest neighbor frequency point is:

[0020] (6),

[0021] Where m is usually chosen as , and This feature takes advantage of the fact that howling signals are typically very narrow-band.

[0022] For the selected ,like If so, then a howling sound is determined to have occurred, and Index of howling frequencies; here This is the decision threshold for this feature parameter.

[0023] 5) Inter-frame Peak Magnitude Persistence (IPMP): This feature is a time-domain feature used to calculate the index of candidate howling frequency points in past Q frames. The percentage of times appears, i.e.:

[0024] (7),

[0025] in, (8),

[0026] This feature is based on the idea that howling frequency indexes typically last longer than the frequency indexes of speech or tone components. If so, then a howling sound is determined to have occurred, and Index of howling frequencies; here This is the decision threshold for this feature parameter.

[0027] 6) Interframe Magnitude Slope Deviation (IMSD): This is a time-domain feature, defined by averaging the amplitude differences of the spectral components corresponding to candidate whistling frequency indices across Q consecutive signal frames. The difference is calculated between older and newer signal frames.

[0028] (9),

[0029] Because the spectral amplitude dB scale corresponding to the howling frequency index increases almost linearly over (frame) time, its corresponding IMSD characteristic value is relatively small, which is an important characteristic of howling components. If so, then a howling sound is determined to have occurred, and Index of howling frequencies; here This is the decision threshold for this feature parameter.

[0030] Using only a single feature for howling detection results in a high false positive rate. Therefore, an intuitive approach is to directly combine multiple of the aforementioned signal feature parameters into a detection decision criterion to achieve better howling detection performance. MP Oster et al. proposed a howling detection decision criterion based on PHPR and IPMP features as follows:

[0031] like (10)

[0032] Then it is determined that a howling sound has occurred, and Index for howling frequency points.

[0033] N. Osmanovic et al. proposed a feature called Feedback Existence Probability (FEP) based on PNPR and IMSD features, and gave a corresponding criterion for howling detection. FEP is defined as:

[0034] (11),

[0035] in, (12)

[0036] (13)

[0037] Therefore, the decision criterion for howling detection based on FEP is:

[0038] like (14) Then it is determined that a howling has occurred, and Index for howling frequency points.

[0039] Toon van Waterschoot and Marc Moonen proposed the following four multi-feature howling detection criteria based on the three signal features PHPR, PNPR, and IMSD: (15)-(18)

[0040] like (15)

[0041] Then it is determined that a howling sound has occurred, and Index for howling frequency points;

[0042] like (16) then it is determined that a howling has occurred, and Index for howling frequency points;

[0043] like (17) then it is determined that a howling has occurred, and Index for howling frequency points;

[0044] like (18)

[0045] Then it is determined that a howling sound has occurred, and Index for howling frequency points.

[0046] The aforementioned feedback detection criteria based on single-signal features typically have a high detection probability, but also a high false detection probability. Feedback detection criteria based on multiple signal features typically have a lower false detection probability, but their overall detection probability is significantly lower than that of single-feature feedback detection criteria. It's important to note that a high false detection probability can cause the system to mistakenly activate subsequent acoustic feedback control processing, leading to a deterioration in signal quality. Conversely, a low detection probability can prevent the system from activating subsequent acoustic feedback control processing, also resulting in deteriorated signal quality due to feedback, or even system malfunction. Summary of the Invention

[0047] Compared with existing technologies, a new sound signal parameter feature is used to detect howling in communication systems, thereby providing a method and device for howling detection in communication systems.

[0048] To address the aforementioned technical problems, the present invention discloses at least one method and apparatus for detecting howling in a communication system.

[0049] In a first aspect, embodiments of the present invention disclose a method for detecting howling in a communication system, comprising:

[0050] Acquire the sound signal to be detected;

[0051] Generate the short-time spectrum of the sound signal to be detected at frame t;

[0052] The presence of a first type of howling frequency index in the short-time spectrum of the target signal at frame t is detected based on the peak power spectrum stability measure. The peak power spectrum stability measure is formed by weighting a specified short-time spectrum amplitude square with a specified weighting coefficient and then taking a decibel scale. The specified weighting coefficient is formed with Euler's number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the target sound signal as the exponent. The specified short-time spectrum amplitude square is the short-time spectrum amplitude square of the target sound signal at the local peak frequency index. The first type of howling frequency index is a local peak frequency index where the peak power spectrum stability measure is greater than a preset peak power spectrum stability measure threshold.

[0053] If the first type of howling frequency index does not exist in the short-time spectrum to be detected at frame t, then the presence of the second type of howling frequency index is detected in the short-time spectrum to be detected at frame t is determined based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio. The peak-to-average amplitude ratio stability measure is formed by weighting the specified short-time spectrum amplitude ratio with the specified weighting coefficient and taking a decibel scale. The specified short-time spectrum amplitude ratio is the ratio of the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index to the average value of the short-time spectrum amplitude across the entire frequency band. The peak-to-harmonic power ratio is formed by taking the square of a specified short-time spectral amplitude in decibels. The specified short-time spectral amplitude square ratio is the ratio of the square of the short-time spectral amplitude of the sound signal to be detected at a local peak frequency index to the square of the short-time spectral amplitude at the corresponding harmonic frequency index of the local peak frequency. The second type of howling frequency index is a local peak frequency index corresponding to a peak-to-average amplitude stability measure that is greater than a preset peak-to-average amplitude stability measure threshold and a peak-to-harmonic power ratio that is greater than a preset peak-to-harmonic power ratio threshold.

[0054] Optionally, before detecting whether a first type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure, the method further includes: calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t; obtaining a preset number of candidate howling frequency indexes from the short-time spectrum to be detected at frame t based on the long-time frame amplitude squared coherence coefficient, wherein the candidate howling frequency indexes are frequency indexes corresponding to long-time frame amplitude squared coherence coefficients at frame t that are greater than a preset long-time frame amplitude squared coherence coefficient threshold parameter; the method further includes: detecting whether a first type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure; and obtaining a preset number of candidate howling frequency indexes from the short-time spectrum to be detected at frame t based on the long-time frame amplitude squared coherence coefficients, wherein the candidate howling frequency indexes are frequency indexes corresponding to long-time frame amplitude squared coherence coefficients at frame t that are greater than a preset long-time frame amplitude squared coherence coefficient threshold parameter; the method further includes: calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure; the method further includes: calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t based on the long-time frame amplitude squared coherence coefficients threshold parameter; the method further includes: calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected based on the peak power spectrum stability measure; the method further includes: calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected based on the long-time frame amplitude squared coherence coefficients threshold parameter; the method further includes: calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected based on the peak power spectrum stability measure; the method further includes: calculating the long-time frame amplitude squared coherence coefficient The presence of a first type of howling frequency index in the short-time spectrum to be detected at frame t is determined by: detecting whether a first type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure and the candidate howling frequency index; the presence of a second type of howling frequency index in the short-time spectrum to be detected at frame t is determined by: detecting whether a second type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio.

[0055] Optionally, obtaining a preset number of candidate howling frequency point indices from the short-time spectrum to be detected at frame t based on the squared coherence coefficient of the long-time frame amplitude includes: obtaining all candidate howling frequency point indices from the short-time spectrum to be detected at frame t based on the squared coherence coefficient of the long-time frame amplitude.

[0056] Obtain a preset number of candidate howling frequency indexes from all the candidate howling frequency indexes, based on the largest long-time frame amplitude squared coherence coefficient.

[0057] Optionally, before calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t, the method further includes: performing Cherrett-Berucic kernel smoothing on the short-time spectrum to be detected at frame t.

[0058] Optionally, the step of obtaining a preset number of candidate howling frequency indexes from the short-time spectrum to be detected at frame t based on the long-time frame amplitude squared coherence coefficient further includes: obtaining the candidate short-time spectrum corresponding to the candidate howling frequency index from the short-time spectrum to be detected at frame t; the step of detecting whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure is as follows: if the candidate howling frequency index exists in the short-time spectrum to be detected at frame t, then the first type of howling frequency index is detected in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure; if the candidate howling frequency index does not exist in the short-time spectrum to be detected at frame t, then the next frame short-time spectrum in the short-time spectrum to be detected at frame t is obtained, until all short-time spectra to be detected at frame t are detected.

[0059] Optionally, the index for whether a type-1 howling frequency point exists in the short-time spectrum to be detected at frame t based on the peak power spectrum stability metric is: using the formula This paper implements a method to detect the existence of a type I howling frequency index in the short-time spectrum of the target device at frame t based on the peak power spectral stability metric. This is the binary detection indication signal output by the first howling detector in frame t; "V" is the logical "OR" operator; (Unit: signal frame) represents the preset decision threshold parameter for the first howling detector; Index of candidate howling frequency points in the first howling detector The counter value at frame t is determined by the index of the candidate howling frequency point at frame t. When the peak power spectrum stability measure at a certain point is greater than a preset peak power spectrum stability measure threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

[0060] Optionally, the index for whether a second type of howling frequency point exists in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio is: using the formula This method implements a mechanism to detect the existence of a second type of howling frequency index in the short-time spectrum of the target device at frame t, based on the peak-to-average amplitude ratio stability metric and the peak-to-harmonic power ratio. This is the binary detection indication signal output by the second howling detector at frame t; (Unit: signal frame) represents the preset decision threshold parameter for the second howling detector; Index of candidate howling frequency points in the second howling detector The counter value at frame t. The determination method is to use the candidate howling frequency index in frame t. When the peak-to-average amplitude ratio stability measure is greater than a preset peak-to-average amplitude ratio stability measure threshold and the peak-to-harmonic power ratio is greater than a preset peak-to-harmonic power ratio threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

[0061] Optionally, the method further includes: determining the detection indication signal of the howling frequency index of the short-time spectrum to be detected at frame t using the following hierarchical final decision expression, wherein the hierarchical final decision expression is:

[0062] ,

[0063] in, The detection indication signal is the howling frequency index of the short-time spectrum to be detected at frame t. The set of candidate howling frequency points for frame t; This is a detection indication signal indicating whether the first type of howling frequency point index exists in the short-time spectrum to be detected at frame t; This is a detection indication signal for whether the second type of howling frequency index exists in the short-time spectrum to be detected at frame t; the output is the detection indication signal of the howling frequency index of the short-time spectrum to be detected at frame t, where 1 indicates that the detection result is true and 0 indicates that the detection result is false.

[0064] Secondly, embodiments of the present invention disclose a communication system howling detection device, comprising:

[0065] The sound signal acquisition module is used to acquire the sound signal to be detected;

[0066] The short-time spectrum generation module is used to generate the short-time spectrum of the sound signal to be detected at frame t.

[0067] The first howling detector is used to detect whether a first type of howling frequency index exists in the short-time spectrum of the target signal at frame t based on the peak power spectrum stability measure. The peak power spectrum stability measure is formed by weighting a specified short-time spectrum amplitude square with a specified weighting coefficient and taking a decibel scale. The specified weighting coefficient is formed with Euler number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the target sound signal as the exponent. The specified short-time spectrum amplitude square is the short-time spectrum amplitude square of the target sound signal at the local peak frequency index. The first type of howling frequency index is a local peak frequency index where the peak power spectrum stability measure is greater than a preset peak power spectrum stability measure threshold.

[0068] The technical solutions provided by the embodiments of the present invention can have the following beneficial effects:

[0069] This paper presents a novel method for detecting howling frequency points in a short-time spectrum based on a new peak power spectrum stability metric of sound signals. Compared to existing technologies, this method offers a novel approach. Furthermore, it employs another new sound feature, the peak-to-average amplitude ratio stability metric, and an existing sound feature, the peak-to-harmonic power ratio, to detect howling frequency points in a short-time spectrum. Finally, it superimposes a long-time frame amplitude squared coherence coefficient to distinguish howling signals from normal speech signals. This approach achieves a high detection probability while reducing false detection probability, effectively improving howling suppression. Attached Figure Description

[0070] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0071] Figure 1 A schematic diagram of the whistling formation process is shown;

[0072] Figure 2 A schematic diagram of the howling formation process in a hands-free communication system is shown.

[0073] Figure 3 A flowchart of an acoustic feedback processing method in a voice communication system provided by an embodiment of the present invention;

[0074] Figure 4 A flowchart of another acoustic feedback processing method in a voice communication system provided by an embodiment of the present invention is shown;

[0075] Figure 5 A schematic diagram of a whistling formation process in an embodiment of the present invention is shown;

[0076] Figure 6 This diagram illustrates an acoustic feedback processing procedure in another voice communication system provided by an embodiment of the present invention.

[0077] Figure 7 A schematic diagram of the acoustic feedback processing device in a voice communication system provided by an embodiment of the present invention is shown. Detailed Implementation

[0078] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present invention as detailed in the summary of the invention. Example 1

[0079] like Figure 3 The flowchart shown is a communication system howling detection method provided in an embodiment of the present invention. The method includes:

[0080] S31: Acquire the sound signal to be detected;

[0081] S32: Generate the short-time spectrum of the sound signal to be detected at frame t;

[0082] S33: Detect whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure. If there is no first type of howling frequency index in the short-time spectrum to be detected at frame t, then execute S34.

[0083] Among them, the peak power spectrum stability measure is formed by weighting the square of a specified short-time spectrum amplitude with a specified weighting coefficient and taking a decibel scale. The specified weighting coefficient is formed with Euler number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the sound signal to be detected as the exponent. The specified short-time spectrum amplitude square is the square of the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index. The first type of howling frequency index is the local peak frequency index where the peak power spectrum stability measure is greater than the preset peak power spectrum stability measure threshold.

[0084] S34: Detect whether there is a second type of howling frequency index in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio.

[0085] Among them, the peak-to-average amplitude ratio stability measure is formed by weighting a specified short-time spectral amplitude ratio with a specified weighting coefficient and then taking a decibel scale. The specified short-time spectral amplitude ratio is the ratio of the short-time spectral amplitude of the sound signal to be detected at the local peak frequency index to the average short-time spectral amplitude across the entire frequency band. The peak-to-harmonic power ratio is formed by taking a decibel scale for the specified short-time spectral amplitude square ratio. The specified short-time spectral amplitude square ratio is the ratio of the square of the short-time spectral amplitude of the sound signal to be detected at the local peak frequency index to the square of the short-time spectral amplitude at the corresponding harmonic frequency index of the local peak frequency. The second type of howling frequency index is the local peak frequency index corresponding to the peak-to-average amplitude stability measure being greater than a preset peak-to-average amplitude stability measure threshold and the peak-to-harmonic power ratio being greater than a preset peak-to-harmonic power ratio threshold.

[0086] It is understood that the technical solution provided in this embodiment, based on a new sound signal parameter feature peak power spectrum stability measure, detects the presence of a first type of howling frequency index in the short-time spectrum to be detected. Compared with existing technologies, this provides a new howling detection method. Furthermore, it employs another new sound feature, the peak-to-average amplitude ratio stability measure, and the existing acoustic feature, the peak-to-harmonic power ratio, to detect the presence of a second type of howling frequency index in the short-time spectrum to be detected. It further superimposes a long-time frame amplitude squared coherence coefficient to distinguish howling signals from normal speech signals. This achieves a high detection probability while reducing the false detection probability, effectively improving howling suppression. Example 2

[0087] like Figure 4 The flowchart shown is another communication system howling detection method provided by an embodiment of the present invention. The method includes:

[0088] S401: Acquire the sound signal to be detected.

[0089] S402: Generate the short-time spectrum of the sound signal to be detected at frame t.

[0090] S403: Perform Cherit-Belukrani (CB) kernel smoothing on the short-time spectrum to be detected at frame t.

[0091] This invention employs a signal spectrum long-time frame amplitude squared coherence coefficient (MSC) estimation method based on short-time spectrum CB kernel smoothing processing. This method overcomes the shortcomings of existing Welch average periodogram algorithm and MVDR algorithm for MSC estimation, further improving the estimation performance of long-time frame MSC and ensuring the reliability of screening candidate howling frequency point sets.

[0092] S404: Calculate the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t.

[0093] S405: Based on the long-time frame amplitude squared coherence coefficient, obtain a preset number of candidate howling frequency indexes from the short-time spectrum to be detected at frame t. The candidate howling frequency indexes are the frequency indexes corresponding to the long-time frame amplitude squared coherence coefficient at frame t being greater than the preset long-time frame amplitude squared coherence coefficient threshold parameter.

[0094] S406: Does the short-time spectrum to be detected at frame t contain a candidate howling frequency index? If the short-time spectrum to be detected at frame t contains a candidate howling frequency index, then execute S407. If the short-time spectrum to be detected at frame t does not contain a candidate howling frequency index, then execute S409.

[0095] S407: Based on the peak power spectrum stability measure and the alternative howling frequency index, detect whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t. If there is no first type of howling frequency index in the short-time spectrum to be detected at frame t, then execute S408. If there is a first type of howling frequency index in the short-time spectrum to be detected at frame t, then execute S409.

[0096] Among them, the peak power spectrum stability measure is formed by weighting the square of a specified short-time spectrum amplitude with a specified weighting coefficient and taking a decibel scale. The specified weighting coefficient is formed with Euler number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the sound signal to be detected as the exponent. The specified short-time spectrum amplitude square is the square of the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index. The first type of howling frequency index is the local peak frequency index where the peak power spectrum stability measure is greater than the preset peak power spectrum stability measure threshold.

[0097] Specifically, the formula can be used:

[0098] This paper implements a method to detect the existence of a type I howling frequency index in the short-time spectrum of the target device at frame t based on the peak power spectral stability metric. This is the binary detection indication signal output by the first howling detector in frame t; "V" is the logical "OR" operator; (Unit: signal frame) represents the preset decision threshold parameter for the first howling detector; Index of candidate howling frequency points in the first howling detector The counter value at frame t is determined by the index of the candidate howling frequency point at frame t. When the peak power spectrum stability measure at a certain point is greater than the preset peak power spectrum stability measure threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

[0099] S408: Based on the peak-to-average amplitude ratio stability measure, peak-to-harmonic power ratio, and candidate howling frequency index, the second type of howling frequency index in the short-time spectrum to be detected at frame t is obtained. The peak-to-average amplitude ratio stability measure is formed by weighting a specified short-time spectrum amplitude ratio with a specified weighting coefficient and then taking a decibel scale. The specified short-time spectrum amplitude ratio is the ratio of the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index to the average value of the short-time spectrum amplitude across the entire frequency band. The peak-to-harmonic power ratio is formed by taking a decibel scale of the square ratio of the specified short-time spectrum amplitude. The specified short-time spectrum amplitude square ratio is the ratio of the square of the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index to the square of the short-time spectrum amplitude at the corresponding harmonic frequency index of the local peak frequency. The second type of howling frequency index is the local peak frequency index corresponding to the peak-to-average amplitude stability measure being greater than a preset peak-to-average amplitude stability measure threshold and the peak-to-harmonic power ratio being greater than a preset peak-to-harmonic power ratio threshold.

[0100] Specifically, using the formula: This method implements a mechanism to detect the existence of a second type of howling frequency index in the short-time spectrum of the target device at frame t, based on the peak-to-average amplitude ratio stability metric and the peak-to-harmonic power ratio. This is the binary detection indication signal output by the second howling detector at frame t; (Unit: signal frame) represents the preset decision threshold parameter for the second howling detector; Index of candidate howling frequency points in the second howling detector The counter value at frame t is determined by the index of the candidate howling frequency point at frame t. When the peak-to-average amplitude ratio stability measure is greater than the preset peak-to-average amplitude ratio stability measure threshold and the peak-to-harmonic power ratio is greater than the preset peak-to-harmonic power ratio threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

[0101] S409: The detection indication signal of the howling frequency index of the short-time spectrum to be detected at frame t is determined by the following hierarchical final decision expression:

[0102] ;

[0103] in, The detection indication signal is the howling frequency index of the short-time spectrum to be detected at frame t. The set of candidate howling frequency points for frame t; This is a detection indication signal indicating whether the first type of howling frequency point index exists in the short-time spectrum to be detected at frame t; This is a detection indication signal for whether the second type of howling frequency index exists in the short-time spectrum to be detected at frame t. 1 indicates that the detection result is true and 0 indicates that the detection result is false.

[0104] This invention employs a hierarchical howling decision rule: First, the estimated signal spectrum long-time frame (MSC) is used to filter a set of candidate howling frequency points. If this set is empty, the algorithm directly proceeds to the next signal frame. Otherwise, the algorithm uses the CB kernel smoothed spectrum selected from the candidate howling frequency points to enter the decision processing of the first howling detector. If the decision result of the first howling detector is "true," its detection result is output and the algorithm proceeds to the next signal frame. Otherwise, the algorithm proceeds to the decision processing of the second howling detector, outputs its detection result, and proceeds to the next signal frame. Applying this hierarchical howling decision rule can further improve the howling detection probability, reduce the false detection probability, and decrease the computational complexity of the detection algorithm.

[0105] This invention proposes two novel howling detectors: a first howling detector and a second howling detector. The first howling detector uses the Peak Power Spectrum Stability (PPS) feature for each candidate howling frequency. If any frequency satisfies the howling detection condition of this feature (i.e., the PPS value exceeds the decision threshold), the frame signal is determined to contain a howling component. This operation, equivalent to a logical "OR" operation of the detection condition, increases the detection probability while reducing the decision complexity. Furthermore, since the first howling detector uses the PPS feature and is used to detect "saturated howling" periods, its false detection probability is also very low. The second howling detector uses both the Peak-to-Average Amplitude Ratio Stability (PAMRS) and Peak-to-Harmonic Power Ratio (PHPR) features for howling detection decision for each candidate howling frequency. This operation, equivalent to a logical "AND" operation of the detection condition, further reduces the false detection probability. In addition, the logical "OR" operation of the combined decision result of the two features for each candidate howling frequency increases the detection probability of the second howling detector while relatively reducing its decision complexity.

[0106] S410: Outputs a detection indication signal for the howling frequency index of the short-time spectrum to be detected at frame t.

[0107] Subsequently, the next frame's short-time spectrum from the short-time spectrum to be detected at frame t is acquired, until all short-time spectra to be detected at frame t are detected.

[0108] In some alternative embodiments, S405 above includes:

[0109] S405-1. Based on the squared coherence coefficient of the long-time frame amplitude, obtain all candidate howling frequency indexes from the short-time spectrum to be detected at frame t.

[0110] S405-2. Obtain a preset number of candidate howling frequency indexes with the largest long-time frame amplitude squared coherence coefficient from all candidate howling frequency indexes.

[0111] S405-3. Obtain the candidate short-time spectrum corresponding to the candidate howling frequency index from the short-time spectrum to be detected at frame t.

[0112] To facilitate understanding, the technical principles and specific implementation methods involved in the embodiments of the present invention will be described in detail below.

[0113] The short-time spectrum of a speech signal has a very small long-time coherence coefficient on the frame time axis, while the spectrum of a howling signal has a very large long-time coherence coefficient on the frame time axis. Therefore, the long-time coherence coefficient of the speech signal spectrum can be used as a feature to distinguish howling signals from normal speech signals (i.e., speech signals without howling components). However, audio signals from instruments such as flutes and suonas also have high long-time coherence coefficients on the frame time axis. Therefore, this application uses the Peak Power Spectrum Stability Measure (PPS), Peak-to-Average Amplitude Ratio Stability Measure (PAMRS), and Peak-to-Harmonic Power Ratio (PHPR) signal features to further distinguish normal audio signals from howling signals, thereby completing the howling detection task. The system block diagram of this howling detection method proposed in this invention is as follows: Figure 5 As shown, the input signal x(n) undergoes frame segmentation and then short-time Fourier transform (STFT) to obtain a short-time spectral signal X(k,t), where n and t are the sample index and frame index of the input signal, respectively; k is the frequency index of the short-time spectrum, k = 0, 1, 2, …, N-1, and N is the window function length in the STFT operation, in units of the number of samples (the same applies below, without further explanation). This short-time spectral signal X(k,t) is sent to the "CB kernel smoothing processor" for smoothing to further reduce crosstalk caused by spectral leakage and improve the resolution of spectral peaks. The "long-time frame amplitude squared coherence coefficient calculator" calculates the coherence coefficient based on the output of the "CB kernel smoothing processor". To calculate the long-time frame amplitude squared coherence coefficient of the spectral signal. This is used to distinguish between normal speech and feedback components. The "alternate feedback frequency selector" is used to select P of the largest (exceeding the preset threshold parameter) long-time frame amplitude squared coherence coefficients. Corresponding frequency index set As an alternative set of whistling frequency indexes Here, it is recommended that P be set to 1-5. The "Alternate Howling Frequency Point Spectrum Selector" is based on the alternative howling frequency point index set. From a smooth spectrum Select P corresponding values The signal is then sent to the subsequent howling detector. Note that the howling process can be decomposed into two periods: "pre-saturation howling" and "saturation howling." During the "saturation howling" period, the howling frequency time trajectory is relatively stable, and the power spectrum amplitude is very high. Therefore, the "first howling detector based on peak power spectrum stability measurement characteristics" can easily detect the "saturation howling" period. During the "pre-saturation howling" period, the "second howling detector based on peak-to-average amplitude ratio stability measurement and peak-to-harmonic power ratio characteristics" can easily detect it. It should be noted that the activation and invocation of the first and second howling detectors are controlled by the "hierarchical howling decision rule" module, as detailed in the subsequent discussion. Finally, by applying the hierarchical howling decision rule proposed in this invention, the final detection result is obtained. It should be noted that selecting candidate howling point sets based on the long-time frame amplitude squared coherence coefficient and performing howling detection based on a logical "AND" operation using multiple features in the candidate howling point sets both further reduce the probability of false howling detection. Furthermore, the application of hierarchical howling decision rules (equivalent to a logical "OR" operation of each detection result) further improves and increases the howling detection probability while reducing the computational complexity of the howling detection algorithm. Therefore, compared with existing howling detection technologies, the howling detector proposed in this invention has a higher detection probability, a lower false detection probability, and lower computational complexity.

[0114] Now Figure 5 The working principles of the main modules in the software are briefly described below:

[0115] I. Working principle of the long-term (frame) amplitude squared coherence coefficient calculator and CB core smoothing processor module:

[0116] The temporal coherence function (TCF) is a fundamental physical quantity first defined in optics in the early 20th century, used to measure the correlation between a light wave and its delayed version. In fact, the TCF is a special manifestation of the generalized self-spectral coherence function of a signal, measuring the time-varying coherence of two spectra of a random process. Noting the fact that howling components have relatively long coherence times while speech signals have relatively short coherence times, researchers have used it to detect and estimate howling frequencies in closed-loop speech reinforcement systems. In this application, the TCF estimation was obtained using the traditional Welch average periodogram algorithm, resulting in limited frequency domain resolution and non-negligible spectral leakage (although the Hanning window function, which has some resistance to spectral leakage, was applied in the STFT variation). Therefore, researchers have suggested using the Minimum Variance Distortionless Response (MVDR) technique proposed by J. Benesty et al. to estimate the TCF, in order to overcome these shortcomings. However, the power spectrum estimated using MVDR technology is actually equivalent to the output of a set of matched bandpass filters designed on a finite uniform sampling analysis frequency grid. If the signal frequency does not match any of the analysis frequencies, then the signal frequency components will be suppressed and cannot be detected from the spectrum; this is the common defect of MVDR spectrum estimation technology known as the "signal adaptation problem". In recent years, the application of Kernel with Compact Support (KCS) in time-frequency distribution has been very widespread, and it has been proven that KCS time-frequency distribution has reduced spectral leakage cross-interference terms and high-resolution measurement characteristics at its instantaneous frequency. As one of the compact support kernels, the Cheriet-Belouchrani kernel (hereinafter referred to as the CB kernel) has been successfully applied in image and video processing. By controlling the parameters of the CB kernel, it can effectively suppress spectral leakage cross-interference and improve time-frequency resolution.

[0117] In this embodiment of the invention, the present application proposes to apply the CB kernel function to smooth the short-time spectrum of the signal, and then use the smoothed spectrum to estimate the TCF of the signal, thereby improving and enhancing the estimation performance of TCF.

[0118] Specifically, given a short-time spectrum of signal x(n) denoted as X(k,t), the CB kernel function is applied to smooth X(k,t) in the following manner:

[0119] (19)

[0120] Where M is a preset positive integer, representing a frequency window length of 2M+1; It is a preset positive integer representing the frame time window length. ; This is a two-dimensional linear convolution operator; This is a CB kernel function, defined as follows:

[0121] (20),

[0122] Here, C>0 and B>0 are preset parameters that control the width and peak value of the CB kernel window function; m = -M, -M+1, …, M-1, M; n = , , …, , .

[0123] remember The delayed version of the long-time frame q (where q >> 1) is ,So and The long-term (frame) amplitude squared coherence coefficient can be calculated by the following formula:

[0124] (twenty one),

[0125] The superscript here is " " is the complex conjugate operator; q>>1 is a preset integer long-time frame parameter, which is recommended to be a value in the range of 160 ms to 320 ms; for example, if the signal sampling rate of the system to be processed is Fs = 16 kHz, and the number of skip samples used when dividing the signal into frames is L = 64, then the long-time frame parameter corresponding to 200 ms is... ,here It represents the smallest integer not less than x.

[0126] II. Working principle of alternative howling frequency points and their spectrum selector modules:

[0127] The long-term (frame) amplitude squared coherence coefficient calculated by formula (21) Select P frequency indexes that exceed the threshold maximum value as the candidate frequency index set for howling, in the following manner. ,Right now:

[0128] (twenty two),

[0129] in, (23), here The preset decision threshold parameter, P, typically ranges from 1 to 5. The candidate frequency point index set... The corresponding set of smoothed spectra The selected features are sent to the subsequent howling detector for howling feature extraction and decision processing.

[0130] III. Working principle of the howling detector:

[0131] This section first defines two new howling signal detection features, and then combines them with the existing howling feature PHPR for howling detection.

[0132] III-a. New howling detection features:

[0133] Spectrum smoothed from the CB kernel In terms of its alternative whistling frequency index Peak power spectrum stability measure Peak-to-average amplitude ratio stability measure They are defined as follows:

[0134] (twenty four),

[0135] (25)

[0136] in, The relative rate of change of spectral amplitude between frames; The average spectral amplitude of frame t, here For short-time spectrum The amplitude spectrum, where N is the length of the window function in the STFT. When When indexing the frequency of howling, its position is relatively stable in time and has a high spectral amplitude, therefore... The value is very small, therefore its and The value is very large at this point. Clearly, and All of these belong to the time-frequency domain type of signal characteristics.

[0137] III-b. Working principle of the first howling detector based on PPS features:

[0138] The detector is based on an index of alternative howling frequency points. Peak power spectrum stability measure The criteria for detecting "saturation howling" during the detection period are as follows:

[0139] (26)

[0140] in, This is the binary detection indication signal output by the first howling detector at frame t; "V" is the logical "OR" operator (the same applies below, without further explanation); (Unit: signal frame) represents the preset decision threshold parameter for the first howling detector; Candidate howling frequencies in the first howling detector The counter value at frame t is defined by the following formula:

[0141] (27)

[0142] here (Unit: dB) represents the preset decision threshold parameter for the first howling detector; The set of candidate howling frequency indexes for frame t.

[0143] III-c. Working principle of the second howling detector based on PAMRS and PHPR features:

[0144] The detector is based on an index of alternative howling frequency points. Peak-to-average amplitude ratio stability measure Peak-to-harmonic power ratio These two features are used to detect the "pre-whistling" period of a howling signal, and the decision criterion for howling detection is as follows:

[0145] (28)

[0146] here This is the binary detection indication signal output by the second howling detector at frame t; (Unit: signal frame) represents the preset decision threshold parameter for the second howling detector. Candidate howling frequencies in the second howling detector The counter value at frame t is defined by the following formula:

[0147] (29),

[0148] in, (Unit: dB) and (Unit: dB) These are the preset decision threshold parameters for the second howling detector; The set of alternative howling frequency indexes for frame t; Candidate howling frequencies in the second howling detector The peak-to-harmonic power ratio at frame t is defined by the following formula:

[0149] (30)

[0150] in, .

[0151] IV. Hierarchical decision rules for the howling detector and a flowchart illustrating its engineering implementation:

[0152] The final decision expression for the howling detector proposed in this invention is a hierarchical decision rule, which can be expressed as follows:

[0153] (31),

[0154] in The set of alternative howling frequency indexes for frame t; This is the binary detection indication signal output by the first howling detector at frame t; This is the binary detection indication signal output by the second howling detector at frame t.

[0155] Specifically, the working principle of this hierarchical howling decision rule proposed in this invention is as follows:

[0156] First, examine the candidate howling frequency point index set for frame t obtained by filtering based on the squared coherence coefficient of the long-term frame amplitude. Is it an empty set? If the set is empty, then the signal in frame t is determined to have no howling component, and the output indication of the final howling detector in frame t is set to "false" (i.e., hdFlag(t) = 0). Simultaneously, the howling decision processing for this signal frame ends, and the process moves to the howling decision processing for the next signal frame. If If the set is non-empty, then the howling decision immediately proceeds to the first howling detector for howling detection; if the output of the first howling detector is "true" (i.e.: If the signal in frame t is determined to have a howling component, then the output indication of the final howling detector in frame t is set to "true" (i.e., hdFlag(t) = 1), and the howling decision processing for this signal frame ends, transitioning to the howling decision processing for the next signal frame; if the output indication of the first howling detector is "false" (i.e., ... If so, the howling decision will immediately proceed to the second howling detector for howling detection, and the output indication of the second howling detector will be used as the output indication of the final howling detector (i.e.: At the same time, the howling decision processing of this signal frame ends and the howling decision processing of the next signal frame begins.

[0157] Therefore, this hierarchical howling decision rule can reduce the computational complexity and false detection probability of the howling detection algorithm, while improving its detection probability. A detailed flowchart of the engineering implementation of the howling detection algorithm proposed in this invention can be found in [link to relevant documentation]. Figure 6 As shown.

[0158] It is understood that the technical solution provided in this embodiment proposes two time-frequency features for howling detection, namely "Peak Power Spectrum Stability Measure (PPS)" and "Peak-to-Average Amplitude Ratio Stability Measure (PAMRS)". Compared with existing howling detection features, these two feature parameters have better detection performance and stronger robustness to the operating environment.

[0159] First, based on a novel sound signal parameter feature, the peak power spectrum stability measure, the presence of a first-type howling frequency index in the short-time spectrum to be detected is provided, offering a new howling detection method compared to existing technologies. Further, another novel sound feature, the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio, are employed to detect the presence of a second-type howling frequency index in the short-time spectrum to be detected. Furthermore, a long-time frame amplitude squared coherence coefficient is superimposed to distinguish howling signals from normal speech signals. This achieves a high detection probability while reducing the false detection probability, effectively improving howling suppression. This method first uses the long-time frame amplitude squared coherence coefficient (MSC) of the signal spectrum to screen "candidate howling frequencies," thereby reducing the false detection probability of howling caused by speech signal formants. Within the candidate howling frequency set, the peak power spectrum stability measure and the peak-to-average amplitude ratio stability measure—two time-frequency features of howling signals proposed in this invention—are applied, combined with the frequency domain features of the peak-to-harmonic ratio howling signal, for howling detection, further reducing the false detection probability of howling caused by musical instrument audio signals. Example 3

[0160] like Figure 7 As shown in the schematic diagram, a communication system howling detection device provided in an embodiment of the present invention includes:

[0161] The sound signal acquisition module 71 is used to acquire the sound signal to be detected.

[0162] The short-time spectrum generation module 72 is used to generate the short-time spectrum of the sound signal to be detected at frame t.

[0163] The first howling detector 73 is used to detect whether a first type of howling frequency index exists in the short-time spectrum of the target signal at frame t based on the peak power spectrum stability measure. The peak power spectrum stability measure is formed by weighting a specified short-time spectrum amplitude square with a specified weighting coefficient and taking a decibel scale. The specified weighting coefficient is formed with Euler number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the target sound signal as the exponent. The specified short-time spectrum amplitude square is the short-time spectrum amplitude square of the target sound signal at the local peak frequency index. The first type of howling frequency index is a local peak frequency index where the peak power spectrum stability measure is greater than a preset peak power spectrum stability measure threshold.

[0164] The second howling detector 74, if the first type of howling frequency index does not exist in the short-time spectrum to be detected at frame t, then detects whether the second type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio. The peak-to-average amplitude ratio stability measure is formed by weighting the specified short-time spectrum amplitude ratio with the specified weighting coefficient and taking a decibel scale. The specified short-time spectrum amplitude ratio is the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index and the short-time spectrum amplitude across the entire frequency band. The ratio of average values; the peak-to-harmonic power ratio is formed by taking the decibel scale of the specified short-time spectral amplitude square ratio, which is the ratio of the square of the short-time spectral amplitude of the sound signal to be detected at the local peak frequency index to the square of the short-time spectral amplitude at the corresponding harmonic frequency index of the local peak frequency; the second type of howling frequency index is the local peak frequency index corresponding to the peak-to-average amplitude stability measure being greater than a preset peak-to-average amplitude stability measure threshold and the peak-to-harmonic power ratio being greater than a preset peak-to-harmonic power ratio threshold.

[0165] In some alternative embodiments, such as Figure 7 As shown by the dashed line, the aforementioned communication system howling detection device may further include:

[0166] The long-time frame amplitude squared coherence coefficient calculation module 75 is used to calculate the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t.

[0167] The alternative howling frequency index acquisition module 76 is used to acquire a preset number of alternative howling frequency indexes from the short-time spectrum to be detected at frame t based on the long-time frame amplitude square coherence coefficient. The alternative howling frequency index is the frequency index corresponding to the long-time frame amplitude square coherence coefficient at frame t being greater than the preset long-time frame amplitude square coherence coefficient threshold parameter.

[0168] First howling detector 73: Based on the peak power spectrum stability measure and the candidate howling frequency index, detect whether there is a first type of howling frequency index in the short-time spectrum to be detected when frame t;

[0169] Specifically, the first howling detector 73 can utilize the formula This paper implements a method to detect the existence of a type I howling frequency index in the short-time spectrum of the target device at frame t based on the peak power spectral stability metric. This is the binary detection indication signal output by the first howling detector in frame t; "V" is the logical "OR" operator; (Unit: signal frame) represents the preset decision threshold parameter for the first howling detector; Index of candidate howling frequency points in the first howling detector The counter value at frame t is determined by the index of the candidate howling frequency point at frame t. When the peak power spectrum stability measure at a certain point is greater than a preset peak power spectrum stability measure threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

[0170] Second howling detector 74: Based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio and the candidate howling frequency index, detect whether there is a second type of howling frequency index in the short-time spectrum to be detected at frame t.

[0171] Specifically, the second howling detector 74 can utilize the formula This method implements a mechanism to detect the existence of a second type of howling frequency index in the short-time spectrum of the target device at frame t, based on the peak-to-average amplitude ratio stability metric and the peak-to-harmonic power ratio. This is the binary detection indication signal output by the second howling detector at frame t; (Unit: signal frame) represents the preset decision threshold parameter for the second howling detector; Index of candidate howling frequency points in the second howling detector The counter value at frame t is determined by the index of the candidate howling frequency point at frame t. When the peak-to-average amplitude ratio stability measure is greater than a preset peak-to-average amplitude ratio stability measure threshold and the peak-to-harmonic power ratio is greater than a preset peak-to-harmonic power ratio threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

[0172] In some alternative embodiments, such as Figure 7 As shown by the dashed line, the aforementioned alternative howling frequency index acquisition module 76 may include:

[0173] The submodule 761 for obtaining all candidate howling frequency indexes is used to obtain all candidate howling frequency indexes from the short-time spectrum to be detected at frame t based on the long-time frame amplitude squared coherence coefficient.

[0174] The candidate howling frequency index acquisition submodule 762 is used to acquire a preset number of candidate howling frequency indexes with the largest long-time frame amplitude square coherence coefficient from all the candidate howling frequency indexes.

[0175] The alternative short-time spectrum acquisition submodule 76 3 is used to acquire the alternative short-time spectrum corresponding to the alternative howling frequency index from the short-time spectrum to be detected at frame t; if the alternative howling frequency index exists in the short-time spectrum to be detected at frame t, the first howling detector 73 detects whether the first type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure and the alternative howling frequency index;

[0176] If the candidate howling frequency index is not present in the short-time spectrum to be detected at frame t, then the next frame of short-time spectrum in the short-time spectrum to be detected at frame t is obtained, until all short-time spectrum to be detected at frame t is detected.

[0177] In some alternative embodiments, such as Figure 7 As shown by the dashed line, the aforementioned communication system howling detection device may further include:

[0178] The smoothing module 77 is used to perform Cherrett-Beruclani kernel smoothing on the short-time spectrum to be detected at frame t.

[0179] In some alternative embodiments, such as Figure 7 As shown by the dashed line, the aforementioned communication system howling detection device may further include:

[0180] Hierarchical howling decision rule 78 is used to determine the detection indication signal of the howling frequency point index of the short-time spectrum to be detected at frame t using the following hierarchical final decision expression: ,

[0181] in, The detection indication signal is the howling frequency index of the short-time spectrum to be detected at frame t. The set of candidate howling frequency points for frame t; This is a detection indication signal indicating whether the first type of howling frequency point index exists in the short-time spectrum to be detected at frame t; This is a detection indication signal for whether the second type of howling frequency index exists in the short-time spectrum to be detected at frame t; the output is the detection indication signal for the howling frequency index of the short-time spectrum to be detected at frame t.

[0182] It is understood that the technical solution provided in this embodiment detects the presence of a first type of howling frequency index in the short-time spectrum of the target audio signal based on a new peak power spectrum stability measure of audio signal parameters. Compared with existing technologies, this provides a new howling detection method. Furthermore, it employs another new audio feature, the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio, to detect the presence of a second type of howling frequency index in the short-time spectrum of the target audio signal. It also further superimposes the long-time frame amplitude squared coherence coefficient to distinguish howling signals from normal speech signals. This achieves a high detection probability while reducing the false detection probability, effectively improving howling suppression.

[0183] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention. Example 4

[0184] Based on the same technical concept, this application also provides a computer device, including a memory 1 and a processor 2, wherein the memory 1 stores a computer program, and the processor 2 executes the computer program to implement the acoustic feedback processing method in the voice communication system described above.

[0185] The memory 1 includes at least one type of readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 1 can be an internal storage unit of the howling detection system, such as a hard disk. In other embodiments, the memory 1 can be an external storage device of the howling detection system, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Furthermore, the memory 1 can include both internal storage units and external storage devices of the howling detection system. The memory 1 can be used not only to store application software and various types of data installed in the howling detection system, such as the code of the howling detection program, but also to temporarily store data that has been output or will be output.

[0186] In some embodiments, processor 2 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program code stored in memory 1 or process data, such as executing a howling detection program.

[0187] It is understood that the technical solution provided in this embodiment detects the presence of a first type of howling frequency index in the short-time spectrum of the target audio signal based on a new peak power spectrum stability measure of audio signal parameters. Compared with existing technologies, this provides a new howling detection method. Furthermore, it employs another new audio feature, the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio, to detect the presence of a second type of howling frequency index in the short-time spectrum of the target audio signal. It also further superimposes the long-time frame amplitude squared coherence coefficient to distinguish howling signals from normal speech signals. This achieves a high detection probability while reducing the false detection probability, effectively improving howling suppression.

[0188] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the acoustic feedback processing method in the voice communication system described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0189] The computer program product of the acoustic feedback processing method in the voice communication system provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the acoustic feedback processing method in the voice communication system described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0190] The present invention also discloses a computer program that, when executed by a processor, implements any of the methods described in the foregoing embodiments. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0191] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0192] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.

[0193] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0194] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0195] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0196] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0197] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.

[0198] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0199] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method of howl detection for a communication system, characterized by, include: Acquire the sound signal to be detected; Generate the short-time spectrum of the sound signal to be detected at frame t; The presence of a first-type howling frequency index in the short-time spectrum of the target signal at frame t is detected based on the peak power spectrum stability measure. The peak power spectrum stability measure is formed by weighting a specified short-time spectrum amplitude square with a specified weighting coefficient and then taking a decibel scale. The specified weighting coefficient is formed with Euler's number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the target sound signal as the exponent. The specified short-time spectrum amplitude square is the short-time spectrum amplitude square of the target sound signal at the local peak frequency index. The first-type howling frequency index is a local peak frequency index where the peak power spectrum stability measure is greater than a preset peak power spectrum stability measure threshold. If the first type of howling frequency index does not exist in the short-time spectrum to be detected at frame t, then the presence of the second type of howling frequency index is detected in the short-time spectrum to be detected at frame t is determined based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio. The peak-to-average amplitude ratio stability measure is formed by weighting the specified short-time spectrum amplitude ratio with the specified weighting coefficient and taking a decibel scale. The specified short-time spectrum amplitude ratio is the ratio of the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index to the average value of the short-time spectrum amplitude across the entire frequency band. The peak-to-harmonic power ratio is formed by taking the square of a specified short-time spectral amplitude in decibels. The specified short-time spectral amplitude square ratio is the ratio of the square of the short-time spectral amplitude of the sound signal to be detected at a local peak frequency index to the square of the short-time spectral amplitude at the corresponding harmonic frequency index of the local peak frequency. The second type of howling frequency index is a local peak frequency index corresponding to a peak-to-average amplitude stability measure that is greater than a preset peak-to-average amplitude stability measure threshold and a peak-to-harmonic power ratio that is greater than a preset peak-to-harmonic power ratio threshold.

2. The communication system howl detection method according to claim 1, characterized in that, Before detecting whether a first type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak power spectral stability metric, the method further includes: Calculate the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t; Based on the long-time frame amplitude squared coherence coefficient, a preset number of candidate howling frequency indexes are obtained from the short-time spectrum to be detected at frame t. The candidate howling frequency indexes are the frequency indexes corresponding to the long-time frame amplitude squared coherence coefficient at frame t being greater than the preset long-time frame amplitude squared coherence coefficient threshold parameter. The method of detecting whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure is: whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure and the candidate howling frequency index; The method for detecting whether a second type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio is as follows: Detecting whether a second type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure, the peak-to-harmonic power ratio, and the candidate howling frequency index.

3. The communication system howl detection method according to claim 2, characterized in that, The step of obtaining a preset number of candidate howling frequency point indices from the short-time spectrum to be detected at frame t based on the squared coherence coefficient of the long-time frame amplitude includes: Based on the long-time frame amplitude squared coherence coefficient, obtain all candidate howling frequency indexes from the short-time spectrum to be detected at frame t. Obtain a preset number of candidate howling frequency indexes from all the candidate howling frequency indexes, based on the largest long-time frame amplitude squared coherence coefficient.

4. The communication system howl detection method according to claim 3, characterized by, Before calculating the long-time frame amplitude squared coherence coefficient of the short-time spectrum to be detected at frame t, the method further includes: The short-time spectrum to be detected at frame t is smoothed using the Cherit-Beruclani kernel.

5. The communication system howl detection method according to claim 4, characterized in that, The step of obtaining a preset number of candidate howling frequency point indices from the short-time spectrum to be detected at frame t based on the long-time frame amplitude squared coherence coefficient further includes: Obtain the candidate short-time spectrum corresponding to the candidate howling frequency point index from the short-time spectrum to be detected at frame t; The method of detecting whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t based on the peak power spectrum stability measure is as follows: if the candidate howling frequency index exists in the short-time spectrum to be detected at frame t, then the method of detecting whether there is a first type of howling frequency index in the short-time spectrum to be detected at frame t is based on the peak power spectrum stability measure and the candidate howling frequency index. If the candidate howling frequency index is not present in the short-time spectrum to be detected at frame t, then the next frame of short-time spectrum in the short-time spectrum to be detected at frame t is obtained, until all short-time spectrum to be detected at frame t is detected.

6. The communication system howling detection method according to claim 5, characterized in that, The index for whether a type I howling frequency point exists in the short-time spectrum to be detected at frame t based on the peak power spectral stability metric is: using the formula... This paper implements a method to detect the existence of a type I howling frequency index in the short-time spectrum of the target device at frame t based on the peak power spectral stability metric. "V" represents the binary detection indication signal output by the first howling detector in frame t; "V" is the logical "OR" operator. (Unit: signal frame) represents the preset decision threshold parameter for the first howling detector; Index of candidate howling frequency points in the first howling detector The counter value at frame t. The determination method is to use the candidate howling frequency index in frame t. When the peak power spectrum stability measure at a certain point is greater than a preset peak power spectrum stability measure threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

7. The communication system howling detection method according to claim 6, characterized in that, The index for whether a second type of howling frequency point exists in the short-time spectrum to be detected at frame t, based on the peak-to-average amplitude ratio stability metric and the peak-to-harmonic power ratio, is: using the formula... This method implements a mechanism to detect the existence of a second type of howling frequency index in the short-time spectrum of the target device at frame t, based on the peak-to-average amplitude ratio stability metric and the peak-to-harmonic power ratio. This is the binary detection indication signal output by the second howling detector at frame t; (Unit: signal frame) represents the preset decision threshold parameter for the second howling detector; Index of candidate howling frequency points in the second howling detector The counter value at frame t. The determination method is to use the candidate howling frequency index in frame t. When the peak-to-average amplitude ratio stability measure is greater than a preset peak-to-average amplitude ratio stability measure threshold and the peak-to-harmonic power ratio is greater than a preset peak-to-harmonic power ratio threshold, It will automatically increment by one, or it will automatically decrement by one until it reaches zero.

8. The communication system howling detection method according to claim 7, characterized in that, The method further includes: The detection indication signal of the howling frequency point index of the short-time spectrum to be detected at frame t is determined by the following hierarchical final decision expression: , in, The detection indication signal is the howling frequency index of the short-time spectrum to be detected at frame t. The set of candidate howling frequency points for frame t; This is a detection indication signal indicating whether the first type of howling frequency point index exists in the short-time spectrum to be detected at frame t; This is a detection indication signal for whether the index of the second type of howling frequency point exists in the short-time spectrum to be detected at frame t. 1 indicates that the detection result is true and 0 indicates that the detection result is false. Output the detection indication signal of the howling frequency index of the short-time spectrum to be detected at frame t.

9. A communication system howling detection device, characterized in that, include: The sound signal acquisition module is used to acquire the sound signal to be detected; The short-time spectrum generation module is used to generate the short-time spectrum of the sound signal to be detected at frame t. The first howling detector is used to detect whether a first type of howling frequency index exists in the short-time spectrum of the target signal at frame t based on the peak power spectrum stability measure. The peak power spectrum stability measure is formed by weighting a specified short-time spectrum amplitude square with a specified weighting coefficient and taking a decibel scale. The specified weighting coefficient is formed with Euler number e as the base and the negative of the absolute value of the relative change rate of the amplitude spectrum between frames at the local peak frequency index of the short-time spectrum amplitude of the target sound signal as the exponent. The specified short-time spectrum amplitude square is the short-time spectrum amplitude square of the target sound signal at the local peak frequency index. The first type of howling frequency index is a local peak frequency index where the peak power spectrum stability measure is greater than a preset peak power spectrum stability measure threshold. The second howling detector, if the first type of howling frequency index does not exist in the short-time spectrum to be detected at frame t, then detects whether the second type of howling frequency index exists in the short-time spectrum to be detected at frame t based on the peak-to-average amplitude ratio stability measure and the peak-to-harmonic power ratio. The peak-to-average amplitude ratio stability measure is formed by weighting a specified short-time spectrum amplitude ratio with the specified weighting coefficient and taking a decibel scale. The specified short-time spectrum amplitude ratio is the short-time spectrum amplitude of the sound signal to be detected at the local peak frequency index and the short-time spectrum amplitude across the entire frequency band. The ratio of the mean; the peak-harmonic power ratio is formed by taking the decibel scale of the specified short-time spectral amplitude square ratio, which is the ratio of the square of the short-time spectral amplitude of the sound signal to be detected at the local peak frequency index to the square of the short-time spectral amplitude at the corresponding harmonic frequency index of the local peak frequency; the second type of howling frequency index is the local peak frequency index corresponding to the peak-average amplitude stability measure being greater than a preset peak-average amplitude stability measure threshold and the peak-harmonic power ratio being greater than a preset peak-harmonic power ratio threshold.

Citation Information

Patent Citations

  • Howling processing method and device of voice broadcast sound amplification system

    CN116312592A