Howling detection method, apparatus, device, and medium
By analyzing the amplitude changes of audio frequency points and the characteristics of human ear perception, combined with features such as spectral sparsity, the system identifies and distinguishes howling signals, solving the problems of misjudgment and insufficient generalization ability in howling detection, and achieving high-precision and stable howling detection.
Patent Information
- Application Number
- CN202411737342.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing methods for detecting howling have a high false positive rate and insufficient generalization ability, making them difficult to adapt to complex and changing environments. They also rely on a large amount of labeled data and frequent threshold adjustments.
By analyzing the amplitude variation information of audio at various frequency points and the characteristics of human ear perception, combined with the amplitude variation trend and distribution characteristics, candidate howling frequency points are identified. Furthermore, by utilizing features such as amplitude variation acceleration and spectral sparsity, the howling type and detection results are determined, reducing the reliance on labeled data.
It achieves high-precision whistling detection in complex environments, reduces false detection and false negative rates, improves the adaptability and stability of detection, and reduces resource consumption.
Smart Images

Figure CN119763598B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to a method, apparatus, device and medium for detecting howling. Background Technology
[0002] In audio equipment, feedback is a significant factor affecting sound quality and equipment stability. Feedback is typically caused by excessive gain in the acoustic feedback path or amplification of specific frequencies of the audio signal. It is characterized by a sharp, single-frequency sound and can, in severe cases, lead to equipment failure.
[0003] Traditional howling detection methods typically rely on setting multiple thresholds to determine howling frequency points. However, setting these thresholds requires extensive verification and debugging, and they are prone to failure in complex or changing environments. Different environments and device configurations may require different threshold adjustments. Deep learning methods, through neural network models, can better handle complex scenarios and improve detection accuracy, but they require a large amount of labeled training data. The data acquisition and labeling process is both labor-intensive and resource-intensive, and is also limited by the data scenario and environment, resulting in insufficient generalization ability of the model.
[0004] Therefore, there is an urgent need for a whistling detection method that reduces resource consumption, avoids complex verification, and adapts to changing environments. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for detecting howling, in order to solve the defects of high false positive rate and insufficient generalization ability in related technologies.
[0006] This invention provides a method for detecting howling, comprising:
[0007] Obtain the audio to be detected;
[0008] Based on the amplitude variation information of the audio at each frequency point, candidate howling frequency points in the audio are determined;
[0009] Based on the amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, the howling detection result of the audio is determined.
[0010] According to a feedback detection method provided by the present invention, determining candidate feedback frequency points in the audio based on amplitude variation information of the audio at various frequency points includes:
[0011] Based on the amplitude variation trend and amplitude variation acceleration of the audio at each frequency point, candidate howling frequency points and the type of the candidate howling frequency points are determined from each frequency point.
[0012] According to the feedback detection method provided by the present invention, when the type of the candidate feedback frequency point is attenuated feedback, the amplitude distribution information of the candidate feedback frequency point in the audio includes spectral sparsity.
[0013] The determination of the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude values of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear, includes:
[0014] If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
[0015] According to the feedback detection method provided by the present invention, the amplitude distribution information of the candidate feedback frequency points in the audio also includes amplitude dispersion;
[0016] When the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human auditory perception, determining that the howling detection result includes attenuated howling includes:
[0017] If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, the amplitude dispersion of the candidate howling frequency point is less than the dispersion threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
[0018] According to a feedback detection method provided by the present invention, when the type of the candidate feedback frequency point is an increasing feedback, the step of determining the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency point in the audio, and / or the relationship between the amplitude value of the candidate feedback frequency point in the audio and the amplitude range perceived by the human ear, includes:
[0019] If the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, it is determined that the howling detection result includes an increasing howling.
[0020] According to the feedback detection method provided by the present invention, when the audio is acquired based on a loudspeaker system, it further includes:
[0021] Based on the feedback detection results of the audio and the current gain of the amplification system, determine the critical gain of the amplification system;
[0022] When the current gain is greater than the critical gain, the loudspeaker system is subjected to howling control.
[0023] This invention provides a training method for a howling suppression model, comprising:
[0024] Candidate audio is synthesized based on near-end speech and noisy audio;
[0025] Based on the aforementioned feedback detection method, feedback detection is performed on the candidate audio to obtain the feedback detection result of the candidate audio;
[0026] Based on the feedback detection results of the candidate audio, sample audio is selected from the candidate audio;
[0027] A howling suppression model is trained based on the sample audio and the near-end speech used to synthesize the sample audio.
[0028] According to a training method for a howling suppression model provided by the present invention, the step of selecting sample audio from the candidate audio based on the howling detection results of the candidate audio includes:
[0029] The severity of the feedback of the candidate audio is determined based on the number of feedback frequency points in the feedback detection results.
[0030] Based on the proportion of howling audio required for training and the howling severity of the candidate audio, sample audio is selected from the candidate audio.
[0031] The present invention also provides a howling detection device, comprising the following modules:
[0032] The acquisition unit is used to acquire the audio to be detected;
[0033] An evaluation unit is used to determine candidate howling frequency points in the audio based on the amplitude variation information of the audio at each frequency point;
[0034] The detection unit is used to determine the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude value of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear.
[0035] The present invention also provides a training device for a howling suppression model, comprising the following modules:
[0036] An audio synthesis unit is used to synthesize candidate audio based on near-end speech and noisy audio;
[0037] The feedback detection unit is used to perform feedback detection on the candidate audio based on the feedback detection method, and obtain the feedback detection result of the candidate audio;
[0038] A sample filtering unit is used to filter sample audio from the candidate audio based on the feedback detection results of the candidate audio;
[0039] The model training unit is used to train a howling suppression model based on the sample audio and the near-end speech used to synthesize the sample audio.
[0040] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the training method of any of the above-described howling detection methods or howling suppression models.
[0041] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for a howling detection method or a howling suppression model as described above.
[0042] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a training method for any of the above-described howling detection methods or howling suppression models.
[0043] The feedback detection method, apparatus, device, and medium provided by this invention detect feedback frequencies by analyzing the amplitude characteristics of frequency points and combining this with the characteristics of human auditory perception. Through dynamic analysis of amplitude characteristics, it can adapt to complex or changing environments without the need for frequent threshold adjustments. Furthermore, it does not rely on large amounts of labeled data, but rather achieves feedback detection through feature analysis, thereby reducing data dependence. This method can accurately distinguish between feedback signals and non-feedback signals, reducing false positives and false negatives, and possesses high detection accuracy and environmental adaptability, providing a stable and efficient feedback detection solution for audio equipment. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a flowchart illustrating the whistling detection method provided by the present invention.
[0046] Figure 2 This is a flowchart illustrating the training method of the howling suppression model provided by the present invention.
[0047] Figure 3 This is a schematic diagram of the acoustic feedback closed-loop simulation system provided by the present invention.
[0048] Figure 4 This is a schematic diagram of the whistling detection device provided by the present invention.
[0049] Figure 5 This is a schematic diagram of the training device for the howling suppression model provided by the present invention.
[0050] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0052] In audio equipment, feedback is a significant factor affecting sound quality and equipment stability. Feedback is typically caused by excessive gain in the acoustic feedback path or amplification of specific frequencies of the audio signal. It is characterized by a sharp, single-frequency sound, and in severe cases, can lead to equipment malfunction, negatively impact user experience, and reduce equipment reliability. Feedback is common in various audio devices such as public address systems, hearing aids, and teleconferencing systems, and is particularly pronounced in complex acoustic environments or when equipment configurations are suboptimal.
[0053] Traditional feedback detection methods typically rely on setting multiple thresholds to determine feedback frequencies. These methods monitor characteristics such as the amplitude and frequency response of audio signals to determine the presence of feedback. However, setting these thresholds requires extensive experimental verification and debugging, and they are prone to failure in complex or changing environments. Furthermore, threshold settings have a significant impact on device stability; improper threshold settings can easily lead to false positives or false negatives, affecting the reliability of the device's operation.
[0054] With the development of deep learning technology, neural network-based howling detection methods have gradually attracted attention. Deep learning methods can extract complex features of signals through self-learning, providing higher detection accuracy in more complex and dynamic environments. These methods optimize models using training data, enabling them to cope with changes in different environments and improve model adaptability. However, a major problem with deep learning methods is their dependence on large amounts of labeled data. The data collection and labeling process is not only costly in terms of manpower and resources, but also limited by data scenarios and environments, resulting in high costs for acquiring labeled data. Furthermore, due to the limitations of training data, the trained models may lack sufficient generalization ability and cannot effectively adapt to all real-world application scenarios.
[0055] To address the above problems, embodiments of the present invention provide a method for detecting howling. Figure 1 This is a flowchart illustrating the whistling detection method provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0056] Step 110: Obtain the audio to be detected.
[0057] Specifically, the audio to be detected can originate from various devices, such as real-time speech signals captured by a microphone, audio files read from a storage device, or real-time signals transmitted from a loudspeaker. The audio to be detected may contain multiple components, including the target signal (such as speech or music), ambient noise, and howling signals caused by acoustic feedback paths.
[0058] Step 120: Based on the amplitude variation information of the audio at each frequency point, determine the candidate howling frequency points in the audio.
[0059] Specifically, after acquiring the audio to be detected, the first step is to perform spectral analysis. This step typically involves transforming the audio from the time domain to the frequency domain using a short-time Fourier transform or other time-frequency analysis methods (such as wavelet transform). This process decomposes the audio into different frequency components and displays the amplitude changes at each frequency point within each time frame. In this way, a spectrum of the audio at different time frames and frequency points can be obtained, thus revealing the frequency characteristics of the audio and their changes over time.
[0060] In a spectrum analyzer, the amplitude variation at each frequency point reflects the characteristics of the audio, especially when the audio contains feedback components. Feedback typically manifests as a sharp increase or decrease in amplitude at certain frequency points. In the case of increasing feedback, the amplitude increases rapidly and remains at a high level; while in the case of decreasing feedback, the amplitude change may show a gradual decline, although the decline is often relatively gentle or stable. Both increasing and decreasing feedback will exhibit obvious amplitude variation characteristics in the spectrum.
[0061] To accurately identify howling frequencies, one can calculate the rate of change of amplitude for each frequency in the audio, paying particular attention to frequencies with large amplitude variations. For example, by calculating the rate of change of amplitude for each frequency between adjacent time frames, frequencies with sharp increases or decreases in amplitude can be identified. These frequencies are often associated with howling. Next, an amplitude variation threshold can be set to help filter out potential howling frequencies. For instance, if the amplitude variation of a frequency exceeds a preset threshold, or its increase or decrease exceeds the average rate of change of amplitude of the signal, then these frequencies are likely howling frequencies. Besides simple amplitude variation thresholds, other methods can be considered. For example, the standard deviation of amplitude can help identify frequencies with large amplitude fluctuations. By calculating the standard deviation of amplitude for each frequency within a certain time window, frequencies with large amplitude variations can be identified; these frequencies may be indicators of howling. Another method to consider is the acceleration of the rate of change of amplitude, i.e., calculating the second derivative (acceleration) of the amplitude change. A large acceleration indicates a more drastic amplitude change at that frequency, potentially indicating howling, whether it is an increasing or decreasing type.
[0062] Based on the amplitude variation information of each frequency point, it is possible to effectively identify and distinguish between increasing and decreasing howling frequency points. Ultimately, the selected frequency points will serve as candidate howling frequency points, acting as the basis for subsequent howling identification. After further verification and screening, these candidate frequency points will help accurately locate the howling source and effectively suppress it.
[0063] Step 130: Based on the amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, determine the howling detection result of the audio.
[0064] Specifically, after identifying candidate howling frequencies, it's necessary to further determine whether they are actual howling frequencies based on their amplitude distribution information. Amplitude distribution information refers to the temporal distribution characteristics of the amplitude values of these frequencies throughout the entire audio signal. Although these candidate frequencies exhibit drastic amplitude changes, amplitude variation alone is insufficient to definitively confirm them as actual howling frequencies. By analyzing amplitude distribution information, frequencies that better match howling characteristics can be further filtered out. The amplitude distribution of howling frequencies typically exhibits energy concentration and stability within a certain time range. For example, the amplitude distribution of increasing howling frequencies maintains a high value over multiple time frames, even showing a gradual increase followed by stabilization. In contrast, while the amplitude of decreasing howling frequencies gradually decreases, its amplitude distribution still exhibits regularity, showing a smooth attenuation trend over time. In comparison, background noise or other non-howling frequencies, even with drastic amplitude changes, often have a more dispersed amplitude distribution, exhibiting random fluctuations and lacking stability.
[0065] By statistically analyzing the amplitude distribution information of these candidate frequency points, potential howling frequencies can be further screened. One approach is to analyze whether the amplitude changes of the frequency points exhibit a certain degree of temporal stability. For example, the amplitude of frequency points in an increasing howling pattern should fluctuate stably within a certain range, while the amplitude of frequency points in an decreasing howling pattern should show a regular decreasing trend. A further approach is to divide the audio into multiple time windows and calculate the amplitude distribution characteristics of the frequency points within each time window. If a frequency point exhibits high energy and sustained concentration across all time windows, it further supports the possibility of being a howling frequency point.
[0066] Furthermore, because the human ear has a weak perception of low-loudness signals, some frequency points, although showing amplitude changes in the spectrum, may not be perceptible in actual hearing due to their low loudness. Even if the amplitude changes of these frequency points are significant in the spectrum, they may not have a significant impact on the user and therefore cannot be considered actual feedback components. To improve the accuracy of feedback detection, candidate frequency points with low loudness that are difficult for the human ear to perceive can be eliminated. Specifically, for each candidate feedback frequency point, its loudness value in the audio needs to be calculated. If the loudness of some frequency points is below the human ear's perception threshold (e.g., below a certain phon value), even if their amplitude changes are large, they should be considered to have little impact on the user's auditory experience and thus excluded as valid feedback frequency points. This screening process helps reduce misjudgments caused by environmental noise or non-feedback sources, ensuring that the finally detected feedback frequency points accurately reflect the feedback components that actually affect sound quality and stability.
[0067] This method effectively eliminates low-loudness frequencies that are difficult for the human ear to detect, ensuring more accurate feedback detection results and reducing the risk of false positives and false negatives. As a result, the final feedback detection results will better reflect auditory perception in real-world usage scenarios, effectively improving the performance of audio devices and the user experience.
[0068] The feedback detection method provided in this invention detects feedback frequencies by analyzing the amplitude characteristics of frequency points and combining this with the characteristics of human auditory perception. Through dynamic analysis of amplitude characteristics, it can adapt to complex or changing environments without frequent threshold adjustments. Furthermore, it does not rely on large amounts of labeled data, but rather achieves feedback detection through feature analysis, thereby reducing data dependence. This method can accurately distinguish between feedback signals and non-feedback signals, reducing false positives and false negatives, and possesses high detection accuracy and environmental adaptability, providing a stable and efficient feedback detection solution for audio devices.
[0069] Based on the above embodiments, determining candidate howling frequency points in the audio based on the amplitude variation information of the audio at each frequency point includes:
[0070] Based on the amplitude variation trend and amplitude variation acceleration of the audio at each frequency point, candidate howling frequency points and the type of the candidate howling frequency points are determined from each frequency point.
[0071] The amplitude variation trend at each frequency point refers to the direction and rate of change of the amplitude at each frequency point in the audio over time. For example, it can be represented by the sign and magnitude of the first derivative of the amplitude at each frequency point. The formula for calculating the first derivative is:
[0072]
[0073] in, Frequency point In time frame The amplitude value, Frequency point In time frame The amplitude value, The derivative of the time frame, Frequency point In time frame The first derivative of the amplitude value. The acceleration due to amplitude change represents the rate of change of amplitude with time, and can be expressed as the magnitude of the second derivative of the amplitude at each frequency point. The formula for calculating the second derivative is:
[0074]
[0075] in, Frequency point In time frame The second derivative of the amplitude value. By analyzing the amplitude variation trend and acceleration of the audio signal at various frequency points, candidate howling frequency points can be effectively identified, and the type of these frequency points can be further determined. Specifically, if the first derivative of the amplitude of a certain frequency point with respect to a time frame is constant over consecutive time frames, and the value of the second derivative is close to zero (e.g., below a certain set threshold), it can be determined that the amplitude variation rate of that frequency point is relatively gentle, which is usually an important characteristic of howling frequency points. The stability of this amplitude variation rate reflects the persistence of the howling signal in the time dimension, which is significantly different from the violent fluctuations of non-howling frequency points.
[0076] Furthermore, the sign of the first derivative can be used to determine the type of candidate howling frequency points. When the first derivative is positive and constant, it indicates that the amplitude of the frequency point is increasing at a stable rate, and it can be classified as a candidate howling frequency point for an increasing type. The amplitude of these frequency points continuously increases over time and exhibits high energy concentration. Conversely, when the first derivative is negative and constant, it indicates that the amplitude of the frequency point is decreasing at a stable rate, and it can be classified as a candidate howling frequency point for a decreasing type. The amplitude of these frequency points continuously decreases over time, but still exhibits high energy concentration.
[0077] By jointly analyzing the amplitude variation trend (first derivative) and amplitude variation acceleration (second derivative) at each frequency point, not only can candidate howling frequency points be accurately identified, but also howling increasing and howling decreasing can be distinguished based on the stability and rate direction of amplitude variation. This mathematical feature-based analysis method not only enhances detection accuracy but also improves the method's adaptability in complex environments, providing a reliable data foundation for subsequent howling suppression.
[0078] Based on the above embodiments, when the type of the candidate howling frequency point is attenuated howling, the amplitude distribution information of the candidate howling frequency point in the audio includes spectral sparsity.
[0079] The determination of the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude values of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear, includes:
[0080] If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
[0081] Specifically, when the selected howling frequency is attenuated, its amplitude distribution in the audio can be analyzed to further confirm whether it is the actual howling frequency. Amplitude distribution information includes the important characteristic of spectral sparsity. Spectral sparsity refers to the degree of concentration of energy distribution in the frequency domain. For signals with high sparsity, the sparsity is small, and its energy is mainly concentrated on a few frequency points. Conversely, for signals with low sparsity, the sparsity is large, resulting in a more uniform energy distribution spread across more frequency points. In attenuated howling, although the amplitude gradually decreases, the energy is often concentrated on a few frequency points, leading to high spectral sparsity. This high sparsity characteristic is a crucial basis for distinguishing attenuated howling from other non-howling signals.
[0082] There are several methods to measure spectral sparsity. For example, the ratio of the second-order norm to the fourth-order norm can be used to reflect the concentration of energy at a frequency point. Specifically, the formula for calculating the ratio of the second-order norm to the fourth-order norm is:
[0083]
[0084]
[0085] in, Indicates the number of time frames. It is the ratio of the second-order norm to the fourth-order norm. It is a normalized sparsity feature. When the sparsity is minimized, At frequency point The probability of a howling sound is highest when sparsity is at its maximum. At frequency point The probability of feedback is lowest at these frequency points. Energy distribution can also be assessed by calculating the ratio of the total energy at the highest frequency points to the overall energy. Furthermore, entropy can also be used to measure spectral sparsity; a lower entropy value indicates more concentrated energy, while a higher entropy value indicates more dispersed energy distribution. These methods can analyze the frequency distribution characteristics of signal energy from different perspectives, thereby accurately determining its sparsity.
[0086] During the detection process, when the spectral sparsity of a candidate howling frequency point is higher than the set sparsity threshold, and its amplitude value is within the range of human hearing perception, the frequency point can be confirmed as an actual attenuated howling signal. The sparsity condition effectively eliminates signals with relatively dispersed energy distribution, while the condition of human hearing perception range further excludes frequency points with low loudness that do not affect auditory perception. Through this joint judgment, the real howling signal can be accurately screened from the candidate frequency points.
[0087] Based on the above embodiments, the amplitude distribution information of the candidate howling frequency points in the audio also includes amplitude dispersion;
[0088] When the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human auditory perception, determining that the howling detection result includes attenuated howling includes:
[0089] If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, the amplitude dispersion of the candidate howling frequency point is less than the dispersion threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
[0090] Specifically, when the candidate howling frequency point is of the attenuated howling type, further analysis of the amplitude distribution information of the candidate howling frequency point in the audio is used to comprehensively determine whether it is the actual howling frequency point. Besides spectral sparsity, amplitude distribution information also includes amplitude dispersion. Amplitude dispersion is an important indicator for measuring the degree of fluctuation of signal amplitude within a time frame, used to assess whether the amplitude of a frequency point is stable. For howling frequency points, especially attenuated howling, their amplitude distribution characteristics are characterized by small fluctuation amplitudes that gradually decrease over time, thus the amplitude dispersion is usually low. Ordinary noise or other non-howling signals, on the other hand, have more random amplitude fluctuations and higher amplitude dispersion; this characteristic helps distinguish attenuated howling from other signals. Amplitude dispersion can be represented by calculating parameters such as the standard deviation or variance of the frequency point's amplitude value within a time frame. For example, the formula for calculating the standard deviation is:
[0091]
[0092]
[0093] in, It is a frequency point In time frame The standard deviation of the amplitude values. If the amplitude dispersion of a candidate howling frequency point is lower than the set dispersion threshold, it indicates that the amplitude change at that frequency point is relatively stable, consistent with the distribution characteristics of attenuated howling. Conversely, if the amplitude dispersion is high, it indicates that the amplitude fluctuation at that frequency point is large, and it is more likely to be background noise or other non-howling signals. For example, when the standard deviation... When it is less than 3, the current frequency point There may be a whistling phenomenon.
[0094] During the detection process, three conditions need to be considered together: spectral sparsity, amplitude dispersion, and the range of human hearing. When the spectral sparsity of a candidate howling frequency point is greater than the sparsity threshold, it indicates that its energy distribution has significant concentration; simultaneously, if the amplitude dispersion is lower than the dispersion threshold, it indicates that its amplitude distribution is relatively stable; and if the amplitude value is within the range of human hearing, it indicates that the signal has a real impact on auditory perception. If these conditions are met, the candidate frequency point can be determined as an actual attenuated howling and included in the howling detection results.
[0095] By combining multi-feature analysis of spectral sparsity, amplitude dispersion, and human auditory perception range, the detection results of attenuated feedback are more accurate. This comprehensive judgment method not only improves the accuracy and robustness of detection but also adapts to varying audio environments, providing important support for subsequent feedback suppression.
[0096] Based on the above embodiments, when the type of the candidate howling frequency point is an increasing howling, the...
[0097] The amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, are used to determine the howling detection result of the audio, including:
[0098] If the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, it is determined that the howling detection result includes an increasing howling.
[0099] Specifically, the typical characteristic of growing feedback is that its amplitude continuously increases over time and remains highly stable across multiple time frames. Unlike decaying feedback, the amplitude of growing feedback gradually increases at frequency points, with a higher concentration of energy, and it usually has a more significant impact on audio quality.
[0100] The identification of growing feedback primarily relies on the relationship between amplitude values and the range of human hearing, without requiring further analysis of amplitude distribution information. This is because the core characteristic of growing feedback is a significant increase in the amplitude of a frequency point; this sustained increase is sufficient to distinguish the feedback signal from background noise or other non-feedback signals. In contrast, decaying feedback gradually weakens in amplitude and may have dispersed energy, thus requiring the integration of amplitude distribution information (such as frequency sparsity or amplitude dispersion) to confirm its compliance with feedback characteristics. Growing feedback, due to its prominent signal characteristics, allows the amplitude changes themselves to fully reflect its pattern, eliminating the need for further distribution information. Therefore, determining whether a candidate frequency point's amplitude value is within the perceptual range is sufficient to effectively confirm whether it is an actual feedback frequency point. Specifically, when the amplitude value of a candidate frequency point exceeds a set perceptual threshold, it can be directly identified as a growing feedback frequency point; if its amplitude value is below the perceptual range, even a significant increase in amplitude may not substantially affect audio quality, and it should not be identified as a growing feedback frequency point.
[0101] The method for identifying growing howling directly confirms the presence of growing howling in the detection result by combining the amplitude value with the range of human hearing. This method effectively utilizes the prominent characteristics of growing howling signals, simplifying the judgment process and improving detection efficiency while ensuring detection accuracy.
[0102] Based on the above embodiments, when the audio is acquired by a loudspeaker system, it further includes:
[0103] Based on the feedback detection results of the audio and the current gain of the amplification system, determine the critical gain of the amplification system;
[0104] When the current gain is greater than the critical gain, the loudspeaker system is subjected to howling control.
[0105] Specifically, critical gain is a key threshold for the gain value in a public address system, representing the maximum gain allowed without triggering feedback. When the actual gain of the public address system exceeds the critical gain, gain accumulation in the acoustic feedback path will cause feedback, affecting audio quality and system stability.
[0106] The determination of the critical gain relies on the feedback detection method. By analyzing the feedback frequency points and their energy distribution characteristics in the current audio signal, combined with the current gain value of the amplifier system, the critical gain range of the amplifier system can be accurately identified. The feedback detection method identifies the current feedback signal, determines whether the current gain is close to the boundary state that triggers feedback, and accordingly confirms the maximum permissible gain value without triggering feedback.
[0107] When the current gain exceeds the critical gain, feedback control of the loudspeaker system is necessary. Feedback control methods include reducing the loudspeaker system gain below the critical gain, adjusting the feedback path of the loudspeaker equipment, or using adaptive feedback cancellation technology to reduce the impact of the feedback signal in real time. Reducing the gain is the most direct method; by making the current gain less than the critical gain, feedback can be effectively avoided. Simultaneously, to maintain audio output clarity and gain while controlling feedback, fine-tuning can be done by considering the acoustic characteristics of the feedback path, such as optimizing the microphone and speaker layout, or adjusting delay parameters and room impulse response in the feedback path.
[0108] By dynamically determining the critical gain and adjusting the amplifier system gain in real time through feedback detection methods, stable system operation can be ensured in varying acoustic environments, avoiding feedback interference with audio quality. This mechanism enhances the stability and adaptability of the amplifier system, while also improving the user experience in complex scenarios.
[0109] This invention provides a training method for a howling suppression model. Figure 2 This is a flowchart illustrating the training method for the howling suppression model provided by this invention, as shown below. Figure 2 As shown, the method includes:
[0110] Step 210: Synthesize candidate audio based on near-end speech and noisy audio.
[0111] Near-end speech refers to the raw speech signal captured by the target microphone, which is usually the main audio content that needs to be processed and enhanced in a public address system. Noise audio simulates background noise that may exist in the public address environment, such as environmental noise, equipment noise, or feedback signal interference. By synthesizing near-end speech and noise audio, the complex audio signals that a public address system may encounter in real-world operating scenarios can be simulated, providing diverse data samples for training feedback suppression models.
[0112] In practice, the near-end speech signal can be mixed with noisy audio at a preset signal-to-noise ratio (SNR) to control the relative clarity of the target speech in a noisy environment. For example, at a high SNR, the energy of the near-end speech dominates, while at a low SNR, the noisy audio significantly interferes with the overall audio. By adjusting the SNR, candidate audio with different interference intensities can be generated to cover the amplification needs of various real-world scenarios.
[0113] Step 220: Based on the aforementioned feedback detection method, feedback detection is performed on the candidate audio to obtain the feedback detection result of the candidate audio.
[0114] The aforementioned howling detection method is used to comprehensively analyze the synthesized candidate audio, identify the presence of howling signals, and generate corresponding detection results. Candidate audio typically contains near-end speech, noise interference, and potential howling components. The howling detection method can accurately extract howling features from these audio samples, providing foundational data for subsequent processing. Specifically, the howling detection method analyzes the spectral characteristics of the audio signal, comprehensively judging the amplitude variation trend, amplitude distribution information, and other features (such as spectral sparsity or amplitude dispersion) of each frequency point. Typical characteristics of howling signals are drastic amplitude changes and concentrated energy; therefore, the howling detection method first extracts potential candidate howling frequencies from the candidate audio. The amplitude variation information of these candidate frequencies is matched with howling characteristics to determine if they conform to the features of a howling signal. Finally, the detection results clearly identify the howling frequencies present in the candidate audio and the type of howling signal (growing or decaying). These detection results can not only be used to screen training data but also provide a reference for optimizing the subsequent howling suppression model.
[0115] Step 230: Based on the feedback detection results of the candidate audio, select sample audio from the candidate audio.
[0116] Specifically, in this step, by analyzing the feedback detection results of candidate audio files, sample audio files that meet specific conditions are selected for use in the subsequent training of the feedback suppression model. The purpose of selecting sample audio files is to ensure that the training dataset has sufficient diversity and representativeness, while avoiding sample audio files with excessive feedback that may interfere with the model training effect.
[0117] In candidate audio samples, amplification gain and room impulse response significantly influence the occurrence and severity of feedback. When the amplification gain is high or the cumulative gain of the signal is high due to feedback path characteristics, both the probability and severity of feedback increase significantly. While such severely feedback-laden audio samples can demonstrate the extreme processing capabilities of the feedback suppression model, if the proportion of such data in the training set is too high, the model may struggle to recover the target speech signal from these samples during training, leading to difficulties in converging the training loss function. Therefore, it is necessary to statistically analyze the feedback detection results of candidate audio samples and select sample audio samples with appropriate feedback levels based on the number of feedback frequency points. For example, sample audio samples with a large number of feedback frequency points are considered severe feedback data, and selecting such sample audio samples helps in better training the model.
[0118] Besides the number of howling frequency points, key information in howling detection results includes the proportion of howling signals to the total audio signal (howling audio percentage) and the type of howling (growing or decaying). Analyzing this information provides more diverse sample audio for model training. This selection process generates a representative set of sample audio, including howling samples of appropriate proportion and severity, and a reasonable distribution of growing and decaying howling samples. Such a dataset can effectively improve the training performance of the howling suppression model in complex environments, ensuring its strong generalization ability.
[0119] Step 240: Train a howling suppression model based on the sample audio and the near-end speech used to synthesize the sample audio.
[0120] Specifically, the model is trained using selected sample audio as input and the corresponding near-end speech as the target label. The goal of the training is to learn the howling features in the audio and accurately distinguish the howling signal from the target speech signal, thereby effectively suppressing the howling signal.
[0121] The sample audio contains rich feedback characteristics and noise interference. After the aforementioned screening steps, these samples contain a moderate proportion of feedback audio and cover different characteristics of both increasing and decreasing feedback, providing diverse training data for the model. During training, the sample audio is input into the feedback suppression model. The model learns the feature structure of the audio through layer-by-layer processing, including the frequency distribution of the feedback signal, amplitude variation patterns, and the differences between the feedback and the target speech.
[0122] Near-end speech serves as a label signal, providing a reference for the model and ensuring that it preserves the target speech signal as completely as possible while suppressing feedback. Near-end speech is typically a clean target speech signal, unaffected by feedback or noise, and can therefore guide the model in learning how to remove feedback components while preserving the frequency characteristics, content, and quality of the speech. The model's training loss function usually employs error metrics from speech signal reconstruction, such as mean square error in the time domain or energy spectrum error in the frequency domain. It can also incorporate Bark frequency domain loss, which is perceived by the human ear, to ensure that the model's suppression effect matches actual auditory perception.
[0123] In practical implementation, deep learning models (such as convolutional neural networks or time-frequency hybrid networks) can be used to extract and reconstruct features from the input signal. For example, the mixed signal, which combines near-end speech, noisy audio, and feedback signals, can be processed first. Converted to a complex spectrum via short-time Fourier transform. And calculate the energy spectrum To better align with the auditory characteristics of the human ear, the mixed signal energy is mapped from the frequency domain to the Bark domain, and the logarithmic decibel value is taken to obtain the Bark domain energy spectrum. ,in This represents the frequency points in the Bark domain. During feature extraction, a feedforward sequence memory neural network model is used for training. The training objective is to generate a mask matrix (Mask) for howling suppression, calculated using the following formula:
[0124]
[0125] in, and These represent the energy spectra of the noise audio and the feedback signal in the Bark domain, respectively. This mask matrix is used to estimate the enhancement process of the target signal under howling interference. To optimize the model, the loss function is defined as the mean squared error between the model's predicted mask matrix and the true mask matrix for each frame:
[0126]
[0127] in, It is the total number of frequency points in the Bark domain. is the true mask matrix. By minimizing this loss function, the model can learn how to accurately distinguish between near-end speech, noisy audio, and feedback signals, thereby effectively suppressing howling and recovering a clear target speech signal.
[0128] Through this training process, the howling suppression model can effectively grasp the characteristic differences between the howling signal and the target speech signal, thereby recognizing and suppressing howling in real time in practical applications while preserving the integrity and naturalness of the speech signal. Ultimately, the trained model can be applied to various sound reinforcement scenarios, providing stable and reliable howling suppression capabilities for audio equipment.
[0129] Based on the above embodiments, the step of selecting sample audio from the candidate audio based on the feedback detection results of the candidate audio includes:
[0130] The severity of the feedback of the candidate audio is determined based on the number of feedback frequency points in the feedback detection results.
[0131] Based on the proportion of howling audio required for training and the howling severity of the candidate audio, sample audio is selected from the candidate audio.
[0132] Specifically, the severity of feedback in candidate audio samples is determined by analyzing the number of feedback frequency points in the feedback detection results. The number of feedback frequency points reflects the coverage and intensity of the feedback signal in the audio; a higher number indicates a more severe feedback signal. However, audio samples with excessively severe feedback may severely mask the target speech signal, making it difficult for the model to accurately learn effective feedback suppression characteristics during training, and may also cause the training loss function to fail to converge. Therefore, audio samples with a set upper limit for the number of feedback frequency points need to be removed during the screening process. This removal operation can effectively filter out overly extreme audio samples, avoiding negative impacts on training results.
[0133] While removing audio with severe feedback, it's also necessary to further filter the remaining candidate audio based on the proportion of feedback audio required for training. The proportion of feedback audio is the percentage of feedback signal audio in the training dataset, and it's a crucial indicator for controlling the distribution and characteristics of the training data. A reasonable proportion ensures that the model can fully learn feedback suppression characteristics while maintaining good reproduction of the target speech signal. In practice, the proportion of feedback audio in the sample audio can be controlled by setting the amplification gain range. As shown in Table 1 below, different gain ranges correspond to different proportions of feedback audio.
[0134] Table 1: Different gain ranges correspond to different proportions of howling audio.
[0135]
[0136] This screening process removes audio samples with severe feedback while constructing a sample audio set covering different proportions of feedback audio. This method ensures the quality and diversity of the model training data, enabling the feedback suppression model to exhibit strong adaptability in complex scenarios.
[0137] The training method for the howling suppression model provided by this invention analyzes the number and proportion of howling frequency points in candidate audio samples to reasonably control the quality of the sample audio, enabling the model to have stronger generalization ability and robustness in complex scenarios. Simultaneously, it does not rely on a large amount of real-world scene data, significantly reducing data acquisition and annotation costs. Through a training loss function optimized for howling characteristics, the model can accurately suppress howling signals while preserving the sound quality and integrity of the target speech, providing stable and efficient howling suppression capabilities for audio devices in practical applications.
[0138] Based on the above embodiments, Figure 3 This is a schematic diagram of the acoustic feedback closed-loop simulation system provided by the present invention. Figure 3As shown, the input signal consists of a mixture of near-end audio and noise audio, used to simulate the actual input signal in a loudspeaker environment. The mixing of the two can be adjusted by setting the signal-to-noise ratio to generate audio signals with different interference intensities. The adaptive feedback cancellation module simulates the real-time suppression process of acoustic feedback by loudspeaker equipment. By analyzing the feedback signal of the previous frame, it adaptively adjusts the current audio, thereby partially eliminating the influence of feedback, which is a key function in the system to reduce howling. The delay module simulates the time delay caused by spatial propagation or equipment processing in the acoustic feedback path. The existence of delay causes a phase difference between the feedback signal and the original signal, which may lead to amplification of certain frequencies. The gain module simulates the amplification effect of loudspeaker equipment. The magnitude of the gain value determines the energy intensity of the signal in the feedback path. When the gain is too high, the feedback signal is continuously amplified, which may cause howling. The room impulse response module simulates the propagation characteristics of the audio signal in physical space, such as acoustic effects such as reflection and absorption. Its characteristics determine the frequency response and attenuation characteristics of the feedback path. After being processed by the above modules, the feedback signal is returned to the input, forming a closed-loop feedback path. When the gain is set too high and the accumulated energy of the feedback path exceeds the system's stability condition, howling may occur in the audio signal. This simulation system can generate howling audio with different intensities and characteristic distributions, providing diverse simulation data for training the howling suppression model. As shown in Table 2, after acoustic feedback control is enabled, the probability of howling at the same gain decreases.
[0139] Table 2: Percentage of howling audio frequencies corresponding to different gains before and after acoustic feedback control is enabled
[0140]
[0141] The following describes the howling detection device and the training device for the howling suppression model provided by the present invention. The howling detection device and the training device for the howling suppression model described below can be referred to in correspondence with the howling detection method and the howling suppression model method described above.
[0142] Figure 4 This is a schematic diagram of the whistling detection device provided by the present invention, as shown below. Figure 4 As shown, the device includes:
[0143] Acquisition unit 410 is used to acquire the audio to be detected;
[0144] Evaluation unit 420 is used to determine candidate howling frequency points in the audio based on the amplitude change information of the audio at each frequency point;
[0145] The detection unit 430 is used to determine the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude value of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear.
[0146] The feedback detection device provided in this invention detects feedback frequencies by analyzing the amplitude characteristics of frequency points and combining this with the characteristics of human auditory perception. Through dynamic analysis of amplitude characteristics, it can adapt to complex or changing environments without the need for frequent threshold adjustments. Furthermore, it does not rely on large amounts of labeled data, but rather achieves feedback detection through feature analysis, thereby reducing data dependence. This method can accurately distinguish between feedback signals and non-feedback signals, reducing false positives and false negatives, and possesses high detection accuracy and environmental adaptability, providing a stable and efficient feedback detection solution for audio devices.
[0147] Based on any of the above embodiments, the evaluation unit is specifically used for:
[0148] Based on the amplitude variation trend and amplitude variation acceleration of the audio at each frequency point, candidate howling frequency points and the type of the candidate howling frequency points are determined from each frequency point.
[0149] Based on any of the above embodiments, the detection unit is specifically used for:
[0150] When the type of the candidate howling frequency point is attenuated howling, the amplitude distribution information of the candidate howling frequency point in the audio includes spectral sparsity;
[0151] The determination of the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude values of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear, includes:
[0152] If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
[0153] The amplitude distribution information of the candidate howling frequency points in the audio also includes amplitude dispersion.
[0154] When the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human auditory perception, determining that the howling detection result includes attenuated howling includes:
[0155] If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, the amplitude dispersion of the candidate howling frequency point is less than the dispersion threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
[0156] The amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, are used to determine the howling detection result of the audio, including:
[0157] If the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, it is determined that the howling detection result includes an increasing howling.
[0158] Based on the feedback detection results of the audio and the current gain of the amplification system, determine the critical gain of the amplification system;
[0159] When the current gain is greater than the critical gain, the loudspeaker system is subjected to howling control.
[0160] Figure 5 This is a schematic diagram of the training device for the howling suppression model provided by the present invention, as shown below. Figure 5 As shown, the device includes:
[0161] The audio synthesis unit 510 is used to synthesize candidate audio based on near-end speech and noisy audio;
[0162] The feedback detection unit 520 is used to perform feedback detection on the candidate audio based on the feedback detection method, and obtain the feedback detection result of the candidate audio.
[0163] The sample filtering unit 530 is used to filter sample audio from the candidate audio based on the howling detection results of the candidate audio;
[0164] The model training unit 540 is used to train a howling suppression model based on the sample audio and the near-end speech used to synthesize the sample audio.
[0165] The training device for the feedback suppression model provided by this invention analyzes the number and proportion of feedback frequency points in candidate audio, and reasonably controls the quality of sample audio, enabling the model to have stronger generalization ability and robustness in complex scenarios. Simultaneously, it does not rely on a large amount of real-world scene data, significantly reducing data acquisition and annotation costs. Through a training loss function optimized for feedback characteristics, the model can accurately suppress feedback signals while preserving the sound quality and integrity of the target speech, providing stable and efficient feedback suppression capabilities for audio devices in practical applications.
[0166] Based on any of the above embodiments, the sample screening unit is specifically used for:
[0167] The severity of the feedback of the candidate audio is determined based on the number of feedback frequency points in the feedback detection results.
[0168] Based on the proportion of howling audio required for training and the howling severity of the candidate audio, sample audio is selected from the candidate audio.
[0169] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. The processor 610, communication interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a howling detection method or a howling suppression model training method. The howling detection method includes:
[0170] Obtain the audio to be detected;
[0171] Based on the amplitude variation information of the audio at each frequency point, candidate howling frequency points in the audio are determined;
[0172] Based on the amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, the howling detection result of the audio is determined.
[0173] Training methods for howling suppression models include:
[0174] Candidate audio is synthesized based on near-end speech and noisy audio;
[0175] Based on the aforementioned feedback detection method, feedback detection is performed on the candidate audio to obtain the feedback detection result of the candidate audio;
[0176] Based on the feedback detection results of the candidate audio, sample audio is selected from the candidate audio;
[0177] A howling suppression model is trained based on the sample audio and the near-end speech used to synthesize the sample audio.
[0178] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0179] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute a howling detection method or a howling suppression model training method. The howling detection method includes:
[0180] Obtain the audio to be detected;
[0181] Based on the amplitude variation information of the audio at each frequency point, candidate howling frequency points in the audio are determined;
[0182] Based on the amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, the howling detection result of the audio is determined.
[0183] Training methods for howling suppression models include:
[0184] Candidate audio is synthesized based on near-end speech and noisy audio;
[0185] Based on the aforementioned feedback detection method, feedback detection is performed on the candidate audio to obtain the feedback detection result of the candidate audio;
[0186] Based on the feedback detection results of the candidate audio, sample audio is selected from the candidate audio;
[0187] A howling suppression model is trained based on the sample audio and the near-end speech used to synthesize the sample audio.
[0188] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the training method for the howling detection method or howling suppression model provided by the above methods. The howling detection method includes:
[0189] Obtain the audio to be detected;
[0190] Based on the amplitude variation information of the audio at each frequency point, candidate howling frequency points in the audio are determined;
[0191] Based on the amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, the howling detection result of the audio is determined.
[0192] Training methods for howling suppression models include:
[0193] Candidate audio is synthesized based on near-end speech and noisy audio;
[0194] Based on the aforementioned feedback detection method, feedback detection is performed on the candidate audio to obtain the feedback detection result of the candidate audio;
[0195] Based on the feedback detection results of the candidate audio, sample audio is selected from the candidate audio;
[0196] A howling suppression model is trained based on the sample audio and the near-end speech used to synthesize the sample audio.
[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting howling, characterized in that, include: Obtain the audio to be detected; Based on the amplitude variation information of the audio at each frequency point, candidate howling frequency points in the audio are determined; Based on the amplitude distribution information of the candidate howling frequency points in the audio, and / or the relationship between the amplitude values of the candidate howling frequency points in the audio and the range of amplitude perceived by the human ear, the howling detection result of the audio is determined. The step of determining candidate howling frequency points in the audio based on the amplitude variation information of the audio at each frequency point includes: Based on the amplitude variation trend and amplitude variation acceleration of the audio at each frequency point, candidate howling frequency points and the type of the candidate howling frequency points are determined from each frequency point. The amplitude change trend refers to the direction and rate of change of the amplitude of the frequency point over time; wherein, the frequency point where the amplitude increases at a stable rate is a candidate howling frequency point of the growth type, and the frequency point where the amplitude decreases at a stable rate is a candidate howling frequency point of the attenuation type.
2. The whistling detection method according to claim 1, characterized in that, When the type of the candidate howling frequency point is attenuated howling, the amplitude distribution information of the candidate howling frequency point in the audio includes spectral sparsity; The determination of the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude values of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear, includes: If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
3. The whistling detection method according to claim 2, characterized in that, The amplitude distribution information of the candidate howling frequency points in the audio also includes amplitude dispersion. When the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human auditory perception, determining that the howling detection result includes attenuated howling includes: If the frequency sparsity of the candidate howling frequency point is greater than the sparsity threshold, the amplitude dispersion of the candidate howling frequency point is less than the dispersion threshold, and the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, then the howling detection result is determined to include attenuated howling.
4. The whistling detection method according to claim 1, characterized in that, When the type of the candidate howling frequency point is an increasing howling, determining the howling detection result of the audio based on the amplitude distribution information of the candidate howling frequency point in the audio, and / or the relationship between the amplitude value of the candidate howling frequency point in the audio and the range of amplitude perceived by the human ear, includes: If the amplitude value of the candidate howling frequency point in the audio is within the range of human ear perception amplitude, it is determined that the howling detection result includes an increasing howling.
5. The whistling detection method according to any one of claims 1 to 4, characterized in that, When the audio is acquired based on a loudspeaker system, it also includes: Based on the feedback detection results of the audio and the current gain of the amplification system, determine the critical gain of the amplification system; When the current gain is greater than the critical gain, the loudspeaker system is subjected to howling control.
6. A training method for a howling suppression model, characterized in that, include: Candidate audio is synthesized based on near-end speech and noisy audio; Based on the feedback detection method as described in any one of claims 1 to 5, feedback detection is performed on the candidate audio to obtain the feedback detection result of the candidate audio; Based on the feedback detection results of the candidate audio, sample audio is selected from the candidate audio; A howling suppression model is trained based on the sample audio and the near-end speech used to synthesize the sample audio.
7. The training method for a howling suppression model according to claim 6, characterized in that, The feedback detection result based on the candidate audio, selecting sample audio from the candidate audio, includes: The severity of the feedback of the candidate audio is determined based on the number of feedback frequency points in the feedback detection results. Based on the proportion of howling audio required for training and the howling severity of the candidate audio, sample audio is selected from the candidate audio.
8. A whistling detection device, characterized in that, include: The acquisition unit is used to acquire the audio to be detected; An evaluation unit is used to determine candidate howling frequency points in the audio based on the amplitude variation information of the audio at each frequency point; The detection unit is used to determine the feedback detection result of the audio based on the amplitude distribution information of the candidate feedback frequency points in the audio, and / or the relationship between the amplitude value of the candidate feedback frequency points in the audio and the range of amplitude perceived by the human ear; The evaluation unit is specifically used for: Based on the amplitude variation trend and amplitude variation acceleration of the audio at each frequency point, candidate howling frequency points and the type of the candidate howling frequency points are determined from each frequency point. The amplitude change trend refers to the direction and rate of change of the amplitude of the frequency point over time; wherein, the frequency point where the amplitude increases at a stable rate is a candidate howling frequency point of the growth type, and the frequency point where the amplitude decreases at a stable rate is a candidate howling frequency point of the attenuation type.
9. A training device for a howling suppression model, characterized in that, include: An audio synthesis unit is used to synthesize candidate audio based on near-end speech and noisy audio; The feedback detection unit is used to perform feedback detection on the candidate audio based on the feedback detection method as described in any one of claims 1 to 5, and obtain the feedback detection result of the candidate audio; A sample filtering unit is used to filter sample audio from the candidate audio based on the feedback detection results of the candidate audio; The model training unit is used to train a howling suppression model based on the sample audio and the near-end speech used to synthesize the sample audio.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the howling detection method as described in any one of claims 1 to 5, or the training method for the howling suppression model as described in claim 6 or 7.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the howling detection method as described in any one of claims 1 to 5, or the training method for the howling suppression model as described in claim 6 or 7.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the howling detection method as described in any one of claims 1 to 5, or the training method for the howling suppression model as described in claim 6 or 7.
Citation Information
Patent Citations
Squeaking detection method and device
CN107645696A
Howling detection and suppression method
CN111402911A