An audio howling suppression method based on intelligent algorithm

By performing short-time frequency domain analysis on the microphone input signal and the monitor speaker drive signal, and combining it with an intelligent discrimination model to identify early feedback frequencies, the problem of inaccurate early feedback identification in stage performance scenarios is solved, and effective suppression of early feedback and protection of sound quality are achieved.

CN122493820APending Publication Date: 2026-07-31NANJING BAIYIN INTELLIGENT SOUND TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610920906.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing audio feedback suppression methods have difficulty accurately identifying early feedback components attached to the frequency band of human voice overtones in stage performance scenarios, resulting in effective suppression of vocal timbre errors or missed detection of early feedback.

Method used

By collecting microphone input signals and monitor speaker drive signals for short-time frequency domain decomposition, the human voice overtone attachment degree and monitor feedback consistency are determined. Combined with the frequency point energy growth state and narrowband concentration state, an intelligent discrimination model is used to identify early howling frequencies and perform adaptive narrowband suppression.

Benefits of technology

It improves the accuracy and targeting of early feedback detection, reduces the attenuation of normal singing sound, and enhances the sound stability and listening fidelity during stage sound reinforcement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493820A_ABST
    Figure CN122493820A_ABST
Patent Text Reader

Abstract

This invention discloses an audio feedback suppression method based on intelligent algorithms, relating to the field of speech and audio signal processing technology. The method includes: acquiring the microphone input signal from a performer's handheld microphone and the monitor speaker drive signal, and performing short-time frequency domain decomposition; determining the vocal overtone attachment degree based on the microphone input signal, and determining the monitor feedback consistency degree based on the microphone frequency domain signal and the monitor speaker frequency domain signal; determining candidate feedback frequencies and candidate frequency point feature vectors by combining the frequency point energy growth state and narrowband concentration state; inputting the candidate frequency point feature vectors into a trained intelligent discrimination model to obtain early feedback confidence; performing adaptive narrowband suppression based on the candidate feedback frequencies and early feedback confidence, and outputting the suppressed audio signal. This invention can identify early feedback components attached to the frequency band adjacent to the vocal overtones, reducing vocal timbre error suppression and the probability of missed early feedback detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech and audio signal processing technology, and in particular to an audio howling suppression method based on intelligent algorithms. Background Technology

[0002] In audio amplification scenarios such as stage performances, theater concerts, live events, and conference sound reinforcement, it is usually necessary to use handheld microphones to capture the voices of performers, hosts, or speakers, and then amplify the output through mixing consoles, audio processing equipment, power amplifiers, and speakers. In order for performers to hear their own singing status and accompaniment rhythm in real time, monitor speakers are usually set up in the stage area. The monitor speakers play back the performer's voice, accompaniment, or mixed monitoring signal to the area where the performer is located. Because the monitor speakers are close to the handheld microphones, and the output sound of the monitor speakers may be picked up again by the handheld microphones and enter the sound reinforcement link, when the backflow sound of a certain frequency band is continuously amplified, it is easy to form audio feedback, which affects the stability of the live sound reinforcement, the timbre of the singing, and the audience's listening experience.

[0003] During actual stage performances, the position and direction of the microphone held by the performer are not fixed, but change with the singing movements, such as tilting the head, turning the wrist, turning the body, moving closer to the monitor speaker, or holding the microphone head. Under normal setup, the monitor speaker is usually positioned as close as possible to the low-sensitivity pickup direction of the handheld microphone to reduce the proportion of the monitor speaker output sound picked up by the handheld microphone. However, when the performer moves or changes the posture of holding the microphone, the output sound of the monitor speaker may shift from the original low-sensitivity pickup direction to a higher-sensitivity pickup direction, causing the monitor sound backflow ratio to increase temporarily. At the same time, the long notes, high notes, lingering notes, and vibrato produced by the performer have a stable fundamental frequency and multiple order of vocal overtones. The critical feedback frequency band is easily excited by the current vocal overtones, so that the early howling does not appear as an isolated spike completely independent of the vocal sound, but rather as a narrow band component that slowly increases, gradually becomes sharper, and slightly drifts on the frequency band adjacent to the vocal overtones.

[0004] Existing audio feedback suppression methods typically identify and suppress feedback through fixed frequency point detection, energy threshold judgment, spectral peak search, adaptive notch filtering, or conventional adaptive filtering. These methods are effective when feedback has formed obvious sharp peaks. However, in stage performance scenarios, normal vocal overtones and early feedback have similar time-frequency structures, and both may exhibit continuous narrow-band energy. If judgment is based solely on the magnitude of frequency energy or fixed frequency peaks, normal vocal overtones may be misjudged as feedback and suppressed, resulting in a weakening of the effective vocal timbre. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies, such as the difficulty in accurately identifying early feedback components attached to the frequency band of human voice overtones, which leads to effective suppression of human voice timbre errors or missed detection of early feedback. Therefore, this invention proposes an audio feedback suppression method based on intelligent algorithms.

[0006] To address the problems existing in the prior art, the present invention adopts the following technical solution: An audio feedback suppression method based on intelligent algorithms includes: S1. Collect the microphone input signal of the performer's handheld microphone and the monitor speaker drive signal, and perform short-time frequency domain decomposition on the two to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal. S2. Determine the human voice overtone attachment degree, which characterizes the degree of similarity between the frequency point and the human voice overtone trajectory, based on the microphone input signal; S3. Determine the monitor feedback consistency degree, which is used to characterize the monitor feedback component in the microphone input, based on the microphone frequency domain signal and the monitor speaker frequency domain signal. S4. Based on the human voice overtone attachment degree, the feedback feedback consistency degree, and the frequency point energy growth state and narrowband concentration state determined by the microphone frequency domain signal, determine the candidate howling frequency and the candidate frequency point feature vector. S5. Input the candidate frequency feature vector into the trained intelligent discrimination model to obtain the early howling confidence level; S6. Based on the candidate howling frequency and the early howling confidence level, adaptive narrowband suppression is performed on the microphone input signal to obtain the suppressed audio signal.

[0007] Preferably, the microphone input signal from the performer's handheld microphone and the monitor speaker drive signal are collected, and short-time frequency domain decomposition is performed on both to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal, including: The microphone input signal and the monitor speaker drive signal are framed with the same frame length and frame shift; Windowing processing is applied to the microphone input signal and the monitor speaker drive signal after framing. Short-time Fourier transforms are performed on the windowed microphone input signal and the monitor speaker drive signal to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal with the same frame number and frequency point number.

[0008] Preferably, determining the vocal overtone attachment degree, which characterizes the proximity of a frequency point to the vocal overtone trajectory, based on the microphone input signal includes: Perform normalized autocorrelation calculation on the current frame signal in the microphone input signal to determine the fundamental frequency of the human voice in the current frame signal; The trajectory of human voice overtones in the current frame signal is determined based on the fundamental frequency of the human voice. Determine the minimum frequency distance between each frequency point and the human voice overtone trajectory; Based on the minimum frequency distance and the frequency resolution of the short-time frequency domain decomposition, the attachment degree of human voice overtones at each frequency point is determined.

[0009] Preferably, the monitor-speaker backfeed consistency, used to characterize the monitor-speaker backfeed component in the microphone input, is determined based on the microphone frequency domain signal and the monitor speaker frequency domain signal, including: The cross power spectrum of each frequency point is determined based on the microphone frequency domain signal and the monitor speaker frequency domain signal. The microphone self-power spectrum at each frequency point is determined based on the microphone frequency domain signal, and the monitor speaker self-power spectrum at each frequency point is determined based on the monitor speaker frequency domain signal. Short-time smoothing is performed on the cross-power spectrum, the microphone self-power spectrum, and the monitor speaker self-power spectrum. Based on the cross-power spectrum, microphone self-power spectrum, and monitor speaker self-power spectrum after short-time smoothing, the amplitude squared coherence coefficient of each frequency point is determined, and the amplitude squared coherence coefficient is used as the monitor feedback consistency.

[0010] Preferably, determining the candidate howling frequencies and candidate frequency point feature vectors includes: Based on the microphone frequency domain signal, determine the short-term energy increase of each frequency point relative to the same frequency point in the previous frame; Based on the microphone frequency domain signal, determine the narrowband concentration of each frequency point relative to its adjacent frequency points; The early feedback score of vocal overtone attachment is determined based on the vocal overtone attachment degree, feedback feedback consistency, short-term energy increase and narrowband concentration at the same frequency point. The frequency point with the highest early howling score of the human voice overtone attachment is determined as the candidate howling frequency point, and the candidate howling frequency is determined based on the candidate howling frequency point; The candidate frequency point feature vector is composed of the human voice overtone attachment degree, the echo feedback consistency degree, the short-time energy increase and the narrowband concentration corresponding to the candidate howling frequency point.

[0011] Preferably, the short-term energy increase is determined by the logarithmic difference between the energy of the current frame frequency point and the energy of the same frequency point in the previous frame, and negative values ​​are set to zero; the narrowband concentration is determined by the ratio between the energy of the current frequency point and the average energy of two adjacent frequency points.

[0012] Preferably, the trained intelligent discrimination model is a logistic regression discrimination model, the input of which is the candidate frequency feature vector, and the output is the early feedback confidence score; the training samples of the logistic regression discrimination model include normal singing long tone samples, samples where the performer holds a microphone close to the monitor speaker but does not produce feedback, and samples where the performer holds a microphone close to the monitor speaker and produces early feedback, and candidate frequency feature vectors are extracted for each training sample.

[0013] Preferably, the logistic regression discriminant model uses cross-entropy loss to train the model weights, and the model weights include weights that act on vocal overtone attachment, feedback consistency, short-term energy increase, and narrowband concentration, respectively.

[0014] Preferably, adaptive narrowband suppression is performed on the microphone input signal based on the candidate howling frequencies and the early howling confidence level, including: Using the candidate howling frequency as the center, determine the half-power points on both sides of the candidate peak in the microphone frequency domain signal, and determine the half-power bandwidth based on the half-power points; Construct a second-order adaptive notch filter based on the candidate howling frequency and the half-power bandwidth; The microphone input signal is input into the second-order adaptive notch filter to obtain a notch-processed signal. The early feedback confidence level is used as the mixing weight of the notch filter signal, and the value obtained by subtracting the early feedback confidence level from the unit value is used as the mixing weight of the microphone input signal. The microphone input signal and the notch filter signal are mixed to obtain the suppressed audio signal.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention synchronously acquires microphone input signals and monitor speaker drive signals, and calculates vocal overtone attachment, monitor feedback consistency, short-time energy increase, and narrowband concentration under the same short-time frequency domain coordinates. This makes the determination of candidate feedback frequency points no longer solely dependent on the energy magnitude of a single frequency point or a fixed spectral peak value, but also considers whether the frequency point is close to the current vocal overtone trajectory, whether it is consistent with the output content of the monitor speaker, whether it increases over time, and whether it exhibits a narrowband concentration state. When the performer holds the microphone close to the monitor speaker and a monitor rejection angle mismatch occurs, it can identify early feedback components attached to the frequency band of vocal overtones, reducing the probability of misjudging normal singing long notes, high notes, vocal embellishments, or vibrato as feedback, and improving the targeting and accuracy of early feedback detection. 2. This invention further inputs the candidate frequency feature vector into the trained intelligent discrimination model to obtain the early howling confidence level, and performs adaptive narrowband suppression based on the candidate howling frequency and the early howling confidence level, so that the suppression center can be adjusted with the change of the candidate howling frequency, and the suppression intensity can be continuously changed with the risk level of the early howling; when the candidate frequency is closer to the normal human voice overtones, the system reduces the participation of notch processing, thereby reducing the weakening of the effective human voice timbre; when the candidate frequency simultaneously exhibits human voice overtone attachment, consistent monitor feedback, energy growth and narrowband concentration, the system enhances the suppression of the narrowband component, thereby controlling the howling before it develops into a clearly audible spike, improving the sound stability and listening fidelity during stage monitor sound reinforcement. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating an audio feedback suppression method based on an intelligent algorithm, provided as an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Example: This example provides an audio feedback suppression method based on intelligent algorithms, which is suitable for stage performances, live singing, theater sound reinforcement, and event sound reinforcement scenarios; In the above scenario, the performer uses a handheld microphone to sing, and the monitor speaker is located in front of or slightly below the performer to return the performer's own singing voice and accompaniment sound. During the performance, the performer may move closer to the monitor speaker to hear the monitor sound and may make movements such as lowering their head, turning their wrist, turning their body to the side, or holding the microphone head. This causes the output sound of the monitor speaker to shift from the original low-sensitivity pickup direction of the handheld microphone to a higher-sensitivity pickup direction, resulting in an increased proportion of the monitor speaker output sound re-entering the handheld microphone. Since the performer is singing, the microphone input signal contains a stable fundamental frequency of human voice and multiple order human voice overtones. The critical feedback frequency band is easily excited by the current human voice overtones, causing the early howling to manifest as a narrow band component that slowly intensifies, gradually becomes sharper, and slightly drifts on the frequency band adjacent to the human voice overtones. It should be noted that the monitor rejection angle mismatch in this embodiment refers to the acoustic state in which, during a stage performance, after the performer lowers their head, turns their wrist, turns to the side, moves closer to the monitor speaker, or wraps their hand around the microphone head, the output sound of the monitor speaker is no longer mainly in the low-sensitivity pickup direction of the handheld microphone, but instead enters the higher-sensitivity pickup direction of the handheld microphone for a short period of time, thus increasing the proportion of the monitor speaker output sound being picked up again by the handheld microphone. This state is not directly measured by an additional angle sensor, but is characterized by the frequency domain consistency between the microphone input signal and the monitor speaker drive signal. It should be noted that the early feedback of vocal overtones in this embodiment refers to the early feedback frequency band not appearing as an isolated frequency point completely independent of the singing voice, but rather as a narrow band component with gradually increasing energy, gradually concentrated spectrum, and consistent feedback within the fundamental frequency of the current singing voice and its adjacent frequency bands of multiple overtones; this term is used to describe the audio state in stage performance scenarios where the early feedback and effective vocal overtones are close to each other and easily confused. In this embodiment, the early feedback of human voice overtones under monitor rejection angle mismatch refers to the early feedback state in which, after a short-term change in the relative position and relative direction of the handheld microphone and the monitor speaker, the proportion of the output sound from the monitor speaker entering the handheld microphone increases, and the critical feedback frequency band close to the current human voice overtone trajectory shows energy growth and narrowband concentration.

[0019] In this embodiment, the microphone input signal refers to the time-domain audio signal collected by the performer holding the microphone; the monitor speaker drive signal refers to the time-domain audio signal sent to the monitor speaker, which can characterize the actual content played by the monitor speaker; the microphone frequency domain signal refers to the frequency domain representation obtained after short-time frequency domain decomposition of the microphone input signal; the monitor speaker frequency domain signal refers to the frequency domain representation obtained after short-time frequency domain decomposition of the monitor speaker drive signal; the vocal overtone trajectory refers to the set of multi-order vocal overtone frequencies determined according to the fundamental frequency of the vocal voice in the current frame of the microphone input signal; the vocal overtone attachment degree refers to the degree to which the frequency point is close to the current vocal overtone trajectory; the monitor feedback consistency degree refers to the degree of consistency between the microphone frequency domain signal and the monitor speaker frequency domain signal at the same frequency point; the candidate frequency point feature vector refers to the feature vector used to input the intelligent discrimination model; the early feedback confidence degree refers to the probabilistic output that the candidate frequency point belongs to the vocal overtone attachment type early feedback under the monitor rejection angle mismatch; It should be noted that the frame-level processing in this embodiment refers to using the framed microphone input signal and monitor speaker drive signal as the processing objects, and determining the vocal overtone attachment, monitor feedback consistency, candidate howling frequency, and early howling confidence in each frame. For continuous audio output, adjacent frames are processed continuously using frame shift consistent with short-time frequency domain decomposition, and are superimposed and added after each frame is processed to obtain a continuous suppressed audio signal. Therefore, the candidate howling frequency and early howling confidence obtained in the m-th frame only apply to the microphone input frame signal corresponding to the m-th frame and will not be mixed with the processing parameters of other frames. It should be noted that in this embodiment... A uniform stability constant is set to prevent the denominator from being zero, logarithmic operations from being abnormal, or numerical instability. Based on the effective audio energy level, a positive number less than that effective audio energy level is selected; in the calculation formulas of this embodiment, Keep the same value so that it is only used for numerical stability without changing the relative magnitude between different frequency points.

[0020] The method in this embodiment includes a training phase and an application phase; the training phase is used to obtain the model weights of the intelligent discrimination model, and the application phase is used to perform real-time or near-real-time feedback suppression processing on the microphone input signals during stage performances. During the training phase, three types of training samples were collected. The first type of training sample consisted of normal singing long notes, in which the performer continuously produced long notes, high notes, sustained notes, or vibrato, but did not approach the monitor speaker to form early feedback. The second type of training sample consisted of the performer holding the microphone close to the monitor speaker but without forming feedback. In this type of training sample, there was a situation where the output sound from the monitor speaker entered the microphone, but there was no continuously enhanced narrowband component attached to the frequency band adjacent to the vocal overtones. The third type of training sample consisted of the performer holding the microphone close to the monitor speaker and forming early feedback. In this type of training sample, there was enhanced monitor feedback, and a continuously enhanced and gradually concentrated narrowband component appeared in the frequency band adjacent to the current vocal overtones. The first and second types of training samples were labeled with a sample label of 0, and the third type of training samples were labeled with a sample label of 1. The early feedback state in the third type of training samples was determined through a combination of manual listening and time-frequency plot verification. Manual listening was used to confirm the presence of a sharp feedback precursor or a change in the initial timbre of the feedback caused by monitor backflow in the training sample. Time-frequency plot verification was used to confirm that the candidate frequency points in the training sample simultaneously met the following conditions: the candidate frequency point is close to the current vocal overtone trajectory, the candidate frequency point has frequency domain consistency with the monitor speaker drive signal, and the candidate frequency point is within a certain range in consecutive frames. The training samples show an energy growth trend relative to the previous frame, and the candidate frequency points are narrow-band concentrated relative to adjacent frequency points. Although the first type of training samples contains stable long notes, high notes, vocal embellishments, or vibrato, they do not simultaneously have consistent monitor feedback and a continuous narrow-band enhancement state. Although the second type of training samples have situations where the performer holds the microphone close to the monitor speaker, they do not have a continuous energy growth and narrow-band concentration state attached to the frequency band of the vocal overtones. Through the above labeling method, the training sample labels correspond to the early feedback state of vocal overtones attached under monitor rejection angle mismatch. For each training sample, process it according to S1 to S4 in the following usage phase to extract candidate frequency point feature vectors; use the candidate frequency point feature vectors as input to the logistic regression discriminant model, use the sample labels as the supervised output of the logistic regression discriminant model, and train to obtain the model weights; The output of the logistic regression discriminant model is: ; ; Where i is the training sample number. Let be the model output probability corresponding to the i-th training sample. Let i be the linear combination value corresponding to the i-th training sample. Let be the attachment degree of the human voice overtones corresponding to the candidate howling frequency point of the i-th training sample. Let be the echo feedback consistency corresponding to the candidate howling frequency point of the i-th training sample. Let be the short-term energy increase corresponding to the candidate howling frequency point of the i-th training sample. Let represent the narrowband concentration corresponding to the candidate howling frequency point of the i-th training sample. For bias weights, As the weights applied to the attachment degree of human voice overtones, As a weight that applies to the consistency of the feedback loop, As the weight applied to the short-term energy increase, The weights applied to narrowband concentration. It is a natural constant; It should be noted that during the training phase , , and These are the four features obtained after extracting candidate frequency point feature vectors from the i-th training sample according to usage stages S1 to S4, respectively, and are compared with the features in the usage stages. , , and Representation of similar features at different processing stages.

[0021] The model training uses the cross-entropy loss function: ; in, For training loss, Let be the sample label of the i-th training sample. Let be the model output probability corresponding to the i-th training sample. This represents the summation of training samples in the training sample set; the model weights are obtained by minimizing the cross-entropy loss function. , , , and After training, the logistic regression discriminant model is used as a trained intelligent discriminant model to output the early howling confidence level in the usage phase.

[0022] Usage phase: S1. Collect the microphone input signal of the performer's handheld microphone and the monitor speaker drive signal, and perform short-time frequency domain decomposition on the two to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal. Specifically, during a stage performance, the audio processing equipment simultaneously acquires the microphone input signal collected by the performer's handheld microphone and sends the monitor speaker drive signal to the monitor speaker. It should be noted that the microphone input signal in this embodiment refers to the time-domain audio signal actually collected by the performer holding the microphone during the stage performance. It may also include the performer's vocals, accompaniment leakage, stage ambient sound, and re-entry components formed by the sound output from the monitor speaker re-entering the handheld microphone. The monitor speaker drive signal in this embodiment refers to the time-domain audio signal sent from the monitor output channel of the mixing console, the monitor bus of the digital audio processor, or the input terminal of the power amplifier to the monitor speaker. It is used to characterize the actual content played by the monitor speaker and serves as a reference signal for determining whether the microphone input signal contains monitor speaker re-entry components. The microphone input signal and the monitor speaker drive signal are divided into frames with the same frame length N and the same frame shift H to obtain the m-th frame microphone input frame signal. and the m-th frame monitor speaker drive frame signal Where m is the frame number, q is the intra-frame sampling point number, N is the number of sampling points per frame, and H is the frame shift. , ; For each frame, the microphone input frame signal and each frame of monitor speaker drive frame signal Using window functions respectively Windowing is applied; short-time Fourier transforms are performed on the windowed microphone input frame signal and the monitor speaker drive frame signal respectively to obtain the microphone frequency domain signal corresponding to the m-th frame and the k-th frequency point. and monitor speaker frequency domain signal ; Microphone frequency domain signal The calculation method is as follows: ; in, This represents the microphone frequency domain signal corresponding to the k-th frequency point in the m-th frame. The value of the microphone input frame signal in the m-th frame is taken at the sampling point in the q-th frame. Let be the value of the window function at the sampling point in the q-th frame, N be the number of sampling points per frame, k be the frequency point number, m be the frame number, q be the sampling point number within the frame, j be the imaginary unit, and e be the natural constant. Monitor speaker frequency domain signal The calculation method is as follows: ; in, This represents the frequency domain signal of the monitor speaker corresponding to the k-th frequency point in the m-th frame. The value of the monitor speaker driver frame signal in frame m is taken at the sampling point in frame q. The meanings of the other parameters are the same as those of the microphone frequency domain signal. The corresponding parameters are the same.

[0023] S2. Determine the human overtone attachment degree, which is used to characterize the proximity of the frequency point to the human voice overtone trajectory, based on the microphone input signal; Specifically, for the m-th frame microphone input frame signal Perform normalized autocorrelation calculations to obtain the normalized autocorrelation values ​​of the current frame signal at different delay points. The calculation method is as follows: ; in, For the m-th frame, the microphone input frame signal has a delay of [number] points. Normalized autocorrelation value at , The value of the microphone input frame signal in the m-th frame is taken at the sampling point in the q-th frame. For the m-th frame, the microphone input frame signal is delayed. The value at the corresponding position after each sampling point Where q is the delay point number, and q is the intra-frame sampling point number. This indicates summation over the intra-frame samples involved in the calculation; To avoid interference from zero-delay points and non-vocal periods on the estimation of the fundamental frequency of human voice, this embodiment performs a normalized autocorrelation peak search within a preset fundamental frequency range for human voice. The preset fundamental frequency range for human voice is determined based on the fundamental frequency range of the stage singer's voice and is determined by the sampling frequency. Converted to delay point search range; assuming the lower limit of the preset human voice fundamental frequency range is... The upper limit is The search range for delay points is then... to Furthermore, the search range does not include zero-latency points; within this search range of latency points, the latency point with the largest normalized autocorrelation value is determined as... and according to Determine the fundamental frequency of the human voice in the current frame signal. : ; ; in, This is the lower limit of the preset fundamental frequency range for human voices. To preset the upper limit of the fundamental frequency range of human voice, This represents the number of fundamental frequency period points of the human voice corresponding to the m-th frame. Indicates to make The delay points to obtain the maximum value. The fundamental frequency of the human voice corresponding to the m-th frame. The sampling frequency is used; by performing peak search within the preset fundamental frequency range of human voice, the influence of environmental noise, low-frequency components of accompaniment, and zero-delay autocorrelation peaks on the estimation of the fundamental frequency of human voice can be reduced. It should be noted that the vocal overtone trajectory in this embodiment refers to the set of multi-order vocal overtone frequencies determined based on the fundamental frequency of the current vocal voice in the same frame of microphone input signal; this vocal overtone trajectory is used to characterize the fundamental frequency harmonic structure that is stable in the current singing vocal voice, and serves as the basis for judging whether a certain frequency point is close to the effective vocal overtone adjacent frequency band.

[0024] Based on the fundamental frequency of human voice Determine the trajectory of the human voice overtones in the current frame signal; the h-th order human voice overtone frequency is: ; in, Let h be the overtone frequency of the human voice in the m-th frame, where h is a positive integer, and The frequency does not exceed the Nyquist frequency of the audio signal; the human voice overtone trajectory is composed of the human voice overtone frequencies of each order that satisfy the above conditions within the current frame; The actual frequency corresponding to the kth frequency point for: ; in, This represents the actual frequency corresponding to the k-th frequency point. Frequency point number Sampling frequency, The number of sampling points per frame; Determine the minimum frequency distance between the k-th frequency point and the vocal overtone trajectory of the current frame. : ; in, This represents the minimum frequency distance between the k-th frequency point in the m-th frame and the vocal overtone trajectory in the current frame. This indicates that the minimum value is taken among all human voice overtone orders h that meet the conditions; It should be noted that, in this embodiment, the vocal overtone attachment degree refers to the degree of closeness of a certain frequency point to the vocal overtone trajectory of the current frame. The smaller the minimum frequency distance between the frequency point and any order of vocal overtone frequencies in the current frame, the higher the vocal overtone attachment degree of the frequency point, indicating that the frequency point is more likely to be confused with normal singing vocal overtones. The larger the minimum frequency distance between the frequency point and each order of vocal overtone frequencies in the current frame, the lower the vocal overtone attachment degree of the frequency point.

[0025] Based on minimum frequency distance and frequency resolution of short-time frequency domain decomposition Determine the attachment degree of the human voice overtones at the k-th frequency. : ; ; in, Assign a value to the vocal overtone at the k-th frequency point in the m-th frame. This represents the frequency resolution of the short-time Fourier transform. The larger the value, the closer the k-th frequency point is to the human voice overtone trajectory of the current frame; The smaller the value, the more the k-th frequency point deviates from the vocal overtone trajectory of the current frame.

[0026] S3. Determine the monitor feedback consistency based on the microphone frequency domain signal and the monitor speaker frequency domain signal, which is used to characterize the monitor speaker feedback component in the microphone input. Specifically, based on the microphone frequency domain signal and the monitor speaker frequency domain signal, the cross power spectrum, microphone self power spectrum and monitor speaker self power spectrum of the k-th frequency point of the m-th frame are determined, and the above power spectra are smoothed for a short time. The short-time smoothing can be achieved by averaging the current frame with its adjacent historical frames, so as to reduce the impact of random fluctuations in a single frame on the consistency of monitor feedback. In one specific implementation, short-time smoothing adopts an exponential smoothing method; for the first frame, the cross power spectrum, microphone self-power spectrum, and monitor speaker self-power spectrum are respectively taken as the instantaneous cross power spectrum, instantaneous microphone self-power spectrum, and instantaneous monitor speaker self-power spectrum of the corresponding frame as the initial smoothed power spectrum; for each frame after the first frame, the smoothed power spectrum of the current frame is obtained by weighting the instantaneous power spectrum of the current frame and the smoothed power spectrum of the previous frame; The smoothed power spectrum initialization method for the first frame is as follows: ; ; ; in, The cross-power spectrum at the k-th frequency point in the first frame. The microphone autopower spectrum at the k-th frequency point in the first frame. The self-power spectrum of the monitor speaker at the k-th frequency point in the first frame. This is the microphone frequency domain signal at the k-th frequency point in the first frame. This is the frequency domain signal of the monitor speaker at the k-th frequency point in the first frame. for Conjugate; For each frame after the first frame, the short-time smoothing method for the cross-power spectrum is as follows: ; The short-time smoothing methods for the microphone power spectrum and the monitor speaker power spectrum are as follows: ; ; in, The cross-power spectrum of the k-th frequency point in the m-th frame. The microphone auto-power spectrum at the k-th frequency point in the m-th frame. The self-power spectrum of the monitor speaker at the k-th frequency point in the m-th frame. For the conjugate of the frequency domain signal R(k,m) of the monitor speaker, For smoothing coefficients, Greater than 0 and less than 1, and consistent throughout the same processing; through exponential smoothing, the impact of single-frame noise on the amplitude squared coherence coefficient can be reduced while preserving the feedback variation trend; It should be noted that the monitor feedback consistency in this embodiment refers to the degree of consistency between the microphone frequency domain signal and the monitor speaker frequency domain signal at the same frequency point. It is used to characterize whether the frequency component in the microphone input signal may originate from the feedback of the monitor speaker output sound. When there is a short-term mismatch between the pickup direction of the performer's handheld microphone and the output direction of the monitor speaker, the proportion of the monitor speaker output sound picked up by the handheld microphone increases, and the consistency between the microphone frequency domain signal and the monitor speaker frequency domain signal at the corresponding frequency point is enhanced accordingly.

[0027] Based on the short-time smoothed cross-power spectrum, microphone self-power spectrum, and monitor speaker self-power spectrum, the amplitude squared coherence coefficient at the k-th frequency point in the m-th frame is determined, and this amplitude squared coherence coefficient is used as the monitor feedback consistency. : ; in, The echo feedback consistency at the k-th frequency point of the m-th frame. It is the square of the cross-power spectral amplitude. To prevent the use of a unified stability constant with a denominator of zero; Feedback consistency The physical meaning is that if the microphone input signal at a certain frequency point has a high degree of consistency with the monitor speaker drive signal, then the microphone input signal at that frequency point is more likely to contain the re-entry component of the monitor speaker output sound into the microphone; when the performer holds the microphone close to the monitor speaker and makes movements such as looking down, turning the wrist, turning to the side, or covering the microphone with their head, a short-term mismatch occurs between the actual pickup direction of the handheld microphone and the output direction of the monitor speaker, the proportion of the monitor speaker output sound picked up by the microphone increases, and the consistency of the monitor re-entry at the corresponding frequency point increases.

[0028] S4. Based on the human voice overtone attachment degree, the consistency of the feedback and backfeed, and the frequency point energy growth state and narrowband concentration state determined by the microphone frequency domain signal, determine the candidate howling frequency and the candidate frequency point feature vector. It should be noted that the effective audio frequency point set in this embodiment refers to the set of frequency points that, after removing DC frequency points and Nyquist boundary frequency points from the frequency points obtained by short-time frequency domain decomposition, can simultaneously obtain the left and right adjacent frequency points and participate in the candidate howling frequency search; boundary frequency points that cannot simultaneously obtain the left and right adjacent frequency points are not included in the narrowband concentration calculation, nor are they considered as candidate howling frequency points. It should be noted that, in this embodiment, the short-time energy increase refers to the degree of energy enhancement of the same frequency point in the current frame relative to the previous frame, used to characterize whether the frequency point is in a gradually increasing state; the narrowband concentration in this embodiment refers to the degree of energy concentration of the current frequency point relative to the energy of adjacent frequency points, used to characterize whether the frequency point exhibits a narrowband peak state; the short-time energy increase and the narrowband concentration are used together to distinguish between ordinary human voice overtones and the emerging early narrowband howling components; This embodiment determines candidate howling frequencies within the effective audio frequency point set; since there is no energy at the same frequency point in the previous frame in the first frame, this embodiment calculates the short-term energy increase and determines candidate howling frequencies starting from the second frame; For the second frame and all subsequent frames, based on the microphone frequency domain signal, determine the short-time energy increase of the k-th frequency point in the m-th frame relative to the same frequency point in the previous frame; the short-time energy increase is used to characterize whether the energy at the same frequency point is increasing, and its calculation method is as follows: ; in, This represents the short-time energy increase at the k-th frequency point in the m-th frame. The microphone frequency domain signal at the k-th frequency point in the m-th frame. This refers to the microphone frequency domain signal at the same frequency point in the previous frame. To prevent abnormal logarithmic operations, a uniform stability constant is needed. Represents the natural logarithm; Based on the microphone frequency domain signal, determine the narrowband concentration of the k-th frequency point in the m-th frame relative to its adjacent frequency points; the narrowband concentration is used to characterize whether the current frequency point exhibits a more concentrated narrowband peak state relative to its adjacent frequency points, and its calculation method is as follows: ; in, The narrowband concentration of the k-th frequency point in the m-th frame. This represents the microphone frequency domain signal at the (k-1)th frequency point in the m-th frame. For the microphone frequency domain signal at the (k+1)th frequency point in the m-th frame, a unified stability constant is used to prevent the denominator from being zero; It should be noted that the early howling score based on vocal overtone attachment in this embodiment refers to a comprehensive score obtained by simultaneously considering the degree to which the frequency point is close to the current vocal overtone trajectory, the degree of consistency between the frequency point and the monitor speaker drive signal, the degree of frequency point energy increase over time, and the degree of frequency point concentration relative to adjacent frequency points. This score is not used alone to determine the magnitude of frequency point energy, but rather to screen candidate howling frequency points that simultaneously possess the characteristics of vocal overtone proximity, monitor feedback, energy growth, and narrowband concentration.

[0029] Based on the attachment of human vocal overtones at the same frequency Consistency of Feedback and Refeedback Short-term energy increase and narrowband concentration Determine the early feedback score of the human voice overtone at the k-th frequency point in the m-th frame. : ; in, The score is the early feedback score for the human voice overtones at the k-th frequency point in the m-th frame. The larger the score, the more the frequency point simultaneously meets the following conditions: close to the current human voice overtone trajectory, consistent with the monitor speaker drive signal, increasing energy, and exhibiting a narrowband concentrated state. The frequency point with the highest early feedback score for human voice overtones was determined as the candidate feedback frequency point in the m-th frame. : ; in, Let m be the candidate howling frequency point in the m-th frame. Indicates to make The frequency index where the maximum value is obtained; Based on candidate howling frequency points Determine candidate howling frequencies : ; in, Let be the candidate howling frequencies for the m-th frame. Sampling frequency, The number of sampling points per frame; It should be noted that the candidate frequency feature vector in this embodiment refers to the feature combination extracted for the candidate howling frequency and input into the intelligent discrimination model; the candidate frequency feature vector includes at least the human voice overtone attachment degree, the echo feedback consistency degree, the short-time energy growth amount and the narrowband concentration degree corresponding to the candidate howling frequency, so that the intelligent discrimination model can simultaneously obtain human voice overtone proximity information, echo feedback information, time growth information and frequency domain concentration information.

[0030] The candidate frequency point feature vector is composed of the vocal overtone attachment, monitor feedback consistency, short-time energy increase, and narrowband concentration corresponding to the candidate feedback frequency point. : ; in, Let be the candidate frequency point feature vector of the m-th frame. Assign a value to the human voice overtones corresponding to the candidate howling frequencies. The consistency of the feedback and backfeedback corresponding to the candidate howling frequency points. This represents the short-term energy increase corresponding to the candidate howling frequency points. This represents the narrowband concentration corresponding to the candidate howling frequency points.

[0031] S5. Input the candidate frequency point feature vector into the trained intelligent discrimination model to obtain the early howling confidence. Specifically, the candidate frequency feature vectors are input into the logistic regression discriminant model obtained during the training phase; the candidate frequency feature vectors include the attachment degree of human voice overtones corresponding to the candidate howling frequency points. Consistency of Feedback and Refeedback Short-term energy increase and narrowband concentration ; It should be noted that the early howling confidence in this embodiment refers to the probabilistic discrimination result output by the intelligent discrimination model based on the feature vector of the candidate frequency point; the early howling confidence is used to characterize the probability that the candidate howling frequency point belongs to the early howling state of human voice overtone attachment under the mismatch of the feedback rejection angle, and serves as the source of mixed weights for the notch processing signal and the microphone input signal in the subsequent adaptive narrowband suppression process. The logistic regression discriminant model calculates the early howling confidence score of the m-th frame based on the candidate frequency point feature vectors. : ; ; in, Let be the confidence level of the early howling in the m-th frame. The linear combination value corresponding to the feature vector of the candidate frequency points in the m-th frame. These are the bias weights obtained during the training phase. The weights applied to the attachment degree of vocal overtones, obtained during the training phase. The weights obtained during the training phase that affect the consistency of the feedback loop. The weights obtained during the training phase that affect short-term energy growth. The weights applied to narrowband concentration are obtained during the training phase; Early howling confidence The value ranges from 0 to 1; early howling confidence level The larger the value, the more closely the candidate feedback frequency point matches the early feedback state of human voice overtones under the condition of feedback rejection angle mismatch; early feedback confidence level The smaller the value, the more likely the candidate feedback frequency point is to belong to normal human voice overtones or audio components that do not require significant suppression.

[0032] S6. Based on the candidate howling frequency and early howling confidence, perform adaptive narrowband suppression on the microphone input signal and output the suppressed audio signal. It should be noted that the half-power bandwidth in this embodiment refers to the frequency width between the positions on both sides of the candidate peak value around the candidate howling frequency that reach the half-power level; this half-power bandwidth is used to determine the bandwidth of the second-order adaptive notch filter so that the suppression range of the notch filter corresponds to the actual width of the narrowband peak value near the candidate howling frequency point. By candidate howling frequency Centered on the microphone frequency domain signal, half-power points are determined on both sides of the candidate peak, and the half-power bandwidth is determined based on these half-power points; specifically, the candidate howling frequency points are... The energy at a specific frequency point is determined as the candidate peak energy, and half of the candidate peak energy is determined as the half-power level; from the candidate howling frequency point The search proceeds in both low-frequency and high-frequency directions, identifying the points where half-power level is first reached on both sides as the left and right half-power points. If no point reaching half-power level is found within the effective audio frequency set in either the low-frequency or high-frequency direction, the boundary of the effective audio frequency set in that direction is identified as the corresponding half-power point. When the left and right half-power points coincide, or the bandwidth determined by the two is less than a frequency resolution, a frequency resolution is identified as the half-power bandwidth. Thus, the bandwidth parameters of the second-order adaptive notch filter can still be determined even when the candidate peak is close to the boundary of the effective audio frequency set or when the region adjacent to the candidate peak has not dropped to half-power level. It should be noted that the adaptive narrowband suppression in this embodiment refers to a narrowband audio processing method that determines the suppression center frequency based on the candidate howling frequency and determines the degree of suppression participation based on the early howling confidence. This processing method does not uniformly attenuate the entire audio frequency band, but suppresses the narrowband components near the candidate howling frequency and controls the mixing ratio between the original microphone input frame signal and the notch-processed frame signal through the early howling confidence.

[0033] Based on candidate howling frequencies and half-power bandwidth Constructing a second-order adaptive notch filter : ; ; ; in, Let z be the second-order adaptive notch filter corresponding to the m-th frame, and z be the Z-transform variable. Candidate howling frequencies The corresponding digital angular frequency, Let B(m) be the pole radius of the second-order adaptive notch filter, and B(m) be the half-power bandwidth. Sampling frequency, Pi The cosine function is represented; the zeros of this second-order adaptive notch filter are determined by the candidate howling frequency, and its pole radius is determined by the half-power bandwidth, so that the notch center can correspond to the candidate howling frequency and the notch width can correspond to the bandwidth of the candidate peak. The microphone input frame signal corresponding to the m-th frame. Input the second-order adaptive notch filter corresponding to the m-th frame. The notch-processed frame signal of the m-th frame is obtained. ;in, This is the frame signal after suppressing the narrowband components near the candidate howling frequency in the m-th frame; The early howling confidence level P(m) of the m-th frame is used as the notch processing frame signal of the m-th frame. The mixed weights are calculated by subtracting the early howling confidence of the m-th frame from the unit value. The obtained value is used as the microphone input frame signal for the m-th frame. The mixed weights are applied to the microphone input frame signal of the m-th frame. and the notch-filtered frame signal of the mth frame The signals are mixed to obtain the suppressed frame signal of the m-th frame. : ; in, To suppress the signal of the following frame in the m-th frame, The microphone input frame signal for the m-th frame. For the m-th frame, notch filtering is applied to the frame signal. Suppressed frame signal of consecutive frames The frame shifts corresponding to the framing and windowing in S1 are overlapped and added to obtain a continuous suppressed audio signal. Through this processing method, the candidate howling frequency, half-power bandwidth and early howling confidence determined at the frame level are all applied to the corresponding m-th frame microphone input frame signal, and a continuous audio output is formed by overlapping and adding them.

[0034] When the confidence level of early feedback is low, it indicates that the candidate feedback frequency point is more likely to belong to normal human voice overtones or the feedback component that does not form early feedback. At this time, the microphone input frame signal of the m-th frame is... Suppress the post-frame signal in the m-th frame The proportion in the m-th frame is relatively high, and the notch-filtered frame signal in the m-th frame is processed. Suppress the post-frame signal in the m-th frame The proportion of [something] is relatively low, thus reducing the impact on normal singing timbre; when the confidence level of early feedback is high, it indicates that the candidate feedback frequency points simultaneously have the states of vocal overtone attachment, consistent monitor feedback, energy growth, and narrowband concentration. At this time, the notch-filtered frame signal of the m-th frame is [something]. Suppress the post-frame signal in the m-th frame The higher proportion of this component enhances the suppression of early howling components attached to the frequency bands adjacent to the human voice overtones.

[0035] The method in this embodiment can identify feedback not only based on the frequency energy magnitude when a performer holds a microphone close to the monitor speaker and a rejection angle mismatch occurs, but also simultaneously identify candidate feedback frequencies using vocal overtone attachment, monitor feedback consistency, short-term energy increase, and narrowband concentration. It outputs early feedback confidence through a trained intelligent discrimination model and then performs adaptive narrowband suppression based on the candidate feedback frequencies and early feedback confidence. This method can suppress early feedback components attached to the frequency band of the current vocal overtone and reduce the probability of normal vocal overtones being falsely suppressed.

[0036] In an alternative implementation, the fundamental frequency estimation of human voice can be achieved by using the cepstral method or the YIN algorithm instead of the normalized autocorrelation method. As long as the output can be used to determine the trajectory of human voice overtones in the current frame signal and further obtain the human voice overtone attachment degree, the human voice overtone attachment recognition purpose of this embodiment can be achieved. In another alternative implementation, the intelligent discrimination model can use support vector machine, decision tree model or gradient boosting model to replace the logistic regression discrimination model, but its inputs include the human voice overtone attachment degree, echo feedback consistency degree, short-term energy growth amount and narrowband concentration degree corresponding to the candidate howling frequency point, and the output is the early howling confidence degree or the discrimination result that can be converted into the early howling confidence degree. In another alternative implementation, adaptive narrowband suppression can use frequency domain gain attenuation or parametric equalization filtering instead of a second-order adaptive notch filter, but the suppression center frequency should be determined based on the candidate howling frequency, and the suppression intensity should be determined based on the early howling confidence level, so as to maintain the correspondence with the early howling state of vocal overtone attachment under the monitor rejection angle mismatch.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An audio feedback suppression method based on intelligent algorithms, characterized in that, Includes the following steps: S1. Collect the microphone input signal of the performer's handheld microphone and the monitor speaker drive signal, and perform short-time frequency domain decomposition on the two to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal. S2. Determine the human voice overtone attachment degree, which characterizes the degree of similarity between the frequency point and the human voice overtone trajectory, based on the microphone input signal; S3. Determine the monitor feedback consistency degree, which is used to characterize the monitor feedback component in the microphone input, based on the microphone frequency domain signal and the monitor speaker frequency domain signal. S4. Based on the human voice overtone attachment degree, the feedback feedback consistency degree, and the frequency point energy growth state and narrowband concentration state determined by the microphone frequency domain signal, determine the candidate howling frequency and the candidate frequency point feature vector. S5. Input the candidate frequency point feature vector into the trained intelligent discrimination model to obtain the early howling confidence level; S6. Based on the candidate howling frequency and the early howling confidence level, adaptive narrowband suppression is performed on the microphone input signal to obtain the suppressed audio signal.

2. The audio feedback suppression method based on intelligent algorithms according to claim 1, characterized in that, The microphone input signal from the performer's handheld microphone and the monitor speaker drive signal are collected. Short-time frequency domain decomposition is performed on both to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal, including: The microphone input signal and the monitor speaker drive signal are framed with the same frame length and frame shift; Windowing processing is applied to the microphone input signal and the monitor speaker drive signal after framing. Short-time Fourier transforms are performed on the windowed microphone input signal and the monitor speaker drive signal to obtain the microphone frequency domain signal and the monitor speaker frequency domain signal with the same frame number and frequency point number.

3. The audio feedback suppression method based on intelligent algorithms according to claim 1, characterized in that, Determining the vocal overtone attachment degree, which characterizes the proximity of a frequency point to the vocal overtone trajectory, based on the microphone input signal, includes: Perform normalized autocorrelation calculation on the current frame signal in the microphone input signal to determine the fundamental frequency of the human voice in the current frame signal; The trajectory of human voice overtones in the current frame signal is determined based on the fundamental frequency of the human voice. Determine the minimum frequency distance between each frequency point and the human voice overtone trajectory; Based on the minimum frequency distance and the frequency resolution of the short-time frequency domain decomposition, the attachment degree of human voice overtones at each frequency point is determined.

4. The audio feedback suppression method based on intelligent algorithms according to claim 1, characterized in that, Based on the microphone frequency domain signal and the monitor speaker frequency domain signal, determine the monitor-speaker backfeed consistency, which characterizes the monitor speaker backfeed component in the microphone input, including: The cross power spectrum of each frequency point is determined based on the microphone frequency domain signal and the monitor speaker frequency domain signal. The microphone self-power spectrum at each frequency point is determined based on the microphone frequency domain signal, and the monitor speaker self-power spectrum at each frequency point is determined based on the monitor speaker frequency domain signal. Short-time smoothing is performed on the cross-power spectrum, the microphone self-power spectrum, and the monitor speaker self-power spectrum. Based on the cross-power spectrum, microphone self-power spectrum, and monitor speaker self-power spectrum after short-time smoothing, the amplitude squared coherence coefficient of each frequency point is determined, and the amplitude squared coherence coefficient is used as the monitor feedback consistency.

5. The audio feedback suppression method based on intelligent algorithms according to claim 1, characterized in that, Determine candidate howling frequencies and candidate frequency point feature vectors, including: Based on the microphone frequency domain signal, determine the short-term energy increase of each frequency point relative to the same frequency point in the previous frame; Based on the microphone frequency domain signal, determine the narrowband concentration of each frequency point relative to its adjacent frequency points; The early feedback score of human voice overtone attachment, feedback feedback consistency, short-term energy increase and narrowband concentration at the same frequency point are determined. The frequency point with the highest early howling score of the human voice overtone attachment is determined as the candidate howling frequency point, and the candidate howling frequency is determined based on the candidate howling frequency point; The candidate frequency point feature vector is composed of the human voice overtone attachment degree, the echo feedback consistency degree, the short-time energy increase and the narrowband concentration corresponding to the candidate howling frequency point.

6. The audio feedback suppression method based on intelligent algorithms according to claim 5, characterized in that, The short-term energy increase is determined by the logarithmic difference between the energy of the current frame frequency point and the energy of the same frequency point in the previous frame, and negative values ​​are set to zero; the narrowband concentration is determined by the ratio between the energy of the current frequency point and the average energy of two adjacent frequency points.

7. The audio feedback suppression method based on intelligent algorithms according to claim 1, characterized in that, The trained intelligent discrimination model is a logistic regression discrimination model. The input of the logistic regression discrimination model is the candidate frequency feature vector, and the output is the early feedback confidence. The training samples of the logistic regression discrimination model include normal singing long tone samples, samples where the performer holds a microphone close to the monitor speaker but does not produce feedback, and samples where the performer holds a microphone close to the monitor speaker and produces early feedback. The candidate frequency feature vector is extracted for each training sample.

8. The audio feedback suppression method based on intelligent algorithms according to claim 7, characterized in that, The logistic regression discriminant model uses cross-entropy loss to train the model weights, which include weights that act on vocal overtone attachment, feedback consistency, short-term energy increase, and narrowband concentration, respectively.

9. The audio feedback suppression method based on intelligent algorithms according to claim 7, characterized in that, Based on the candidate howling frequencies and the early howling confidence levels, adaptive narrowband suppression is performed on the microphone input signal, including: Using the candidate howling frequency as the center, determine the half-power points on both sides of the candidate peak in the microphone frequency domain signal, and determine the half-power bandwidth based on the half-power points; Construct a second-order adaptive notch filter based on the candidate howling frequency and the half-power bandwidth; The microphone input signal is input into the second-order adaptive notch filter to obtain a notch-processed signal. The early feedback confidence level is used as the mixing weight of the notch filter signal, and the value obtained by subtracting the early feedback confidence level from the unit value is used as the mixing weight of the microphone input signal. The microphone input signal and the notch filter signal are mixed to obtain the suppressed audio signal.