Howling suppression method, device, equipment and medium

Howling is eliminated by using a self-built beamformer and adaptive filter using multiple microphones, solving the high cost and distortion problems caused by hardware dependence and achieving efficient and low-cost howling suppression.

CN115831140BActive Publication Date: 2025-10-03AI TING ZHI NENG KE JI (SHEN ZHEN) YOU XIAN GONG SI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211393760.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-10-03
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

Existing howling suppression methods rely on hardware, resulting in high system cost, high power consumption, poor adaptability, and sound distortion.

Method used

A self-built beamformer and adaptive filter based on multiple microphones are used to achieve howling suppression by eliminating interference from the direction of the speaker.

Benefits of technology

It reduces sound distortion while suppressing howling, saves deployment costs, and has good adaptability, suitable for a variety of software and hardware platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831140B_ABST
    Figure CN115831140B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech processing technology and provides a howling suppression method, apparatus, device, and medium. These methods utilize a beamformer built from multiple microphones to eliminate interference from the speaker's direction, while simultaneously utilizing an adaptive filter built from path vectors to eliminate any remaining interference, thereby achieving howling suppression. This howling suppression process not only reduces sound distortion but also saves deployment costs by eliminating the need for additional hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a howling suppression method, device, equipment and medium. Background Art

[0002] In acoustic scenarios involving public address systems, particularly in conferences, classrooms, and Karaoke (Karaoke TV), howling can occur when a closed acoustic feedback loop forms. Howling is a form of positive feedback, caused by the speaker's sound being picked up again by the microphone, generating self-excitation. Furthermore, the likelihood of howling increases with increasing sound system volume.

[0003] Howling not only affects hearing, but also burns out audio equipment. Therefore, howling suppression is widely used in daily life.

[0004] Currently, howling suppression is mainly achieved through the following hardware methods:

[0005] (1) Use a microphone with low sensitivity and high directivity.

[0006] By reducing the microphone's perception of sound intensity, the sound from the speaker is blocked from being transmitted to the microphone, thereby reducing the probability of howling.

[0007] (2) Use a hardware frequency shifter.

[0008] By increasing or decreasing the frequency component of the microphone input audio signal (such as increasing or decreasing a certain frequency point by 5 to 10 Hz), the conditions for howling are destroyed and the frequency output signal is changed. In this way, when the audio is transmitted to the microphone and amplifier system again, it will not be superimposed on the original signal frequency, thereby achieving howling suppression.

[0009] The aforementioned howling suppression methods all rely heavily on hardware, resulting in high system cost, high power consumption, and large size. In particular, using frequency shifters to suppress howling can cause significant distortion, resulting in a stiff and less smooth vocal sound and poor adaptability. Summary of the Invention

[0010] In view of the above, it is necessary to provide a howling suppression method, apparatus, device and medium that can reduce sound distortion when performing howling suppression while saving deployment costs.

[0011] A howling suppression method, the howling suppression method comprising:

[0012] In response to a howling suppression instruction for a target device, acquiring audio data of the target device;

[0013] Preprocessing the audio data to obtain data to be processed;

[0014] Creating a beamformer, and eliminating first interference from a direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data;

[0015] Creating an adaptive filter, and eliminating second interference in the first output data based on the adaptive filter to obtain target speech;

[0016] The target speech is outputted using the target device.

[0017] According to a preferred embodiment of the present invention, obtaining audio data of the target device includes:

[0018] Acquire multiple microphones of the target device, and acquire a speaker of the target device;

[0019] collecting audio signals from the plurality of microphones and the loudspeaker as the audio data;

[0020] Wherein, the multiple microphones include an in-ear microphone.

[0021] According to a preferred embodiment of the present invention, preprocessing the audio data to obtain data to be processed includes:

[0022] Performing frame processing on the audio data with a preset number of sampling points as one frame to obtain first data;

[0023] Performing overlap-addition on the first data according to a preset window length to obtain second data;

[0024] Determining the preset window length as the frame length;

[0025] Performing a discrete Fourier transform on the second data with the frame length, the preset frame shift, and the preset window function as parameters to obtain the data to be processed.

[0026] According to a preferred embodiment of the present invention, eliminating first interference from the direction of the speaker in the data to be processed based on the beamformer to obtain first output data includes:

[0027] Create a beamforming coefficient matrix;

[0028] Obtaining feedback paths between the plurality of microphones and the loudspeaker by conducting a test in a free field;

[0029] Calculating a product of a transposed matrix of the beamforming coefficient matrix and the feedback path to obtain third data;

[0030] obtaining estimated paths between the plurality of microphones and the speaker;

[0031] Calculating a difference between the third data and the estimated path to obtain fourth data;

[0032] Create a forward path gain function;

[0033] Calculating a product of the forward path gain function and the fourth data to obtain an intermediate function;

[0034] Calculating the difference between 1 and the intermediate function to obtain a closed-loop transfer function;

[0035] When the value of the fourth data is 0, optimizing the variable corresponding to the fourth data in the closed-loop transfer function by using a least squares method to obtain an estimated value of the beamforming coefficient matrix;

[0036] Acquire a sum of input signals and feedback signals collected by the multiple microphones as fifth data;

[0037] calculating a transpose of an estimated value of the beamforming coefficient matrix as sixth data;

[0038] Calculating the product of the fifth data and the sixth data to obtain the first output data;

[0039] The estimated value of the beamforming coefficient matrix is ​​expressed as follows:

[0040]

[0041] Among them, B LS represents the estimated value of the beamforming coefficient matrix, represents the convolution matrix between the feedback paths corresponding to other microphones except the in-ear microphone, Indicates the feedback path corresponding to the in-ear microphone.

[0042] According to a preferred embodiment of the present invention, creating a beamforming coefficient matrix includes:

[0043] Obtaining a beamforming coefficient submatrix for each microphone in the plurality of microphones;

[0044] Construct a matrix using the beamforming coefficient submatrix of each microphone as an element to obtain an intermediate matrix;

[0045] A transposed matrix of the intermediate matrix is ​​calculated to obtain the beamforming coefficient matrix.

[0046] According to a preferred embodiment of the present invention, eliminating the second interference in the first output data based on the adaptive filter to obtain the target speech includes:

[0047] Create a whitening filter;

[0048] performing whitening processing on the first output data using the whitening filter to obtain a first output signal;

[0049] Obtaining a sampling signal from the speaker;

[0050] Using the whitening filter to perform whitening processing on the sampled signal of the loudspeaker to obtain a second output signal;

[0051] Obtaining estimated paths between the plurality of microphones and the loudspeaker in a previous frame as current estimated paths;

[0052] Calculating a product of the current estimated path and the second output signal as a first product;

[0053] calculating a difference between the first output signal and the first product to obtain seventh data;

[0054] Calculating a product of the second output signal and the seventh data to obtain eighth data;

[0055] Get the change step size and configuration constants;

[0056] Calculating the product of the change step length and the eighth data to obtain ninth data;

[0057] calculating a product of a transposed signal of the second output signal and the second output signal to obtain a second product;

[0058] Calculating the sum of the configuration constant and the second product to obtain a first sum value;

[0059] Calculating a quotient of the ninth data and the first sum to obtain a tenth data;

[0060] Calculating a sum of the current estimated path and the tenth data to obtain an estimated path between the plurality of microphones and the speaker in a current frame, and using the estimated path as a target estimated path;

[0061] Calculating a product of the target estimated path and the echo signal of the speaker to obtain a third product;

[0062] Calculating a difference between the first output data and the third product to obtain the target speech;

[0063] The value range of the configuration constant is (0, 1).

[0064] According to a preferred embodiment of the present invention, before obtaining the change step size and the configuration constant, the method further includes:

[0065] Obtaining the frequency response value of the estimated path of each frame in the previous preset frame;

[0066] Get the frequency response value of the estimated path of the current frame;

[0067] Calculating the Euclidean distance of each frame according to the frequency response value of the estimated path of each frame and the frequency response value of the estimated path of the current frame;

[0068] Get the distance threshold and initial step length;

[0069] comparing the Euclidean distance of each frame with the distance threshold;

[0070] When the Euclidean distance of any frame is greater than the distance threshold, determining that the adjustment condition for the initial step size is met;

[0071] The value of the initial step size is updated based on the adjustment condition until convergence to obtain the changed step size.

[0072] A howling suppression device, comprising:

[0073] an acquiring unit, configured to acquire audio data of a target device in response to a howling suppression instruction to the target device;

[0074] A processing unit, configured to pre-process the audio data to obtain data to be processed;

[0075] an elimination unit, configured to create a beamformer, and eliminate first interference from a direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data;

[0076] The elimination unit is further configured to create an adaptive filter and eliminate the second interference in the first output data based on the adaptive filter to obtain the target speech;

[0077] An output unit is configured to output the target speech using the target device.

[0078] A computer device, comprising:

[0079] a memory storing at least one instruction; and

[0080] The processor executes the instructions stored in the memory to implement the howling suppression method.

[0081] A computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the howling suppression method.

[0082] As can be seen from the above technical solution, the present invention can eliminate interference from the speaker direction using a beamformer built from multiple microphones, while simultaneously eliminating any residual interference using an adaptive filter built from path vectors, thereby achieving howling suppression. This howling suppression process not only reduces sound distortion but also saves deployment costs by eliminating the need for additional hardware. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is a flow chart of a preferred embodiment of the howling suppression method of the present invention.

[0084] Figure 2 It is a functional module diagram of a preferred embodiment of the howling suppression device of the present invention.

[0085] Figure 3 It is a structural diagram of a computer device according to a preferred embodiment of the present invention for implementing the howling suppression method. DETAILED DESCRIPTION

[0086] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0087] like Figure 1 FIG. 1 is a flow chart of a preferred embodiment of the howling suppression method of the present invention. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.

[0088] The howling suppression method is applied to one or more computer devices, which are devices that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Their hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0089] The computer device may be any electronic product that can interact with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.

[0090] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0091] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0092] Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0093] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0094] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0095] S10 , in response to a howling suppression instruction to a target device, acquiring audio data of the target device.

[0096] In this embodiment, the target device is an audio device. For example, the target device may include, but is not limited to, wearable devices such as headphones and hearing aids, and other voice communication devices.

[0097] In this embodiment, the howling suppression instruction may be automatically triggered when it is detected that the target device has input data, so as to achieve real-time suppression of the howling phenomenon.

[0098] In this embodiment, obtaining the audio data of the target device includes:

[0099] Acquire multiple microphones of the target device, and acquire a speaker of the target device;

[0100] collecting audio signals from the plurality of microphones and the loudspeaker as the audio data;

[0101] Wherein, the multiple microphones include an in-ear microphone.

[0102] For example, if the target device is a headset, the headset structure may include an in-ear microphone, two external microphones, and a speaker. Each microphone may have a sampling frequency of 16 kHz and a sampling bit rate of 16 bits. In other words, the audio data of the target device includes the input signals of the three microphones and the output signal of the one speaker.

[0103] In this embodiment, in order to better achieve howling suppression, a howling signal can be simulated first. A person holds a device such as a hearing aid in his hand so that a strong howling phenomenon can be generated when the microphone is close to the speaker. At this time, the serial communication interface can be used to import the data into a designated terminal (such as a personal computer) and save it as a general audio file.

[0104] Furthermore, the stored analog howling signal was analyzed. Specifically, according to the Nyquist-Shannon sampling theorem, at a 16kHz sampling rate, the effective frequency of the audio signal is 0-8kHz. This analysis shows that howling can occur at all frequencies. Howling is a self-excited process, where speech energy at a certain frequency is transferred from the speaker to the microphone and then amplified by the speaker. Therefore, the speech energy of the howling increases, and unless the sound field environment changes, the howling will persist.

[0105] Based on the above analysis of the simulated howling signal, we can better determine how to suppress the howling. That is, starting from the conditions that cause howling, destroying any of the conditions that cause howling can effectively suppress howling.

[0106] S11, pre-processing the audio data to obtain data to be processed.

[0107] In this embodiment, preprocessing the audio data to obtain data to be processed includes:

[0108] Performing frame processing on the audio data with a preset number of sampling points as one frame to obtain first data;

[0109] Performing overlap-addition on the first data according to a preset window length to obtain second data;

[0110] Determining the preset window length as the frame length;

[0111] Performing a discrete Fourier transform on the second data with the frame length, the preset frame shift, and the preset window function as parameters to obtain the data to be processed.

[0112] The preset number may be 256. Then the processing time for each frame is: 1000ms / (16000 / 256)=16ms, where 16000 means 16000 points are sampled in 1s, and 256 means 256 points are sampled per frame.

[0113] The preset window length may be 512.

[0114] The preset frame shift may be 256.

[0115] The preset window function may include, but is not limited to, a Hanning window, a Hamming window, a flat-top window, an exponential window, etc., which is not limited in the present invention.

[0116] In the preprocessing described above, the overlap-add method avoids spectral leakage and aliasing caused by speech framing, achieving better spectral resolution. This helps the howling suppression algorithm distinguish between howling and human voices, achieving better suppression without compromising speech intelligibility. Furthermore, the Fourier transform allows data processing in the frequency domain, reducing algorithm computational overhead.

[0117] S12: Create a beamformer, and eliminate first interference from the direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data.

[0118] In this embodiment, the beamformer can be used to eliminate interference from the direction of the speaker.

[0119] In this embodiment, eliminating first interference from the direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data includes:

[0120] Create the beamforming coefficient matrix B(q);

[0121] Obtaining feedback paths H(q,n) between the plurality of microphones and the loudspeaker by conducting a test in a free field;

[0122] Calculate the transposed matrix B of the beamforming coefficient matrix T The product of (q) and the feedback path H(q,n) is used to obtain the third data B T (q)H(q,n);

[0123] Obtaining estimated paths between the multiple microphones and the speaker

[0124] Calculate the third data B T (q)H(q,n) and the estimated path The difference between

[0125] Creating a forward path gain function G(q,n); wherein the forward path gain function is a pre-configured system function obtained by superimposing a series of functions, representing the comprehensive processing of multiple algorithms (such as noise reduction (NS), wide dynamic range compression (WDRC), and EQ (equalize)). The system function acts on the microphone input signal, which is equivalent to multiplying the input signal by several sets of linear and nonlinear coefficient matrices with the same length as the input signal;

[0126] Calculate the forward path gain function G(q,n) and the fourth data The product of , we get the intermediate function

[0127] Calculate 1 with the intermediate function The closed-loop transfer function is obtained by

[0128] When the value of the fourth data is 0, that is, , optimizing the variables corresponding to the fourth data in the closed-loop transfer function using the least squares method to obtain an estimated value B of the beamforming coefficient matrix LS ;

[0129] Acquire a sum of input signals and feedback signals collected by the plurality of microphones as fifth data y[n];

[0130] Calculate the transpose of the estimated value of the beamforming coefficient matrix as the sixth data B LS T ;

[0131] Calculate the product of the fifth data and the sixth data to obtain the first output data

[0132]

[0133] The estimated value of the beamforming coefficient matrix is ​​expressed as follows:

[0134]

[0135] Among them, B LS represents the estimated value of the beamforming coefficient matrix, represents the convolution matrix between the feedback paths corresponding to other microphones except the in-ear microphone, Indicates the feedback path corresponding to the in-ear microphone.

[0136] Among them, the feedback path can be measured in advance according to the hardware structure.

[0137] The step of creating a beamforming coefficient matrix includes:

[0138] Obtaining a beamforming coefficient submatrix for each microphone in the plurality of microphones;

[0139] Construct a matrix using the beamforming coefficient submatrix of each microphone as an element to obtain an intermediate matrix;

[0140] A transposed matrix of the intermediate matrix is ​​calculated to obtain the beamforming coefficient matrix.

[0141] For example, the beamforming coefficient matrix can be expressed as: B(q)=[B1(q)…B M (q)] T .

[0142] Wherein, B(q) represents the beamforming coefficient matrix, B M (q) represents the beamforming coefficient submatrix of the Mth microphone, [B1(q)…B M (q)] represents the intermediate matrix.

[0143] In all the above formulas, q represents the delay of the filter at the corresponding order, and n represents the corresponding frame.

[0144] S13: Create an adaptive filter, and eliminate second interference in the first output data based on the adaptive filter to obtain target speech.

[0145] Ideally, a beamformer based on multiple microphones can steer the beam suppression direction toward the speaker located in the inner ear, thereby completely eliminating the howling feedback in the ear without affecting the input signal.

[0146] However, changes in the acoustic feedback path will affect the performance of the fixed-steering beamformer, so an adaptive filter needs to be added after the beamformer to eliminate possible residual howling signals.

[0147] Specifically, this embodiment creates an adaptive filter and eliminates residual feedback signals from the first output data obtained after processing by the beamformer. However, due to the strong correlation between the first output data and the speaker signal, according to the principle of the adaptive filter, the magnitude of the correlation matrix between the in-ear microphone signal and the speaker signal determines the magnitude of the estimated bias. To minimize the estimation bias caused by the correlation between the microphone and speaker, a whitening filter is required to whiten the signal.

[0148] Specifically, eliminating the second interference in the first output data based on the adaptive filter to obtain the target speech includes:

[0149] Create a whitening filter The whitening filter may be estimated using the Levinson-Durbin method (fast recursive method) in the Prediction-Error-Method (PEM), for example, the whitening filter may be an all-pole filter.

[0150] Using the whitening filter The first output data Perform whitening processing to obtain the first output signal

[0151] Obtaining the sampling signal u[n] of the loudspeaker;

[0152] Using the whitening filter The speaker's collected signal u[n] is whitened to obtain a second output signal

[0153] Obtain the estimated paths between the multiple microphones and the speaker in the previous frame as the current estimated path

[0154] Calculate the current estimated path With the second output signal u pw The product of [n] is taken as the first product

[0155] Calculate the first output signal Multiplying the first The difference between

[0156] Calculate the second output signal u pw [n] and the seventh data e pw The product of [n] is the eighth data u pw [n]e pw [n];

[0157] Get the change step size μ and configuration constant α;

[0158] Calculate the change step μ and the eighth data u pw [n]e pw The product of [n] is the ninth data μ*u pw [n]e pw [n];

[0159] Calculate the transposed signal of the second output signal With the second output signal u pw [n] to get the second product

[0160] Calculate the configuration constant α and the second product The sum of , get the first sum value

[0161] Calculate the ninth data μ*u pw [n]e pw [n] and the first sum value The quotient of

[0162] Calculate the current estimated path With the tenth data The sum of the two is used to obtain the estimated path between the multiple microphones and the speaker in the current frame, and the estimated path is used as the target estimated path

[0163] Calculate the estimated path to the target The product of the sampling signal u[n] of the speaker is obtained as the third product

[0164] Calculate the first output data Multiplying the third The target speech is obtained by

[0165] The value range of the configuration constant is (0, 1).

[0166] Through the self-constructed adaptive filter, the interference not eliminated by the beamformer can be further eliminated, thereby effectively suppressing howling with less distortion.

[0167] In this embodiment, before obtaining the change step size and the configuration constant, the method further includes:

[0168] Get the frequency response value of the estimated path of each frame in the previous preset frame; for example, you can get the frequency response value X of the estimated path of each frame in the first 10 frames i , i = 10;

[0169] Get the frequency response value Y of the estimated path of the current frame i ;

[0170] The Euclidean distance of each frame is calculated based on the frequency response value of the estimated path of each frame and the frequency response value of the estimated path of the current frame.

[0171] Get the distance threshold T limit , and initial step length D;

[0172] Compare the Euclidean distance d(x,y) of each frame with the distance threshold T limit ;

[0173] When the Euclidean distance d(x,y) of any frame is greater than the distance threshold T limit , determining that the adjustment condition for the initial step length D is met;

[0174] The value of the initial step size D is updated based on the adjustment condition until convergence, thereby obtaining the change step size μ.

[0175] The distance threshold can be obtained by analyzing and processing historical data. For example, after the shape of a headset is determined, a better distance threshold can be determined based on the range of continuously calculated Euclidean distances.

[0176] It can be understood that when the Euclidean distance d(x,y) of any frame is greater than the distance threshold T limit When , it means that the current path has undergone a significant change. Therefore, the step size is continuously updated to speed up convergence and realize the modification of the change step size.

[0177] S14: Output the target voice using the target device.

[0178] Specifically, the self-built beamformer and adaptive filter are deployed to the target device, which not only has good adaptability and a wider frequency domain, but also has a smaller size and is easy to deploy on most software and hardware platforms with low deployment cost.

[0179] Moreover, when performing howling suppression, since this embodiment has low requirements on memory computing power, it also has a very positive effect on reducing latency, has high accuracy, and is more suitable for use in devices including but not limited to hearing aids and wearable devices.

[0180] The target speech finally obtained through this embodiment effectively suppresses howling.

[0181] As can be seen from the above technical solution, the present invention can eliminate interference from the speaker direction using a beamformer built from multiple microphones, while simultaneously eliminating any residual interference using an adaptive filter built from path vectors, thereby achieving howling suppression. This howling suppression process not only reduces sound distortion but also saves deployment costs by eliminating the need for additional hardware.

[0182] like Figure 2, which is a functional block diagram of a preferred embodiment of a howling suppression device according to the present invention. The howling suppression device 11 includes an acquisition unit 110, a processing unit 111, a cancellation unit 112, and an output unit 113. As used herein, a module / unit refers to a series of computer program segments that can be executed by a processor and perform fixed functions, and is stored in a memory. The functions of each module / unit in this embodiment will be described in detail in subsequent embodiments.

[0183] The acquiring unit 110 is configured to acquire audio data of a target device in response to a howling suppression instruction for the target device.

[0184] In this embodiment, the target device is an audio device. For example, the target device may include, but is not limited to, wearable devices such as headphones and hearing aids, and other voice communication devices.

[0185] In this embodiment, the howling suppression instruction may be automatically triggered when it is detected that the target device has input data, so as to achieve real-time suppression of the howling phenomenon.

[0186] In this embodiment, the acquiring unit 110 acquires the audio data of the target device including:

[0187] Acquire multiple microphones of the target device, and acquire a speaker of the target device;

[0188] collecting audio signals from the plurality of microphones and the loudspeaker as the audio data;

[0189] Wherein, the multiple microphones include an in-ear microphone.

[0190] For example, if the target device is a headset, the headset structure may include an in-ear microphone, two external microphones, and a speaker. Each microphone may have a sampling frequency of 16 kHz and a sampling bit rate of 16 bits. In other words, the audio data of the target device includes the input signals of the three microphones and the output signal of the one speaker.

[0191] In this embodiment, in order to better achieve howling suppression, a howling signal can be simulated first. A person holds a device such as a hearing aid in his hand so that a strong howling phenomenon can be generated when the microphone is close to the speaker. At this time, the serial communication interface can be used to import the data into a designated terminal (such as a personal computer) and save it as a general audio file.

[0192] Furthermore, the stored analog howling signal was analyzed. Specifically, according to the Nyquist-Shannon sampling theorem, at a 16kHz sampling rate, the effective frequency of the audio signal is 0-8kHz. This analysis shows that howling can occur at all frequencies. Howling is a self-excited process, where speech energy at a certain frequency is transferred from the speaker to the microphone and then amplified by the speaker. Therefore, the speech energy of the howling increases, and unless the sound field environment changes, the howling will persist.

[0193] Based on the above analysis of the simulated howling signal, we can better determine how to suppress the howling. That is, starting from the conditions that cause howling, destroying any of the conditions that cause howling can effectively suppress howling.

[0194] The processing unit 111 is configured to pre-process the audio data to obtain data to be processed.

[0195] In this embodiment, the processing unit 111 pre-processes the audio data to obtain data to be processed, including:

[0196] Performing frame processing on the audio data with a preset number of sampling points as one frame to obtain first data;

[0197] Performing overlap-addition on the first data according to a preset window length to obtain second data;

[0198] Determining the preset window length as the frame length;

[0199] Performing a discrete Fourier transform on the second data with the frame length, the preset frame shift, and the preset window function as parameters to obtain the data to be processed.

[0200] The preset number may be 256. Then the processing time for each frame is: 1000ms / (16000 / 256)=16ms, where 16000 means 16000 points are sampled in 1s, and 256 means 256 points are sampled per frame.

[0201] The preset window length may be 512.

[0202] The preset frame shift may be 256.

[0203] The preset window function may include, but is not limited to, a Hanning window, a Hamming window, a flat-top window, an exponential window, etc., which is not limited in the present invention.

[0204] In the preprocessing described above, the overlap-add method avoids spectral leakage and aliasing caused by speech framing, achieving better spectral resolution. This helps the howling suppression algorithm distinguish between howling and human voices, achieving better suppression without compromising speech intelligibility. Furthermore, the Fourier transform allows data processing in the frequency domain, reducing algorithm computational overhead.

[0205] The elimination unit 112 is configured to create a beamformer and eliminate first interference from the direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data.

[0206] In this embodiment, the beamformer can be used to eliminate interference from the direction of the loudspeaker.

[0207] In this embodiment, the elimination unit 112 eliminates the first interference from the direction of the speaker in the data to be processed based on the beamformer, and obtains the first output data including:

[0208] Create the beamforming coefficient matrix B(q);

[0209] Obtaining feedback paths H(q,n) between the plurality of microphones and the loudspeaker by conducting a test in a free field;

[0210] Calculate the transposed matrix B of the beamforming coefficient matrix T The product of (q) and the feedback path H(q,n) is used to obtain the third data B T (q)H(q,n);

[0211] Obtaining estimated paths between the multiple microphones and the speaker

[0212] Calculate the third data B T (q)H(q,n) and the estimated path The difference between

[0213] Creating a forward path gain function G(q,n); wherein the forward path gain function is a pre-configured system function obtained by superimposing a series of functions, representing the comprehensive processing of multiple algorithms (such as noise reduction (NS), wide dynamic range compression (WDRC), and EQ (equalize)). The system function acts on the microphone input signal, which is equivalent to multiplying the input signal by several sets of linear and nonlinear coefficient matrices with the same length as the input signal;

[0214] Calculate the forward path gain function G(q,n) and the fourth data The product of , we get the intermediate function

[0215] Calculate 1 with the intermediate function The closed-loop transfer function is obtained by

[0216] When the value of the fourth data is 0, that is, , optimizing the variables corresponding to the fourth data in the closed-loop transfer function using the least squares method to obtain an estimated value B of the beamforming coefficient matrix LS ;

[0217] Acquire a sum of input signals and feedback signals collected by the plurality of microphones as fifth data y[n];

[0218] Calculate the transpose of the estimated value of the beamforming coefficient matrix as the sixth data B LS T ;

[0219] Calculate the product of the fifth data and the sixth data to obtain the first output data

[0220]

[0221] The estimated value of the beamforming coefficient matrix is ​​expressed as follows:

[0222]

[0223] Among them, B LS represents the estimated value of the beamforming coefficient matrix, represents the convolution matrix between the feedback paths corresponding to other microphones except the in-ear microphone, Indicates the feedback path corresponding to the in-ear microphone.

[0224] Among them, the feedback path can be measured in advance according to the hardware structure.

[0225] The step of creating a beamforming coefficient matrix includes:

[0226] Obtaining a beamforming coefficient submatrix for each microphone in the plurality of microphones;

[0227] Construct a matrix using the beamforming coefficient submatrix of each microphone as an element to obtain an intermediate matrix;

[0228] A transposed matrix of the intermediate matrix is ​​calculated to obtain the beamforming coefficient matrix.

[0229] For example, the beamforming coefficient matrix can be expressed as: B(q)=[B1(q)…B M (q)] T .

[0230] Wherein, B(q) represents the beamforming coefficient matrix, B M (q) represents the beamforming coefficient submatrix of the Mth microphone, [B1(q)…B M (q)] represents the intermediate matrix.

[0231] In all the above formulas, q represents the delay of the filter at the corresponding order, and n represents the corresponding frame.

[0232] The elimination unit 112 is further configured to create an adaptive filter and eliminate the second interference in the first output data based on the adaptive filter to obtain the target speech.

[0233] Ideally, a beamformer based on multiple microphones can steer the beam suppression direction toward the speaker located in the inner ear, thereby completely eliminating the howling feedback in the ear without affecting the input signal.

[0234] However, changes in the acoustic feedback path will affect the performance of the fixed-steering beamformer, so an adaptive filter needs to be added after the beamformer to eliminate possible residual howling signals.

[0235] Specifically, this embodiment creates an adaptive filter and eliminates residual feedback signals from the first output data obtained after processing by the beamformer. However, due to the strong correlation between the first output data and the speaker signal, according to the principle of the adaptive filter, the magnitude of the correlation matrix between the in-ear microphone signal and the speaker signal determines the magnitude of the estimated bias. To minimize the estimation bias caused by the correlation between the microphone and speaker, a whitening filter is required to whiten the signal.

[0236] Specifically, the elimination unit 112 eliminates the second interference in the first output data based on the adaptive filter to obtain the target speech, including:

[0237] Create a whitening filter The whitening filter may be estimated using the Levinson-Durbin method (fast recursive method) in the Prediction-Error-Method (PEM), for example, the whitening filter may be an all-pole filter.

[0238] Using the whitening filter The first output data Perform whitening processing to obtain the first output signal

[0239] Obtaining the sampling signal u[n] of the loudspeaker;

[0240] Using the whitening filter The speaker's collected signal u[n] is whitened to obtain a second output signal

[0241] Obtain the estimated paths between the multiple microphones and the speaker in the previous frame as the current estimated path

[0242] Calculate the current estimated path With the second output signal u pw The product of [n] is taken as the first product

[0243] Calculate the first output signal Multiplying the first The difference between

[0244] Calculate the second output signal u pw [n] and the seventh data e pw The product of [n] is the eighth data u pw [n]e pw [n];

[0245] Get the change step size μ and configuration constant α;

[0246] Calculate the change step μ and the eighth data u pw [n]e pw The product of [n] is the ninth data μ*u pw [n]e pw [n];

[0247] Calculate the transposed signal of the second output signal With the second output signal u pw [n] to get the second product

[0248] Calculate the configuration constant α and the second product The sum of , get the first sum value

[0249] Calculate the ninth data μ*u pw [n]e pw [n] and the first sum value The quotient of

[0250] Calculate the current estimated path With the tenth data The sum of the two is used to obtain the estimated path between the multiple microphones and the speaker in the current frame, and the estimated path is used as the target estimated path

[0251] Calculate the estimated path to the target The product of the sampling signal u[n] of the speaker is obtained as the third product

[0252] Calculate the first output data Multiplying the third The target speech is obtained by

[0253] The value range of the configuration constant is (0, 1).

[0254] Through the self-constructed adaptive filter, the interference not eliminated by the beamformer can be further eliminated, thereby effectively suppressing howling with less distortion.

[0255] In this embodiment, before obtaining the change step size and the configuration constant, the frequency response value of the estimated path of each frame in the previous preset frame is obtained; for example, the frequency response value X of the estimated path of each frame in the previous 10 frames can be obtained. i , i = 10;

[0256] Get the frequency response value Y of the estimated path of the current frame i ;

[0257] The Euclidean distance of each frame is calculated based on the frequency response value of the estimated path of each frame and the frequency response value of the estimated path of the current frame.

[0258] Get the distance threshold T limit , and initial step length D;

[0259] Compare the Euclidean distance d(x,y) of each frame with the distance threshold T limit ;

[0260] When the Euclidean distance d(x,y) of any frame is greater than the distance threshold T limit , determining that the adjustment condition for the initial step length D is met;

[0261] The value of the initial step size D is updated based on the adjustment condition until convergence, thereby obtaining the change step size μ.

[0262] The distance threshold can be obtained by analyzing and processing historical data. For example, after the shape of a headset is determined, a better distance threshold can be determined based on the range of continuously calculated Euclidean distances.

[0263] It can be understood that when the Euclidean distance d(x,y) of any frame is greater than the distance threshold T limit When , it means that the current path has undergone a significant change. Therefore, the step size is continuously updated to speed up convergence and realize the modification of the change step size.

[0264] The output unit 113 is configured to output the target speech using the target device.

[0265] Specifically, the self-built beamformer and adaptive filter are deployed to the target device, which not only has good adaptability and a wider frequency domain, but also has a smaller size and is easy to deploy on most software and hardware platforms with low deployment cost.

[0266] Moreover, when performing howling suppression, since this embodiment has low requirements on memory computing power, it also has a very positive effect on reducing latency, has high accuracy, and is more suitable for use in devices including but not limited to hearing aids and wearable devices.

[0267] The target speech finally obtained through this embodiment effectively suppresses howling.

[0268] As can be seen from the above technical solution, the present invention can eliminate interference from the speaker direction using a beamformer built from multiple microphones, while simultaneously eliminating any residual interference using an adaptive filter built from path vectors, thereby achieving howling suppression. This howling suppression process not only reduces sound distortion but also saves deployment costs by eliminating the need for additional hardware.

[0269] like Figure 3 FIG. 1 is a schematic diagram of the structure of a computer device according to a preferred embodiment of the howling suppression method of the present invention.

[0270] The computer device 1 may include a memory 12 , a processor 13 , and a bus, and may further include a computer program stored in the memory 12 and executable on the processor 13 , such as a howling suppression program.

[0271] Those skilled in the art will understand that the schematic diagram is merely an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 may have either a bus structure or a star structure. The computer device 1 may also include more or less other hardware or software than shown in the figure, or a different arrangement of components. For example, the computer device 1 may also include input and output devices, network access devices, etc.

[0272] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that are suitable for the present invention should also be included in the scope of protection of the present invention and included here by reference.

[0273] The memory 12 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 12 may be an internal storage unit of the computer device 1, such as a mobile hard disk of the computer device 1. In other embodiments, the memory 12 may also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 1. Furthermore, the memory 12 may include both an internal storage unit and an external storage device of the computer device 1. The memory 12 may be used not only to store application software and various types of data installed in the computer device 1, such as the code of the howling suppression program, but also to temporarily store data that has been output or is to be output.

[0274] In some embodiments, the processor 13 may be comprised of an integrated circuit, such as a single packaged integrated circuit or a combination of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core (Control Unit) of the computer device 1, connecting the various components of the entire computer device 1 using various interfaces and circuits. It executes or runs programs or modules stored in the memory 12 (e.g., executing a howling suppression program) and accesses data stored in the memory 12 to perform various functions and process data.

[0275] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-mentioned various howling suppression method embodiments, for example Figure 1 Steps shown.

[0276] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, a processing unit 111, an elimination unit 112, and an output unit 113.

[0277] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module stored in a storage medium includes a number of instructions for causing a computer device (which can be a personal computer, computer equipment, or network equipment, etc.) or a processor to execute the portion of the howling suppression method described in various embodiments of the present invention.

[0278] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the processes in the above-mentioned method embodiments by instructing relevant hardware devices through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments.

[0279] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory, etc.

[0280] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0281] Blockchain, as used in this article, refers to a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each block contains information about a batch of online transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.

[0282] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The figure shows that only one straight line is used, but it does not mean that there is only one bus or one type of bus. The bus is configured to realize the connection and communication between the memory 12 and at least one processor 13.

[0283] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be detailed here.

[0284] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.

[0285] Optionally, the computer device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the computer device 1 and to display a visual user interface.

[0286] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0287] Figure 3 Only the computer device 1 having components 12-13 is shown, and it can be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0288] Combine Figure 1 The memory 12 in the computer device 1 stores a plurality of instructions to implement a howling suppression method, and the processor 13 can execute the plurality of instructions to implement:

[0289] In response to a howling suppression instruction for a target device, acquiring audio data of the target device;

[0290] Preprocessing the audio data to obtain data to be processed;

[0291] Creating a beamformer, and eliminating first interference from a direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data;

[0292] Creating an adaptive filter, and eliminating second interference in the first output data based on the adaptive filter to obtain target speech;

[0293] The target speech is outputted using the target device.

[0294] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.

[0295] It should be noted that the data involved in this case were all obtained legally.

[0296] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division, and actual implementation may employ other division methods.

[0297] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0298] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0299] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0300] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0301] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0302] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the present invention may also be implemented by a single unit or device through software or hardware. Terms such as first and second are used to indicate names and do not imply any particular order.

[0303] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A howling suppression method, characterized in that: The howling suppression method includes: In response to a howling suppression instruction to a target device, the step of acquiring audio data of the target device includes: Acquire multiple microphones of the target device, and acquire a speaker of the target device; collecting audio signals from the plurality of microphones and the loudspeaker as the audio data; Wherein, the plurality of microphones includes an in-ear microphone; Preprocessing the audio data to obtain data to be processed; Creating a beamformer, and eliminating first interference from a direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data; Creating an adaptive filter, and eliminating second interference in the first output data based on the adaptive filter to obtain target speech; The target speech is outputted using the target device.

2. The howling suppression method according to claim 1, wherein: The preprocessing of the audio data to obtain data to be processed comprises: Performing frame processing on the audio data with a preset number of sampling points as one frame to obtain first data; Performing overlap-addition on the first data according to a preset window length to obtain second data; Determining the preset window length as the frame length; Performing a discrete Fourier transform on the second data with the frame length, the preset frame shift, and the preset window function as parameters to obtain the data to be processed.

3. The howling suppression method according to claim 1, wherein: Eliminating first interference from a direction of a loudspeaker in the data to be processed based on the beamformer to obtain first output data includes: Create a beamforming coefficient matrix; Obtaining feedback paths between the plurality of microphones and the loudspeaker by conducting a test in a free field; Calculating a product of a transposed matrix of the beamforming coefficient matrix and the feedback path to obtain third data; obtaining estimated paths between the plurality of microphones and the speaker; Calculating a difference between the third data and the estimated path to obtain fourth data; Create a forward path gain function; Calculating a product of the forward path gain function and the fourth data to obtain an intermediate function; Calculating the difference between 1 and the intermediate function to obtain a closed-loop transfer function; When the value of the fourth data is 0, optimizing the variable corresponding to the fourth data in the closed-loop transfer function by using a least squares method to obtain an estimated value of the beamforming coefficient matrix; Acquire a sum of input signals and feedback signals collected by the multiple microphones as fifth data; calculating a transpose of an estimated value of the beamforming coefficient matrix as sixth data; Calculating the product of the fifth data and the sixth data to obtain the first output data; The estimated value of the beamforming coefficient matrix is ​​expressed as follows: in, represents the estimated value of the beamforming coefficient matrix, represents the convolution matrix between the feedback paths corresponding to other microphones except the in-ear microphone, Indicates the feedback path corresponding to the in-ear microphone.

4. The howling suppression method according to claim 3, wherein: The creating of the beamforming coefficient matrix includes: Obtaining a beamforming coefficient submatrix for each microphone in the plurality of microphones; Construct a matrix using the beamforming coefficient submatrix of each microphone as an element to obtain an intermediate matrix; A transposed matrix of the intermediate matrix is ​​calculated to obtain the beamforming coefficient matrix.

5. The howling suppression method according to claim 1, wherein: Eliminating the second interference in the first output data based on the adaptive filter to obtain the target speech includes: Create a whitening filter; performing whitening processing on the first output data using the whitening filter to obtain a first output signal; Obtaining a sampling signal from the speaker; Using the whitening filter to perform whitening processing on the sampled signal of the loudspeaker to obtain a second output signal; Obtaining estimated paths between the plurality of microphones and the loudspeaker in a previous frame as current estimated paths; Calculating a product of the current estimated path and the second output signal as a first product; calculating a difference between the first output signal and the first product to obtain seventh data; Calculating a product of the second output signal and the seventh data to obtain eighth data; Get the change step size and configuration constants; Calculating the product of the change step length and the eighth data to obtain ninth data; calculating a product of a transposed signal of the second output signal and the second output signal to obtain a second product; Calculating the sum of the configuration constant and the second product to obtain a first sum value; Calculating a quotient of the ninth data and the first sum to obtain a tenth data; Calculating a sum of the current estimated path and the tenth data to obtain an estimated path between the plurality of microphones and the speaker in a current frame, and using the estimated path as a target estimated path; Calculating a product of the target estimated path and the echo signal of the speaker to obtain a third product; Calculating a difference between the first output data and the third product to obtain the target speech; The value range of the configuration constant is (0, 1).

6. The howling suppression method according to claim 5, wherein: Before obtaining the change step size and the configuration constant, the method further includes: Obtaining the frequency response value of the estimated path of each frame in the previous preset frame; Get the frequency response value of the estimated path of the current frame; Calculating the Euclidean distance of each frame according to the frequency response value of the estimated path of each frame and the frequency response value of the estimated path of the current frame; Get the distance threshold and initial step length; comparing the Euclidean distance of each frame with the distance threshold; When the Euclidean distance of any frame is greater than the distance threshold, determining that the adjustment condition for the initial step size is met; The value of the initial step size is updated based on the adjustment condition until convergence to obtain the changed step size.

7. A howling suppression device, characterized in that: The howling suppression device comprises: an acquiring unit, configured to acquire audio data of the target device in response to a howling suppression instruction to the target device, wherein the acquiring unit is further configured to: Acquire multiple microphones of the target device, and acquire a speaker of the target device; collecting audio signals from the plurality of microphones and the loudspeaker as the audio data; Wherein, the plurality of microphones includes an in-ear microphone; A processing unit, configured to pre-process the audio data to obtain data to be processed; an elimination unit, configured to create a beamformer, and eliminate first interference from a direction of the loudspeaker in the data to be processed based on the beamformer to obtain first output data; The elimination unit is further configured to create an adaptive filter and eliminate the second interference in the first output data based on the adaptive filter to obtain the target speech; An output unit is configured to output the target speech using the target device.

8. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the howling suppression method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the howling suppression method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio signal processing method and device, electronic equipment and storage medium

    CN114664321A