Audio signal processing method, device, readable storage medium and electronic device
By using a combination of high and low delay adaptive filters in adaptive echo cancellation to judge and adjust the filtering quality, the problem of adaptive filter overfitting is solved and the processing accuracy and quality of audio signals are improved.
Patent Information
- Application Number
- CN202211334465.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Existing adaptive echo cancellation methods require a strong convergence speed when the scene environment changes, which leads to overfitting of the adaptive filter, resulting in poor audio signal quality and even oscillation.
A first adaptive filter with high signal processing delay and a second adaptive filter with low signal processing delay are used. By comparing the difference in signal transfer gain between the two filters, the filtering quality is judged. If it is unqualified, the parameters of the second adaptive filter are adjusted to reduce the risk of overfitting.
This effectively reduces the risk of overfitting of the adaptive filter, improves the processing accuracy and quality of the audio signal, and improves the audio effect of speaker playback.
Smart Images

Figure CN115691542B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method, device, computer-readable storage medium, and electronic device for processing an audio signal. Background Art
[0002] Adaptive echo cancellation, with its excellent adaptive tracking capabilities and strong suppression capabilities, offers significant advantages in preventing howling during audio playback. The basic concept of adaptive echo cancellation is to estimate the characteristic parameters of the echo path to generate a simulated echo signal, which is then subtracted from the received signal to achieve echo cancellation. Because the echo path is often unknown and time-varying, this technology typically uses an adaptive filter as an echo canceller to simulate the echo path. The primary goal is to accurately estimate the echo path parameters while rapidly tracking changes in the echo path.
[0003] Conventional adaptive echo cancellation methods require a high convergence speed to quickly track changes in the impulse response function within a scene (such as a car interior) when the scene environment changes (for example, due to occlusion between the microphone and speaker, or relative movement between the microphone and speaker). However, a high convergence speed can lead to overfitting of the adaptive filter when the amplitude of the sound frequency captured by the microphone is small, resulting in poor audio signal quality and even "oscillation" of the sound energy, where it fluctuates. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a method, apparatus, computer-readable storage medium, and electronic device for processing an audio signal.
[0005] An embodiment of the present disclosure provides a method for processing an audio signal, the method comprising: acquiring a target audio signal, wherein the target audio signal comprises a microphone signal and an echo reference signal; adaptively filtering the target audio signal using a first adaptive filter and a second adaptive filter, respectively, wherein the signal processing delay of the first adaptive filter is greater than the signal processing delay of the second adaptive filter; converting the first adaptive filter based on the signal processing delay of the second adaptive filter to obtain a converted first adaptive filter; determining a first signal transfer gain of the converted first adaptive filter and a second signal transfer gain of the second adaptive filter; determining a first filtering quality inspection result of the second adaptive filter based on a difference between the first signal transfer gain and the second signal transfer gain; in response to the first filtering quality inspection result being unqualified, adjusting parameters of the second adaptive filter based on the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality; and processing the target audio signal based on the second adaptive filter with qualified filtering quality.
[0006] According to another aspect of an embodiment of the present disclosure, a device for processing an audio signal is provided, the device comprising: an acquisition module for acquiring a target audio signal, wherein the target audio signal comprises a microphone signal and an echo reference signal; a first filtering module for adaptively filtering the target audio signal using a first adaptive filter and a second adaptive filter, respectively, wherein a signal processing delay of the first adaptive filter is greater than a signal processing delay of the second adaptive filter; a conversion module for converting the first adaptive filter based on the signal processing delay of the second adaptive filter to obtain a converted first adaptive filter; a first determination module for determining a first signal transfer gain of the converted first adaptive filter and a second signal transfer gain of the second adaptive filter; a second determination module for determining a first filtering quality inspection result of the second adaptive filter based on a difference between the first signal transfer gain and the second signal transfer gain; a first adjustment module for adjusting parameters of the second adaptive filter based on parameters of the converted first adaptive filter in response to an unqualified first filtering quality inspection result to obtain a second adaptive filter with qualified filtering quality; and a second filtering module for processing the target audio signal based on the second adaptive filter with qualified filtering quality.
[0007] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program is used to execute the above-mentioned audio signal processing method.
[0008] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the above-mentioned audio signal processing method.
[0009] Based on the audio signal processing method, device, computer-readable storage medium and electronic device provided by the above-mentioned embodiments of the present disclosure, the target audio signal is adaptively filtered by a first adaptive filter with high signal processing delay and a second adaptive filter with low signal processing delay, and then the first adaptive filter is converted based on the signal processing delay of the second adaptive filter, and the difference between the first signal transfer gain and the second signal transfer gain corresponding to the converted first adaptive filter and the second adaptive filter is determined, and the first filtering quality inspection result of the second adaptive filter is determined according to the search; in response to the first filtering quality inspection result being unqualified, the parameters of the second adaptive filter are adjusted based on the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality, and finally the target audio signal is processed based on the second adaptive filter with qualified filtering quality. The disclosed embodiment implements the addition of a high-latency first adaptive filter on the basis of echo cancellation through a low-latency second adaptive filter. By comparing the difference in signal transfer gains of the two filters, the filtering quality of the second adaptive filter can be effectively judged. If the filtering quality is unqualified, the parameters of the second adaptive filter are adjusted to obtain a second filter with qualified filtering quality, thereby effectively reducing the risk of overfitting of the second adaptive filter, and then performing high-precision filtering processing on the target audio signal, thereby improving the quality of the audio played by the speaker.
[0010] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 is a system diagram to which the present disclosure is applicable.
[0013] Figure 2 It is a flowchart of a method for processing an audio signal provided by an exemplary embodiment of the present disclosure.
[0014] Figure 3FIG. 4 is a flowchart of a method for processing an audio signal provided by another exemplary embodiment of the present disclosure.
[0015] Figure 4 FIG. 4 is a flowchart of a method for processing an audio signal provided by another exemplary embodiment of the present disclosure.
[0016] Figure 5 FIG. 4 is a flowchart of a method for processing an audio signal provided by another exemplary embodiment of the present disclosure.
[0017] Figure 6 FIG. 4 is a flowchart of a method for processing an audio signal provided by another exemplary embodiment of the present disclosure.
[0018] Figure 7 FIG. 4 is an exemplary schematic diagram of a beam enhancement direction according to an exemplary embodiment of the present disclosure.
[0019] Figure 8 FIG. 4 is a flowchart of a method for processing an audio signal provided by another exemplary embodiment of the present disclosure.
[0020] Figure 9 FIG. 4 is a structural diagram of an audio signal processing apparatus provided by an exemplary embodiment of the present disclosure.
[0021] Figure 10 FIG. 4 is a structural diagram of an audio signal processing apparatus provided by another exemplary embodiment of the present disclosure.
[0022] Figure 11 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0024] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0025] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.
[0026] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0027] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0028] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.
[0029] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.
[0030] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0031] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0032] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0033] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0034] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.
[0035] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.
[0036] Application Overview
[0037] In the field of audio processing, there are scenarios where real-time audio processing is crucial, such as when a user sings or speaks into a microphone. In these scenarios, the block length of the audio signal processed by the adaptive filtering algorithm must be kept very short to reduce the delay caused by audio analysis and synthesis.
[0038] In some scenarios, the fundamental frequency and harmonics of the audio signal collected by the microphone remain largely unchanged for an extended period. For example, when a user prolongs a sound, the signal processed by the adaptive filter at the previous moment and played back through the speaker may not differ much from the current sound produced by the user. This indicates that the reference signal and the near-end collected signal have the same source, leading to algorithm overfitting. Overfitting can degrade the sound quality of the sound played back by the speaker, especially by causing energy oscillations in low-frequency signals and harsh howling in high-frequency signals.
[0039] To solve this problem, an embodiment of the present disclosure provides a method for audio signal processing, by setting two adaptive filters with high signal processing delay and low signal processing delay, comparing the difference in signal transfer gain of the two filters to determine whether overfitting occurs, and adjusting the parameters of the low signal processing delay filter when overfitting occurs, thereby reducing the risk of overfitting.
[0040] Exemplary Systems
[0041] Figure 1 An exemplary system architecture 100 is shown to which the audio signal processing method or the audio signal processing apparatus according to the embodiments of the present disclosure can be applied.
[0042] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, a server 103, a microphone 104, and a speaker 105. Network 102 is used to provide a medium for a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0043] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as audio playback applications, singing applications, instant messaging tools, etc.
[0044] The microphone 104 is used to collect audio signals including the user's voice, and the speaker is used to play audio signals obtained by filtering the audio signals collected by the microphone.
[0045] The terminal device 101 can be a variety of electronic devices, including but not limited to mobile terminals such as vehicle-mounted terminals, mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and fixed terminals such as smart home appliances, digital TVs, and desktop computers. The terminal device 101 can obtain audio signals collected from the microphone 104 and filter the collected audio signals, or send the collected audio signals to the server 103, which then filters the audio signals. The terminal device 101 can then send the filtered audio signals to the speaker 105 for playback.
[0046] Server 103 can be a server that provides various services, such as a background audio processing server that filters the audio signal uploaded by terminal device 101. The background audio processing server can process the received audio signal and perform other operations such as filtering quality inspection on the adaptive filter. Server 103 can also feed back the filtered audio signal to terminal device 101.
[0047] It should be noted that the audio signal processing method provided in the embodiment of the present disclosure can be executed by the server 103 or by the terminal device 101. Accordingly, the audio signal processing device can be set in the server 103 or in the terminal device 101.
[0048] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above description is merely illustrative. Any number of terminal devices, networks, and servers may be used as needed. If the audio signal does not need to be acquired remotely, the above system architecture may include only terminal devices, microphones, and speakers, without including the network and servers.
[0049] Exemplary Methods
[0050] Figure 2 This is a flowchart of a method for processing an audio signal provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices (such as Figure 1On the terminal device 101 or server 103 shown in FIG. Figure 2 As shown, the method includes the following steps:
[0051] Step 201: Acquire a target audio signal.
[0052] In this embodiment, the electronic device can obtain the target audio signal locally or remotely, wherein the target audio signal includes a microphone signal and an echo reference signal.
[0053] The microphone signal can be as follows Figure 1 The audio signal collected by the microphone 104 shown in the figure, the echo reference signal can be sent to the target time period before the current moment. Figure 1 The audio signal played by speaker 105, i.e., the echo reference signal, is the audio signal obtained after filtering by the second adaptive filter during the target time period. When speaker 105 plays the audio signal, microphone 104 simultaneously captures the user's voice and the sound played by the speaker. The collected sound played by the speaker is the echo signal. The purpose of this embodiment is to estimate the echo signal component in the microphone signal through the second adaptive filter, remove this echo signal component from the microphone signal, and obtain the user's voice signal.
[0054] Step 202: Use a first adaptive filter and a second adaptive filter to adaptively filter the target audio signal.
[0055] In this embodiment, the electronic device may use a first adaptive filter and a second adaptive filter to adaptively filter the target audio signal, respectively, wherein the signal processing delay of the first adaptive filter is greater than the signal processing delay of the second adaptive filter.
[0056] Signal processing delay refers to the time it takes to process a frame of microphone signal using an adaptive filter. This time can include sampling time and analysis time. The sampling time can be the time it takes for the microphone to collect a new section of microphone signal (also known as a frame-shifted signal), i.e., the recording delay time; the analysis time can be the time required to synthesize, analyze, and perform other processing on the current frame of microphone signal. Typically, a frame of microphone signal includes a section of signal collected within the above-mentioned sampling time (i.e., the frame-shifted signal), and also includes a historical signal collected over a period of time (i.e., a signal overlapping with the previous frame of microphone signal). During the signal processing delay corresponding to the frame, the electronic device completes signal collection and signal analysis, synthesis, and other processing. For example, a frame of microphone signal with a frame length of 32ms includes a new 8ms frame-shifted signal and a 24ms historical signal. The sum of the time to collect the 8ms signal and the time to process the 32ms frame of microphone signal is the signal processing delay. That is, under the condition of a fixed frame shift and frame length ratio, the longer the frame length of the microphone signal, the longer the signal processing delay.
[0057] When the microphone signal is filtered using the adaptive filter, since the signal processing delay of the first adaptive filter is greater than the signal processing delay of the second adaptive filter, the first adaptive filter processes more audio data each time.
[0058] The process of filtering the microphone signal using the first adaptive filter is shown in the following formula (1):
[0059]
[0060] Where N = Q*L, e is the filtered audio signal, d is the microphone signal, x is the echo reference signal, and w is the weight parameter of the first adaptive filter. To ensure sufficient resolution, the first adaptive filter usually uses a relatively large Q value. For example, the number of samples corresponding to a microphone signal with a frame length of 256ms, that is, the number of samples after converting the microphone signal in the time domain to the frequency domain, corresponds to a frequency resolution of 1000 / 256 = 3.90625Hz. Therefore, according to formula (1), the echo reference signal is processed in blocks, and the amount of data contained in each block is Q. The frame length of the microphone signal processed by the first adaptive filter each time is 256ms, and the amount of data contained is also Q.
[0061] The process of filtering the microphone signal using the second adaptive filter is shown in the following formula (2):
[0062]
[0063] Where N = P * M. To ensure a sufficiently short output delay, the second adaptive filter typically uses a relatively small value for P. For example, the number of samples corresponding to an 8ms frame length microphone signal corresponds to a frequency resolution of 1000 / 8 = 125Hz. Therefore, according to equation (2), the echo reference signal is processed in blocks, with each block containing P data. The frame length of the microphone signal processed by the first adaptive filter is 256ms, and the amount of data contained in each block is also P.
[0064] The above equations (1) and (2) are formulas in the time domain, and the summation terms in equations (1) and (2) can be treated as convolution. Generally, in order to reduce the amount of computation, the convolution in the time domain can be converted to the frequency domain for multiplication. That is, during filtering, the target audio signal and the parameters of the first adaptive filter and the second adaptive filter are converted from the time domain to the frequency domain by algorithms such as FFT (Fast Fourier Transform), filtered in the frequency domain, and finally the filtered signal is converted from the frequency domain back to the time domain by algorithms such as IFFT (Inverse Fast Fourier Transform).
[0065] It should be noted that, according to the principle of the adaptive filter, the adaptive filter will perform a parameter update operation after each filtering. Therefore, usually after step 202 is completed, the first adaptive filter and the second adaptive filter will perform an iterative update operation of the parameter w.
[0066] Step 203: Based on the signal processing delay of the second adaptive filter, the first adaptive filter is converted to obtain a converted first adaptive filter.
[0067] In this embodiment, the electronic device may determine the converted first adaptive filter based on the signal processing delay of the second adaptive filter.
[0068] Specifically, the signal processing delay of the first adaptive filter may be adjusted to be the same as or close to the signal processing delay of the second adaptive filter, thereby adjusting the parameters of the first adaptive filter to the same dimension as the parameters of the second adaptive filter.
[0069] Step 204 : Determine the first signal transfer gain of the first adaptive filter and the second signal transfer gain of the second adaptive filter after conversion.
[0070] In this embodiment, the electronic device may determine the converted first signal transfer gain of the first adaptive filter and the converted second signal transfer gain of the second adaptive filter.
[0071] The signal transfer gain is used to measure the degree to which the adaptive filter filters the echo signal. Typically, the sum of the squares of the parameters of the first and second adaptive filters can be used as the signal transfer gain. For example, the sum of the squares of the parameters w in equations (1) and (2) above is the signal transfer gain.
[0072] Step 205: Determine a first filtering quality inspection result of the second adaptive filter based on the difference between the first signal transfer gain and the second signal transfer gain.
[0073] In this embodiment, the electronic device may determine the first filtering quality inspection result of the second adaptive filter based on the difference between the first signal transfer gain and the second signal transfer gain.
[0074] Specifically, if the difference between the first signal transmission gain and the second signal transmission gain is greater than or equal to the preset gain value, it can be determined that the second adaptive filter is overfitting, that is, the first filtering quality inspection result of the second adaptive filter is unqualified, and the reason for the failure is that the second adaptive filter is overfitting.
[0075] Typically, the resolution of the second adaptive filter with low latency is low (i.e., the frequency span between the two distinguishable frequency signals is large), and the frequency shift range of the low-frequency signal does not exceed the single resolution range, so it is easy to overfit; after the high-frequency signal is frequency shifted, the frequency shift range is easy to cross the resolution, and the coherence of the corresponding frequency domain is weakened, so it is not easy to overfit. The resolution of the first adaptive filter with high latency is high (i.e., the frequency span between the two distinguishable frequency signals is small), and after the low-frequency signal and the high-frequency signal are frequency shifted, the frequency shift range is easy to cross the resolution, so it is not easy to overfit.
[0076] When the echo reference signal and the near-end human voice signal collected by the microphone are highly similar (for example, when the user makes a prolonged sound, the current near-end human voice signal and the echo reference signal at the previous moment are highly similar), the adaptive filter can not only estimate the transfer function between the echo reference signal and the sound signal emitted by the speaker collected by the microphone, but also simultaneously estimate the transfer function between the echo reference signal and the near-end human voice. The high similarity between the echo reference signal and the near-end human voice signal will cause the two estimated transfer functions to be strongly correlated, and the absolute value of the superposition of the two transfer functions will be close to the superposition of the absolute values of the two transfer functions. When the similarity between the echo reference signal and the near-end human voice signal is low, due to the phase difference between the two transfer functions, the absolute value of the superposition of the two transfer functions is much smaller than the superposition of the absolute values of the two transfer functions. When the echo reference signal and the near-end human voice signal are highly similar, the second adaptive filter with low latency is prone to overfitting, while the first adaptive filter with high latency is not prone to overfitting. Each parameter value of the filter corresponds to different gains and phase information of the sound transmission process. The sum of the squares of the parameters reflects the overall gain level. The manifestation of overfitting is an increase in the signal transmission gain value. Therefore, when the difference between the first signal transmission gain and the second signal transmission gain is too large, it can be determined that the second adaptive filter has overfitted.
[0077] Step 206 : In response to the first filter quality inspection result being unqualified, adjusting the parameters of the second adaptive filter based on the converted parameters of the first adaptive filter to obtain a second adaptive filter with qualified filtering quality.
[0078] In this embodiment, in response to the first filter quality inspection result being unqualified, the electronic device may adjust the parameters of the second adaptive filter based on the converted parameters of the first adaptive filter to obtain a second adaptive filter with qualified filtering quality.
[0079] Specifically, since the parameters of the converted first adaptive filter have the same dimension as the parameters of the second adaptive filter, the parameters of the converted first adaptive filter can be directly replaced by the parameters of the second adaptive filter, or a set value can be added to or subtracted from each parameter of the second adaptive filter to make each parameter of the second adaptive filter close to the corresponding parameter of the converted first adaptive filter.
[0080] Step 207 : Process the target audio signal based on the second adaptive filter with qualified filtering quality.
[0081] In this embodiment, the electronic device may process the target audio signal based on the second adaptive filter with qualified filtering quality.
[0082] Specifically, according to the calculation method shown in the above formula (2), the parameter w of the second adaptive filter is multiplied by the echo reference signal x to obtain an estimated echo signal, and then the estimated echo signal is subtracted from the microphone signal d to obtain the audio signal to be sent to the speaker for playback.
[0083] The method provided by the above-mentioned embodiments of the present disclosure adaptively filters a target audio signal using a first adaptive filter with a high signal processing delay and a second adaptive filter with a low signal processing delay, then converts the first adaptive filter based on the signal processing delay of the second adaptive filter, and determines the difference between the first signal transfer gain and the second signal transfer gain corresponding to the converted first adaptive filter and the second adaptive filter, and determines a first filter quality inspection result of the second adaptive filter based on the search; in response to the first filter quality inspection result being unqualified, the parameters of the second adaptive filter are adjusted based on the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality, and finally the target audio signal is processed based on the second adaptive filter with qualified filtering quality. The embodiment of the present disclosure realizes that on the basis of echo cancellation using the low-latency second adaptive filter, a high-latency first adaptive filter is added. By comparing the difference in the signal transfer gains of the two filters, the filtering quality of the second adaptive filter can be effectively judged. If the filtering quality is unqualified, the parameters of the second adaptive filter are adjusted to obtain a second filter with qualified filtering quality, thereby effectively reducing the risk of overfitting of the second adaptive filter, thereby performing high-precision filtering processing on the target audio signal and improving the quality of the audio played by the speaker.
[0084] In some optional implementations, such as Figure 3 As shown, step 203 includes:
[0085] Step 2031: Adjust the signal processing delay of the first adaptive filter based on the signal processing delay of the second adaptive filter.
[0086] Specifically, the parameters of each segment of the first adaptive filter in the frequency domain can be converted to the time domain through an algorithm such as IFFT, and the parameters of each segment in the time domain are merged. Then, the first adaptive filter is re-segmented according to the above formula (2), so that the signal processing delay of the first adaptive filter is the same as the signal processing delay of the second adaptive filter, even if the first adaptive filter processes more segmented signals M and a smaller signal length P.
[0087] Step 2032: Based on the adjusted signal processing delay of the first adaptive filter, convert the parameters of the first adaptive filter to obtain a converted first adaptive filter.
[0088] Specifically, according to the method described above, after the first adaptive filter is re-segmented in the time domain, the re-segmented first adaptive filter is converted to the frequency domain to obtain a converted first adaptive filter. The parameters of the converted first adaptive filter have the same dimension as the parameters of the second adaptive filter, so that the signal transmission gain can be compared.
[0089] In this embodiment, by adjusting the delay of the first adaptive filter to be the same as that of the second adaptive filter, the dimension of the parameters of the first adaptive filter can be made the same as the dimension of the parameters of the second adaptive filter, thereby helping to more accurately compare the first signal transmission gain and the second signal transmission gain, and further accurately judge the filtering quality of the second adaptive filter.
[0090] In some optional implementations, step 206 may be performed as follows:
[0091] The parameters of the second adaptive filter are replaced by the converted parameters of the first adaptive filter to obtain a second adaptive filter with qualified filtering quality.
[0092] Since the parameters of the converted first adaptive filter and the parameters of the second adaptive filter have the same dimensions, the parameters of the second adaptive filter can be directly replaced with the parameters of the first adaptive filter. Furthermore, since the first adaptive filter generally does not overfit, the second adaptive filter after the parameter replacement will also not overfit, resulting in a second adaptive filter with acceptable filtering quality.
[0093] In this embodiment, by replacing the parameters of the second adaptive filter with the converted parameters of the first adaptive filter, a second adaptive filter with qualified filtering quality can be quickly obtained, thereby reducing the risk of overfitting of the second adaptive filter.
[0094] In some optional implementations, such as Figure 4 As shown, after step 203, the method further includes:
[0095] Step 208 : Use the converted first adaptive filter to adaptively filter the target audio signal to obtain a first filtered signal.
[0096] Step 209 : Determine a second filtering quality test result of the second adaptive filter based on the difference between the energy of the first filtered signal and the energy of the second filtered signal, and the first filtering quality test result.
[0097] The second filtered signal is obtained by adaptively filtering the target audio signal with the second adaptive filter. That is, when executing step 202 , the second filtered signal can be obtained. This step can directly obtain the second filtered signal.
[0098] Typically, the energy of the filtered signal can be calculated by calculating the sum of the squares of each sampling point of the filtered signal. If the first filter quality test result indicates that the signal transfer function estimated by the second adaptive filter has undergone a sudden change, resulting in unstable energy of the second filtered signal, the second filter quality test result is determined to be unqualified.
[0099] Step 210 : In response to the second filter quality inspection result being unqualified, adjusting the step size of the second adaptive filter to obtain a second adaptive filter with qualified filtering quality.
[0100] The step size of the second adaptive filter is the step size during iterative update. As an example, when the second adaptive filter adopts the LMS (Least Mean Square) algorithm, the iterative update formula of the second adaptive filter is shown in the following formula (3):
[0101] w(n+1)=w(n)+2μe(n)x(n) (3)
[0102] Where w(n+1) is the updated parameter, w(n) is the parameter before the update, μ is the iterative update step size, e(n) is the filtered signal e as shown in equation (2), and x(n) is the echo reference signal. When the energy of the second filtered signal is unstable, the update rate of the second adaptive filter can be increased, that is, the step size μ can be increased. The increase in the step size μ can be preset.
[0103] This embodiment determines whether the filtering quality of the second adaptive filter is acceptable by comparing the difference in signal energy after filtering by the first adaptive filter and the second adaptive filter. The filtering quality of the second adaptive filter is restored by adjusting the step size of the iterative update. This enables the filtering result of the first adaptive filter to be used as a reference, the filtering quality of the second adaptive filter to be monitored in real time, and the update rate of the second adaptive filter to be adjusted in a timely manner, thereby making the filtering quality of the second adaptive filter more stable.
[0104] In some optional implementations, such as Figure 5 As shown, step 207 includes:
[0105] Step 2071 : Use the second adaptive filter with qualified filtering quality to adaptively filter the target audio signal to obtain a third filtered signal.
[0106] Specifically, the target audio signal may be adaptively filtered according to the above formula (2), where e in the formula (2) is the third filtered signal.
[0107] Step 2072: Determine the signal to be divided based on the third filtered signal.
[0108] Optionally, the third filtered signal may be determined as the signal to be divided, or the beam enhancement direction may be determined and the beam enhancement may be performed on the third filtered signal according to the method described in the following optional embodiment to obtain the signal to be divided.
[0109] Step 2073: perform frequency division processing on the signal to be divided to obtain at least three frequency division signals.
[0110] The at least three frequency-divided signals correspond to different frequency bands. As an example, the signal to be divided can be decomposed into three non-overlapping frequency bands, corresponding to signals below 200 Hz, signals between 200 Hz and 558 Hz, and signals above 558 Hz, respectively. Alternatively, the at least three frequency-divided signals can also have overlapping portions, for example, decomposing the signal to be divided into signals below 200 Hz, signals below 558 Hz, and signals above 558 Hz.
[0111] Step 2074: Perform frequency conversion on the at least three frequency-divided signals according to corresponding frequency conversion methods to obtain at least three frequency-converted signals.
[0112] The frequency conversion method may include but is not limited to at least one of the following: frequency doubling, frequency shifting, etc.
[0113] Step 2075: Superimpose at least three frequency-converted signals to obtain an audio signal to be played.
[0114] Specifically, if the frequency ranges of at least three frequency-converted signals do not overlap, the at least three frequency-converted signals can be directly spliced into a complete audio signal to be played; if there are overlapping frequency bands in at least three frequency-converted signals, the signals of the overlapping frequency bands can be superimposed, and the signals of the remaining non-overlapping frequency bands can be spliced to obtain the audio signal to be played.
[0115] Typically, in some scenarios, the fundamental frequency and harmonics of the audio signal collected by the microphone remain largely unchanged for an extended period. For example, when a user prolongs their voice, the signal processed by the adaptive filter at the previous moment and played back through the speaker may not differ much from the user's current voice. This indicates that the echo reference signal and the signal collected by the near-end microphone originate from the same source, leading to algorithm overfitting.
[0116] This embodiment filters the target audio signal by using a second adaptive filter with qualified filtering quality, generates at least three frequency-divided signals based on the third filtered signal, and performs frequency conversion processing on the at least three frequency-divided signals. This allows the audio signal to be played back by the speaker and the microphone to capture the sound played back by the speaker, so that the signal captured by the microphone and the reference signal have a certain difference in each frequency band, thereby further reducing the risk of overfitting of the second adaptive filter and improving the quality of the played audio.
[0117] In some optional implementations, such as Figure 6 As shown, step 2072 includes:
[0118] Step 20721: Determine the beam enhancement direction based on the sound source position of the target audio signal.
[0119] Specifically, the beam boosting direction is defined as the angle between the line connecting the microphone's reference point and the sound source location and the microphone's reference line. The sound source location can be determined in real time based on a sound source localization method, and the beam boosting direction can be determined simultaneously. Alternatively, a fixed sound source location can be pre-set, and a fixed beam boosting direction can be determined simultaneously.
[0120] like Figure 7 As shown, when this embodiment is applied to a vehicle, a microphone array can be used, including MIC1 and MIC2 in the figure. The line connecting the two microphones is the reference line, and the midpoint of the reference line is the reference point. Usually, MIC1 and MIC2 are arranged above the roof, and the reference point is located between the driver's seat and the co-driver's seat. Therefore, the beam enhancement direction can be selected as 45° and 135° as shown in the figure. Figure 7 BEAM1 and BEAM2 shown in the figure represent beams in these two directions. These two directions point to the driver's seat and the front passenger seat, so the sound of the people in the driver's seat and the front passenger seat can be collected.
[0121] Step 20722: Based on the beam enhancement direction, perform beam enhancement on the third filtered audio signal to obtain a signal to be divided.
[0122] Specifically, based on a beamforming algorithm, beam enhancement processing may be performed on the above-mentioned beam enhancement direction, while suppressing sounds collected in other directions.
[0123] This embodiment can specifically collect the sound emitted from the sound source position by determining the beam enhancement direction, thereby improving the quality of the audio played by the speaker.
[0124] In some optional implementations, such as Figure 8 As shown, step 2073 includes:
[0125] Step 20731: Perform frequency division processing on the signal to be divided to obtain a first frequency division signal, a second frequency division signal, and a third frequency division signal with frequency bands increasing in sequence.
[0126] As an example, the first frequency-divided signal is a signal with a frequency below 200 Hz, the second frequency-divided signal is a signal with a frequency below 558 Hz, that is, the first frequency-divided signal and the second frequency-divided signal have an overlapping part, and the third frequency-divided signal is a signal with a frequency above 558 Hz.
[0127] Based on step 20731, step 2074 includes:
[0128] Step 20741: perform frequency multiplication processing on the first frequency-divided signal to obtain a first frequency-converted signal.
[0129] Typically, the first frequency-divided signal may be frequency-multiplied by an integer multiple, such as by 2 times. Optionally, an algorithm such as SOLA (Synchronized Overlap-Add) may be used for frequency multiplication.
[0130] Step 20742: Perform frequency shift processing on the second frequency-divided signal according to a preset frequency shift amount to obtain a second frequency-converted signal.
[0131] That is, all frequency components included in the second frequency-divided signal are shifted by the same frequency shift amount. Optionally, the time domain signal can be multiplied by e -jwt The principle of frequency shift is to multiply the real and imaginary parts of the second frequency-divided signal by e in the frequency domain. ±jf The frequency is shifted up and down, and then the frequency-shifted second frequency-divided signal is inversely transformed from the frequency domain to the time domain.
[0132] Step 20743: Scale the third frequency-divided signal according to a preset frequency scaling factor to obtain a third frequency-converted signal.
[0133] Specifically, the method for scaling the frequency of the third frequency-divided signal can be implemented by using the frequency multiplication method described in step 20741. Typically, the frequency scaling factor is set close to 1, such as 1.007 frequency multiplication or 1 / 1.007 frequency multiplication.
[0134] In order to ensure a stable auditory experience when listening to audio played by speakers and not to experience large fluctuations in the energy of the played audio, the frequency variation of the spectrum needs to be as small as possible. However, after the low-frequency signal undergoes a small frequency conversion, the frequency variation of the signal is very weak, which is not very helpful in reducing overfitting. Therefore, in the frequency band below 558Hz (i.e., the second frequency division signal), frequency shifting can be used to shift all frequency components in this band by the same frequency to increase the amplitude of the frequency variation without affecting the user's hearing experience. At the same time, since the first resonance peak is located in the frequency band above 200Hz, the frequency band below 200Hz (i.e., the first frequency division signal) can be doubled to increase the frequency variation amplitude and compensate for the loss of the baseband signal. In the high frequency band above 558Hz (i.e., the third frequency division signal), in order to ensure that the frequency variation amplitude is sufficient and does not affect the user's hearing experience, a frequency scaling factor of no more than 1.007 can be selected.
[0135] This embodiment performs high-multiplication frequency multiplication processing, frequency shift processing, and low-multiplication frequency scaling processing on the three-way frequency-divided signals, thereby achieving targeted frequency conversion processing of audio signals in different frequency bands based on the user's hearing perception and the characteristics of parameter fitting of the adaptive filter. This further reduces the risk of overfitting of the second adaptive filter without affecting the user's hearing perception, thereby improving the quality of the filtered audio signal.
[0136] Exemplary devices
[0137] Figure 9 FIG. 1 is a schematic diagram of a structure of an audio signal processing apparatus provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices such as Figure 9As shown, the apparatus for processing an audio signal includes: an acquisition module 901 for acquiring a target audio signal, wherein the target audio signal includes a microphone signal and an echo reference signal; a first filtering module 902 for adaptively filtering the target audio signal using a first adaptive filter and a second adaptive filter, respectively, wherein the signal processing delay of the first adaptive filter is greater than the signal processing delay of the second adaptive filter; a conversion module 903 for converting the first adaptive filter based on the signal processing delay of the second adaptive filter to obtain a converted first adaptive filter; a first determination module 904 for determining a first signal transfer gain of the converted first adaptive filter and a second signal transfer gain of the second adaptive filter; a second determination module 905 for determining a first filtering quality inspection result of the second adaptive filter based on a difference between the first signal transfer gain and the second signal transfer gain; a first adjustment module 906 for adjusting parameters of the second adaptive filter based on the parameters of the converted first adaptive filter in response to an unqualified first filtering quality inspection result to obtain a second adaptive filter with qualified filtering quality; and a second filtering module 907 for processing the target audio signal based on the second adaptive filter with qualified filtering quality.
[0138] In this embodiment, the acquisition module 901 can acquire the target audio signal locally or remotely, wherein the target audio signal includes a microphone signal and an echo reference signal.
[0139] The microphone signal can be as follows Figure 1 The audio signal collected by the microphone 104 shown in the figure, the echo reference signal can be sent to the target time period before the current moment. Figure 1 The audio signal played by speaker 105, i.e., the echo reference signal, is the audio signal obtained after filtering by the second adaptive filter during the target time period. When speaker 105 plays the audio signal, microphone 104 simultaneously captures the user's voice and the sound played by the speaker. The collected sound played by the speaker is the echo signal. The purpose of this embodiment is to estimate the echo signal component in the microphone signal through the second adaptive filter, remove this echo signal component from the microphone signal, and obtain the user's voice signal.
[0140] In this embodiment, the first filtering module 902 may adaptively filter the target audio signal using a first adaptive filter and a second adaptive filter, wherein the signal processing delay of the first adaptive filter is greater than the signal processing delay of the second adaptive filter.
[0141] Signal processing delay refers to the length of each segment of an echo reference signal, divided into multiple segments of equal length. It should be understood that the length here refers to the time dimension. When filtering the microphone signal using an adaptive filter, the adaptive filter processes the echo reference signal in segments. Because the signal processing delay of the first adaptive filter is greater than that of the second adaptive filter, the first adaptive filter processes more sampled data each time.
[0142] In this embodiment, the conversion module 903 may determine the converted first adaptive filter based on the signal processing delay of the second adaptive filter.
[0143] Specifically, the signal processing delay of the first adaptive filter may be adjusted to be the same as or close to the signal processing delay of the second adaptive filter, thereby adjusting the parameters of the first adaptive filter to the same dimension as the parameters of the second adaptive filter.
[0144] In this embodiment, the first determining module 904 may determine the converted first signal transfer gain of the first adaptive filter and the converted second signal transfer gain of the second adaptive filter.
[0145] The signal transfer gain is used to measure the degree to which the adaptive filter filters the echo signal. Generally, the square sum of the parameters of the first adaptive filter and the second adaptive filter can be used as the signal transfer gain.
[0146] In this embodiment, the second determination module 905 may determine the first filtering quality inspection result of the second adaptive filter based on the difference between the first signal transfer gain and the second signal transfer gain.
[0147] Specifically, if the difference between the first signal transmission gain and the second signal transmission gain is greater than or equal to the preset gain value, it can be determined that the second adaptive filter is overfitting, that is, the first filtering quality inspection result of the second adaptive filter is unqualified, and the reason for the failure is that the second adaptive filter is overfitting.
[0148] In this embodiment, in response to the first filtering quality inspection result being unqualified, the first adjustment module 906 may adjust the parameters of the second adaptive filter based on the converted parameters of the first adaptive filter to obtain a second adaptive filter with qualified filtering quality.
[0149] Specifically, since the parameters of the converted first adaptive filter have the same dimension as the parameters of the second adaptive filter, the parameters of the converted first adaptive filter can be directly replaced by the parameters of the second adaptive filter, or a set value can be added to or subtracted from each parameter of the second adaptive filter to make each parameter of the second adaptive filter close to the corresponding parameter of the converted first adaptive filter.
[0150] In this embodiment, the second filtering module 907 may process the target audio signal based on the second adaptive filter with qualified filtering quality.
[0151] Specifically, according to the calculation method shown in the above formula (2), the parameter w of the second adaptive filter is multiplied by the echo reference signal x to obtain an estimated echo signal, and then the estimated echo signal is subtracted from the microphone signal d to obtain the audio signal to be sent to the speaker for playback.
[0152] Reference Figure 10 , Figure 10 FIG. 4 is a structural diagram of an audio signal processing apparatus provided by another exemplary embodiment of the present disclosure.
[0153] In some optional implementations, the conversion module 903 includes: an adjustment unit 9031, which is used to adjust the signal processing delay of the first adaptive filter based on the signal processing delay of the second adaptive filter; a conversion unit 9032, which is used to convert the parameters of the first adaptive filter based on the adjusted signal processing delay of the first adaptive filter to obtain a converted first adaptive filter.
[0154] In some optional implementations, the first adjustment module 906 is further configured to replace the parameters of the second adaptive filter with the converted parameters of the first adaptive filter to obtain a second adaptive filter with qualified filtering quality.
[0155] In some optional implementations, the device also includes: a third filtering module 908, which is used to use the converted first adaptive filter to adaptively filter the target audio signal to obtain a first filtered signal; a third determination module 909, which is used to determine the second filtering quality inspection result of the second adaptive filter based on the difference between the energy of the first filtered signal and the energy of the second filtered signal, and the first filtering quality inspection result, wherein the second filtered signal is obtained by adaptively filtering the target audio signal by the second adaptive filter; a second adjustment module 910, which is used to adjust the step size of the second adaptive filter in response to the second filtering quality inspection result being unqualified, to obtain a second adaptive filter with qualified filtering quality.
[0156] In some optional implementations, the second filtering module 907 includes: a filtering unit 9071, which is used to use a second adaptive filter with qualified filtering quality to adaptively filter the target audio signal to obtain a third filtered signal; a determination unit 9072, which is used to determine the signal to be divided based on the third filtered signal; a frequency division unit 9073, which is used to perform frequency division processing on the signal to be divided to obtain at least three frequency division signals, wherein the at least three frequency division signals correspond to different frequency bands; a frequency conversion unit 9074, which is used to perform frequency conversion processing on the at least three frequency division signals according to corresponding frequency conversion methods to obtain at least three frequency-converted signals; and a superposition unit 9075, which is used to superimpose at least three frequency-converted signals to obtain the audio signal to be played.
[0157] In some optional implementations, the determination unit 9072 includes: a determination subunit 90721, used to determine the beam enhancement direction based on the sound source position of the target audio signal; an enhancement subunit 90722, used to perform beam enhancement on the third filtered audio signal based on the beam enhancement direction to obtain the signal to be divided.
[0158] In some optional implementations, the frequency division unit 9073 is further used to: perform frequency division processing on the signal to be divided to obtain a first frequency division signal, a second frequency division signal and a third frequency division signal with successively increasing frequency bands; the frequency conversion unit 9074 includes: a frequency multiplication subunit 90741, used to perform frequency multiplication processing on the first frequency division signal to obtain a first frequency-converted signal; a frequency shifting subunit 90742, used to perform frequency shift processing on the second frequency division signal according to a preset frequency shift amount to obtain a second frequency-converted signal; and a scaling subunit 90743, used to scale the third frequency division signal according to a preset frequency scaling factor to obtain a third frequency-converted signal.
[0159] The audio signal processing device provided by the above-mentioned embodiment of the present disclosure adaptively filters the target audio signal using a first adaptive filter with a high signal processing delay and a second adaptive filter with a low signal processing delay, then converts the first adaptive filter based on the signal processing delay of the second adaptive filter, and determines the difference between the first signal transfer gain and the second signal transfer gain corresponding to the converted first adaptive filter and the second adaptive filter, and determines the first filter quality inspection result of the second adaptive filter based on the search; in response to the first filter quality inspection result being unqualified, the parameters of the second adaptive filter are adjusted based on the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality, and finally processes the target audio signal based on the second adaptive filter with qualified filtering quality. The embodiment of the present disclosure realizes that on the basis of echo cancellation using the low-latency second adaptive filter, a high-latency first adaptive filter is added. By comparing the difference in the signal transfer gains of the two filters, the filtering quality of the second adaptive filter can be effectively determined. If the filtering quality is unqualified, the parameters of the second adaptive filter are adjusted to obtain a second filter with qualified filtering quality, thereby effectively reducing the risk of overfitting of the second adaptive filter, thereby performing high-precision filtering processing on the target audio signal and improving the quality of audio played by the speaker.
[0160] Exemplary electronic devices
[0161] Below, reference Figure 11 To describe the electronic device according to the embodiment of the present disclosure. The electronic device may be as follows Figure 1 Any one or both of the terminal device 101 and the server 103 shown, or a stand-alone device independent of them, can communicate with the terminal device 101 and the server 103 to receive the collected input signals from them.
[0162] Figure 11 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0163] like Figure 11 As shown, the electronic device 1100 includes one or more processors 1101 and a memory 1102 .
[0164] The processor 1101 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1100 to perform desired functions.
[0165] The memory 1102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1101 may execute the program instructions to implement the audio signal processing methods of the various embodiments of the present disclosure described above and / or other desired functions. Various content, such as a target audio signal, may also be stored in the computer-readable storage medium.
[0166] In one example, the electronic device 1100 may further include an input device 1103 and an output device 1104 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0167] For example, when the electronic device is the terminal device 101 or the server 103 , the input device 1103 may be a microphone, a mouse, a keyboard, a touch screen, or the like, for inputting target audio signals, various commands, etc. When the electronic device is a standalone device, the input device 1103 may be a communication network connector for receiving input target audio signals, various commands, etc. from the terminal device 101 and the server 103 .
[0168] The output device 1104 can output various information to the outside, including the filtered audio signal. The output device 1104 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto.
[0169] Of course, to simplify, Figure 11 Only some of the components related to the present disclosure in the electronic device 1100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1100 may further include any other appropriate components.
[0170] Exemplary computer program products and computer-readable storage media
[0171] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, and when the computer program instructions are executed by a processor, the processor performs the steps of the method of audio signal processing according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0172] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0173] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps of the method for audio signal processing according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0174] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0175] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0176] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.
[0177] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0178] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0179] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0180] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0181] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for audio signal processing, comprising: Acquire a target audio signal, wherein the target audio signal includes a microphone signal and an echo reference signal; Adaptively filtering the target audio signal using a first adaptive filter and a second adaptive filter, respectively, wherein a signal processing delay of the first adaptive filter is greater than a signal processing delay of the second adaptive filter; converting the first adaptive filter based on a signal processing delay of the second adaptive filter to obtain a converted first adaptive filter; determining a first signal transfer gain of the converted first adaptive filter and a second signal transfer gain of the second adaptive filter; determining a first filtering quality inspection result of the second adaptive filter based on a difference between the first signal transfer gain and the second signal transfer gain; In response to the first filtering quality inspection result being unqualified, adjusting the parameters of the second adaptive filter based on the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality; The target audio signal is processed based on the second adaptive filter with qualified filtering quality.
2. The method according to claim 1, wherein The converting of the first adaptive filter based on the signal processing delay of the second adaptive filter to obtain a converted first adaptive filter includes: adjusting a signal processing delay of the first adaptive filter based on a signal processing delay of the second adaptive filter; Based on the adjusted signal processing delay of the first adaptive filter, parameters of the first adaptive filter are converted to obtain a converted first adaptive filter.
3. The method according to claim 1, wherein The adjusting the parameters of the second adaptive filter based on the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality includes: The parameters of the second adaptive filter are replaced by the parameters of the converted first adaptive filter to obtain a second adaptive filter with qualified filtering quality.
4. The method according to claim 1, wherein After the signal processing delay based on the second adaptive filter is performed on the first adaptive filter to obtain a converted first adaptive filter, the method further includes: Adaptively filtering the target audio signal using the converted first adaptive filter to obtain a first filtered signal; determining a second filtering quality test result of the second adaptive filter based on a difference between energy of the first filtered signal and energy of a second filtered signal, and the first filtering quality test result, wherein the second filtered signal is obtained by adaptively filtering the target audio signal by the second adaptive filter; In response to the second filtering quality inspection result being unqualified, the step size of the second adaptive filter is adjusted to obtain a second adaptive filter with qualified filtering quality.
5. The method according to any one of claims 1 to 4, wherein: The processing of the target audio signal by the second adaptive filter based on the qualified filtering quality includes: Adaptively filtering the target audio signal using the second adaptive filter with qualified filtering quality to obtain a third filtered signal; determining a signal to be frequency-divided based on the third filtered signal; Performing frequency division processing on the signal to be divided to obtain at least three frequency-divided signals, wherein the at least three frequency-divided signals correspond to different frequency bands; Performing frequency conversion processing on the at least three frequency-divided signals according to corresponding frequency conversion methods to obtain at least three frequency-converted signals; The at least three frequency-converted signals are superimposed to obtain an audio signal to be played.
6. The method according to claim 5, wherein: The step of determining the signal to be frequency-divided based on the third filtered signal includes: Determining a beam enhancement direction based on a sound source position of the target audio signal; Based on the beam boosting direction, beam boosting is performed on the third filtered signal to obtain a signal to be frequency-divided.
7. The method according to claim 5, wherein: The frequency division processing is performed on the signal to be divided to obtain at least three frequency-divided signals, including: Performing frequency division processing on the signal to be divided to obtain a first frequency division signal, a second frequency division signal and a third frequency division signal with successively higher frequency bands; The performing frequency conversion processing on the at least three frequency-divided signals according to corresponding frequency conversion modes to obtain at least three frequency-converted signals includes: Performing frequency multiplication processing on the first frequency-divided signal to obtain a first frequency-converted signal; Performing frequency shift processing on the second frequency-divided signal according to a preset frequency shift amount to obtain a second frequency-converted signal; The third frequency-divided signal is scaled according to a preset frequency scaling factor to obtain a third frequency-converted signal.
8. An apparatus for audio signal processing, comprising: An acquisition module, configured to acquire a target audio signal, wherein the target audio signal includes a microphone signal and an echo reference signal; a first filtering module, configured to adaptively filter the target audio signal using a first adaptive filter and a second adaptive filter, respectively, wherein a signal processing delay of the first adaptive filter is greater than a signal processing delay of the second adaptive filter; a conversion module, configured to convert the first adaptive filter based on a signal processing delay of the second adaptive filter to obtain a converted first adaptive filter; a first determining module, configured to determine a first signal transfer gain of the first adaptive filter and a second signal transfer gain of the second adaptive filter after the conversion; a second determining module, configured to determine a first filtering quality inspection result of the second adaptive filter based on a difference between the first signal transfer gain and the second signal transfer gain; a first adjustment module, configured to adjust parameters of the second adaptive filter based on the parameters of the converted first adaptive filter in response to the first filtering quality inspection result being unqualified, to obtain a second adaptive filter with qualified filtering quality; The second filtering module is configured to process the target audio signal based on the second adaptive filter with qualified filtering quality.
9. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the method according to any one of claims 1 to 7 when executed by a processor.
10. An electronic device, comprising: processor; a memory for storing executable instructions for the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Hearing aid system with feedback arrangement to predict and cancel acoustic feedback, method and use
CN101874412A
Signal processing method, signal processing device, and signal processing program
CN102165709A