Decorrelation processing method and apparatus, device, medium, and product
By using time-domain decorrelation processing, frequency, phase, and amplitude adjustments are made to reduce the signal correlation between the microphone and the speaker, thus solving the feedback problem in the in-vehicle karaoke system and achieving efficient feedback suppression and sound quality improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-05-20
- Publication Date
- 2026-07-31
AI Technical Summary
In in-car karaoke systems, closed-loop feedback between the microphone and speaker causes howling. Existing howling suppression methods are ineffective, especially the AFC adaptive filter, which eliminates part of the desired speaker's voice signal without decorrelation processing.
A time-domain decorrelation processing method is adopted to reduce the correlation between the human voice signal captured by the microphone and the audio signal played by the speaker through frequency adjustment, phase adjustment and amplitude adjustment, including signal processing using Hilbert transform and dynamic range controller.
It improves howling suppression, reduces computational load, meets real-time processing requirements, and is independent of AFC and notch filtering methods, thus improving sound quality and howling suppression.
Smart Images

Figure CN122496754A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of signal processing technology, and in particular relates to a decorrelation processing method, apparatus, device, medium and product. Background Technology
[0002] With the development of automotive technology, in-car entertainment scenarios have become increasingly diversified, and in-car karaoke, especially microphone-free karaoke, has gradually become a popular application in in-car entertainment systems. In karaoke scenarios, the basic sound reinforcement system consists of in-car microphones (handheld or in-car microphones) and speakers (and their amplifiers). The sound of passengers singing is picked up by the in-car microphones, amplified by the amplifier, and then played through the in-car speakers. The signal picked up by the microphones is amplified and played by the speakers, and the played sound signal is picked up by the microphones again, resulting in a closed-loop positive feedback loop between the microphones and speakers. If the sound is directly amplified without processing, the signal will be continuously fed back and amplified, eventually causing howling, which will greatly affect the experience of the local sound reinforcement scenario. Therefore, appropriate howling suppression methods are needed.
[0003] Feedback suppression methods all require additional de-processing equipment to function properly. Without such equipment, feedback suppression methods will eliminate some of the desired speaker voice signal, resulting in poor feedback suppression performance. Summary of the Invention
[0004] This application provides a decorrelation processing method, apparatus, device, medium, and product that can improve the howling suppression effect.
[0005] In a first aspect, embodiments of this application provide a decorrelation processing method, the method comprising: The time-domain signal to be processed at the current sampling point is obtained. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the sound pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the sound pickup device is located. The time-domain signal to be processed is subjected to decorrelation processing in the time domain to obtain the target time-domain signal. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment.
[0006] Secondly, embodiments of this application provide a decorrelation processing apparatus, the apparatus comprising: The acquisition module is used to acquire the time-domain signal to be processed at the current sampling point. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. The processing module is used to perform time-domain decorrelation processing on the time-domain signal to be processed to obtain the target time-domain signal. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment.
[0007] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the decorrelation processing method as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the decorrelation processing method as described in the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the decorrelation processing method as described in the first aspect.
[0010] In this embodiment of the application, the correlation between the first audio signal (i.e., the audio signal played by the playback device) and the second audio signal (the human voice signal collected by the microphone) can be reduced through the above-described decorrelation process, thereby improving the effect of howling suppression. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of a local sound reinforcement scenario in the related technologies provided in the embodiments of this application; Figure 2 This is a schematic diagram of the decorrelation processing method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the howling suppression process based on AFC and decorrelation module provided in the embodiments of this application; Figure 4 This is a schematic diagram of the decorrelation processing apparatus provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0015] In all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. Additionally, when embodiments of this application require access to sensitive personal information, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments obtained.
[0016] With the development of automotive technology, in-car entertainment scenarios have become increasingly diversified, and in-car karaoke, especially microphone-free karaoke, has gradually become a popular application in in-car entertainment systems. In a karaoke scenario, the basic sound reinforcement system consists of an in-car microphone (handheld or in-car microphone), speakers (and their amplifiers). The voices of passengers singing are picked up by the in-car microphones, amplified by the amplifier, and then played back through the in-car speakers. Figure 1 As shown.
[0017] Figure 1 As shown, the smiley face represents the speaker, the dotted line segment represents the sound emitted by the speaker, the horizontal line segment represents the speaker's voice, the solid line represents the entire audio loop system, G represents the audio system (including the power amplifier), and F represents the acoustic path between the speaker and the microphone. Figure 1 As can be seen, this is a typical local sound reinforcement system. In-car karaoke is just one application scenario; in-vehicle communication (ICC) systems can also be used in in-vehicle applications. Figure 1 This is for display purposes. Of course, large classrooms, conference rooms, karaoke rooms, and other similar applications all fall under the category of local sound reinforcement and can also be represented using this technology. Figure 1 express.
[0018] The signal picked up by the microphone is amplified and played by the speaker, and the played sound signal is then picked up by the microphone again, creating a closed-loop positive feedback loop between the microphone and the speaker. If the signal is amplified directly without processing, it will be continuously fed back and amplified, eventually causing howling, which will greatly affect the experience in local sound reinforcement scenarios. Therefore, appropriate howling suppression methods are needed.
[0019] There are three main types of howling suppression methods.
[0020] The first method is frequency shifting, which shifts the frequency of the signal collected by the microphone. The greater the frequency shift, the better the feedback suppression effect, but the greater the loss of sound quality.
[0021] The second method is the notch filter method, which uses a notch filter to process the detected howling frequency band, thereby achieving howling suppression. This method requires first detecting the howling sound and accurately locating the frequency band where the howling occurs before a suitable notch filter can be used for howling suppression. No matter how precise the howling detection and howling frequency location are, this method can only work after the howling has occurred.
[0022] The third method is Acoustic Feedback Cancellation (AFC). AFC uses an adaptive filter to estimate the acoustic path between the microphone and the loudspeaker, thereby eliminating the sound signal emitted by the loudspeaker from the microphone signal. If the adaptive filter can perfectly estimate the acoustic path between the microphone and the loudspeaker, the howling problem will not occur; that is, the AFC method is the theoretically optimal solution among howling suppression schemes.
[0023] There are many design methods for adaptive filtering in AFC (Augmented Feedback Control), such as Least Mean Squares (LMS), Normalized Least Mean Squares (NLMS), Recursive Least Squares (RLS), and Kalman Filter. However, these methods all require additional decorrelation processing to function properly. This is because the signal played by the speaker is actually the signal of the speaker's voice captured by the microphone after passing through the audio loop system, and the two are highly correlated. Without decorrelation processing, the adaptive filtering in AFC will eliminate some of the desired speaker's voice signal, resulting in poor feedback suppression.
[0024] To address the problems of the prior art, embodiments of this application provide a decorrelation processing method, apparatus, device, medium, and product. The decorrelation processing method provided in this application embodiment will be described first below.
[0025] Figure 2 A flowchart illustrating a decorrelation processing method according to an embodiment of this application is shown. Figure 2 As shown, the decorrelation processing method provided in this application embodiment is applied to an electronic device and includes the following steps 101-102, wherein: Step 101: Obtain the time-domain signal to be processed at the current sampling point. The time-domain signal to be processed is obtained by performing howling suppression processing on the initial signal collected by the pickup device. Howling suppression processing can employ Acoustic Feedback Cancellation (AFC) or notch filtering. In this embodiment, decorrelation operation can be performed on the signal output by the adaptive filter in AFC or the notch filter in notch filtering (i.e., the time-domain signal to be processed).
[0026] The sound pickup device can refer to a microphone. The initial signals collected by the sound pickup device include a first audio signal and a second audio signal. The first audio signal is obtained after the signal to be played is output through a playback device (e.g., a speaker); the second audio signal is the human voice signal in the environment where the sound pickup device is located.
[0027] For example, step 101 specifically includes: processing the signal to be played and the initial signal through AFC to obtain the acoustic path between the playback device and the pickup device; estimating the first audio signal based on the acoustic path to obtain an audio estimation signal; and subtracting the audio estimation signal from the initial signal to obtain the time-domain signal to be processed at the current sampling point.
[0028] like Figure 3As shown, x is the signal to be played by the speaker. The acoustic path F between the speaker and the microphone is convolved on x to obtain the speaker's broadcast signal r. This broadcast signal r is collected by the microphone. The total signal (i.e. the initial signal) y collected by the microphone includes the broadcast signal r and the human voice signal s. The overall signal y and the signal to be played x are first processed by AFC to estimate the acoustic path between the speaker and the microphone. Then based on The signal r emitted by the loudspeaker is estimated to obtain Subtract from the overall signal The output signal e after AFC is obtained, and this output signal e is the time domain signal to be processed. Figure 3 In this context, G represents the audio system (including the power amplifier), and "decorrelated" represents the decorrelation module that uses the decorrelation processing method provided in the embodiments of this application.
[0029] Will Figure 3 After replacing the AFC method with the notch filter method, the overall signal y is first detected to find the frequency band where the howling occurs. A suitable notch filter is then used to suppress the howling in this frequency band to obtain the output signal e.
[0030] Step 102: Perform time-domain decorrelation processing on the time-domain signal to be processed to obtain the target time-domain signal. The decorrelation processing includes frequency adjustment, phase adjustment, and amplitude adjustment. The decorrelation processing is used to reduce the correlation between the first audio signal and the second audio signal. In this embodiment, the time-domain signal to be processed is decorrelated in the time domain. On the one hand, since the decorrelation process is implemented in the time domain, it does not involve related processing such as Fourier transform and inverse Fourier transform, resulting in low computational load and high efficiency, which can meet the real-time processing requirements of the time-domain signal to be processed. On the other hand, through the above decorrelation process, the correlation between the first audio signal (i.e., the audio signal played by the playback device) and the second audio signal (the human voice signal collected by the microphone) can be reduced, thereby improving the effect of howling suppression.
[0031] The decorrelation processing method provided in this application can be used as a plug-in, independent of the design of the adaptive filter in the AFC scheme and the design of the notch filter in the notch method. It has great flexibility and can be efficiently integrated into the in-vehicle karaoke and in-vehicle communication system in the vehicle cabin.
[0032] In one embodiment of this application, the time-domain signal to be processed is suppressed if its amplitude is too large or too small. Specifically, the time-domain signal to be processed if its amplitude is greater than a first threshold or less than a second threshold is subjected to frequency adjustment, phase adjustment, and amplitude adjustment. Specifically, step 102 involves performing time-domain decorrelation processing on the time-domain signal to be processed to obtain the target time-domain signal, including: If the amplitude of the time-domain signal to be processed is greater than a first threshold, or if the amplitude of the time-domain signal to be processed is less than a second threshold, then decorrelation processing is performed on the time-domain signal to be processed to obtain the target time-domain signal, wherein the first threshold is greater than the second threshold. On the one hand, when howling occurs, the energy of the time-domain signal to be processed with an amplitude greater than the first threshold will increase significantly. Adjusting the amplitude can provide additional amplitude control for the time-domain signal to be processed, thereby improving the effect of howling suppression. On the other hand, adjusting the amplitude (i.e., amplifying the amplitude) of the time-domain signal to be processed with an excessively small amplitude can improve the sound effect of the target time-domain signal.
[0033] In another embodiment of this application, the time-domain decorrelation processing of the time-domain signal to be processed is performed to obtain the target time-domain signal, including: Based on the time-domain signal to be processed, a first analytic signal represented in rectangular coordinates is constructed, wherein the time-domain signal to be processed is the real part of the first analytic signal, and the Hilbert transform of the time-domain signal to be processed is the imaginary part of the first analytic signal; The first analytical signal is transformed to obtain a second analytical signal represented in polar coordinates. The second analytical signal includes instantaneous amplitude and instantaneous phase. The instantaneous phase of the second analytical signal is adjusted using the target phase increment to obtain the first time-domain signal; The first time-domain signal is transformed to obtain a second time-domain signal represented in rectangular coordinates; The real part of the second time-domain signal is extracted to obtain the target time-domain signal.
[0034] For example, if the number n corresponding to the current sampling point (i.e., the sampling point corresponding to the time-domain signal to be processed) is odd, the Hilbert transform of the time-domain signal to be processed is: The time-domain signal to be processed is compared with The product of these is used as the imaginary part of the first analytic signal; if the number n corresponding to the current sampling point is even, the Hilbert transform of the time-domain signal to be processed is... The time-domain signal to be processed is compared with The product of is used as the imaginary part of the first analytic signal, which can be found in Equation (5) below.
[0035] The first analytical signal is transformed to obtain a second analytical signal characterized by instantaneous amplitude and instantaneous phase. The second analytical signal can be found in equation (7) below. The first analytical signal and the second analytical signal are the same analytical signal. The distinction between "first" and "second" is only to differentiate the ways in which the analytical signal is expressed in different coordinate systems.
[0036] The instantaneous phase of the second analytical signal is adjusted by using the target phase increment to obtain the first time domain signal. Then, the first time domain signal is transformed to obtain the second time domain signal characterized by the real part and the imaginary part. Taking the real part of the second time domain signal can obtain the target time domain signal, which is the time domain signal after phase adjustment or frequency adjustment.
[0037] In this embodiment, a first analytical signal represented in rectangular coordinates is constructed based on the time-domain signal to be processed; the first analytical signal is transformed to obtain a second analytical signal represented in polar coordinates; the instantaneous phase of the second analytical signal is adjusted using a target phase increment to obtain a first time-domain signal; the first time-domain signal is transformed to obtain a second time-domain signal represented in rectangular coordinates; the real part of the second time-domain signal is taken to obtain a target time-domain signal. This achieves decorrelation processing of the time-domain signal to be processed in the time domain, thereby reducing the correlation between the first audio signal and the second audio signal and improving the effect of howling suppression.
[0038] In one embodiment of this application, the amplitude adjustment includes: adjusting the real part of the second time-domain signal; The phase adjustment includes: adjusting the angle value when the target phase increment is determined based on the angle value; The frequency adjustment includes adjusting the moving frequency when the target phase increment is determined based on the moving frequency.
[0039] Specifically, if the target phase increment is an angle value Then by adjusting It can achieve phase adjustment of the time-domain signal to be processed, for example, adding the target phase increment to the instantaneous phase of the second analytic signal. By adjusting This is to achieve phase adjustment of the time-domain signal to be processed.
[0040] If Represented as a form that changes over time, i.e. According to the properties of the Fourier transform, in the time domain... This is equivalent to shifting the spectrum of the time-domain signal to be processed to the right in the frequency domain. ( For example, adding the target phase increment to the instantaneous phase of the second analyzed signal. By adjusting the movement frequency This is used to adjust the frequency of the time-domain signal being processed.
[0041] The amplitude, frequency, and phase of the time-domain signal to be processed can be adjusted according to the actual situation to reduce the correlation between the first audio signal and the second audio signal, thereby improving the effect of howling suppression and making the sound effect of the target time-domain signal played better.
[0042] This application embodiment also provides a method for phase adjustment of a time-domain signal to be processed. Specifically, the phase adjustment includes: delaying the time-domain signal to be processed. Second, It is an integer multiple of the sampling period, or, It is not an integer multiple of the sampling period.
[0043] The time-domain signal to be processed is a discrete signal, delayed by several sampling points, and the corresponding delay time is also discrete, i.e., an integer multiple of the sampling period. When the time interval is not an integer multiple of the sampling period, interpolation is performed on the time-domain signal to be processed based on historical sampling points to obtain the delay. The time-domain signal to be processed corresponds to seconds. For example, interpolation can be performed using linear interpolation or nonlinear interpolation (for details on linear interpolation, please refer to the relevant records of equations (14)-(16) below, which will not be elaborated here), thereby obtaining the delay. The time-domain signal to be processed corresponds to a given second. Through the interpolation method described above, it is possible to achieve a delay of arbitrary duration (i.e.,...) of the time-domain signal to be processed. It can take any value (not limited to an integer multiple of the sampling period), thus enabling arbitrary phase adjustment of the time-domain signal to be processed in the frequency domain.
[0044] The following examples illustrate the de-relation processing method provided in this application.
[0045] like Figure 3 As shown, x is the signal to be played by the loudspeaker. The acoustic path F between the loudspeaker and the microphone is convolved on x to obtain the loudspeaker's broadcast signal r. This broadcast signal r is collected by the microphone. The total signal y collected by the microphone includes the broadcast signal r and the human voice signal s, as shown in equation (1): (1) The overall signal y and the signal to be played x are first processed by AFC to estimate the acoustic path between the speaker and the microphone. Then based on The signal r emitted by the loudspeaker is estimated to obtain Subtract from the overall signal The output signal e after AFC is obtained, and this output signal e is the time-domain signal to be processed at the current sampling point. As shown in equation (2): (2) In this embodiment, frequency and phase adjustments are implemented in the time domain using the Hilbert transform. Hilbert transform It is a convolution operator, whose discrete representation is shown in equation (3), where n represents the sampling point number, and the corresponding discrete time can be calculated.
[0046] (3) The time-domain signal to be processed at the current sampling point n The Hilbert transform can be expressed as equation (4): (4) Based on equation (4), the time-domain signal to be processed can be constructed. Analyzed signal (The complex signal, also known as the first analytic signal) is shown in equation (5): (5) At this point, the spectrum Z(k) of z(n) after discrete Fourier transform and the spectrum E(k) of the time-domain signal e(n) after discrete Fourier transform are consistent, except that the amplitude is doubled. Their relationship is as shown in equation (6): (6) From equation (6), it can be seen that the analytic signal The phase information in the frequency domain is consistent with the time-domain signal e(n) to be processed, and the amplitude information is also a simple multiple relationship. Equation (5) can be expressed in the form of instantaneous amplitude and instantaneous phase (called the second analytic signal), as shown in equation (7): (7) Multiply equation (7) by Phase adjustment can then be achieved in this case. The target phase increment is shown in equation (8): (8) Then to Taking the real part yields the phase-adjusted output signal of the time-domain signal e(n), as shown in equation (9), where Re represents the real part taking operation: (9) In equation (8) Represented in time-varying form, i.e. This allows for frequency adjustment of the analytic signal z(n), as shown in equation (10): (10) Based on the properties of the Fourier transform, in the time domain This is equivalent to shifting the spectrum of the time-domain signal to be processed to the right in the frequency domain. f ( f= In this case (w / 2π), This represents the target phase increment. For Taking the real part yields the frequency-adjusted output signal of the original signal e(n), as shown in equation (11): (11) According to the properties of Fourier transform, the spectrum of the time-domain signal e(t) to be processed can be obtained in the frequency domain. Multiply by phase increment This is equivalent to delaying the time-domain signal e(t) to be processed by τ seconds in the time domain, as shown in equation (12): (12) In the above formula, F represents the Fourier transform. This indicates that the time-domain signal e(t) is delayed by τ seconds. Since the time-domain signal e(n) to be processed is a discrete signal, the delay time corresponding to several sampling points is also discrete, that is, an integer multiple of the sampling period. Assume the sampling rate is... Hertz (Hz), the duration of one sampling period is shown in equation (13): (13) Assumption , The delay at the sampling point level can only result in a delay that is an integer multiple of 0.0000625s, and it is impossible to achieve an arbitrary delay of τ seconds. If an arbitrary delay of τ seconds could be achieved, it can be seen from equation (12) that arbitrary phase increment adjustment can be realized in the frequency domain.
[0047] To achieve an arbitrary time delay for e(n), interpolation can be used for fitting. For example, linear interpolation is employed, simply assuming that the signal changes linearly between two sampling points. Suppose a time delay is required. Seconds, and no For integer multiples of the specified value, first calculate the delay. The number of sampling points corresponding to a second is shown in equation (14): (14) Then S is decomposed into an integer part N and a fractional part μ, as shown in equation (15): in This indicates rounding down. Delay e(n). The output data after one second can be expressed as equation (16): (16) Equation (16) can delay e(n) by any τ seconds, thereby adjusting the arbitrary phase increment of e(n) in the frequency domain. However, Equation (16) is not very accurate, and the linear assumption is too "coarse" for speech signals, which will introduce distortion. To improve accuracy, higher-order polynomials (such as quadratic or cubic) can be used to fit the curve between sampling points, such as Lagrange interpolation or cubic spline interpolation. If a higher precision delay is still required, a fractional delay filter can be used, such as the Farrow filter. The Farrow filter is a fractional delay filter based on polynomial interpolation. Its core idea is to achieve a continuously variable fractional delay through parameterized polynomial coefficients.
[0048] In the time domain, the general approach to adjusting the amplitude of the input signal is to suppress the time-domain signal with excessive amplitude, while keeping the time-domain signal at other sampling points unchanged or appropriately amplified. The following explanation uses a Dynamic Range Compressor (DRC) as an example to illustrate amplitude adjustment: right The magnitude of the logarithmic domain is calculated as shown in equation (17): (17) in This represents the linear amplitude of the input signal. This represents the dB value in the logarithmic field. It can also be calculated using the root mean square (RMS) value of the input signal.
[0049] Will The amplitude of the time-domain signal to be processed at sampling points exceeding a preset amplitude threshold T (i.e., the first threshold, in decibels, dB) is suppressed. A transition bandwidth (knee width) can be introduced to achieve a smooth gain change near the threshold T. The transition bandwidth is denoted as W, also in dB. The target gain for adjusting the amplitude of the time-domain signal to be processed at each sampling point is... (Unit: dB) is achieved through formula (18): (18) In equation (18), R represents the compression ratio of the DRC input parameter, which is a preset value. The envelope detector in the DRC can be implemented using equation (19).
[0050] (19) in , and The attack time and release time are obtained from the DRC input parameters respectively through corresponding calculation formulas (see relevant technologies for details, which will not be elaborated here), and both are preset values. The gain factor in the linear domain can be obtained, as shown in equation (20). This is a protection measure applied to prevent the gain adjustment from being too small.
[0051] (20) right The amplitude adjustment in the time domain can be expressed as equation (21): (twenty one) It should be noted that amplitude adjustment can only be based on amplitude information in the time domain. Although modules such as noise reduction can also achieve amplitude adjustment, these modules utilize other information and cannot be simply classified as amplitude adjustment. The DRC described above is only one time-domain amplitude adjustment scheme, and other time-domain amplitude adjustment schemes are also applicable, such as side-chain compression technology.
[0052] In summary, the decorrelation process that simultaneously adjusts the frequency, phase, and amplitude of the time-domain signal e(n) can be expressed as equation (22): (twenty two) Equation (21) is only one scheme for amplitude adjustment in the time domain. There are many schemes for time domain amplitude adjustment, such as time domain equalizers (EQ) based on Infinite Impulse Response (IIR), and time domain automatic gain control (AGC) schemes. Regarding the DRC implementation used as an example in this invention, there are other variations, such as multiband dynamic range compression. Multiband dynamic range compression can achieve greater amplitude compression for bands that are more prone to feedback, while other bands do not undergo or undergo less amplitude compression, which is beneficial to the overall sound quality improvement.
[0053] The decorrelation processing method in the above embodiments is a time-domain-based decorrelation processing method. It uses frequency adjustment, phase adjustment, and amplitude adjustment to perform decorrelation operations on the output signal of the adaptive filter in the AFC scheme or the output signal of the notch filter in the notch method. Since it is implemented in the time domain, it can achieve small delay and realize real-time processing. The module implemented using the decorrelation processing method in the above embodiments can be used as a plug-in, independent of the design of the adaptive filter in the AFC scheme and the design of the notch filter in the notch method.
[0054] The decorrelation processing method in the above embodiments is based on frequency adjustment, phase adjustment and amplitude adjustment in the time domain, rather than decorrelation based on data-trained neural networks. This can avoid the data requirements during model training and reduce the computation and storage requirements during cockpit chip deployment.
[0055] Figure 4 A structural diagram of the decorrelation processing apparatus provided in an embodiment of this application is shown. Figure 4 As shown, a decorrelation processing device 400 includes: The acquisition module 401 is used to acquire the time-domain signal to be processed at the current sampling point. The time-domain signal to be processed is obtained by performing howling suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. The processing module 402 is used to perform time-domain decorrelation processing on the time-domain signal to be processed to obtain the target time-domain signal. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment.
[0056] The decorrelation processing apparatus 400 provided in this application embodiment can implement the various processes implemented in the aforementioned decorrelation processing method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0057] Figure 5 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0058] The electronic device may include a processor 601 and a memory 602 storing computer program instructions.
[0059] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0060] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 602 is non-volatile solid-state memory.
[0061] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to the first or second aspect of this disclosure.
[0062] The processor 601 reads and executes computer program instructions stored in the memory 602 to implement any of the decorrelation processing methods in the above embodiments.
[0063] In one example, the electronic device may also include a communication interface 603 and a bus 610. For example, Figure 5 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 610 and complete communication with each other.
[0064] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0065] Bus 610 includes hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0066] Furthermore, in conjunction with the decorrelation processing methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the decorrelation processing methods in the above embodiments.
[0067] This application provides a computer program product in which the instructions are executed by the processor of an electronic device, causing the electronic device to perform any of the decorrelation processing methods described in the above embodiments.
[0068] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described as examples. However, the method process of this application is not limited to the specific steps described. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0069] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0070] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0071] The foregoing flowcharts and / or block diagrams of methods, apparatus (systems) according to embodiments of the present disclosure have described various aspects of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0072] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A decorrelation processing method, characterized in that, The method includes: The time-domain signal to be processed at the current sampling point is obtained. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the sound pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the sound pickup device is located. The time-domain signal to be processed is subjected to decorrelation processing in the time domain to obtain the target time-domain signal. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment.
2. The decorrelation processing method according to claim 1, characterized in that, Performing decorrelation processing in the time domain on the time-domain signal to be processed to obtain the target time-domain signal includes: If the amplitude of the time-domain signal to be processed is greater than the first threshold, or if the amplitude of the time-domain signal to be processed is less than the second threshold, then the time-domain signal to be processed is subjected to decorrelation processing to obtain the target time-domain signal.
3. The decorrelation processing method according to claim 1, characterized in that, Performing decorrelation processing in the time domain on the time-domain signal to be processed to obtain the target time-domain signal includes: Based on the time-domain signal to be processed, a first analytic signal represented in rectangular coordinates is constructed, wherein the time-domain signal to be processed is the real part of the first analytic signal, and the Hilbert transform of the time-domain signal to be processed is the imaginary part of the first analytic signal; The first analytical signal is transformed to obtain a second analytical signal represented in polar coordinates. The second analytical signal includes instantaneous amplitude and instantaneous phase. The instantaneous phase of the second analytical signal is adjusted using the target phase increment to obtain the first time-domain signal; The first time-domain signal is transformed to obtain a second time-domain signal represented in rectangular coordinates; The real part of the second time-domain signal is extracted to obtain the target time-domain signal.
4. The decorrelation processing method according to claim 3, characterized in that, The amplitude adjustment includes: adjusting the real part of the second time-domain signal; The phase adjustment includes: adjusting the angle value when the target phase increment is determined based on the angle value; The frequency adjustment includes adjusting the moving frequency when the target phase increment is determined based on the moving frequency.
5. The decorrelation processing method according to claim 1, characterized in that, The phase adjustment includes: delaying the time-domain signal to be processed. Second, It is an integer multiple of the sampling period, or, It is not an integer multiple of the sampling period.
6. The decorrelation processing method according to claim 5, characterized in that, exist When the time interval is not an integer multiple of the sampling period, interpolation is performed on the time-domain signal to be processed based on historical sampling points to obtain the delay. The time-domain signal to be processed corresponding to a second.
7. The decorrelation processing method according to claim 1, characterized in that, Obtain the time-domain signal to be processed at the current sampling point, including: The acoustic feedback elimination AFC process is used to process the signal to be played and the initial signal to obtain the acoustic path between the playback device and the pickup device. The first audio signal is estimated based on the acoustic path to obtain an estimated audio signal; The audio estimation signal is subtracted from the initial signal to obtain the time-domain signal to be processed at the current sampling point.
8. A decorrelation processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the time-domain signal to be processed at the current sampling point. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. The processing module is used to perform time-domain decorrelation processing on the time-domain signal to be processed to obtain the target time-domain signal. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment.
9. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the decorrelation processing method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the decorrelation processing method as described in any one of claims 1-7.
11. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the decorrelation processing method as described in any one of claims 1-7.