Decorrelation processing method and apparatus, device, medium, and product

CN122554762APending Publication Date: 2026-08-11ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

如果没有去相关处理装置,啸叫抑制方法会消除部分期望的说话人声信号,造成啸叫抑制效果较差

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554762A_ABST
    Figure CN122554762A_ABST
Patent Text Reader

Abstract

This application discloses a decorrelation processing method, apparatus, device, medium, and product. The method includes: acquiring a time-domain signal to be processed, which is obtained by performing howling suppression processing on an initial signal acquired by a pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting a signal to be played through a playback device; the second audio signal is a human voice signal in the environment where the pickup device is located. The method involves performing STFT on the time-domain signal to be processed to obtain multiple first frequency domain signals corresponding to multiple time frames. Decorrelation processing is then performed on the first frequency domain signals corresponding to each time frame to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment, and amplitude adjustment. Finally, ISTFT is performed on each second frequency domain signal to obtain a target time-domain signal. Through the above processing, the correlation between the first audio signal and the second audio signal can be reduced, thereby improving the howling suppression effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of signal processing technology, and in particular relates to a decorrelation processing method, apparatus, device, medium and product. Background Technology

[0002] With the development of vehicle technology, in-car entertainment scenarios have become increasingly diversified, and in-car karaoke, especially microphone-free karaoke, has gradually become a popular application in in-car entertainment systems. In karaoke scenarios, the basic sound reinforcement system consists of in-car microphones (handheld or in-car microphones) and speakers (and their amplifiers). The sound of passengers singing is picked up by the in-car microphones, amplified by the amplifier, and then played through the in-car speakers. The signal picked up by the microphones is amplified and played by the speakers, and the played sound signal is picked up by the microphones again, resulting in a closed-loop positive feedback loop between the microphones and speakers. If the sound is directly amplified without processing, the signal will be continuously fed back and amplified, eventually causing feedback, which will greatly affect the experience of the local sound reinforcement scenario. Therefore, appropriate feedback suppression methods are needed.

[0003] Feedback suppression methods all require additional de-processing equipment to function properly. Without such equipment, feedback suppression methods will eliminate some of the desired speaker voice signal, resulting in poor feedback suppression performance. Summary of the Invention

[0004] This application provides a decorrelation processing method, apparatus, device, medium, and product that can improve the howling suppression effect.

[0005] In a first aspect, embodiments of this application provide a decorrelation processing method, the method comprising: The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. Perform a short-time Fourier transform (STFT) on the time-domain signal to be processed to obtain first frequency domain signals corresponding to multiple time frames; The first frequency domain signal corresponding to each time frame is subjected to decorrelation processing to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment. Perform an inverse short-time Fourier transform (ISTFT) on each of the second frequency domain signals to obtain the target time domain signal.

[0006] Secondly, embodiments of this application provide a decorrelation processing apparatus, the apparatus comprising: The acquisition module is used to acquire a time-domain signal to be processed. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. The first conversion module is used to perform a short-time Fourier transform (STFT) on the time-domain signal to be processed to obtain a first frequency domain signal corresponding to multiple time frames. The decorrelation module is used to perform decorrelation processing on the first frequency domain signal corresponding to each time frame to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment. The second conversion module is used to perform an inverse short-time Fourier transform (ISTFT) on each of the second frequency domain signals to obtain the target time domain signal.

[0007] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the decorrelation processing method as described in the first aspect.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the decorrelation processing method as described in the first aspect.

[0009] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the decorrelation processing method as described in the first aspect.

[0010] In this embodiment of the application, the correlation between the first audio signal (i.e., the audio signal played by the playback device) and the second audio signal (the human voice signal collected by the microphone) can be reduced through the above-described decorrelation process, thereby improving the effect of howling suppression. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of a local sound reinforcement scenario in the related technologies provided in the embodiments of this application; Figure 2 This is a schematic diagram of the decorrelation processing method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the howling suppression process based on AFC and decorrelation module provided in the embodiments of this application; Figure 4 This is a schematic diagram of the decorrelation processing apparatus provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0013] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0015] In all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. Additionally, when embodiments of this application require access to sensitive personal information, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments obtained.

[0016] With the development of vehicle technology, in-car entertainment scenarios have become increasingly diversified, and in-car karaoke, especially microphone-free karaoke, has gradually become a popular application in in-car entertainment systems. In a karaoke scenario, an in-car microphone (handheld or in-car microphone), speakers (and their amplifiers) form a basic sound reinforcement system. The voices of passengers singing are picked up by the in-car microphones, amplified by the amplifier, and then played back through the in-car speakers. Figure 1 As shown.

[0017] Figure 1 As shown, the smiley face represents the speaker, the dotted line segment represents the sound emitted by the speaker, the horizontal line segment represents the speaker's voice, the solid line represents the entire audio loop system, G represents the audio system (including the power amplifier), and F represents the acoustic path between the speaker and the microphone. Figure 1 As can be seen, this is a typical local sound reinforcement system. In-car karaoke is just one application scenario; in-vehicle communication (ICC) systems can also be used in in-vehicle applications. Figure 1 This is for display purposes. Of course, large classrooms, conference rooms, karaoke rooms, and other similar applications all fall under the category of local sound reinforcement and can also be represented using this technology. Figure 1 express.

[0018] The signal picked up by the microphone is amplified and played by the speaker, and the played sound signal is then picked up by the microphone again, creating a closed-loop positive feedback loop between the microphone and the speaker. If the signal is amplified directly without processing, it will be continuously fed back and amplified, eventually causing howling, which will greatly affect the experience in local sound reinforcement scenarios. Therefore, appropriate howling suppression methods are needed.

[0019] There are three main types of howling suppression methods.

[0020] The first method is frequency shifting, which shifts the frequency of the signal collected by the microphone. The greater the frequency shift, the better the feedback suppression effect, but the greater the loss of sound quality.

[0021] The second method is the notch filter method, which uses a notch filter to process the detected howling frequency band, thereby achieving howling suppression. This method requires first detecting the howling sound and accurately locating the frequency band where the howling occurs before a suitable notch filter can be used for howling suppression. No matter how precise the howling detection and howling frequency location are, this method can only work after the howling has occurred.

[0022] The third method is Acoustic Feedback Cancellation (AFC). AFC uses an adaptive filter to estimate the acoustic path between the microphone and the loudspeaker, thereby eliminating the sound signal emitted by the loudspeaker from the microphone signal. If the adaptive filter can perfectly estimate the acoustic path between the microphone and the loudspeaker, the howling problem will not occur; that is, the AFC method is the theoretically optimal solution among howling suppression schemes.

[0023] There are many design methods for adaptive filtering in AFC (Augmented Feedback Control), such as Least Mean Squares (LMS), Normalized Least Mean Squares (NLMS), Recursive Least Squares (RLS), and Kalman Filter. However, these methods all require additional decorrelation processing to function properly. This is because the signal played by the speaker is actually the signal of the speaker's voice captured by the microphone after passing through the audio loop system, and the two are highly correlated. Without decorrelation processing, the adaptive filtering in AFC will eliminate some of the desired speaker's voice signal, resulting in poor feedback suppression.

[0024] To address the problems of the prior art, embodiments of this application provide a decorrelation processing method, apparatus, device, medium, and product. The decorrelation processing method provided in this application embodiment will be described first below.

[0025] Figure 2 A flowchart illustrating a decorrelation processing method according to an embodiment of this application is shown. Figure 2 As shown, the decorrelation processing method provided in this application embodiment is applied to an electronic device and includes the following steps 101-104, wherein: Step 101: Obtain the time-domain signal to be processed. The time-domain signal to be processed is obtained by performing acoustic feedback cancellation (AFC) processing on the initial signal collected by the pickup device.

[0026] The sound pickup device can refer to a microphone. The initial signal (y) collected by the sound pickup device includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through a playback device (e.g., a speaker) (r); the second audio signal is the human voice signal (s) in the environment where the sound pickup device is located.

[0027] For example, step 101 specifically includes: processing the signal to be played and the initial signal using AFC to obtain the acoustic path between the playback device and the pickup device; estimating the first audio signal based on the acoustic path to obtain an audio estimation signal; and subtracting the audio estimation signal from the initial signal to obtain the time-domain signal to be processed.

[0028] like Figure 3 As shown, x is the signal to be played by the speaker. The acoustic path F between the speaker and the microphone is convolved on x to obtain the speaker's broadcast signal r. This broadcast signal r is collected by the microphone. The total signal (i.e. the initial signal) y collected by the microphone includes the broadcast signal r and the human voice signal s. The overall signal y and the signal to be played x are first processed by AFC to estimate the acoustic path between the speaker and the microphone. Then based on The signal r emitted by the loudspeaker is estimated to obtain Subtract from the overall signal The output signal e after AFC is obtained, and this output signal e is the time domain signal to be processed. Figure 3 In this context, G represents the audio system (including the power amplifier), and "decorrelated" represents the decorrelation module that uses the decorrelation processing method provided in the embodiments of this application.

[0029] Will Figure 3 After replacing the AFC method with the notch filter method, the overall signal y is first detected to find the frequency band where the howling occurs. The notch filter is then used to suppress the howling in this frequency band to obtain the output signal e.

[0030] Step 102: Perform a Short-Time Fourier Transform (STFT) on the time-domain signal to be processed to obtain the first frequency domain signal corresponding to multiple time frames. The time-domain signal to be processed is converted to the frequency domain via STFT to obtain the first frequency domain signal corresponding to each time frame. Each time frame corresponds to a time period. For details, please refer to the relevant documentation on STFT, which will not be elaborated upon here.

[0031] Step 103: Perform decorrelation processing on the first frequency domain signal corresponding to each time frame to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment.

[0032] In this embodiment of the application, the first frequency domain signal corresponding to each time frame is subjected to decorrelation processing, which includes frequency adjustment, phase adjustment and amplitude adjustment of the first frequency signal.

[0033] Step 104: Perform an inverse short-time Fourier transform (ISTFT) on each of the second frequency domain signals to obtain the target time domain signal.

[0034] In this embodiment, a time-domain signal to be processed is acquired; a Short-Time Fourier Transform (STFT) is performed on the time-domain signal to obtain first frequency-domain signals corresponding to multiple time frames; decorrelation processing is performed on the first frequency-domain signals corresponding to each time frame to obtain multiple second frequency-domain signals, the decorrelation processing including frequency adjustment, phase adjustment, and amplitude adjustment; an ISTFT is performed on each second frequency-domain signal to obtain a target time-domain signal. Through the above decorrelation processing, the correlation between the first audio signal (i.e., the audio signal played by the playback device) and the second audio signal (the human voice signal collected by the microphone) can be reduced, thereby improving the effect of feedback suppression.

[0035] The decorrelation processing method provided in this application can be used as a plug-in, independent of the design of the adaptive filter in the AFC scheme and the design of the notch filter in the notch method. It has great flexibility and can be efficiently integrated into the in-vehicle karaoke and in-vehicle communication system in the vehicle cabin.

[0036] In one embodiment of this application, frequency bands with excessively large amplitudes are suppressed, that is, frequency adjustment, phase adjustment, and amplitude adjustment are performed on the first frequency domain signal with an amplitude greater than a first threshold. Specifically, step 103 involves decorrelation processing on the first frequency domain signal corresponding to each time frame to obtain multiple second frequency domain signals, including: For each of the first frequency domain signals, if the amplitude of the first frequency domain signal is greater than a first threshold, or if the amplitude of the first frequency domain signal is less than a second threshold, then the first frequency domain signal is decorrelated to obtain a second frequency domain signal. The first threshold is greater than the second threshold. On the one hand, when feedback occurs, the energy of the first frequency domain signal with an amplitude greater than the first threshold will increase significantly. Adjusting the amplitude can have an additional amplitude control effect on the first frequency domain signal, thereby improving the feedback suppression effect. On the other hand, adjusting the amplitude of the first frequency domain signal with an amplitude less than the second threshold, for example, by amplifying the amplitude, can improve the sound quality.

[0037] In one embodiment of this application, a short-time Fourier transform (STFT) is performed on the time-domain signal to be processed to obtain a first frequency-domain signal corresponding to multiple time frames, including: The time-domain signal to be processed is divided according to time to obtain the first time-domain signal corresponding to multiple time frames; For each time frame, the first time-domain signal corresponding to the time frame is frequency-shifted according to a first frequency to obtain a second time-domain signal; and... Perform a Fast Fourier Transform on the second time-domain signal to obtain the first frequency-domain signal corresponding to the time frame.

[0038] Through the above processing, the frequency of the first time-domain signal corresponding to each time frame can be shifted to prepare for subsequent decorrelation processing, thereby improving the effect of howling suppression.

[0039] The higher the first frequency, the better the decorrelation effect. However, considering other factors, the first frequency should be controlled to be less than the third threshold.

[0040] In one embodiment of this application, performing a Fast Fourier Transform on the second time-domain signal to obtain a first frequency-domain signal corresponding to the time frame includes: Perform a Fast Fourier Transform on the second time-domain signal to obtain the intermediate frequency-domain signal; If the first frequency is less than the third threshold, the amplitude of the intermediate frequency domain signal is taken as the amplitude of the first frequency domain signal, and the third threshold is determined according to the spectral resolution after STFT. The first frequency-related parameter in the intermediate frequency domain signal is replaced by the target phase increment to obtain the phase of the first frequency domain signal corresponding to the time frame. The target phase increment is used to perform the frequency adjustment and phase adjustment on the first frequency domain signal.

[0041] For example, the third threshold could be the ratio of the sampling rate to the STFT length. First frequency-related parameter Please refer to the relevant records below. It is the first frequency. This refers to the actual duration corresponding to the time frame, which will not be elaborated upon here.

[0042] For example, the target phase increment is the value obtained by weighted summation of the frequency adjustment and the phase adjustment; The frequency adjustment amount is determined based on the first frequency and the actual duration corresponding to the time frame, and the value of the frequency adjustment amount is located within a preset phase range. The phase adjustment amount is determined based on the modulation amplitude and modulation frequency.

[0043] For example, the frequency adjustment amount is ,in, It is the actual duration corresponding to the time frame; if c=0; otherwise c= Although the frequency adjustment is within the preset phase range [ , The cycle changes periodically, but the rate of change decreases from the first frequency. Control. The preset phase range can also be [0, ... The specific settings can be adjusted according to the actual situation, and are not limited here. The frequency adjustment amount of the first frequency domain signal can be adjusted by adjusting the first frequency.

[0044] Phase adjustment amount is Where a is the modulation amplitude. For modulation frequency, the sine function can be replaced by other periodic functions, such as the cosine function. The modulation amplitude and modulation frequency should not be too large, otherwise it will affect the sound quality. For example, the modulation frequency... All values ​​do not exceed 10 Hz; the modulation amplitude 'a' is set in a frequency band manner. For example, as the frequency band increases, the value of 'a' increases. For high-frequency bands, a larger value of 'a' results in better decorrelation. Of course, the modulation amplitude can also be fixed, which is not limited here. The phase adjustment of the first frequency domain signal can be adjusted by adjusting the modulation amplitude and modulation frequency.

[0045] The amplitude adjustment includes: adjusting the amplitude of the first frequency domain signal corresponding to the time frame; the frequency adjustment includes: adjusting the frequency adjustment amount of the first frequency domain signal corresponding to the time frame; the phase adjustment includes: adjusting the phase adjustment amount of the first frequency domain signal corresponding to the time frame.

[0046] The amplitude, frequency, and phase of the first frequency domain signal can be adjusted according to the actual situation to reduce the correlation between the first audio signal and the second audio signal, thereby improving the effect of howling suppression and making the target time domain signal played more clean.

[0047] The amplitude of the first frequency domain signal corresponding to the time frame can be adjusted by multiplying the amplitude by an adjustment factor. If the amplitude of the first frequency domain signal is less than a first threshold, the amplitude of the first frequency domain signal can be left unadjusted; in this case, the amplitude adjustment factor is 1.

[0048] The following examples illustrate the de-relation processing method provided in this application.

[0049] like Figure 3 As shown, x is the signal to be played by the loudspeaker. The acoustic path F between the loudspeaker and the microphone is convolved on x to obtain the loudspeaker's broadcast signal r. This broadcast signal r is collected by the microphone. The total signal y collected by the microphone includes the broadcast signal r and the human voice signal s, as shown in equation (1): (1) The overall signal y and the signal to be played x are first processed by AFC to estimate the acoustic path between the speaker and the microphone. Then based on The signal r emitted by the loudspeaker is estimated to obtain Subtract from the overall signal The output signal e after AFC is obtained, and this output signal e is the time-domain signal to be processed. As shown in equation (2): (2) E represents the result of e after passing through STFT. In the frequency domain, E can represent the amplitude and phase, as shown in equation (3), where k represents the discrete frequency band and t represents the time frame.

[0050] (3) Where k represents the frequency band, not the frequency point, and assuming the length of the STFT is N and the sampling rate is f... s Hertz (Hz), and the bandwidth of the frequency band corresponds to the spectral resolution after STFT, as shown in equation (4): (4) The first time-domain signal corresponding to time frame t according to (i.e., the first frequency) is shifted to obtain the second time-domain signal. It can be expressed as equation (5): (5) right Perform frequency domain transformation to obtain the intermediate frequency domain signal. , expressed as equation (6): (6) Note here that k and in equation (6) The meanings are different; k represents a frequency band with a certain bandwidth, while... This indicates the frequency to be shifted, measured in Hz. Here, it is assumed that the frequency to be shifted is... Compared to spectral resolution F r When the amplitude is reduced by an order of magnitude, it has no effect on the amplitude within this frequency band, as expressed in equation (7): (7) Combining equations (6) and (7), we obtain equation (8): (8) When the first frequency When the spectral resolution is an order of magnitude smaller, the first frequency This can be viewed as a phase increment, using the target phase increment. replace The first frequency domain signal corresponding to time frame t is obtained, which is expressed as equation (9): (9) For ease of explanation, assume that the frame length of the microphone data acquisition is N / 2 (the frame length and the length N of the STFT can have other relationships), then the time time (in seconds) corresponding to the time frame t can be expressed as equation (10): (10) Combining equations (9) and (10), the target phase increment increases with time t. It will get bigger and bigger. Modulo operation is needed to control the phase value within the range [ π, π], as shown in equation (11): (11) It can be a fixed value, or it can be increased linearly or non-linearly according to the increase of the frequency band. If fixed... =1 Hz, for a frequency point of 10 Hz, the shift ratio is 10%; for a frequency point of 100 Hz, the shift ratio is 1%; for a frequency point of 1000 Hz, the shift ratio is 0.1%. As the frequency increases, the shift frequency... The impact is decreasing. This means that higher frequency bands can use larger mobile frequencies.

[0051] if In the case of c = 0, otherwise c = 2π. Although the phase increment is in the interval [ π, π] changes periodically, but The rate of change within the interval is determined by control.

[0052] Target phase increment Phase modulation can also be represented by a periodic function. The form is as shown in equation (12): (12) Where a is the modulation amplitude. For modulation frequency, the sine function can be replaced by other periodic functions, such as the cosine function. The modulation amplitude and modulation frequency should not be too large, otherwise it will affect the sound quality. For example, the modulation frequency... All frequencies are below 10Hz; the modulation amplitude 'a' is set in a frequency band manner. For example, as the frequency band increases, the value of 'a' increases. For high-frequency bands, a larger value of 'a' results in better decorrelation. Of course, the modulation amplitude can also be fixed, and this is not limited here.

[0053] In addition to choosing a fixed value and increasing linearly with frequency, the modulation amplitude 'a' can also be increased non-linearly, such as increasing it from a low frequency band and then maintaining it constant after reaching a certain high frequency band, like the part of the tanh function where x>0.

[0054] For amplitude The overall approach to adjustment is to suppress frequency bands with excessive amplitude, while keeping other frequency bands unchanged or appropriately amplifying them. One implementation method is a frequency-domain-based Dynamic Range Compressor (DRC).

[0055] The energy in each frequency band k is used as input for gain calculation, as shown in equation (13): (13) Will The gain in the frequency band exceeding the preset amplitude threshold T (i.e., the first threshold, in decibels, dB) is suppressed. Referring to the time-domain DRC scheme, a transition bandwidth (knee width) can be introduced to achieve smooth gain changes near the threshold T. The transition bandwidth is denoted as W, in dB. The target gain for amplitude adjustment is... (Unit: dB) is achieved through equation (14): (14) In the formula, R represents the compression ratio of the DRC input parameter, which is a preset value. The envelope detector in the DRC can be implemented using equation (15): (15) Where D , and The attack time and release time are determined by the DRC input parameters, respectively, and both are preset values. The amplitude adjustment level of the envelope detector output, measured in dB. Indicates the previous frame .Depend on The linear amplitude adjustment coefficient can be obtained, as shown in equation (16). This is a protection measure applied to prevent the gain adjustment from being too small.

[0056] (16) Equation (16) is only one method of amplitude adjustment. Other methods, such as automatic gain control in the frequency domain, can also be used for amplitude adjustment in the frequency band. It is worth noting that amplitude adjustment can only be based on amplitude information in the frequency band. Although modules such as noise reduction can also achieve amplitude adjustment, these modules utilize other information and cannot be simply classified as amplitude adjustment. Regarding frequency band amplitude... The adjustment method is based only on the amplitude information in the frequency band. The equalizer (EQ) that adjusts the frequency band in the frequency domain can also be used for amplitude control.

[0057] In summary, regarding formula (3) The final result of simultaneously adjusting the frequency, phase, and amplitude can be expressed as equation (17).

[0058] (17) Among them, the target phase increment , .

[0059] in Used to control phase increment The proportion of mid-frequency adjustment and phase adjustment. Finally, By using the inverse short-time Fourier transform (ISTFT), the processed frequency domain data is transformed back to the time domain, thereby achieving decorrelation of the AFC output signal, as shown in equation (18): (18) It needs to be explained that... Figure 3 The AFC method can be replaced with the notch method, and the relevant positions and methods are equally applicable.

[0060] The decorrelation processing method in the above embodiments is a frequency domain-based decorrelation processing method. It performs decorrelation operation on the output signal of the adaptive filter in the AFC scheme or the output signal of the notch filter in the notch method, and is a method that can achieve real-time processing. The module implemented using the decorrelation processing method in the above embodiments can be used as a plug-in, independent of the design of the adaptive filter in the AFC scheme and the design of the notch filter in the notch method.

[0061] The decorrelation processing method in the above embodiments is based on frequency adjustment, phase adjustment and amplitude adjustment in the frequency domain, rather than decorrelation based on data-trained neural networks. This can avoid the data requirements during model training and reduce the computation and storage requirements during cockpit chip deployment.

[0062] Figure 4 A structural diagram of the decorrelation processing apparatus provided in an embodiment of this application is shown. Figure 4 As shown, a decorrelation processing device 400 includes: The acquisition module 401 is used to acquire a time-domain signal to be processed. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. The first conversion module 402 is used to perform a short-time Fourier transform (STFT) on the time-domain signal to be processed to obtain a first frequency domain signal corresponding to multiple time frames. The decorrelation module 403 is used to perform decorrelation processing on the first frequency domain signal corresponding to each time frame to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment. The second conversion module 404 is used to perform short-time inverse Fourier transform (ISTFT) on each of the second frequency domain signals to obtain the target time domain signal.

[0063] The decorrelation processing apparatus 400 provided in this application embodiment can implement the various processes implemented in the aforementioned decorrelation processing method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0064] Figure 5 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.

[0065] The electronic device may include a processor 601 and a memory 602 storing computer program instructions.

[0066] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0067] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 602 is non-volatile solid-state memory.

[0068] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to the first or second aspect of this disclosure.

[0069] The processor 601 reads and executes computer program instructions stored in the memory 602 to implement any of the decorrelation processing methods in the above embodiments.

[0070] In one example, the electronic device may also include a communication interface 603 and a bus 610. For example, Figure 5 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 610 and complete communication with each other.

[0071] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0072] Bus 610 includes hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0073] Furthermore, in conjunction with the decorrelation processing methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the decorrelation processing methods in the above embodiments.

[0074] This application provides a computer program product in which the instructions are executed by the processor of an electronic device, causing the electronic device to perform any of the decorrelation processing methods described in the above embodiments.

[0075] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described as examples. However, the method process of this application is not limited to the specific steps described. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0076] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0077] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0078] The foregoing flowcharts and / or block diagrams of methods, apparatus (systems) according to embodiments of the present disclosure have described various aspects of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0079] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A decorrelation processing method, characterized by, The method includes: The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. Perform a short-time Fourier transform (STFT) on the time-domain signal to be processed to obtain first frequency domain signals corresponding to multiple time frames; The first frequency domain signal corresponding to each time frame is subjected to decorrelation processing to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment. Perform an inverse short-time Fourier transform (ISTFT) on each of the second frequency domain signals to obtain the target time domain signal.

2. The decorrelation processing method of claim 1, wherein, The first frequency domain signal corresponding to each time frame is subjected to decorrelation processing to obtain multiple second frequency domain signals, including: For each of the first frequency domain signals, if the amplitude of the first frequency domain signal is greater than a first threshold, or if the amplitude of the first frequency domain signal is less than a second threshold, then the first frequency domain signal is subjected to decorrelation processing to obtain the second frequency domain signal.

3. The decorrelation processing method of claim 1, wherein, Perform a Short-Time Fourier Transform (STFT) on the time-domain signal to be processed to obtain a first frequency domain signal corresponding to multiple time frames, including: The time-domain signal to be processed is divided according to time to obtain the first time-domain signal corresponding to multiple time frames; For each time frame, the first time-domain signal corresponding to the time frame is frequency-shifted according to a first frequency to obtain a second time-domain signal; and... Perform a Fast Fourier Transform on the second time-domain signal to obtain the first frequency-domain signal corresponding to the time frame.

4. The decorrelation processing method of claim 3, wherein, Performing a Fast Fourier Transform on the second time-domain signal to obtain the first frequency-domain signal corresponding to the time frame includes: Perform a Fast Fourier Transform on the second time-domain signal to obtain the intermediate frequency-domain signal; If the first frequency is less than the third threshold, then the amplitude of the intermediate frequency domain signal is taken as the amplitude of the first frequency domain signal, and the third threshold is determined according to the spectral resolution after STFT. The first frequency-related parameter in the intermediate frequency domain signal is replaced by the target phase increment to obtain the phase of the first frequency domain signal corresponding to the time frame. The target phase increment is used to perform the frequency adjustment and phase adjustment on the first frequency domain signal.

5. The decorrelation processing method according to claim 4, characterized in that, The target phase increment is the value obtained by weighted summation of the frequency adjustment and the phase adjustment; The frequency adjustment amount is determined based on the first frequency and the actual duration corresponding to the time frame, and the value of the frequency adjustment amount is located within a preset phase range. The phase adjustment amount is determined based on the modulation amplitude and modulation frequency.

6. The decorrelation processing method according to claim 5, characterized in that, The amplitude adjustment includes: adjusting the amplitude of the first frequency domain signal corresponding to the time frame; The frequency adjustment includes: adjusting the frequency adjustment amount of the first frequency domain signal corresponding to the time frame; The phase adjustment includes adjusting the phase adjustment amount of the first frequency domain signal corresponding to the time frame.

7. The decorrelation processing method according to claim 1, characterized in that, Acquire the time-domain signal to be processed, including: The acoustic feedback elimination AFC process is used to process the signal to be played and the initial signal to obtain the acoustic path between the playback device and the pickup device. The first audio signal is estimated based on the acoustic path to obtain an estimated audio signal; The audio estimation signal is subtracted from the initial signal to obtain the time-domain signal to be processed.

8. A decorrelation processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire a time-domain signal to be processed. The time-domain signal to be processed is obtained by performing feedback suppression processing on the initial signal collected by the pickup device. The initial signal includes a first audio signal and a second audio signal. The first audio signal is obtained by outputting the signal to be played through the playback device. The second audio signal is the human voice signal in the environment where the pickup device is located. The first conversion module is used to perform a short-time Fourier transform (STFT) on the time-domain signal to be processed to obtain a first frequency domain signal corresponding to multiple time frames. The decorrelation module is used to perform decorrelation processing on the first frequency domain signal corresponding to each time frame to obtain multiple second frequency domain signals. The decorrelation processing includes frequency adjustment, phase adjustment and amplitude adjustment. The second conversion module is used to perform an inverse short-time Fourier transform (ISTFT) on each of the second frequency domain signals to obtain the target time domain signal.

9. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the decorrelation processing method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the decorrelation processing method as described in any one of claims 1-7.

11. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the decorrelation processing method as described in any one of claims 1-7.