Method of processing an audio signal, electronic device and vehicle

By determining the acoustic transmission path and inverse filter function of the audio input module inside the vehicle, and utilizing adaptive filtering technology, the problem of poor speech recognition performance in noisy and reverberant environments inside the vehicle is solved, achieving efficient speech enhancement with low computational cost.

CN120108410BActive Publication Date: 2026-03-31BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing in-vehicle voice signal processing algorithms are computationally intensive and rely on high-end chips when dealing with in-vehicle noise and reverberation environments, ignoring the special environment inside the vehicle, resulting in poor voice recognition performance.

Method used

By determining the impulse response function of the acoustic transmission path of multiple audio input modules in the vehicle, the inverse filter function is obtained. Adaptive filtering technology is used to enhance the human voice signal and eliminate noise. The in-vehicle environment is characterized by a simulated sound source and acoustic transfer function.

Benefits of technology

It significantly improves speech enhancement performance, reduces computational load, and enhances speech recognition quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108410B_ABST
    Figure CN120108410B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method of processing an audio signal, an electronic device and a vehicle. The method comprises: determining an impulse response function of a plurality of audio input modules in the vehicle for representing an acoustic transfer path of a sound source to the audio input modules; determining an inverse filter function of a corresponding audio input module in the plurality of audio input modules according to the impulse response function; determining a voice enhancement signal and a noise reference signal according to the inverse filter function; and adaptively filtering an audio signal input to the audio input module according to the voice enhancement signal and the noise reference signal. The method according to the embodiments of the present disclosure can utilize the acoustic environment of the vehicle itself to significantly improve the performance of speech enhancement with a lower computational load, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of vehicles, and particularly to a method for processing audio signals, electronic devices, and vehicles. Background Technology

[0002] More and more vehicles, especially electric vehicles, now feature intelligent voice interaction capabilities. Intelligent voice interaction requires front-end signal processing, after which the signal is sent to back-end speech recognition to understand what the user is saying (i.e., human voice or speech, such as voice commands). The in-vehicle environment is characterized by noise and reverberation; therefore, front-end speech signal processing needs to suppress reverberation and reduce noise to make the speech clearer, resulting in better recognition and improved voice interaction. Currently, multi-microphone pickup is commonly used in vehicles. However, traditional speech enhancement algorithms suffer from varying degrees of complexity. Some algorithms have poor signal processing capabilities, while others, although capable, neglect the unique environment of the in-vehicle space, leading to complex algorithms, high computational demands, and reliance on high-end chips. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for processing audio signals is provided. The method includes: determining an impulse response function for a plurality of audio input modules in a vehicle, representing an acoustic transmission path from a sound source to the audio input module; determining an inverse filter function for a corresponding audio input module among the plurality of audio input modules based on the impulse response function; determining a human voice enhancement signal and a noise reference signal based on the inverse filter function; and adaptively filtering the audio signal input to the audio input modules based on the human voice enhancement signal and the noise reference signal.

[0004] In some embodiments, determining the impulse response function includes: receiving an attenuation-compensated sweep signal output from an analog sound source from the audio input module; and determining the impulse response function based on the sweep signal received from the audio input module and the inverse signal of the sweep signal.

[0005] In some embodiments, determining the inverse filter function includes: determining the inverse filter function based at least on the impulse response function and the impulse signal.

[0006] In some embodiments, determining the voice enhancement signal and the noise reference signal includes: determining the voice enhancement signal and the noise reference signal based on a first audio signal input by a first audio input module and its corresponding inverse filter function, and based on a second audio signal input by a second audio input module and its corresponding inverse filter function.

[0007] In some embodiments, adaptive filtering of the audio signal includes: determining noise transmission path parameters based on the noise reference signal; and determining the human voice signal in the audio signal based on the human voice enhancement signal, the noise reference signal, and the noise transmission path parameters.

[0008] In a second aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing machine-executable instructions that, when executed by the at least one processing unit, cause the device to perform actions including: determining an impulse response function for representing an acoustic transmission path from a sound source to the audio input module of a plurality of audio input modules in a vehicle; determining an inverse filter function for a corresponding audio input module among the plurality of audio input modules based on the impulse response function; determining a human voice enhancement signal and a noise reference signal based on the inverse filter function; and adaptively filtering an audio signal input to the audio input module based on the human voice enhancement signal and the noise reference signal.

[0009] In a third aspect of this disclosure, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0010] In a fourth aspect of this disclosure, a vehicle is provided. The vehicle includes: a plurality of audio input modules; and electronic devices according to the second aspect described above.

[0011] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0013] Figure 1 A simplified schematic diagram of a scenario in which the method according to embodiments of the present disclosure can be applied is shown;

[0014] Figure 2 A schematic flowchart of a method for processing audio signals according to an embodiment of the present disclosure is shown;

[0015] Figure 3 An example control block diagram of an adaptive filtering method for processing audio signals according to an embodiment of the present disclosure is shown;

[0016] Figure 4 An exemplary flowchart of a method for processing audio signals according to embodiments of the present disclosure is shown; and

[0017] Figure 5 A schematic block diagram of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0020] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0021] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.

[0022] As used herein, a “unit,” “operational unit,” or “subunit” can consist of any suitable machine learning model or network. As used herein, a set of elements or similar expressions can include one or more such elements.

[0023] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0024] Currently, there are two main approaches to in-vehicle front-end signal processing systems: beamforming and deep learning. Beamforming methods are primarily divided into fixed beamforming and adaptive beamforming. Fixed beamforming is suitable for applications with fixed voice sources or requiring a specific transmission direction; that is, it receives audio signals in one direction. Adaptive beamforming adjusts the weights and phase settings of the antenna array based on the real-time received voice signal and noise environment to provide better voice quality and microphone array performance.

[0025] Deep learning methods primarily learn the commonalities and differences across frequency bands of different microphones to obtain clean speech. Both beamforming and deep learning methods are currently mature. Beamforming has lower computational cost but lower performance than deep learning. Deep learning performs better than beamforming but is computationally intensive. When ported to digital signal processors (DSPs), it increases the processor's cost. Furthermore, deep learning methods are computationally expensive, increasing resource consumption, and require a large amount of data for initial training.

[0026] The inventors discovered through research that both methods neglect the unique environment of the vehicle interior. While this environment includes reverberation and noise, the driver's position is relatively fixed, and the vehicle's enclosure and interior decoration change little. Leveraging this characteristic, the inventors proposed a method for processing audio signals according to embodiments of this disclosure. This method pre-characterizes the propagation path of ambient sound (including human voice or voice control voice). This characterized propagation path is hereinafter referred to as the acoustic transfer function or impulse response function. By determining the acoustic transfer function from the mouth (i.e., the source of the voice command) to each audio input module (e.g., a microphone), inverting the acoustic transfer function yields the inverse filter function, allowing the reconstruction of the human voice (e.g., voice control voice). Adding the human voice signals received by the microphones yields the enhanced human voice signal; subtracting them yields the pure noise component (noise reference signal). Adaptive filtering then perfectly eliminates the noise. In this way, the performance of voice enhancement can be significantly improved with lower computational load by utilizing the vehicle's own acoustic environment, thereby enhancing the user experience.

[0027] The inventive concept according to this disclosure will be described below with reference to the accompanying drawings. Figure 1 A simplified schematic diagram of a vehicle to which the method according to embodiments of the present disclosure can be applied is shown. Figure 1 As shown, the vehicle according to an embodiment of this disclosure includes a processing unit 101 and an audio input module array (e.g., a microphone array). The processing unit 101 may include, for example, a smart cockpit development platform (CDP). The audio input module array may include multiple audio input modules (i.e., microphones). The following description will primarily use an audio input module array with two microphones as an example to illustrate the concept of this disclosure. It should be understood that the same applies to audio input module arrays with more microphones, and these will not be described in detail below.

[0028] Figure 2 A flowchart illustrating the algorithm according to an embodiment of this disclosure is shown. The method mainly consists of four steps, which will be described below. The first and second steps mentioned below are typically performed only once after the model of the audio input module (i.e., microphone) used in the vehicle is determined. If the model and location of the audio input module used in the vehicle do not change, the determined acoustic transfer function and inverse filter function can be embedded in the processing unit for subsequent use.

[0029] The first step is to measure the acoustic transmission path from the human mouth (i.e., the source of the voice command) to multiple audio input modules (i.e., microphones), that is, the acoustic transfer function or impulse response function.

[0030] In determining the impulse function, a simulated sound source is used to play a frequency sweep signal. In some embodiments, the simulated sound source can be an artificial mouth. An artificial mouth is a loudspeaker device installed in a sealed cavity, with a directivity and radiation pattern similar to the average human mouth, used to simulate human mouth sound production, thus suitable for testing acoustic parameters such as frequency response and distortion of telephone transmitters and microphones. In other words, an artificial mouth is a special speaker with low distortion and high stability, used to obtain a sound source device that more closely resembles the sound produced by a real human mouth, and is generally used for professional acoustic testing.

[0031] A swept frequency signal generally refers to a monophonic sine wave signal whose frequency changes continuously within a certain frequency range according to a linear / logarithmic law. Sweep frequencies, as well as their inverse (discussed later), can be generated by setting various parameters of the controller. Since the artificial mouth itself responds inconsistently to different frequency bands in the swept frequency signal, to ensure the accuracy of the determined acoustic transmission path, the artificial mouth itself can be calibrated first to compensate for the attenuation caused by each frequency band.

[0032] To avoid interference from external noise during calibration, calibration can be performed in an anechoic chamber. For example, in a space simulating the internal sound environment of a vehicle (which could be a sealed model of the entire vehicle), a sweep frequency signal is played using an artificial mouth to measure the attenuation of the sweep frequency signal at various frequency bands, thereby completing the calibration of the artificial mouth.

[0033] The artificial mouthpiece is typically positioned in the driver's seat, and the audio input module is located in the same spot where other audio input modules will be placed in the vehicle. When determining the impact response function, an attenuated, compensated sweep signal is played through the artificial mouthpiece, while the audio input module receives the signal.

[0034] Next, for each audio input module, the impulse response function of the microphone can be determined by convolving the attenuation-compensated sweep signal with the inverse of the sweep signal. For example, the impulse response function can be determined using the following equation (1).

[0035]

[0036] Where h(i) represents the impulse response function, i.e., the acoustic transfer function; S(i) is the signal received by the i-th audio input module; and inv_sweep is the inverse signal of the sweep signal, which, as mentioned above, can be generated by the controller at the same time as or after the sweep signal is generated by adjusting the control parameters.

[0037] After determining the impulse response function for each audio input module, the second step can be performed: calculating the collective inverse path of the acoustic transmission path, which involves inverting the impulse response function to determine the inverse filter function. Based on signal and transmission principles, the impulse response function (i.e., the acoustic transfer function) h(i) and the inverse filter function h... inv (i) The impulse signal δ can be obtained by convolution, as shown in equation (2) below.

[0038]

[0039] The impulse signal δ represents a special function that is nearly infinite at a certain moment, but nothing at other moments.

[0040] In some embodiments, spectral decomposition and Diophantine diagrams can be used together to determine the inverse filter function. Spectral decomposition, also known as eigenvalue decomposition or similar canonical form decomposition, is a method of decomposing a matrix into a product of matrices represented by its eigenvalues ​​and eigenvectors. For this scheme, the equation (3) established based on the spectral decomposition algorithm is shown below. It is solved by establishing an equation relationship between the impulse response function (i.e., the acoustic transfer function) and its conjugate function and the function and its corresponding conjugate function in the minimum phase system.

[0041] β * β=h * h (3)

[0042] Diophantus's formula is a commonly used formula in cybernetics. The specific formula is shown in equation (4) below.

[0043]

[0044] Where h inv H represents the inverse filter function. * β is the conjugate function of the impact response function (i.e., the acoustic transfer function), and β is its minimum phase function. * Let δ be the conjugate function of its minimum phase function, where δ represents the impulse signal function, and Q, q, and L are... * d are intermediate parameters used in the calculation process.

[0045] The inverse filter function can be determined by jointly solving equations (3) and (4) above. It should be understood that the above embodiment using spectral decomposition and Diophantine plots to jointly solve for and determine the inverse filter function is merely illustrative and not intended to limit the scope of this disclosure. Any suitable algorithm or process is possible as long as the impulse response function can be inverted to determine the inverse filter function. For example, in some embodiments, algorithms such as multi-channel least squares can also be used to determine the inverse filter function.

[0046] After determining the inverse filter function, the third step of this method allows the determination of the voice enhancement function and the noise reference signal based on the inverse filter function. Specifically, in a normal in-vehicle usage environment, such as when a user wants to perform intelligent interaction via voice, in this case, such as... Figure 2 As shown, each audio input device can acquire audio signals including control speech and noise. For the audio signal acquired by each audio input device, multiplying it with the corresponding inverse filter function and then adding it to the result of multiplying the audio signal acquired by another audio input device with the corresponding inverse filter function gives the human voice enhancement signal. Subtracting the two gives the noise reference signal, as shown in equation (5) below.

[0047]

[0048] Where Sig represents the voice enhancement signal, Noise represents the noise reference signal, mic1 represents the audio signal acquired by the first microphone, and h inv (1) represents the inverse filter function of the first microphone, mic2 represents the audio signal acquired by the second microphone, and h inv (2) represents the inverse filter function of the second microphone.

[0049] After determining the voice enhancement function and the noise reference signal, in the fourth step of this method, the audio signal input to the audio input module can be adaptively filtered based on the determined voice enhancement function and noise reference signal, thereby effectively eliminating noise and improving speech recognition quality.

[0050] In some embodiments, Kalman adaptive filtering can be used and the audio signal can be processed according to the human voice enhancement function and the noise reference signal. Figure 3 A schematic diagram of a control system using Kalman adaptive filtering to process audio signals is shown. Figure 3 As shown, when using Kalman adaptive filtering to process audio signals, a noise reference signal is used as the reference signal, and the noise propagation path parameter w is continuously updated. Finally, a clean human voice signal is obtained by subtracting the product of the propagation path parameter w and the noise reference signal from the human voice enhancement signal.

[0051] It should be understood that the above description based on an embodiment employing Kalman adaptive filtering is merely illustrative and is not intended to limit the scope of this disclosure. Any suitable adaptive filtering algorithm is possible as long as the audio signal can be processed according to the voice enhancement signal and the noise reference signal. For example, in some alternative embodiments, the least mean square (LMS) adaptive algorithm or the recursive least squares (RLS) adaptive algorithm may also be used.

[0052] As can be seen from the above process description, the method according to the embodiments of this disclosure characterizes the acoustic propagation path (e.g., embodied in the impulse response function) and inverses the acoustic transfer function to obtain the inverse filter function. Finally, adaptive filtering is used to effectively eliminate noise and improve speech recognition quality. In this way, the performance of speech enhancement can be significantly improved with lower computational load by utilizing the acoustic environment of the vehicle itself, thereby enhancing the user experience.

[0053] Figure 4 A flowchart illustrating a method for processing audio signals according to embodiments of the present disclosure is shown. In some embodiments, the method may be implemented by the processing unit, DSP, or any other suitable electronic device mentioned above. For ease of understanding, the specific examples, figures, or values ​​mentioned in the following description are merely exemplary and are not intended to limit the scope of protection of this disclosure.

[0054] like Figure 4 As shown, in the method executed by the processing unit, in block 410, the processing unit determines impulse response functions for multiple audio input modules (e.g., microphones) in the vehicle, representing the acoustic transmission path from the sound source to the audio input module. After determining the impulse response functions, in block 420, the processing unit determines the inverse filter function for the corresponding audio input module among the multiple audio input modules based on the impulse response functions.

[0055] In block 430, the processing unit determines the voice enhancement signal and the noise reference signal based on the inverse filter function. Next, in block 440, the processing unit adaptively filters the audio signal input to the audio input module based on the voice enhancement signal and the noise reference signal to obtain a clean voice control signal, thereby improving the performance of voice enhancement and thus improving the user experience.

[0056] In some embodiments, the processing unit receives an attenuation-compensated sweep signal output from an analog sound source from an audio input module, and determines the impulse response function based on the sweep signal received from the audio input module and the inverse signal of the sweep signal.

[0057] In some embodiments, the processing unit determines the inverse filter function based at least on the impulse response function and the impulse signal. The impulse signal causes the impulse response function (i.e., the acoustic transfer function) h(i) and the inverse filter function h to be determined. inv (i) Obtained by convolution.

[0058] In some embodiments, the processing unit determines the human voice enhancement signal and the noise reference signal based on the first audio signal input by the first audio input module and the corresponding inverse filter function, and based on the second audio signal input by the second audio input module and the corresponding inverse filter function.

[0059] In some embodiments, the processing unit determines noise transmission path parameters (e.g., w mentioned above) based on a noise reference signal, and determines the human voice signal in the audio signal based on the human voice enhancement signal, the noise reference signal, and the noise transmission path parameters.

[0060] Figure 5 A block diagram is shown illustrating an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to implement the electronic device mentioned above.

[0061] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose computing device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0062] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 500.

[0063] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0064] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the electronic device 500 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0065] Input device 550 can be one or more input devices, such as a touchscreen or voice input module. Output device 560 can be one or more output devices, such as a display, speaker, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).

[0066] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0067] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0068] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0069] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0070] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0071] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method of processing an audio signal, comprising: determining an impulse response function for a plurality of audio input modules in a vehicle representing an acoustic transfer path of a sound source to the audio input modules; determining an inverse filter function for a corresponding audio input module of the plurality of audio input modules from the impulse response function; determining a voice enhancement signal and a noise reference signal from audio signals input by any two audio input modules of the plurality of audio input modules and their respective inverse filter functions; determining a noise transfer path parameter from the noise reference signal; and determining a voice signal in the audio signal from the voice enhancement signal, the noise reference signal, and the noise transfer path parameter.

2. The method of claim 1, wherein determining the impulse response function comprises: receiving an attenuated compensated swept frequency signal output by an analog sound source from the audio input modules; and determining the impulse response function from the swept frequency signal received from the audio input modules and an inverse of the swept frequency signal.

3. The method of claim 1, wherein determining an inverse filter function comprises: determining the inverse filter function from at least the impulse response function and an impulse signal.

4. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing machine executable instructions that, when executed by the at least one processing unit, cause the device to perform acts comprising: determining an impulse response function for a plurality of audio input modules in a vehicle representing an acoustic transfer path of a sound source to the audio input modules; determining an inverse filter function for a corresponding audio input module of the plurality of audio input modules from the impulse response function; determining a voice enhancement signal and a noise reference signal from audio signals input by any two audio input modules of the plurality of audio input modules and their respective inverse filter functions; determining a noise transfer path parameter from the noise reference signal; and determining a voice signal in the audio signal from the voice enhancement signal, the noise reference signal, and the noise transfer path parameter.

5. The electronic device of claim 4, wherein determining the impulse response function comprises: receiving an attenuated compensated swept frequency signal output by an analog sound source from the audio input modules; and determining the impulse response function from the swept frequency signal received from the audio input modules and an inverse of the swept frequency signal.

6. The electronic device of claim 4, wherein determining an inverse filter function comprises: determining the inverse filter function from at least the impulse response function and an impulse signal.

7. A computer readable storage medium having stored thereon a computer program, the computer program executable by a processor to implement the method of any one of claims 1-3.

8. A vehicle, comprising: a plurality of audio input modules; and the electronic device of any one of claims 4-6. ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Distributed microphone pickup system and method in complex scene

    CN111161751A

  • Voice signal processing method and device, computer readable medium and electronic equipment

    CN111435598A