Audio signal processing method for vehicle cabin, vehicle and storage medium

By acquiring first and second audio signals inside the vehicle cabin and using short-time Fourier analysis and filtering algorithms to determine the single or dual-talk status, the problem of speaker sound interfering with microphone collection is solved, thus improving the accuracy and security of speech recognition.

CN115881111BActive Publication Date: 2026-04-17GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU XIAOPENG MOTORS TECH CO LTD
Filing Date
2022-12-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Inside the vehicle cabin, the sound emitted by the speakers interferes with the microphones that collect user voice commands, leading to inaccurate voice recognition and potentially causing driving safety hazards. Existing echo cancellation technologies suffer from issues such as low correlation or confusion in judging the correlation or energy difference of far-end signals.

Method used

By acquiring the first and second audio signals from inside the vehicle cabin, short-time Fourier analysis and preset filtering algorithms such as Kalman filtering or normalized least mean square filtering are used to determine the residual signal and residual echo function. The single/dual talk status is determined by combining the ratio of the residual signal function to the residual echo function, thus avoiding deviations in correlation and energy estimation.

Benefits of technology

It improves the accuracy of voice recognition, reduces interference between far-end signals emitted by the speaker and near-end signals, saves performance consumption of voice recognition, and ensures the safety of voice interaction in the vehicle cabin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881111B_ABST
    Figure CN115881111B_ABST
Patent Text Reader

Abstract

The application discloses a kind of audio signal processing methods of vehicle cabin, comprising: obtaining the first audio signal in vehicle cabin and second audio signal;According to first audio signal and second audio signal, determine residual signal;According to residual signal, determine residual signal function;According to second audio signal and residual signal function, determine residual echo function;According to residual echo function and residual signal function, determine the audio signal state in vehicle cabin, to output processed audio signal for speech recognition use.The application adopts new method outside correlation determination and energy determination, avoids the damage of low correlation determination result caused by too large remote signal when correlation determination is used, also avoids the signal confusion caused by energy difference of different loudspeaker devices when energy estimation is used and further causes the damage of energy estimation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, specifically to an audio signal processing method for a vehicle cockpit, a vehicle, and a computer-readable storage medium. Background Technology

[0002] With the development of vehicle intelligence, in-vehicle voice technology allows users to interact within the vehicle cabin via voice, such as controlling vehicle components or interacting with components in the in-vehicle system's user interface. Specifically, voice interaction is achieved by using microphones in the vehicle cabin to collect user voice commands and then performing voice recognition. However, in actual use, the speakers in the vehicle cabin often operate simultaneously during human-machine interaction, interfering with the user's voice commands collected by the microphones. This can lead to inaccurate recognition of user voice commands and even potential safety hazards due to misrecognition. Therefore, how to detect and assess the signal environment within the vehicle cabin has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides an audio signal processing method for a vehicle cockpit, a vehicle, and a storage medium to solve the technical problems described in the background art.

[0004] In current echo cancellation technologies, dual-talk detection primarily relies on the correlation between near-end and far-end signals or the energy difference between the two signals. However, when judging based on the correlation of the two signals, the nonlinear characteristics of the far-end signal have a significant negative impact on the correlation when the far-end volume is high, leading to an underestimation of the correlation and affecting accuracy. When judging based on the energy of the two signals, energy estimation is limited by the influence of different types and standards of equipment, as well as individual differences in the same type of equipment, resulting in excessive interference between the near and far ends, making it impossible to clearly distinguish between the far-end and near-end signals.

[0005] In order to solve the technical problems described in the background art and improve the technical effects in the current related art, the audio signal processing method for a vehicle cockpit according to the embodiments of this application includes:

[0006] Acquire the first and second audio signals from inside the vehicle cabin;

[0007] The residual signal is determined based on the first audio signal and the second audio signal;

[0008] Based on the residual signal, determine the residual signal function;

[0009] The residual echo function is determined based on the second audio signal and the residual signal function;

[0010] Based on the residual echo function and the residual signal function, the audio signal state inside the vehicle cabin is determined, and the processed audio signal is output for use in speech recognition.

[0011] Therefore, this application employs residual echo estimation, and determines the current single / dual-talk status inside the vehicle cabin based on the estimation result of the residual echo and the residual signal function. Simultaneously, it improves the quality of echo estimation by separately determining the residual signal function and the residual echo function. A novel method other than correlation determination and energy determination is used, avoiding the damage to the correlation determination result caused by excessively large far-end signals when using correlation determination, and also avoiding signal confusion caused by energy differences between different speaker devices when using energy estimation, which further damages the energy estimation result. Furthermore, by determining the single / dual-talk status, a portion of the speech recognition input is filtered out, saving performance overhead in speech recognition.

[0012] The acquisition of the second audio signal from the speakers in the vehicle cabin includes:

[0013] Determine the time period for acquiring the first audio signal;

[0014] Obtain the audio file played by the speaker during the stated time period;

[0015] The second audio signal is obtained based on the audio file.

[0016] Thus, this application directly obtains the audio signal emitted from the speaker from the vehicle system instead of from the microphone, separating the near-end signal from the far-end signal, avoiding mutual interference between the two sets of signals, and effectively avoiding the weakening of the correlation between the far-end signal and the near-end signal, as well as the weakening of energy assessment.

[0017] Determining the residual signal based on the first audio signal and the second audio signal includes:

[0018] Based on the first audio signal, a first frequency domain signal is determined at a preset frequency point through short-time Fourier analysis;

[0019] Based on the second audio signal, a second frequency domain signal is determined at a preset frequency point through short-time Fourier analysis;

[0020] Based on the first frequency domain signal and the second frequency domain signal, the residual signal at the current time sampling point is determined through a preset filtering algorithm.

[0021] Thus, this application uses short-time Fourier analysis to process the near-end and far-end signals into frequency domain signals with frequency information, and obtains the residual signal through a preset filtering algorithm, providing a data basis for subsequent estimation of residual echoes.

[0022] The preset filtering algorithm includes the Kalman filter algorithm or the normalized least mean square filter algorithm.

[0023] Thus, this application provides an adaptive linear filtering algorithm that enables the obtained filter signal to have a better effect on the estimation of the residual echo function.

[0024] The step of determining the residual signal function based on the residual signal includes:

[0025] At a preset frequency point, the residual signal function at the current time sampling point is determined based on the value of the residual signal at the current time sampling point, the value of the residual signal function at the previous time sampling point, and the first parameter. The value of the residual signal function at the initial time sampling point is the first frequency domain signal.

[0026] Thus, this application provides a method for determining the residual signal function.

[0027] The step of determining the residual echo function based on the second audio signal and the residual signal function includes:

[0028] At a preset frequency point, the second signal power function at the current time sampling point is determined based on the second audio signal;

[0029] At a preset frequency point, the residual echo function at the current time sampling point is determined based on the filter signal at the current time sampling point, the second signal power function at the current time sampling point, the residual echo function at the previous time sampling point, the residual signal function at the previous time sampling point, and the second parameter. The residual echo function has a value of zero at the initial time sampling point.

[0030] Thus, this application provides a method for determining the residual echo function.

[0031] Determining the sound signal state inside the vehicle cabin based on the residual echo function and the residual signal function includes:

[0032] At the current sampling point, determine the average value of the ratio of the residual echo function and the residual signal function for all possible values ​​of the preset frequency point;

[0033] The sound signal status inside the vehicle cabin is determined based on the relationship between the average value and a preset threshold.

[0034] Thus, this application determines the single / dual talk state by averaging the ratio of the residual echo function to the residual signal function, reducing the nonlinear effects caused by the far-end signal emitted by the loudspeaker, and reducing the data fluctuations and uncertainties caused by frequency changes in the determination process.

[0035] The step of determining the sound signal state inside the vehicle cabin based on the relationship between the average value and a preset threshold includes:

[0036] If the average value is less than a preset threshold, the audio signal status in the vehicle cabin is determined to be dual-talk mode;

[0037] If the average value is greater than or equal to a preset threshold, the audio signal state in the vehicle cabin is determined to be a single-talk state.

[0038] Thus, this application provides a specific method for determining single-talk and dual-talk states.

[0039] The step of determining the sound signal state inside the vehicle cabin based on the residual echo function and the residual signal function further includes:

[0040] When the audio signal status in the vehicle cabin is in dual-talk mode, the first audio signal is input into the voice recognition service for recognition.

[0041] Thus, this application also provides a method for processing audio signals after determining that a dual-talk state has been entered.

[0042] This application also provides a vehicle, the vehicle including an on-board computer, the on-board computer including a memory and a processor; the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the audio signal processing method for the vehicle cabin as described above.

[0043] This application also provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the audio signal processing method for a vehicle cabin as described above.

[0044] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description

[0045] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0046] Figure 1This is a flowchart illustrating the audio signal processing method for the vehicle cockpit provided in this application.

[0047] Figure 2 This is a flowchart illustrating the audio signal processing method for the vehicle cockpit provided in this application.

[0048] Figure 3 This is a flowchart illustrating the audio signal processing method for the vehicle cockpit provided in this application.

[0049] Figure 4 This is a flowchart illustrating the audio signal processing method for the vehicle cockpit provided in this application. Detailed Implementation

[0050] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.

[0051] The method proposed in this application is applied to human-machine voice interaction inside the vehicle cabin. The method proposed in this application will be described in detail below.

[0052] like Figure 1 As shown, the vehicle cockpit audio signal processing method provided in this application includes:

[0053] 01: Acquire the first and second audio signals from inside the vehicle's cabin;

[0054] 02: Determine the residual signal based on the first audio signal and the second audio signal;

[0055] 03: Determine the residual signal function based on the residual signal;

[0056] 04: Determine the residual echo function based on the second audio signal and the residual signal function;

[0057] 05: Based on the residual echo function and residual signal function, determine the audio signal state in the vehicle cabin, and output the processed audio signal for speech recognition.

[0058] This application also provides a vehicle, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to acquire a first audio signal and a second audio signal in the vehicle cabin, determine a residual signal based on the first audio signal and the second audio signal, determine a residual signal function based on the residual signal, determine a residual echo function based on the second audio signal and the residual signal function, and determine the audio signal state in the vehicle cabin based on the residual echo function and the residual signal function.

[0059] Specifically, the first audio signal is acquired by a microphone and mainly consists of the user's voice when interacting with the vehicle system; in current echo cancellation (AEC) technology, this corresponds to the near-end signal. The second audio signal mainly consists of the sound emitted by the speakers in the vehicle cabin when the first audio signal is acquired by the microphone; in current AEC technology, this corresponds to the far-end signal. The second audio signal is not acquired through a microphone to avoid excessive confusion with the first audio signal, which could lead to a decrease in processing effectiveness. After acquiring both sets of audio signals, the residual signals of the two sets of signals are determined based on the two sets of audio signal data. A residual signal function representing the sound power spectrum of the residual signal is then determined based on this residual signal. Simultaneously, estimation iterations are performed using the second audio signal and the residual signal function to determine the residual echo function representing the residual echo power spectrum. Finally, based on the calculation results of the residual echo function and the residual signal function, the audio signal state in the vehicle cabin is determined. The audio signal state generally refers to the single-talk state or dual-talk state of the audio signal in the current environment. Compared to current methods for determining single / dual-talk status based on correlation and energy estimation, this application adopts a completely different technical approach. It does not start from the correlation between near and far-end signals, nor from an energy perspective, but instead determines the single / dual-talk status based on the ratio of the residual echo power spectrum to the residual signal power spectrum. In this way, by separately acquiring the first and second audio signals, and avoiding correlation impairment caused by nonlinear effects due to loud speaker volume, as well as signal confusion caused by device differences, this application directly avoids the drawbacks of current methods for determining single / dual-talk status based on correlation and energy estimation, thus improving the accuracy of single / dual-talk status determination. If the signal status in the vehicle cabin is determined to be dual-talk, it proves that the user has issued a voice command. The first audio signal collected by the microphone is then input into the voice recognition service for voice recognition to achieve human-machine voice interaction. If the signal status in the vehicle cabin is determined to be single-talk, it proves that the user has not issued a voice command, and no voice recognition process is performed. This saves performance consumption in voice recognition by filtering out voice recognition input in single-talk status.

[0060] Therefore, this application employs residual echo estimation, and determines the current single / dual-talk status inside the vehicle cabin based on the estimation result of the residual echo and the residual signal function. Simultaneously, it improves the quality of echo estimation by separately determining the residual signal function and the residual echo function. A novel method other than correlation determination and energy determination is used, avoiding the damage to the correlation determination result caused by excessively large far-end signals when using correlation determination, and also avoiding signal confusion caused by energy differences between different speaker devices when using energy estimation, which further damages the energy estimation result. Furthermore, by determining the single / dual-talk status, a portion of the speech recognition input is filtered out, saving performance overhead in speech recognition.

[0061] Step 01 includes:

[0062] Determine the time period for acquiring the first audio signal;

[0063] Get the audio file played by the speaker within a time period;

[0064] Obtain the second audio signal based on the audio file.

[0065] The processor is used to determine the time period for acquiring the first audio signal, to acquire the audio file played by the speaker during the time period, and to acquire the second audio signal based on the audio file.

[0066] Specifically, to achieve interference-free acquisition of the second audio signal, it can be obtained directly from the vehicle's in-vehicle system. In some examples, while the microphone acquires the first audio signal, the time period of signal acquisition is recorded, and the in-vehicle system retrieves the audio being played by the speaker during that time period, including but not limited to navigation voice, music, and video audio tracks. The corresponding time endpoints are then extracted from these audio tracks to form the second audio signal.

[0067] Thus, this application directly obtains the audio signal emitted from the speaker from the vehicle system instead of from the microphone, separating the near-end signal from the far-end signal, avoiding mutual interference between the two sets of signals, and effectively avoiding the weakening of the correlation between the far-end signal and the near-end signal, as well as the weakening of energy assessment.

[0068] Step 02 includes:

[0069] Based on the first audio signal, a first frequency domain signal is determined at a preset frequency point through short-time Fourier analysis;

[0070] Based on the second audio signal, the second frequency domain signal is determined at a preset frequency point through short-time Fourier analysis;

[0071] Based on the first frequency domain signal and the second frequency domain signal, the residual signal at the current time sampling point is determined through a preset filtering algorithm.

[0072] The processor is used to determine a first frequency domain signal at a preset frequency point based on a first audio signal through short-time Fourier analysis, and to determine a second frequency domain signal at a preset frequency point based on a second audio signal through short-time Fourier analysis, and to determine the residual signal at the current time sampling point based on the first frequency domain signal and the second frequency domain signal through a preset filtering algorithm.

[0073] Specifically, this application first performs frequency domain processing on the acquired first and second audio signals, primarily using short-time Fourier analysis. Both the first and second audio signals inherently contain time sampling point information. After short-time Fourier analysis, the first audio signal is transformed into a first frequency domain signal containing both time sampling point and frequency information, and the second audio signal is transformed into a second frequency domain signal containing both time sampling point and frequency information. Frequency information primarily describes the frequency of the frequency domain signal. Time sampling point information primarily describes the time node when the signal was acquired and is a crucial iterative basis. Finally, based on the first and second frequency domain signals obtained by the above method, a preset filtering algorithm is used to calculate and output a residual signal containing both frequency and time sampling point information.

[0074] Thus, this application uses short-time Fourier analysis to process the near-end signal and the far-end signal into a frequency domain signal with frequency information, and obtains the residual signal through a preset filtering algorithm, providing a data basis for subsequent estimation of residual echo.

[0075] The preset filtering algorithms include the Kalman filter algorithm or the normalized least mean square filter algorithm.

[0076] Specifically, the Kalman filter algorithm and the Normalized Least Mean Square (NLMS) algorithm are linear adaptive filtering algorithms that determine the residual signal through iterative filtering operations based on the input signals in the first and second frequency domains. The following section uses the NLMS algorithm as an example to illustrate the process of obtaining the residual data:

[0077] After obtaining the first frequency domain signal d(k,n) and the second frequency domain signal x(k,n), the NLMS algorithm is applied to both, and the iterative formula is as follows:

[0078]

[0079]

[0080] e(k,n)=d(k,n)-y(k,n)

[0081] Where: w(k,n) is the filter signal provided by the filtering algorithm, e(k,n) is the residual signal, y(k,n) is an intermediate output of the filter, k is the frequency point obtained by short-time Fourier analysis, and N is the order of the filter provided by the filtering algorithm.

[0082] After the combined iterative process of the above three sets of formulas, the residual signal e(k, n) and the filter signal w(k, n) can be obtained. These two data are the basis for subsequently determining the residual signal function and the residual echo function.

[0083] Thus, this application provides an adaptive linear filtering algorithm that enables the obtained filter signal to have a better effect on the estimation of the residual echo function.

[0084] Step 03 includes:

[0085] At a preset frequency point, the residual signal function at the current time sampling point is determined based on the value of the residual signal at the current time sampling point, the value of the residual signal function at the previous time sampling point, and the first parameter. The value of the residual signal function at the initial time sampling point is the first frequency domain signal.

[0086] The processor is used to determine the residual signal function at the current time sampling point based on the value of the residual signal at the current time sampling point, the value of the residual signal function at the previous time sampling point, and the first parameter at a preset frequency point.

[0087] Specifically, the residual signal function is the residual signal power spectral density function, and the method described above for obtaining the residual signal function follows the iterative formula below:

[0088] Φ ee (k, n) = λΦ ee (k, n-1)+(1-λ)e T (k, n)e(k, n)

[0089] Where Φ ee (k, n) is the residual signal power spectral density function, e(k, n) is the residual signal, k is the preset frequency point, n is the current time sampling point, n-1 is the previous time sampling point, and λ is the first parameter, which mainly functions to smooth the residual signal and the output residual signal function. Generally, λ∈(0, 1). According to the relationship between the residual signal and the first frequency domain signal in the aforementioned embodiment, when the frequency point k takes the preset value, when the time sampling point is the initial value, that is, when n is 0, the intermediate output quantity y(k, n) of the filter does not exist. Therefore, the value of the residual signal power spectral density function at this time is the value of the first frequency domain signal d(k, n) when n = 0.

[0090] Thus, this application provides a method for determining the residual signal function.

[0091] Step 04 includes:

[0092] At a preset frequency point, the second signal power function at the current time sampling point is determined based on the second audio signal;

[0093] At a preset frequency point, the residual echo function at the current time sampling point is determined based on the filter signal at the current time sampling point, the second signal power function at the current time sampling point, the residual echo function at the previous time sampling point, the residual signal function at the previous time sampling point, and the second parameter. The residual echo function takes the value of zero at the initial time sampling point.

[0094] The processor is used to determine the second signal power function at the current time sampling point based on the second audio signal at a preset frequency point, and to determine the residual echo function at the current time sampling point based on the filter signal at the current time sampling point, the second signal power function at the current time sampling point, the residual echo function at the previous time sampling point, the residual signal function at the previous time sampling point, and the second parameter at the preset frequency point.

[0095] Specifically, the second signal power function corresponds to the far-end signal power spectral density function in AEC-related techniques, and can be determined based on the second frequency domain signal x(k,n). The residual echo function refers to the residual echo power spectral density function, and the method for determining the residual echo function described above follows the iterative formula:

[0096] Φ bb (k, n) = u 2 (k, n-1)·Φ ee (k, n-1)+Φ bb (k, n-1)·(1-2u(k, n-1))+|α*w(k, n)| 2 Φ xx (k, n)

[0097] Where: Φ bb (k, n) is the residual echo function, w(k, n) is the filter signal provided by the aforementioned NLMS algorithm, and Φ ee (k, n) is the residual signal function, Φ xx (k, n) is the second signal power function, k is the preset frequency point, n is the current time sampling point, n-1 is the previous time sampling point, α is the second parameter, 0 < α << 1; u(k, n) is the ratio of the residual echo function to the residual signal function, i.e.:

[0098]

[0099] Based on the aforementioned NLMS algorithm and the mathematical form of the above formula, it can be seen that at the initial time sampling point, that is, when n=0, the initial value of the residual echo function is the same as w(k,n), both being 0.

[0100] Thus, this application provides a method for determining the residual echo function.

[0101] like Figure 2 As shown, step 05 includes:

[0102] 051: At the current sampling point, determine the average value of the ratio of the residual echo function and the residual signal function for all possible values ​​of the preset frequency point;

[0103] 052: Determine the sound signal status inside the vehicle cabin based on the relationship between the average value and the preset threshold.

[0104] The processor is used to determine, at the current sampling point, the average value of the ratio of the residual echo function and the residual signal function under all possible values ​​of the preset frequency point, and to determine the sound signal state inside the vehicle cabin based on the relationship between the average value and the preset threshold.

[0105] Specifically, both the residual echo function and the residual signal function have two parameter values: frequency point and time sampling point. The time sampling point determines the continuous change of the function itself over time. However, in the aforementioned calculation process, the frequency point is a parameter that is controlled and remains constant. In practical applications, the power spectral density function values ​​corresponding to different frequency points vary greatly, and the residual echo function and residual signal function at any single frequency point cannot be proven to be representative. Therefore, this application uses the average of the ratio of the residual echo function and the residual signal function when taking all possible values ​​of the frequency point at a specific time sampling point to determine the state of the sound signal. This avoids the huge data differences caused by the change of frequency point. To increase the amount of data, the average of the ratio of the residual echo function and the residual signal function when taking all possible values ​​of the frequency point at all time sampling points within a certain time period can also be taken. For example, the average of the ratio u(k,n) when the time sampling points are n-1, n, and n+1 and k is all possible values ​​can be taken. Then, the average value The signal is compared with a preset threshold, and the comparison result determines whether the current sound signal status in the vehicle cabin is single-talk or dual-talk.

[0106] Thus, this application determines the single / dual talk state by averaging the ratio of the residual echo function to the residual signal function, reducing the nonlinear effects caused by the far-end signal emitted by the loudspeaker, and reducing the data fluctuations and uncertainties caused by frequency changes in the determination process.

[0107] like Figure 3 As shown, step 052 includes:

[0108] 0521: Determine if the average value is less than the preset threshold. If yes, proceed to step 0622; otherwise, proceed to step 0623.

[0109] 0522: Confirm that the audio signal status inside the vehicle cabin is in dual-talk mode;

[0110] 0523: Confirm that the audio signal status inside the vehicle cabin is in single-talk mode.

[0111] The processor is used to determine whether the average value is less than a preset threshold, and to determine whether the audio signal status in the vehicle cabin is a dual-talk state, and to determine whether the audio signal status in the vehicle cabin is a single-talk state.

[0112] Specifically, the preset threshold is a boundary value representing balance, used to describe the proportion of various sounds within the current vehicle cabin. The residual signal e(k,n) consists of two parts: residual echo and the sound emitted by the user. According to the description in the section on determining the residual signal function and the method for determining the residual echo function, the residual signal function is the power spectral density function of the residual signal, and the residual echo function is the power spectral density function of the residual echo. Therefore, the ratio of the residual echo function to the residual signal function can be understood analogously through the proportion of the residual echo contained in the residual signal. When the average of the ratio of the residual echo function and the residual signal function is greater than or equal to the boundary value, it can be considered that the proportion of the residual echo in the residual signal is not less than the equilibrium boundary. At this time, it can be considered that the influence of the second audio signal inside the vehicle cabin is very large, far greater than the sound emitted by the user, and it can be identified as a single-talk state. Conversely, when the average of the ratio of the residual echo function and the residual signal function is less than the boundary value, it can be considered that the proportion of the residual echo in the residual signal is less than the equilibrium boundary. At this time, the influence of both the sound emitted by the user and the residual echo cannot be ignored, and it can be identified as a dual-talk state inside the vehicle cabin.

[0113] Thus, this application provides a specific method for determining single-talk and dual-talk states.

[0114] like Figure 4 As shown, step 05 is followed by:

[0115] 06: When the audio signal status in the vehicle cabin is in dual-talk mode, the first audio signal is input into the voice recognition service for recognition.

[0116] The processor is used to input the first audio signal into the voice recognition service for recognition when the audio signal status in the vehicle cabin is in dual-talk mode.

[0117] Specifically, as described in the preceding embodiments, when the average value of the ratio of the residual echo function and the residual signal function is greater than or equal to the boundary value, it can be considered that the proportion of the residual echo in the residual signal is not less than the equilibrium boundary. At this time, it can be considered that the influence of the second audio signal inside the vehicle cabin is very large, far greater than the sound emitted by the user, and it can be identified as being in a one-way state. That is, it can be approximately considered that when the sound signal state inside the vehicle cabin is in a one-way state, the sound emitted by the user does not exist, and it is considered that no voice-based human-computer interaction event has occurred. However, when the average value of the ratio of the residual echo function and the residual signal function is less than the boundary value, it can be considered that the proportion of the residual echo in the residual signal is less than the equilibrium boundary. At this time, the influence of both the sound emitted by the user and the residual echo cannot be ignored, and it can be identified that the inside of the vehicle cabin is in a two-way state. At this time, the sound emitted by the user exists, that is, a voice-based human-computer interaction event has occurred. Therefore, it is necessary to input the first audio signal collected by the microphone into the speech recognition service for speech recognition to further determine the content of the human-computer interaction.

[0118] Thus, this application also provides a method for processing audio signals after determining that a dual-talk state has been entered.

[0119] This application also provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the audio signal processing method for a vehicle cabin as described above.

[0120] In the description of this specification, the references to terms such as "some embodiments," "in one example," "exemplarily," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0121] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.

[0122] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method of audio signal processing for a vehicle cabin, the method comprising: The method includes: Acquire a first audio signal and a second audio signal from inside the vehicle cabin, wherein the first audio signal is collected by a microphone and the second audio signal is an audio signal played by a speaker from the vehicle system; The residual signal is determined based on the first audio signal and the second audio signal; Based on the residual signal, determine the residual signal power spectral density function; At a preset frequency point, the second signal power function at the current time sampling point is determined based on the second audio signal; At a preset frequency point, the residual echo power spectral density function at the current time sampling point is determined based on the filter signal at the current time sampling point, the second signal power function at the current time sampling point, the residual echo power spectral density function at the previous time sampling point, the residual signal power spectral density function at the previous time sampling point, and the second parameter. The residual echo power spectral density function takes the value of zero at the initial time sampling point. Based on the comparison relationship between the residual echo power spectral density function and the residual signal power spectral density function and the magnitude relationship with the preset threshold, the audio signal state in the vehicle cabin is determined, and the processed audio signal is output for speech recognition. The comparison relationship is the average value of the ratio of the residual echo power spectral density function to the residual signal power spectral density function at the current sampling point, where the preset frequency point is the average value of all possible values.

2. The method of claim 1, wherein, The acquisition of the first audio signal and the second audio signal within the vehicle cabin includes: Determine the time period for acquiring the first audio signal; Obtain the audio file played by the speaker during the stated time period; The second audio signal is obtained based on the audio file.

3. The method of claim 1, wherein, Determining the residual signal based on the first audio signal and the second audio signal includes: Based on the first audio signal, a first frequency domain signal is determined at a preset frequency point through short-time Fourier analysis; Based on the second audio signal, a second frequency domain signal is determined at a preset frequency point through short-time Fourier analysis; Based on the first frequency domain signal and the second frequency domain signal, the residual signal at the current time sampling point is determined through a preset filtering algorithm.

4. The method of claim 3, wherein, The preset filtering algorithm includes the Kalman filter algorithm or the normalized least mean square filter algorithm.

5. The method of claim 3, wherein, Determining the power spectral density function of the residual signal based on the residual signal includes: At a preset frequency point, the residual signal power spectral density function at the current time sampling point is determined based on the value of the residual signal at the current time sampling point, the value of the residual signal power spectral density function at the previous time sampling point, and the first parameter. The value of the residual signal power spectral density function at the initial time sampling point is the value of the first frequency domain signal at the initial time sampling point.

6. The method of claim 1, wherein, The step of determining the audio signal state within the vehicle cabin based on the comparison relationship between the residual echo power spectral density function and the residual signal power spectral density function and the magnitude of the comparison relationship with a preset threshold includes: If the average value is less than a preset threshold, the audio signal status in the vehicle cabin is determined to be dual-talk mode; If the average value is greater than or equal to a preset threshold, the audio signal state in the vehicle cabin is determined to be a single-talk state.

7. The method of claim 1, wherein, The step of determining the sound signal state inside the vehicle cabin based on the residual echo power spectral density function and the residual signal power spectral density function further includes: When the audio signal status in the vehicle cabin is in dual-talk mode, the first audio signal is input into the voice recognition service for recognition.

8. A vehicle characterized by comprising: The vehicle includes a memory and a processor; the memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Echo cancellation method and device and computer storage medium

    CN111445917A

  • Echo suppression method and device, equipment and storage medium

    CN111524532A