In-vehicle voice processing method, terminal device, and storage medium

By setting up a microphone array in the car to perform differential beam processing and frequency domain Kalman filtering, the problem of multi-zone interference in in-car voice interaction is solved, and the precise separation of voice signals and the improvement of interaction accuracy are achieved.

CN119446173BActive Publication Date: 2025-09-23ZHEJIANG LEAPMOTOR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411361388.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-09-23
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

In the complex environment inside the car, the audio signals collected by the on-board microphone contain voices from multiple sound zones and the vehicle's own audio, causing the voice interaction system to respond incorrectly and reducing the user experience.

Method used

Differential beam processing and Kalman filtering technology are used to perform differential beam processing by setting up a microphone array to enhance the energy difference of microphone signals, and Kalman filtering is performed in the frequency domain to filter out audio signals in non-target sound areas, thereby achieving accurate separation of voice signals.

Benefits of technology

It improves the accuracy of in-car voice interaction and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446173B_ABST
    Figure CN119446173B_ABST
Patent Text Reader

Abstract

The present application discloses a method for processing in-vehicle voice, a terminal device, and a storage medium. The processing method includes: obtaining a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone, wherein the distance between the first microphone and the second microphone is less than a preset distance; performing differential beam processing on the first audio signal and the second audio signal to obtain a first processed signal and a second processed signal; performing Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point to filter out audio signals generated by sound zones that are not corresponding to the microphone itself in different processed signals or audio signals. The above scheme filters out audio signals that are not generated by the sound zone itself, thereby improving the accuracy of in-vehicle voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice processing technology, and in particular to a method for processing in-vehicle voice, a terminal device, and a storage medium. Background Art

[0002] With the rapid development of technology, voice interaction technology is gradually being applied to in-vehicle connected scenarios. Users are becoming more and more accustomed to interacting with in-vehicle devices through voice, resulting in an increasing demand for in-vehicle voice interaction systems. To meet the needs of voice interaction between individual users and in-vehicle devices, in-vehicle voice interaction systems have launched in-vehicle multi-zone voice interaction services to expand the scope of voice interaction.

[0003] Due to the dense crowds and complex environment inside vehicles, the audio signals collected by the onboard microphones corresponding to a particular audio zone contain not only the speech in that zone, but also speech from other zones and the vehicle's own audio. Directly using audio signals as voice commands can easily result in response errors, degrading the user experience. Summary of the Invention

[0004] The present application provides a method for processing in-vehicle voice, a terminal device, and a storage medium.

[0005] A technical solution adopted by the present application is to provide a method for processing in-vehicle voice. At least a first microphone, a second microphone, and a third microphone are provided inside the vehicle, each microphone corresponding to a sound zone. The method includes:

[0006] Obtaining a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone; wherein a target distance between the first microphone and the second microphone is less than a preset threshold, and a distance between the third microphone and the first microphone and the second microphone is greater than the preset threshold respectively;

[0007] performing differential beam processing on the first audio signal to obtain a first processed signal and performing differential beam processing on the second audio signal to obtain a second processed signal;

[0008] Kalman filtering is performed based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point, filtering out audio signals generated by sound regions not corresponding to the first microphone in the first processed signal, filtering out audio signals generated by sound regions not corresponding to the second microphone in the second processed signal, and filtering out audio signals generated by sound regions not corresponding to the third microphone in the third audio signal.

[0009] Optionally, performing Kalman filtering processing according to the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point includes:

[0010] Performing a Fourier transform on the first processed signal to obtain a first frequency domain signal; performing a Fourier transform on the second processed signal to obtain a second frequency domain signal; and performing a Fourier transform on the third audio signal to obtain a third frequency domain signal;

[0011] Performing Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point, filtering out audio signals generated by a sound region not corresponding to the first microphone in the first processed signal, filtering out audio signals generated by a sound region not corresponding to the second microphone in the second processed signal, and filtering out audio signals generated by a sound region not corresponding to the third microphone in the third audio signal, including:

[0012] According to the energy size of each frequency point, Kalman filtering is performed on the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal to filter out the audio signal generated by the sound area not corresponding to the first microphone in the first frequency domain signal, filter out the audio signal generated by the sound area not corresponding to the second microphone in the second frequency domain signal, and filter out the audio signal generated by the sound area not corresponding to the third microphone in the third frequency domain signal to obtain filtered signals corresponding to the sound areas where different microphones are located.

[0013] Optionally, Kalman filtering is performed on the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal according to the energy of each frequency point to obtain filtered signals corresponding to the sound ranges where different microphones are located, including:

[0014] Determine the target signal from the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal in sequence, and use the remaining signals as reference signals;

[0015] Comparing a first energy of the target signal at each frequency point with a second energy of each reference signal;

[0016] In response to the first energy being less than the second energy, performing Kalman filtering on the target signal using the reference signal to obtain a separated signal corresponding to the target signal at each frequency point;

[0017] In response to the first energy being greater than or equal to the second energy, taking the target signal as the separation signal;

[0018] According to the separated signals at different frequency points, the filtered signals corresponding to the sound ranges where different microphones are located are obtained.

[0019] Optionally, obtaining filtered signals corresponding to different microphones according to the separated signals at different frequency points includes:

[0020] Calculating the first frequency domain coherence value between the target signal and the separated signal at each frequency point;

[0021] Calculating a second frequency domain coherence value between the target signal and each reference signal at each frequency point;

[0022] Determine a suppression factor at each frequency point based on a first frequency domain coherence value and a second frequency domain coherence value corresponding to the same frequency point;

[0023] The suppression factor and the separated signals are used to obtain the filtered signals corresponding to different microphones.

[0024] Optionally, the inhibition factor is obtained according to the following formula:

[0025] α=max(0.05,min(1-cohxd,cohed)), where α is the suppression factor, cohed is the first frequency domain coherence value, and cohxd is the second frequency domain coherence value.

[0026] Optionally, the first microphone corresponds to the first sound range, the second microphone corresponds to the second sound range, and the third microphone corresponds to the third sound range;

[0027] After obtaining the audio signal generated by each sound zone according to the filtered signals corresponding to the different sound zones, the method includes:

[0028] Performing an inverse Fourier transform on the first filtered signal corresponding to the first sound range to obtain a first target audio signal generated by the first sound range;

[0029] Performing an inverse Fourier transform on the second filtered signal corresponding to the second sound range to obtain a second target audio signal generated by the second sound range;

[0030] An inverse Fourier transform is performed on the third filtered signal corresponding to the third sound range to obtain a third target audio signal generated by the third sound range.

[0031] Optionally, the first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range;

[0032] Performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal, including:

[0033] Obtaining a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range;

[0034] The first audio signal is differentially beam processed using a first suppression angle to obtain a first processed signal; and the second audio signal is differentially beam processed using a second suppression angle to obtain a second processed signal.

[0035] Optionally, performing differential beam processing on the first audio signal using the first suppression angle to obtain a first processed signal includes:

[0036] determining a first filter coefficient according to a first suppression angle, a target distance, and a signal angular frequency;

[0037] Performing differential beam processing on the first audio signal using the first filter coefficient to obtain a first processed signal;

[0038] Performing differential beam processing on the second audio signal using the second suppression angle to obtain a second processed signal includes:

[0039] determining a second filter coefficient according to a second suppression angle, a target distance, and a signal angular frequency;

[0040] The second audio signal is subjected to differential beam processing using the second filter coefficient to obtain a second processed signal.

[0041] Another technical solution adopted by the present application is to provide a terminal device, the terminal device including a memory and a processor connected to the memory;

[0042] The memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned in-vehicle voice processing method.

[0043] Another technical solution adopted in the present application is to provide a computer storage medium, which is used to store program data. When the program data is executed by a computer, it is used to implement the above-mentioned in-vehicle voice processing method.

[0044] The beneficial effects of the present application are as follows: when the distance between the first microphone and the second microphone is less than a preset threshold, a first processed signal and a second processed signal are obtained by performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, so as to increase the difference between the audio signals collected by the microphones with smaller spacing. At each frequency point, the audio signal is Kalman filtered based on the energy difference between the audio signals to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0046] Figure 1 This is a flow chart of an embodiment of a method for processing in-vehicle speech provided by the present application;

[0047] Figures 2A-2D This is a schematic diagram of the distribution locations of vehicle-mounted microphones under different embodiments provided by this application;

[0048] Figure 3 This is a flow chart of another embodiment of the method for processing in-vehicle speech provided by the present application;

[0049] Figure 4 This is a schematic structural diagram of an embodiment of a terminal device provided by this application;

[0050] Figure 5 It is a structural diagram of an embodiment of a computer storage medium provided by this application. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0052] To enhance the passenger experience, smart vehicles utilize onboard microphones at various locations within the vehicle. These microphones capture voice commands from passengers in different locations, providing personalized service for each passenger. However, due to the densely populated and complex environment within the vehicle, the voice signals collected by the onboard microphones include not only the target passenger's voice, but also interference signals from passengers in other locations, audio signals from the vehicle itself, and various ambient noises. Accurately sensing the voice signals of passengers in different locations is an urgent challenge.

[0053] Related technologies primarily employ methods such as beamforming and blind source separation. Beamforming algorithms are susceptible to the number and placement of microphones in the array; for example, they are inapplicable to a mix of linear and distributed microphone arrays. Blind source separation also suffers from the problem of permutation uncertainty. In neural network-based speech separation solutions, the simulated room impulse response differs significantly from the actual in-car environment, resulting in distortion in the separated speech.

[0054] This application mainly designs a set of methods to improve the accuracy of audio signal separation. Unlike traditional methods, this application starts from the positional relationship between vehicle-mounted microphones and combines Kalman filtering to filter out audio signals that do not belong to their own sound zones from the signals collected by the vehicle-mounted microphones and mixed with audio generated by different sound zones, thereby accurately identifying users speaking in different sound zones.

[0055] Please refer to the following for details: Figure 1 , Figure 1This is a flow chart of an embodiment of a method for processing in-vehicle voice provided by this application.

[0056] like Figure 1 As shown, the method for processing in-vehicle speech in an embodiment of the present application may specifically include the following steps:

[0057] S1, obtaining a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0058] Among them, at least a first microphone, a second microphone and a third microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.

[0059] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0060] like Figure 2A As shown, mic-1 (microphone 1), mic-2 (microphone 2), and mic-3 correspond to the first sound zone, the second sound zone, and the third sound zone, respectively, wherein mic-1 and mic-2 constitute a linear array microphone array.

[0061] like Figure 2B As shown, mic-1, mic-2, mic-3, and mic-4 correspond to the first sound zone, the second sound zone, the third sound zone, and the fourth sound zone, respectively. Among them, mic-1 and mic-2 form a linear microphone array.

[0062] like Figure 2C As shown, mic-1, mic-2, mic-3, mic-4, mic-5, and mic-6 correspond to the first, second, third, fourth, fifth, and sixth sound zones, respectively. Among them, mic-1 and mic-2 constitute a linear microphone array.

[0063] like Figure 2D As shown, mic-1, mic-2, mic-3, mic-4, mic-5, and mic-6 correspond to the first, second, third, fourth, fifth, and sixth sound zones, respectively. mic-1 and mic-2 form a linear array microphone array, and mic-5 and mic-6 form a linear array microphone array.

[0064] The method for processing in-vehicle speech provided in the present application is mainly performed by a speech processing device. In some embodiments, the speech processing device may be the vehicle-mounted microphone itself. In some embodiments, the distance measuring device may be a device that is communicatively connected to the vehicle-mounted microphone. For example, the device may be a device for monitoring images, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, and any one or more products such as self-driving cars, robots, security systems, glasses, and helmets for augmented reality or virtual reality. In some possible implementations, the method for processing in-vehicle speech may be implemented by a processor calling computer-readable instructions stored in a memory.

[0065] Specifically, the speech processing device obtains a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0066] Exemplarily, the first microphone corresponds to the first sound zone, and the second microphone corresponds to the second sound zone. The distance between the first microphone and the second microphone is less than a preset threshold, that is, the first microphone and the second microphone constitute a linear array microphone array (or simply a microphone array). Since the distance between the first microphone and the second microphone is relatively close, the first audio signal collected by the first microphone includes not only the audio generated in the first sound zone, but also the audio generated in the second sound zone. Similarly, the second audio signal collected by the second microphone includes the audio generated in the first sound zone and the second sound zone. After a period of research, the applicant found that directly performing Kalman filtering operations on the first audio signal and the second audio signal cannot clearly separate the audio signals that do not belong to their own sound zones. Based on this, it is necessary to perform relevant processing on the first audio signal and the second audio signal to amplify the difference between the first audio signal and the second audio signal.

[0067] In some embodiments, the preset threshold is less than 50 cm. In some embodiments, the preset threshold is less than 40 cm. In some embodiments, the preset threshold is less than 30 cm. In some embodiments, the preset threshold is less than 20 cm. In some embodiments, the preset threshold is less than 10 cm. Exemplarily, the preset threshold is 15 cm.

[0068] S2. Perform differential beam processing on the first audio signal to obtain a first processed signal, and perform differential beam processing on the second audio signal to obtain a second processed signal.

[0069] Specifically, the speech processing device performs differential beam processing on the first audio signal to obtain a first processed signal, and the speech processing device performs differential beam processing on the second audio signal to obtain a second processed signal.

[0070] Related technologies primarily employ beamforming. However, this technology is used in an additive microphone array consisting of a first and second microphone. The main lobe shape of the additive microphone array is determined by setting the main lobe itself. For the first and second microphones in a vehicle-mounted microphone array, the spacing between them is small, and the main lobe width is large, resulting in poor directivity and insignificant array gain. This results in weak enhancement of the target direction and poor suppression of interference.

[0071] The differential beamforming algorithm is a microphone array-based sound source localization and sound separation technology that measures differences in spatial sound pressure. Compared to related technologies, the differential microphone array determines the shape of its main lobe by setting the null point direction. By placing the null point in the interference direction, the speech energy ratio in the target direction and the interference direction is increased. The array's sensitivity to different incident directions is represented by the beam pattern. The differential beamform satisfies the following relationship:

[0072] B[h(ω),θ]=d H (ω,cosθ)h(ω)

[0073] Where d(ω, cosθ) is the array's steering vector, ω is the signal angular frequency, θ is the incident direction, h(ω) is the output weight of each microphone, and H is the conjugate transpose. For example, with dual microphones with an end-fire direction of 0° and a suppression direction of 180°, the beam pattern calculation formula for differential beamforming is:

[0074]

[0075] Where h′(ω) is τ is the ratio of the microphone array spacing to the speed of sound.

[0076] In this embodiment, the speech processing device applies a differential beamforming algorithm to audio enhancement for a vehicle's linear microphone array. For example, the speech processing device sets the null point at the angle of the second sound zone relative to the microphone array, effectively suppressing the audio energy in the second sound zone. This emphasizes the audio signal in the first sound zone and increases the energy difference between the audio signals in the first and second sound zones.

[0077] For another example, the speech processing device sets the zero point to the angle of the first sound zone relative to the microphone array, thereby effectively suppressing the audio energy in the direction of the first sound zone and highlighting the audio signal of the second sound zone.

[0078] In this embodiment, the first and second audio signals are processed by the differential beamforming algorithm, resulting in a larger energy difference between the first and second processed signals, and a signal energy distribution similar to that obtained by a distributed microphone. A distributed microphone refers to two microphones separated by a distance greater than or equal to a preset threshold.

[0079] For example, the first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range. Because the first and second microphones are closely spaced, the energy difference between the speech signals in the first and second sound ranges picked up by the speech processing device is small, affecting the effectiveness of subsequent Kalman filtering and causing significant speech impairment. Therefore, differential beamforming is first performed on the first and second audio signals corresponding to the closely spaced first and second microphones, respectively. This directional interference suppression reduces the sound in the second sound range received by the first microphone, and reduces the sound in the first sound range received by the second microphone.

[0080] S3. Perform Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point to filter out audio signals generated by sound regions not corresponding to the first microphone in the first processed signal, filter out audio signals generated by sound regions not corresponding to the second microphone in the second processed signal, and filter out audio signals generated by sound regions not corresponding to the third microphone in the third audio signal.

[0081] Assuming that each frequency point corresponds to only one speaker, it is understandable that the signal energy collected by the microphone closest to the sound area is the largest. Therefore, the energy size at the frequency point is used to determine whether the current sound area is speaking.

[0082] Exemplarily, the speech processing device performs Kalman filtering on the first processed signal, the second processed signal, and the third processed signal based on the energy between the first processed signal, the second processed signal, and the third audio signal at each frequency point, filtering out the audio signal generated by the sound area not corresponding to the first microphone in the first processed signal, filtering out the audio signal generated by the sound area not corresponding to the second microphone in the second processed signal, and filtering out the audio signal generated by the sound area not corresponding to the third microphone in the third audio signal.

[0083] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0084] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0085] S11, obtaining a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0086] Among them, at least a first microphone, a second microphone and a third microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.

[0087] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0088] S12: Perform differential beam processing on the first audio signal to obtain a first processed signal, and perform differential beam processing on the second audio signal to obtain a second processed signal.

[0089] S13, performing Fourier transform on the first processed signal to obtain a first frequency domain signal; performing Fourier transform on the second processed signal to obtain a second frequency domain signal; and performing Fourier transform on the third audio signal to obtain a third frequency domain signal.

[0090] Since the first processed signal, the second processed signal and the third audio signal are time domain signals, in subsequent steps, Kalman filtering needs to be performed based on the energy of the signals in different frequency domains. Therefore, a conversion operation from time domain signals to frequency domain signals is required.

[0091] Exemplarily, the speech processing apparatus may perform short-time Fourier transform (STFT) on the first processed signal, the second processed signal, and the third audio signal, respectively, to obtain a first frequency domain signal, a second frequency domain signal, and a third frequency domain signal.

[0092] S14. Perform Kalman filtering on the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal according to the energy level of each frequency point, filter out the audio signal generated by the sound zone not corresponding to the first microphone in the first frequency domain signal, filter out the audio signal generated by the sound zone not corresponding to the second microphone in the second frequency domain signal, and filter out the audio signal generated by the sound zone not corresponding to the third microphone in the third frequency domain signal, to obtain filtered signals corresponding to the sound zones where different microphones are located.

[0093] Specifically, the speech processing device performs Kalman filtering on the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal according to the energy size of each frequency point, filters out the audio signal generated by the sound area not corresponding to the first microphone in the first frequency domain signal, filters out the audio signal generated by the sound area not corresponding to the second microphone in the second frequency domain signal, and filters out the audio signal generated by the sound area not corresponding to the third microphone in the third frequency domain signal, to obtain filtered signals corresponding to the sound areas where different microphones are located.

[0094] Exemplarily, the first microphone corresponds to the first filtered signal, the second microphone corresponds to the second filtered signal, and the third microphone corresponds to the third filtered signal.

[0095] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0096] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0097] S21: Acquire a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0098] Among them, at least a first microphone, a second microphone and a third microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.

[0099] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0100] S22: Perform differential beam processing on the first audio signal to obtain a first processed signal, and perform differential beam processing on the second audio signal to obtain a second processed signal.

[0101] S23, performing Fourier transform on the first processed signal to obtain a first frequency domain signal; performing Fourier transform on the second processed signal to obtain a second frequency domain signal; and performing Fourier transform on the third audio signal to obtain a third frequency domain signal.

[0102] S24 , determining a target signal from the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal in sequence, and using the remaining signals as reference signals.

[0103] In some embodiments, the speech processing apparatus uses the first frequency domain signal as the target signal, and uses the second frequency domain signal and the third frequency domain signal as reference signals in sequence.

[0104] In some embodiments, the speech processing apparatus uses the second frequency domain signal as a target signal, and uses the first frequency domain signal and the fourth frequency domain signal as reference signals in sequence.

[0105] In some embodiments, the speech processing apparatus uses the third frequency domain signal as the target signal, and uses the first frequency domain signal and the second frequency domain signal as reference signals in sequence.

[0106] S25 , comparing the first energy of the target signal at each frequency point with the second energy of each reference signal.

[0107] Specifically, the speech processing device traverses all frequency points and obtains the energy corresponding to the target signal (i.e., first energy) and the energy corresponding to the reference signal (i.e., second energy) at each frequency point. Further, the speech processing device compares the magnitude relationship between the first energy and the second energy.

[0108] S26 , in response to the first energy being less than the second energy, performing Kalman filtering on the target signal using the reference signal to obtain a separated signal corresponding to the target signal at each frequency point.

[0109] Specifically, at the first target frequency point, in response to the first energy being less than the second energy, it indicates that speech is generated in the sound area corresponding to the reference signal. The speech processing device uses the reference signal to perform Kalman filtering on the target signal to obtain a separated signal corresponding to the target signal at the first target frequency point.

[0110] Among them, the separation signal satisfies the following relationship:

[0111]

[0112] Among them, d n is the target signal, x n is the reference signal, w nis the Kalman filter coefficient, se n The separated signal output after filtering can also be called an error signal.

[0113] In some embodiments, the separated signal may be a single frequency point or a partial frequency band interval.

[0114] S27 , in response to the first energy being greater than or equal to the second energy, taking the target signal as a separation signal.

[0115] Specifically, in response to the first energy being greater than or equal to the second energy, it indicates that speech may have been generated in the sound range corresponding to the target signal, and the speech processing device does not perform any operation.

[0116] S28, obtaining filtered signals corresponding to the sound ranges where different microphones are located according to the separated signals at different frequency points.

[0117] In some embodiments, the first frequency domain signal is a target signal, and the second frequency domain signal is a reference signal. The speech processing device determines a first candidate filtered signal based on the obtained first separated signal (point) at different frequency points. Furthermore, the speech processing device uses the first candidate filtered signal as the target signal and the third frequency domain signal as the reference signal, executes steps S25-S27, obtains the second separated signal (point) at different frequency points, and thereby determines the filtered signal corresponding to the sound range where the first microphone is located, i.e., the filtered signal corresponding to the first sound range.

[0118] Similarly, the filtered signal corresponding to the second sound range and the filtered signal corresponding to the third sound range can also be obtained in the above manner.

[0119] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0120] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0121] S31: Acquire a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0122] Among them, at least a first microphone, a second microphone and a third microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.

[0123] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0124] S32: Perform differential beam processing on the first audio signal to obtain a first processed signal, and perform differential beam processing on the second audio signal to obtain a second processed signal.

[0125] S33, performing Fourier transform on the first processed signal to obtain a first frequency domain signal; performing Fourier transform on the second processed signal to obtain a second frequency domain signal; and performing Fourier transform on the third audio signal to obtain a third frequency domain signal.

[0126] S34, determining a target signal from the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal in sequence, and using the remaining signals as reference signals.

[0127] S35 , comparing the first energy of the target signal at each frequency point with the second energy of each reference signal.

[0128] S36 , in response to the first energy being less than the second energy, performing Kalman filtering on the target signal using the reference signal to obtain a separated signal corresponding to the target signal at each frequency point.

[0129] S37 , in response to the first energy being greater than or equal to the second energy, taking the target signal as a separation signal.

[0130] S38, calculating a first frequency domain coherence value between the target signal and the separated signal at each frequency point.

[0131] Specifically, the speech processing device obtains a target separated signal corresponding to the target microphone, and then calculates a first frequency domain coherence value between the target signal and the separated signal at each frequency point.

[0132] The first frequency domain coherence value is used to describe the similarity between the target signal and the separated signal at the current frequency point. The higher the first frequency domain coherence value, the closer the energy values ​​of the target signal and the separated signal are at the current frequency point.

[0133] In some embodiments, the first frequency domain coherence value satisfies the following relationship:

[0134]

[0135] Where D(f) is the target signal and E(f) is the separated signal.

[0136] S39: Calculate a second frequency domain coherence value between the target signal and each reference signal at each frequency point.

[0137] The second frequency domain coherence value is used to describe the similarity between the target signal and a reference signal at the current frequency. The higher the second frequency domain coherence value, the closer the energy values ​​of the target signal and the reference signal are at that frequency.

[0138] In some embodiments, the second frequency domain coherence value satisfies the following relationship:

[0139]

[0140] Where D(f) is the target signal and X(f) is the reference signal.

[0141] S40 , determining a suppression factor at each frequency point based on the first frequency domain coherence value and the second frequency domain coherence value corresponding to the same frequency point.

[0142] Specifically, the speech processing apparatus calculates the suppression factor corresponding to each frequency point by using the first frequency domain coherence value and the second frequency domain coherence value at the frequency point.

[0143] In some embodiments, the inhibition factor is obtained according to the following formula:

[0144] α=max(0.05,min(1-cohxd,cohed))

[0145] Wherein, α is the suppression factor, cohed is the first frequency domain coherence value, and cohxd is the second frequency domain coherence value.

[0146] S41, using the suppression factor and the separated signal to obtain filtered signals corresponding to different microphones.

[0147] Specifically, the speech processing device calculates the product of the suppression factor and the target separation signal at each frequency point to obtain the target filtered signal corresponding to the target microphone, and so on to obtain the filtered signals corresponding to different microphones.

[0148] It should be noted that the Kalman filter can only eliminate the content in the target signal that is coherent with the reference signal, and there will be some residue in the separated signal. Therefore, it is necessary to further remove the residual interference signal through the above steps.

[0149] In some embodiments, S38 - S41 may also be referred to as nonlinear processing.

[0150] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0151] In addition, considering the limitations of the Kalman filter, the present application further performs post-processing on the result of the Kalman filter processing to further remove the residual interference signal in the speech signal.

[0152] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0153] S51: Acquire a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0154] Among them, at least a first microphone, a second microphone and a third microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.

[0155] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0156] S52: Perform differential beam processing on the first audio signal to obtain a first processed signal, and perform differential beam processing on the second audio signal to obtain a second processed signal.

[0157] S53, performing Fourier transform on the first processed signal to obtain a first frequency domain signal; performing Fourier transform on the second processed signal to obtain a second frequency domain signal; and performing Fourier transform on the third audio signal to obtain a third frequency domain signal.

[0158] S54 , determining a target signal from the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal in sequence, and using the remaining signals as reference signals.

[0159] S55 , comparing the first energy of the target signal at each frequency point with the second energy of each reference signal.

[0160] S56 , in response to the first energy being less than the second energy, performing Kalman filtering on the target signal using the reference signal to obtain a separated signal corresponding to the target signal at each frequency point.

[0161] S57 , in response to the first energy being greater than or equal to the second energy, taking the target signal as a separation signal.

[0162] S58 , obtaining the filtered signals corresponding to the sound ranges where different microphones are located according to the separated signals at different frequency points.

[0163] S59 , performing an inverse Fourier transform on the first filtered signal corresponding to the first sound range to obtain a first target audio signal generated by the first sound range.

[0164] In some embodiments, the first filtered signal is a time domain signal. It is understandable that the time domain signal is difficult to be directly applied to subsequent voice interaction tasks such as voice wake-up and voice recognition. Therefore, it is necessary to convert the first filtered signal belonging to the time domain signal into the first target audio signal belonging to the frequency domain signal.

[0165] Exemplarily, the speech processing device performs short-time Fourier transform on the first processed signal, and the speech processing device performs short-time inverse Fourier transform on the first filtered signal.

[0166] S60 , performing an inverse Fourier transform on the second filtered signal corresponding to the second sound range to obtain a second target audio signal generated by the second sound range.

[0167] In some embodiments, the second filtered signal is a time domain signal. It is understandable that the time domain signal is difficult to be directly applied to subsequent voice interaction tasks such as voice wake-up and voice recognition. Therefore, it is necessary to convert the second filtered signal belonging to the time domain signal into a second target audio signal belonging to the frequency domain signal.

[0168] Exemplarily, the speech processing device performs short-time Fourier transform on the second processed signal, and the speech processing device performs short-time inverse Fourier transform on the second filtered signal.

[0169] S61 , performing an inverse Fourier transform on the third filtered signal corresponding to the third sound range to obtain a third target audio signal generated by the third sound range.

[0170] In some embodiments, the third filtered signal is a time domain signal. It is understandable that the time domain signal is difficult to be directly applied to subsequent voice interaction tasks such as voice wake-up and voice recognition. Therefore, it is necessary to convert the third filtered signal belonging to the time domain signal into a third target audio signal belonging to the frequency domain signal.

[0171] Exemplarily, the speech processing device performs short-time Fourier transform on the third processed signal, and the speech processing device performs short-time inverse Fourier transform on the third filtered signal.

[0172] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0173] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0174] S71: Acquire a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0175] The vehicle is provided with at least a first microphone, a second microphone, and a third microphone, each of which corresponds to a sound zone. The first microphone corresponds to the first sound zone, and the second microphone corresponds to the second sound zone.

[0176] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0177] S72: Acquire a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range.

[0178] Specifically, in response to a target distance between the first microphone and the second microphone being less than a preset threshold, the speech processing device determines that the first microphone and the second microphone constitute a linear array microphone. Furthermore, the speech processing device determines a first suppression angle θ2 of the first microphone relative to the second sound range, and a second suppression angle θ1 of the second microphone relative to the first sound range, based on a large amount of pre-acquired sound pickup data from the first and second microphones collected from real vehicles, as well as theoretical modeling results.

[0179] S73 , performing differential beam processing on the first audio signal using the first suppression angle to obtain a first processed signal; and performing differential beam processing on the second audio signal using the second suppression angle to obtain a second processed signal.

[0180] Specifically, the speech processing device performs differential beam processing on the first audio signal using the first suppression angle θ2 to obtain a first processed signal.

[0181] Furthermore, the speech processing device performs differential beam processing on the second audio signal using the second suppression angle θ1 to obtain a second processed signal.

[0182] S74, performing Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point, filtering out audio signals generated by the sound region not corresponding to the first microphone in the first processed signal, filtering out audio signals generated by the sound region not corresponding to the second microphone in the second processed signal, and filtering out audio signals generated by the sound region not corresponding to the third microphone in the third audio signal.

[0183] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0184] In addition, the in-vehicle speech processing method provided in this application utilizes a differential beam processing method to increase the energy difference of the sound zone signals collected by compactly arranged microphones, thereby improving the filtering effect of subsequent Kalman filtering.

[0185] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0186] S81: Acquire a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, and a third audio signal collected by a third microphone.

[0187] The vehicle is provided with at least a first microphone, a second microphone, and a third microphone, each of which corresponds to a sound zone. The first microphone corresponds to the first sound zone, and the second microphone corresponds to the second sound zone.

[0188] The target distance between the first microphone and the second microphone is smaller than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are larger than the preset thresholds respectively.

[0189] S82: Acquire a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range.

[0190] S83: Determine a first filter coefficient according to the first suppression angle, the target distance, and the signal angular frequency.

[0191] In some embodiments, the speech processing device calculates a ratio of the target distance to the speed of sound to obtain a target ratio.

[0192] Furthermore, the speech processing device determines a first filter coefficient according to the first suppression angle, the target ratio, and the signal angular frequency.

[0193] Among them, the first filter coefficient satisfies the following relationship:

[0194]

[0195] Wherein, C1 is the first filter coefficient.

[0196] S84: Perform differential beam processing on the first audio signal using the first filter coefficient to obtain a first processed signal.

[0197] Specifically, the speech processing device performs differential beam processing on the first audio signal using the first filter coefficient to obtain a first processed signal.

[0198] S85: Determine a second filter coefficient according to the second suppression angle, the target distance, and the signal angular frequency.

[0199] In some embodiments, the speech processing device determines the second filter coefficient based on the second suppression angle, the target coefficient, and the signal angular frequency.

[0200] The second filter coefficient satisfies the following relationship:

[0201]

[0202] Wherein, C2 is the first filter coefficient.

[0203] S86: Perform differential beam processing on the second audio signal using the second filter coefficient to obtain a second processed signal.

[0204] Specifically, the speech processing device performs differential beam processing on the second audio signal using the second filter coefficient to obtain a second processed signal.

[0205] S87, performing Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point, filtering out audio signals generated by the sound range not corresponding to the first microphone in the first processed signal, filtering out audio signals generated by the sound range not corresponding to the second microphone in the second processed signal, and filtering out audio signals generated by the sound range not corresponding to the third microphone in the third audio signal.

[0206] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0207] In addition, the in-vehicle speech processing method provided in this application utilizes a differential beam processing method to increase the energy difference of the sound zone signals collected by compactly arranged microphones, thereby improving the filtering effect of subsequent Kalman filtering.

[0208] See also Figure 3 , Figure 3 This is a flow chart of another embodiment of the method for processing in-vehicle voice provided by the present application.

[0209] like Figure 3 As shown, another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:

[0210] S101: Acquire a first audio signal collected by a first microphone, a second audio signal collected by a second microphone, a third audio signal collected by a third microphone, and a fourth audio signal collected by a fourth microphone.

[0211] The first microphone and the second microphone form a linear microphone array, the third microphone does not form a linear microphone array with other microphones, and the fourth microphone does not form a linear microphone array with other microphones.

[0212] In this embodiment, the layout of the vehicle microphone is as follows: Figure 2B shown.

[0213] S102: Perform differential beam processing on the first audio signal and the second audio signal respectively to obtain a first processed signal corresponding to the first audio signal and a second processed signal corresponding to the second audio signal.

[0214] Since the first microphone and the second microphone are close to each other, the energy difference between the speech signals corresponding to the first sound range and the second sound range picked up by the two microphones is small, which affects the effect of subsequent Kalman filtering and causes more speech damage.

[0215] Therefore, the voice processing device performs corresponding differential beam processing on the first audio signal s1 and the second audio signal s2 corresponding to the compactly arranged first microphone and the second microphone, respectively, and weakens the sound of the second sound zone received by the first microphone and the sound of the first sound zone received by the second microphone by directional interference suppression, thereby obtaining the corresponding first processed signal s′1 and the second processed signal s′2.

[0216] Through the above steps, the energy difference of the signals in different sound zones collected by the vehicle-mounted microphone is increased.

[0217] S103, performing Fourier transform on the first processed signal to obtain a first frequency domain signal; performing Fourier transform on the second processed signal to obtain a second frequency domain signal; performing Fourier transform on the third audio signal to obtain a third frequency domain signal; performing Fourier transform on the fourth audio signal to obtain a fourth frequency domain signal.

[0218] Specifically, the speech processing device performs Fourier transform on the first processed signal s′1, the second processed signal s′2, the third audio signal s3, and the fourth audio signal s4 to obtain the first frequency domain signal X1, the second frequency domain signal X2, the third frequency domain signal X3, and the fourth frequency domain signal X4, thereby realizing the conversion of the time domain signal to the frequency domain signal.

[0219] S104 , performing Kalman filtering on the first frequency domain signal, the second frequency domain signal, the third frequency domain signal, and the fourth frequency domain signal at each frequency point to obtain corresponding first filtered signals, second filtered signals, third filtered signals, and fourth filtered signals.

[0220] Specifically, the Kalman filter performed in this embodiment satisfies the following relationship:

[0221]

[0222] Among them, d n is the target signal, x n is the reference signal, w n is the Kalman filter coefficient, se n is the error signal, that is, the filtered signal output after filtering.

[0223] Assuming there is only one speaker at each frequency, the microphone closest to the speaking area receives the highest signal energy. Therefore, the energy at the frequency is used to determine whether the current speaker is speaking and, therefore, whether filtering should be performed at that frequency. Each microphone signal is filtered once with all other microphone signals to produce the separation result for that area.

[0224] In some application scenarios, S104 may specifically include the following steps:

[0225] The speech processing device uses the first frequency domain signal X1 as the target signal, and the second frequency domain signal X2, the third frequency domain signal X3, and the fourth frequency domain signal X4 as reference signals, respectively, and traverses all frequency points.

[0226] If the energy of the reference signal is greater than the energy of the target signal at a certain frequency, it is considered that the speech is in the corresponding range of the reference signal, and the speech processing device performs a Kalman filter on the target signal. Otherwise, it is considered that the speech is in the corresponding range of the target signal, and the speech processing device does not perform any operation, thereby obtaining the separated signal se1 of the first speech range.

[0227] Similarly, the speech processing device uses the second frequency domain signal X2 as the target signal, and the first frequency domain signal X1, the third frequency domain signal X3, and the fourth frequency domain signal X4 as reference signals, and repeats the above steps to obtain the separated signal se2 of the second sound range.

[0228] Similarly, the speech processing device uses the third frequency domain signal X3 as the target signal, and the first frequency domain signal X1, the second frequency domain signal X2, and the fourth frequency domain signal X4 as reference signals, and repeats the above steps to obtain the separated signal se3 of the third sound zone.

[0229] Similarly, the speech processing device uses the fourth frequency domain signal X4 as the target signal, and the first frequency domain signal X1, the second frequency domain signal X2, and the third frequency domain signal X3 as reference signals respectively, and repeats the above steps to obtain the separated signal se4 of the rear seat of the co-pilot.

[0230] S105 , performing nonlinear processing on the first filtered signal, the second filtered signal, the third filtered signal, and the fourth filtered signal at each frequency point to obtain a speech separation result corresponding to each sound region.

[0231] The Kalman filter can only eliminate the content of the target signal that is relevant to the reference signal, and some residues will remain. To overcome the above problem, the speech processing device calculates the target signal d at each frequency point. n and error signal se n , target signal d n With the reference signal x n The frequency domain coherence values ​​are calculated as follows:

[0232]

[0233] Where D(f), X(f), and E(f) correspond to the target signal, reference signal, and error signal, respectively. The nonlinearity is reflected in the calculation of the suppression factor based on the coherence value, and the calculation formula is as follows:

[0234] α=max(0.05,min(1-cohxd,cohed))

[0235] Furthermore, the speech processing device calculates the product of the suppression factor and the error signal, and determines the speech separation result S1 corresponding to the first sound area, the speech separation result S2 corresponding to the second sound area, the speech separation result S3 corresponding to the third sound area, and the speech separation result S4 corresponding to the fourth sound area, so as to suppress the residual signal.

[0236] Among them, the above-mentioned speech separation results are all time domain signals.

[0237] In the above solution, the distance between the first microphone and the second microphone is less than a preset threshold. By performing differential beam processing on the first audio signal collected by the first microphone and the second audio signal collected by the second microphone, a first processed signal and a second processed signal are obtained, so as to increase the difference between the audio signals collected by the microphones with a smaller distance. At each frequency point, based on the energy difference between the audio signals, the audio signal is Kalman filtered to filter out the signals generated by other sound zones in the audio signal, thereby accurately achieving the separation of the voice signal. Furthermore, the in-car voice processing method provided by the present application can further improve the accuracy of voice interaction for users in the car and enhance the user experience.

[0238] In addition, the in-vehicle speech processing method provided in this application utilizes a differential beam processing method to increase the energy difference of the sound zone signals collected by compactly arranged microphones, thereby improving the filtering effect of subsequent Kalman filtering.

[0239] Please continue to see Figure 4 , Figure 4 The terminal device 500 of the embodiment of the present application includes a processor 51 and a memory 52 .

[0240] The processor 51 and the memory 52 are connected to the bus. The memory 52 stores program data. The processor 51 is used to execute the program data to implement the in-vehicle speech processing method described in the above embodiment.

[0241] In the embodiment of the present application, the processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip having signal processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 51 may be any conventional processor.

[0242] This application also provides a computer storage medium, please continue to refer to Figure 5 , Figure 5 It is a structural diagram of an embodiment of a computer storage medium provided in the present application. The computer storage medium 600 stores program data 61. When the program data 61 is executed by the processor, it is used to implement the in-vehicle voice processing method of the above embodiment.

[0243] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0244] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Equivalent structures or equivalent process changes made by utilizing the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for processing in-vehicle speech, characterized in that: At least a first microphone, a second microphone, and a third microphone are provided inside the vehicle, each microphone corresponding to a sound zone, and the method includes: Obtaining a first audio signal collected by the first microphone, a second audio signal collected by the second microphone, and a third audio signal collected by the third microphone; wherein the target distance between the first microphone and the second microphone is less than a preset threshold, and the distances between the third microphone and the first microphone and the second microphone are respectively greater than the preset threshold; performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal; performing Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point, filtering out audio signals generated in a sound range not corresponding to the first microphone in the first processed signal, filtering out audio signals generated in a sound range not corresponding to the second microphone in the second processed signal, and filtering out audio signals generated in a sound range not corresponding to the third microphone in the third audio signal; The performing Kalman filtering processing according to the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point includes: performing a Fourier transform on the first processed signal to obtain a first frequency domain signal; performing a Fourier transform on the second processed signal to obtain a second frequency domain signal; and performing a Fourier transform on the third audio signal to obtain a third frequency domain signal; The step of performing Kalman filtering on the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal according to the energy of each frequency point to obtain filtered signals corresponding to the sound ranges where different microphones are located includes: determining a target signal from the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal in sequence, and using the remaining signals as reference signals; Comparing a first energy of the target signal at each frequency point with a second energy of each reference signal; In response to the first energy being less than the second energy, performing Kalman filtering on the target signal using the reference signal to obtain a separated signal corresponding to the target signal at each frequency point; In response to the first energy being greater than or equal to the second energy, using the target signal as the separation signal; The filtered signals corresponding to the sound ranges where different microphones are located are obtained according to the separated signals at different frequency points.

2. The method according to claim 1, characterized in that The step of performing Kalman filtering based on the energy of the first processed signal, the second processed signal, and the third audio signal at each frequency point to filter out audio signals generated by a sound region not corresponding to the first microphone in the first processed signal, filter out audio signals generated by a sound region not corresponding to the second microphone in the second processed signal, and filter out audio signals generated by a sound region not corresponding to the third microphone in the third audio signal includes: Kalman filtering is performed on the first frequency domain signal, the second frequency domain signal, and the third frequency domain signal based on the energy level of each frequency point. Audio signals generated by sound regions not corresponding to the first microphone in the first frequency domain signal are filtered out, audio signals generated by sound regions not corresponding to the second microphone in the second frequency domain signal are filtered out, and audio signals generated by sound regions not corresponding to the third microphone in the third frequency domain signal are filtered out to obtain filtered signals corresponding to sound regions where different microphones are located.

3. The method according to claim 1, characterized in that The obtaining, according to the separated signals at different frequency points, the filtered signals corresponding to different microphones includes: Calculating a first frequency domain coherence value between the target signal and the separated signal at each frequency point; Calculating a second frequency domain coherence value between the target signal and each reference signal at each frequency point; Determining a suppression factor at each frequency point based on the first frequency domain coherence value and the second frequency domain coherence value corresponding to the same frequency point; The filtering signals corresponding to different microphones are obtained by using the suppression factor and the separated signal.

4. The method according to claim 3, characterized in that The inhibitory factor is obtained according to the following formula: α=max(0.05,min(1-cohxd,cohed)), where α is the suppression factor, cohed is the first frequency domain coherence value, and cohxd is the second frequency domain coherence value.

5. The method according to claim 1, wherein The first microphone corresponds to the first sound range, the second microphone corresponds to the second sound range, and the third microphone corresponds to the third sound range; After the step of obtaining the filtered signals corresponding to the sound ranges where different microphones are located according to the separated signals at different frequency points, the method further includes: performing an inverse Fourier transform on a first filtered signal corresponding to the first sound region to obtain a first target audio signal generated by the first sound region; performing an inverse Fourier transform on the second filtered signal corresponding to the second sound range to obtain a second target audio signal generated by the second sound range; An inverse Fourier transform is performed on the third filtered signal corresponding to the third sound range to obtain a third target audio signal generated by the third sound range.

6. The method according to claim 1, wherein The first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range; The performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal, includes: Acquire a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range; The first audio signal is differentially beam processed using the first suppression angle to obtain the first processed signal; and the second audio signal is differentially beam processed using the second suppression angle to obtain the second processed signal.

7. The method according to claim 6, characterized in that The performing differential beam processing on the first audio signal by using the first suppression angle to obtain the first processed signal includes: determining a first filter coefficient according to the first suppression angle, the target distance, and the signal angular frequency; performing differential beam processing on the first audio signal using the first filter coefficient to obtain the first processed signal; The performing differential beam processing on the second audio signal by using the second suppression angle to obtain the second processed signal includes: determining a second filter coefficient according to the second suppression angle, the target distance, and the signal angular frequency; Perform differential beam processing on the second audio signal using the second filter coefficient to obtain the second processed signal.

8. A terminal device, characterized in that: The terminal device includes a processor and a memory connected to the processor, wherein: The memory stores program instructions; The processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The storage medium stores program instructions, and when the program instructions are executed, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method of determining noise sound contributions of noise sources of a motorized vehicle

    CN105473988A

  • In-vehicle voice frequency signal processing method and device

    CN109545230A