In-vehicle voice processing method, terminal device, and storage medium
Through differential beam processing and multi-channel speech separation models, the accuracy issues of speech recognition and wake-up in the multi-zone voice interaction system in the car are solved, improving the user experience.
Patent Information
- Application Number
- CN202411356141.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-09-26
AI Technical Summary
In the complex environment inside the car, the audio signals collected by the on-board microphone contain voice and noise in multiple sound zones, resulting in incorrect responses from the voice interaction system and a poor user experience.
Differential beam processing and a multi-channel speech separation model are used to perform differential beamforming through the on-board microphone array to enhance the speech signal in the target sound area and suppress interference signals in other sound areas. The multi-channel speech separation model is then used to further separate the speech signal.
It improves the accuracy of speech recognition and voice wake-up, enhances user experience, and ensures the effectiveness of the voice interaction system.
Smart Images

Figure CN119446172B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice processing technology, and in particular to a method for processing in-vehicle voice, a terminal device, and a storage medium. Background Art
[0002] With the rapid development of technology, voice interaction technology is gradually being applied to in-vehicle connected scenarios. Users are becoming more and more accustomed to interacting with in-vehicle devices through voice, resulting in an increasing demand for in-vehicle voice interaction systems. To meet the needs of voice interaction between individual users and in-vehicle devices, in-vehicle voice interaction systems have launched in-vehicle multi-zone voice interaction services to expand the scope of voice interaction.
[0003] Due to the dense crowds and complex environment inside a vehicle, the audio signal captured by the onboard microphone corresponding to a specific audio zone contains not only the speech in that zone, but also speech from other zones and the vehicle's own audio. Directly using audio signals as voice commands is prone to response errors, resulting in a poor user experience. Summary of the Invention
[0004] The present application provides a method for processing in-vehicle voice, a terminal device, and a storage medium.
[0005] A technical solution adopted by the present application is to provide a method for processing in-vehicle voice, wherein at least a first microphone and a second microphone are provided inside the vehicle, each microphone corresponding to a sound zone; the method comprises:
[0006] Acquire a first audio signal collected by a first microphone and a second audio signal collected by a second microphone;
[0007] In response to a target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal;
[0008] The first processed signal and the second processed signal are input into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone; wherein the first separated signal does not include audio signals generated by sound ranges other than those corresponding to the first microphone, and the second separated signal does not include audio signals generated by sound ranges other than those corresponding to the second microphone.
[0009] Optionally, a third microphone is further provided inside the vehicle, and the distances between the third microphone and the first microphone and the second microphone are respectively greater than a preset threshold;
[0010] Inputting the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone, including:
[0011] The first processed signal, the second processed signal, and the third audio signal collected by the third microphone are input into a multi-channel speech separation model to obtain the speech processing signal generated by each sound zone, and obtain the first separated signal, the second separated signal and the third separated signal, wherein the third separated signal does not include the audio signal generated by the sound zone not corresponding to the third microphone.
[0012] Optionally, inputting the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone includes:
[0013] Performing a short-time Fourier transform on the first processed signal to obtain a first frequency domain signal; and performing a short-time Fourier transform on the second processed signal to obtain a second frequency domain signal;
[0014] Inputting the first frequency domain signal and the second frequency domain signal into a multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal;
[0015] Calculating the product of the first frequency domain signal and the first frequency domain mask to obtain a separated first frequency domain signal; and calculating the product of the second frequency domain signal and the second frequency domain mask to obtain a separated second frequency domain signal;
[0016] An inverse short-time Fourier transform is performed on the separated first frequency domain signal to obtain a first separated signal; and an inverse short-time Fourier transform is performed on the separated second frequency domain signal to obtain a second separated signal.
[0017] Optionally, the multi-channel speech separation model includes an encoder, a time domain convolutional network and a decoder connected in sequence,
[0018] Inputting the first frequency domain signal and the second frequency domain signal into a multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal, including:
[0019] Inputting the first frequency domain signal to the encoder to obtain a first encoding result, and inputting the second frequency domain signal to the encoder to obtain a second encoding result;
[0020] Inputting the first encoding result into a time-domain convolutional network to obtain a first convolution result; and inputting the second encoding result into a time-domain convolutional network to obtain a second convolution result;
[0021] The first convolution result is input to a decoder to obtain a first frequency domain mask; and the second convolution result is input to a decoder to obtain a second frequency domain mask.
[0022] Optionally, the first audio signal is obtained by superimposing a first initial audio signal and a first noise signal, and the second audio signal is obtained by superimposing a second initial audio signal and a second noise signal;
[0023] After the steps of performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain a first separated signal; and performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain a second separated signal, the method includes:
[0024] determining a first loss using a first difference between the first separated signal and the first original audio signal; and determining a second loss using a difference between the second separated signal and the second original audio signal;
[0025] Using the first loss and the second loss, determine the total loss;
[0026] The total loss is used to adjust the parameters of the multi-channel speech separation model.
[0027] Optionally, obtaining a first audio signal collected by a first microphone and a second audio signal collected by a second microphone includes:
[0028] Acquire a first reference audio signal, a second reference audio signal, a first reference noise signal, and a second reference noise signal;
[0029] Determine the room impulse response based on the vehicle's dimensions, microphone placement, and passenger positions;
[0030] Calculating a convolution result of the first reference audio signal and the room impulse response to obtain a first initial audio signal; and calculating a convolution result of the second reference audio signal and the room impulse response to obtain a second initial audio signal;
[0031] Calculating a convolution result of a first reference noise signal and a room impulse response to obtain a first noise signal; and calculating a convolution result of a second reference noise signal and a room impulse response to obtain a second noise signal;
[0032] The first initial audio signal and the first noise signal are superimposed to obtain a first audio signal; and the second initial audio signal and the second noise signal are superimposed to obtain a second audio signal.
[0033] Optionally, the first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range;
[0034] Performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal, including:
[0035] Obtaining a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range;
[0036] The first audio signal is differentially beam processed using a first suppression angle to obtain a first processed signal; and the second audio signal is differentially beam processed using a second suppression angle to obtain a second processed signal.
[0037] Optionally, performing differential beam processing on the first audio signal using the first suppression angle to obtain a first processed signal includes:
[0038] determining a first filter coefficient according to a first suppression angle, a target distance, and a signal angular frequency;
[0039] Performing differential beam processing on the first audio signal using the first filter coefficient to obtain a first processed signal;
[0040] Performing differential beam processing on the second audio signal using the second suppression angle to obtain a second processed signal includes:
[0041] determining a second filter coefficient according to a second suppression angle, a target distance, and a signal angular frequency;
[0042] The second audio signal is subjected to differential beam processing using the second filter coefficient to obtain a second processed signal.
[0043] Another technical solution adopted by the present application is to provide a terminal device, the terminal device comprising a memory and a processor connected to the memory;
[0044] The memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned in-vehicle voice processing method.
[0045] Another technical solution adopted in the present application is to provide a computer storage medium, which is used to store program data. When the program data is executed by a computer, it is used to implement the above-mentioned in-vehicle voice processing method.
[0046] The beneficial effects of the present application are as follows: in response to the target distance between the first microphone and the second microphone being less than a preset threshold, beamforming processing is performed on the first audio signal and the second audio signal respectively, and a first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and a second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sounds generated from other sound zones. The first processed signal and the second processed signal are respectively input into a multi-channel speech separation model to obtain a first separated signal and a second separated signal, and the audio signals that do not correspond to the own sound zone are more effectively filtered out. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided by the present application can more accurately perform speech interaction tasks such as speech recognition and voice wake-up, thereby improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0048] Figure 1 This is a flow chart of an embodiment of a method for processing in-vehicle speech provided by the present application;
[0049] Figures 2A to 2E This is a schematic diagram of the distribution locations of vehicle-mounted microphones under different embodiments provided by this application;
[0050] Figure 3 This is a flow chart of another embodiment of the method for processing in-vehicle speech provided by the present application;
[0051] Figure 4 1 is a schematic structural diagram of an embodiment of a multi-channel speech separation model;
[0052] Figure 5 This is a schematic structural diagram of an embodiment of a terminal device provided by this application;
[0053] Figure 6 It is a structural diagram of an embodiment of a computer storage medium provided by this application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] To enhance the passenger experience, smart vehicles utilize onboard microphones at various locations within the vehicle. These microphones capture voice commands from passengers in different locations, providing personalized service for each passenger. However, due to the densely populated and complex environment within the vehicle, the voice signals collected by the onboard microphones include not only the target passenger's voice, but also interference signals from passengers in other locations, audio signals from the vehicle itself, and various ambient noises. Accurately sensing the voice signals of passengers in different locations is an urgent challenge.
[0056] This application mainly designs a set of methods to improve the accuracy of audio signal separation. Unlike traditional methods, this application starts from the positional relationship between vehicle-mounted microphones and combines a multi-channel speech separation model to separate the speech signals. It can separate the audio signals generated by different sound zones from the signals collected by the vehicle-mounted microphones and mixed with audio generated by different sound zones, and accurately identify users speaking in different sound zones.
[0057] Please refer to the following for details: Figure 1 , Figure 1 This is a flow chart of an embodiment of a method for processing in-vehicle voice provided by this application.
[0058] like Figure 1 As shown, the method for processing in-vehicle speech in an embodiment of the present application may specifically include the following steps:
[0059] S1, obtaining a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0060] Among them, at least a first microphone and a second microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.
[0061] like Figure 2A As shown, mic-1 (microphone 1) and mic-2 (microphone 2) correspond to the first sound zone and the second sound zone respectively, wherein mic-1 and mic-2 constitute a linear array microphone array.
[0062] like Figure 2BAs shown, mic-1, mic-2, mic-3, and mic-4 correspond to the first sound zone, the second sound zone, the third sound zone, and the fourth sound zone, respectively. Among them, mic-1 and mic2 constitute a linear array microphone array, and mic-3 and mic-4 constitute a linear array microphone array.
[0063] like Figure 2C As shown, Figure 2C This is a schematic diagram of an embodiment of a distributed microphone array layout, where mic-1, mic-2, mic-3, and mic-4 correspond to the first, second, third, and fourth sound zones, respectively.
[0064] like Figure 2D As shown, mic-1, mic-2, mic-3, and mic-4 correspond to the first sound zone, the second sound zone, the third sound zone, and the fourth sound zone, respectively. Among them, mic-1 and mic-2 form a linear microphone array.
[0065] like Figure 2E As shown, mic-1, mic-2, mic-3, mic-4, mic-5, and mic-6 correspond to the first, second, third, fourth, fifth, and sixth sound zones, respectively. Among them, mic-1 and mic-2 constitute a linear microphone array.
[0066] Figure 2A-2E In the given embodiments, the number of sound zones is an even number. In some possible application scenarios, the number of sound zones may also be an odd number, as long as the number of sound zones in the car is greater than 1.
[0067] It should be understood that the in-car speech processing method provided in this application is not only applicable to Figure 2A-2E The vehicle has the microphone layout shown. It can also be used in vehicles with similar microphone layouts.
[0068] The method for processing in-vehicle speech provided in the present application is mainly performed by a speech processing device. In some embodiments, the speech processing device may be the vehicle-mounted microphone itself. In some embodiments, the distance measuring device may be a device that is communicatively connected to the vehicle-mounted microphone. For example, the device may be a device for monitoring images, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, and any one or more products such as self-driving cars, robots, security systems, glasses, and helmets for augmented reality or virtual reality. In some possible implementations, the method for processing in-vehicle speech may be implemented by a processor calling computer-readable instructions stored in a memory.
[0069] Specifically, the speech processing device obtains a first audio signal collected by a first microphone, and obtains a second audio signal collected by a second microphone.
[0070] Exemplarily, the first microphone corresponds to the first sound zone, and the second microphone corresponds to the second sound zone. Due to the limited space in the car, the first audio signal inevitably contains the sound produced by the second sound zone (or other sounds that do not belong to the first sound zone). Similarly, the second audio signal inevitably contains the sound produced by the first sound zone (or other sounds that do not belong to the second sound zone). Therefore, the main goal of the present application is to perform a correlation separation operation on the first audio signal, filter out the audio signals that do not belong to the first sound zone in the first audio signal, and finally obtain a first separated signal. Similarly, the present application also performs a correlation separation operation on the second audio signal, filters out the audio signals that do not belong to the second sound zone in the second audio signal, and finally obtains a second separated signal.
[0071] S2, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal.
[0072] Specifically, the condition for two vehicle-mounted microphones to form a linear microphone array is that the relative distance between the two vehicle-mounted microphones is smaller than a preset threshold.
[0073] Similarly, when the relative distance between two vehicle-mounted microphones is greater than or equal to a preset threshold, the two vehicle-mounted microphones do not constitute a linear microphone array. In some embodiments, the preset threshold is less than 50 cm. In some embodiments, the preset threshold is less than 40 cm. In some embodiments, the preset threshold is less than 30 cm. In some embodiments, the preset threshold is less than 20 cm. In some embodiments, the preset threshold is less than 10 cm. Exemplarily, the preset threshold is 15 cm.
[0074] In this embodiment, because the target distance between the first microphone and the second microphone is less than the preset threshold, the first microphone is more likely to obtain audio from the second sound range, and the second microphone is more likely to obtain audio from the first sound range. In other words, the audio processing device may have difficulty distinguishing audio signals that do not belong to the first sound range from the first audio signal, and / or may have difficulty distinguishing audio signals that do not belong to the second sound range from the second audio signal.
[0075] Related technologies primarily employ beamforming. However, this technology is used in an additive microphone array consisting of a first and second microphone. The main lobe shape of the additive microphone array is determined by setting the main lobe itself. For the first and second microphones in a vehicle-mounted microphone array, the spacing between them is small, and the main lobe width is large, resulting in poor directivity and insignificant array gain. This results in weak enhancement of the target direction and poor suppression of interference.
[0076] The differential beamforming algorithm is a microphone array-based sound source localization and sound separation technology that measures differences in spatial sound pressure. Compared to related technologies, the differential microphone array determines the shape of its main lobe by setting the null point direction. By placing the null point in the interference direction, the speech energy ratio in the target direction and the interference direction is increased. The array's sensitivity to different incident directions is represented by the beam pattern. The differential beamform satisfies the following relationship:
[0077] B[h(ω),θ]=d G (ω,cosθ)h(ω)
[0078] Where d(ω, cosθ) is the array's steering vector, ω is the signal angular frequency, θ is the incident direction, h(ω) is the output weight of each microphone, and H is the conjugate transpose. For example, with dual microphones with an end-fire direction of 0° and a suppression direction of 180°, the beam pattern calculation formula for differential beamforming is:
[0079]
[0080] Where h′(ω) is τ is the ratio of the microphone array spacing to the speed of sound.
[0081] In this embodiment, the speech processing device applies a differential beamforming algorithm to audio enhancement for a vehicle's linear microphone array. For example, the speech processing device sets the null point at the angle of the second sound zone relative to the microphone array, effectively suppressing the audio energy in the second sound zone. This emphasizes the audio signal in the first sound zone and increases the energy difference between the audio signals in the first and second sound zones.
[0082] For another example, the speech processing device sets the zero point to the angle of the first sound zone relative to the microphone array, thereby effectively suppressing the audio energy in the direction of the first sound zone and highlighting the audio signal of the second sound zone.
[0083] In this embodiment, the first and second audio signals are processed by the differential beamforming algorithm, resulting in a larger energy difference between the first and second processed signals, and a signal energy distribution similar to that obtained by a distributed microphone. A distributed microphone refers to two microphones separated by a distance greater than or equal to a preset threshold.
[0084] S3: Input the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone.
[0085] The first separated signal does not include audio signals generated by a sound zone not corresponding to the first microphone, and the second separated signal does not include audio signals generated by a sound zone not corresponding to the second microphone.
[0086] Specifically, the multi-channel speech separation model uses audio signals obtained by multiple on-board microphones (such as the first microphone and the second microphone) to distinguish and extract independent sound sources in mixed speech. In a noisy environment, it can accurately separate and recognize speech, distinguish different speaking objects, and has good adaptability to different acoustic environments and noise conditions.
[0087] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0088] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0089] S11, obtaining a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0090] A third microphone is also provided inside the vehicle, and the distances between the third microphone and the first microphone and the second microphone are respectively greater than preset thresholds.
[0091] S12, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal.
[0092] S13, input the first processed signal, the second processed signal, and the third audio signal collected by the third microphone into the multi-channel speech separation model to obtain the speech processing signal generated by each sound zone, and obtain the first separated signal, the second separated signal and the third separated signal.
[0093] The third separated signal does not include audio signals generated by the sound range not corresponding to the third microphone.
[0094] In this embodiment, since the distance between the third microphone and the first microphone is greater than the preset threshold, the distance between the third microphone and the second microphone is greater than the preset threshold, that is, the third microphone does not constitute a linear array microphone array with the first microphone, and the third microphone does not constitute a linear array microphone array with the second microphone.
[0095] Furthermore, because the third microphone does not form a linear microphone array with either the first or second microphone, the speech processing device can more easily perform corresponding enhancement operations on the third audio signal. Therefore, the speech processing device does not need to perform corresponding differential beamforming on the third audio signal collected by the third microphone and can directly input the third audio signal into the multi-channel speech separation model.
[0096] It should be noted that the third microphone and the first microphone constitute a distributed microphone, and the third microphone and the second microphone constitute a distributed microphone. Therefore, there is a relatively obvious energy difference between the third audio signal and the first processed signal, and there is a relatively obvious energy difference between the third audio signal and the second processed signal. Therefore, there is no need to perform corresponding differential beam processing on the third audio signal.
[0097] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0098] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0099] S21: Acquire a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0100] Among them, at least a first microphone and a second microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.
[0101] S22 , in response to a target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal.
[0102] S23, performing a short-time Fourier transform on the first processed signal to obtain a first frequency domain signal; and performing a short-time Fourier transform on the second processed signal to obtain a second frequency domain signal.
[0103] Specifically, the speech processing device performs a short-time Fourier transform (STFT) on the first processed signal, which is a time domain signal, to obtain a first frequency domain signal. Furthermore, the speech processing device performs a short-time Fourier transform (STFT) on the second processed signal, which is a time domain signal, to obtain a second frequency domain signal.
[0104] This step realizes the conversion of time domain signals into frequency domain signals, so that in the subsequent steps, the multi-channel speech separation model can extract the information of the audio signal generated by the sound zone based on the frequency domain characteristics of the signal.
[0105] S24, inputting the first frequency domain signal and the second frequency domain signal into a multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal.
[0106] Specifically, the speech processing device inputs the first frequency domain signal into the multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal, and inputs the second frequency domain signal into the multi-channel speech separation model to obtain a second frequency domain mask corresponding to the second frequency domain signal.
[0107] In some embodiments, the first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range. Further, the first frequency domain mask corresponds to the mask of the first sound range, and the second frequency domain mask corresponds to the mask of the second sound range.
[0108] S25, calculating the product of the first frequency domain signal and the first frequency domain mask to obtain the separated first frequency domain signal; and calculating the product of the second frequency domain signal and the second frequency domain mask to obtain the separated second frequency domain signal.
[0109] In some embodiments, the speech processing device calculates the product of the first frequency domain signal and the first frequency domain mask, and uses the product as the separated first frequency domain signal.
[0110] In some embodiments, the speech processing device calculates the product of the second frequency domain signal and the second frequency domain mask, and uses the product as the separated second frequency domain signal.
[0111] S26, performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain a first separated signal; and performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain a second separated signal.
[0112] It should be noted that the signal output by S25 is a frequency domain signal and cannot be directly used as an input signal for subsequent speech recognition and voice wake-up. Therefore, the signal output by S25 needs to be converted from a frequency domain signal to a time domain signal.
[0113] In some embodiments, the first microphone corresponds to a first sound zone, and the second microphone corresponds to a second sound zone. For example, if no sound is produced in the first sound zone (e.g., no one is speaking in the first sound zone), the first separated signal is a silence signal. For another example, if sound is produced in the first sound zone (e.g., someone is speaking in the first sound zone), the first separated signal is the sound produced in the first sound zone.
[0114] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0115] In addition, the in-car speech processing method provided in this application, based on the preprocessing of the microphone linear array, uses a multi-channel speech separation network to further learn the energy differences of audio signals in different channels and different sound zones, and more accurately separate the speech signals generated by each sound zone.
[0116] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0117] S31: Acquire a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0118] Among them, at least a first microphone and a second microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.
[0119] S32 , in response to a target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal.
[0120] S33, performing short-time Fourier transform on the first processed signal to obtain a first frequency domain signal; and the speech processing device performing short-time Fourier transform on the second processed signal to obtain a second frequency domain signal.
[0121] Among them, the multi-channel speech separation model includes an encoder, a time domain convolutional network and a decoder connected in sequence.
[0122] S34 , inputting the first frequency domain signal into an encoder to obtain a first encoding result, and inputting the second frequency domain signal into an encoder to obtain a second encoding result.
[0123] Specifically, the speech processing apparatus inputs the first frequency domain signal into the encoder to obtain a first encoding result, and the speech processing apparatus inputs the second frequency domain signal into the encoder to obtain a second encoding result.
[0124] S35, inputting the first encoding result into the time domain convolutional network to obtain a first convolution result; and, the speech processing device inputting the second encoding result into the time domain convolutional network to obtain a second convolution result.
[0125] The temporal convolutional network (TCN) used in this embodiment utilizes the characteristics of convolutional neural networks to model long-term dependencies in a shorter time and has better parallel computing capabilities.
[0126] Specifically, the speech processing device inputs the first encoding result into a time domain convolutional network to obtain a first convolution result. Also, the speech processing device inputs the second encoding result into a time domain convolutional network to obtain a second convolution result.
[0127] S36, inputting the first convolution result into a decoder to obtain a first frequency domain mask; and inputting the second convolution result into a decoder to obtain a second frequency domain mask.
[0128] Specifically, the speech processing device inputs the first convolution result into a decoder to obtain a first frequency domain mask. Also, the speech processing device inputs the second convolution result into a decoder to obtain a second frequency domain mask.
[0129] S37, calculating the product of the first frequency domain signal and the first frequency domain mask to obtain the separated first frequency domain signal; and calculating the product of the second frequency domain signal and the second frequency domain mask to obtain the separated second frequency domain signal.
[0130] S38, performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain a first separated signal; and performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain a second separated signal.
[0131] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0132] In addition, the in-car speech processing method provided in this application, based on the preprocessing of the microphone linear array, uses a multi-channel speech separation network to further learn the energy differences of audio signals in different channels and different sound zones, and more accurately separate the speech signals generated by each sound zone.
[0133] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0134] S41: Acquire a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0135] Among them, at least a first microphone and a second microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.
[0136] The first audio signal is obtained by superimposing a first initial audio signal and a first noise signal, and the second audio signal is obtained by superimposing a second initial audio signal and a second noise signal;
[0137] In some embodiments, the first initial audio signal is a clean audio signal generated only in the first sound zone inside the vehicle, that is, the first initial audio signal does not contain noise signals or audio signals generated in other sound zones. Exemplarily, the first initial audio signal is a speech signal.
[0138] Likewise, the second initial audio signal is a clean audio signal generated only in the second sound zone inside the vehicle.
[0139] In some embodiments, the first noise signal may include an audio signal not generated in the first sound zone, or may include ambient noise. Similarly, the second noise signal may include an audio signal not generated in the second sound zone, or may include ambient noise.
[0140] S42 , in response to a target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal.
[0141] S43, performing a short-time Fourier transform on the first processed signal to obtain a first frequency domain signal; and performing a short-time Fourier transform on the second processed signal to obtain a second frequency domain signal.
[0142] S44: Input the first frequency domain signal and the second frequency domain signal into a multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal.
[0143] S45, calculating the product of the first frequency domain signal and the first frequency domain mask to obtain the separated first frequency domain signal; and calculating the product of the second frequency domain signal and the second frequency domain mask to obtain the separated second frequency domain signal.
[0144] S46, performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain a first separated signal; and performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain a second separated signal.
[0145] S47 , determining a first loss using a first difference between the first separated signal and the first original audio signal; and determining a second loss using a difference between the second separated signal and the second original audio signal.
[0146] It can be understood that, during the training process of the multi-channel speech separation model, the smaller the difference between the first separated signal and the first initial audio signal, the better, and the smaller the difference between the second separated signal and the second initial audio signal, the better.
[0147] Exemplarily, the first difference is determined by a first degree of similarity between the first separated signal and the first initial audio signal. For example, the higher the first degree of similarity, the smaller the first difference, and further, the smaller the first loss. Similarly, the second difference is determined by a second degree of similarity between the second separated signal and the second initial audio signal. The higher the second degree of similarity, the smaller the second difference, and further, the smaller the second loss.
[0148] S48, determining the total loss using the first loss and the second loss.
[0149] Specifically, the relationship between the total loss and the first loss and the second loss satisfies the following relationship:
[0150] Loss_t=Loss1*w1+Loss2*w2
[0151] Among them, Loss_t is the total loss, Loss1 is the first loss, w1 is the first weight, Loss2 is the second loss, and w2 is the second weight.
[0152] S49, using the total loss, to adjust the parameters of the multi-channel speech separation model.
[0153] Specifically, the speech processing device performs single or multiple parameter adjustments on the multi-channel speech separation model so that the total loss is less than a preset loss threshold, thereby completing the training of the multi-channel speech separation model.
[0154] In some possible embodiments, the speech processing device performs operations such as pruning and knowledge distillation on the multi-channel speech separation model after training to achieve lightweighting of the model.
[0155] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0156] In addition, the in-car speech processing method provided in this application, based on the preprocessing of the microphone linear array, uses a multi-channel speech separation network to further learn the energy differences of audio signals in different channels and different sound zones, and more accurately separate the speech signals generated by each sound zone.
[0157] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0158] S51 : Acquire a first reference audio signal, a second reference audio signal, a first reference noise signal, and a second reference noise signal.
[0159] Specifically, the speech processing apparatus obtains a first reference audio signal, a second reference audio signal, a first reference noise signal, and a second reference noise signal from a relevant audio data set.
[0160] Illustratively, the audio dataset may be DNS-Challenge2, or other audio datasets, which are not limited here.
[0161] S52 , determining a room impulse response based on the size of the vehicle, the position of the microphone, and the position of the passengers.
[0162] Specifically, the speech processing device calculates the room impulse response (RIR) based on parameters such as the actual interior size of the vehicle, the installation position of the microphone inside the vehicle, and the positions of the driver and passengers.
[0163] S53 , calculating a convolution result of the first reference audio signal and the room impulse response to obtain a first initial audio signal; and calculating a convolution result of the second reference audio signal and the room impulse response to obtain a second initial audio signal.
[0164] Specifically, the speech processing device calculates a convolution result of a first reference audio signal and a room impulse response to obtain a first initial audio signal. Furthermore, the speech processing device calculates a convolution result of a second reference audio signal and a room impulse response to obtain a second initial audio signal.
[0165] It should be noted that due to the complex environment inside the vehicle, if the first reference audio signal and the second reference audio signal obtained from the dataset are directly used as the dataset of the multi-channel speech separation model, the trained model will have poor speech separation effect during the actual vehicle verification phase. For example, the separated speech will be severely distorted.
[0166] In this embodiment, the initial audio signal obtained by calculating the convolution result of the reference audio signal and the room impulse response can better simulate the clean signal collected in the vehicle, effectively improving the performance of the model.
[0167] S54 , calculating a convolution result of the first reference noise signal and the room impulse response to obtain a first noise signal; and calculating a convolution result of the second reference noise signal and the room impulse response to obtain a second noise signal.
[0168] Specifically, the speech processing device calculates a convolution result of a first reference noise signal and a room impulse response to obtain a first noise signal, and calculates a convolution result of a second reference noise signal and a room impulse response to obtain a second noise signal.
[0169] S55 , superimposing the first initial audio signal and the first noise signal to obtain a first audio signal; and superimposing the second initial audio signal and the second noise signal to obtain a second audio signal.
[0170] Specifically, the speech processing device superimposes a first initial audio signal and a first noise signal to obtain a first audio signal, and the speech processing device superimposes a second initial audio signal and a second noise signal to obtain a second audio signal.
[0171] Furthermore, the speech processing device uses the first audio signal and the second audio signal to establish a training set, a verification set, and a test set.
[0172] It should be noted that steps S51 to S55 can more effectively simulate the audio signal collected by the vehicle microphone in an actual vehicle.
[0173] S56 , in response to the target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal.
[0174] S57, performing a short-time Fourier transform on the first processed signal to obtain a first frequency domain signal; and performing a short-time Fourier transform on the second processed signal to obtain a second frequency domain signal.
[0175] S58: Input the first frequency domain signal and the second frequency domain signal into a multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal.
[0176] S59, calculating the product of the first frequency domain signal and the first frequency domain mask to obtain the separated first frequency domain signal; and calculating the product of the second frequency domain signal and the second frequency domain mask to obtain the separated second frequency domain signal.
[0177] S60, performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain a first separated signal; and performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain a second separated signal.
[0178] S61 , determining a first loss using a first difference between the first separated signal and the first original audio signal; and determining a second loss using a difference between the second separated signal and the second original audio signal.
[0179] S62, determining a total loss using the first loss and the second loss.
[0180] S63, using the total loss, to adjust parameters of the multi-channel speech separation model.
[0181] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0182] In addition, the in-car speech processing method provided in this application, based on the preprocessing of the microphone linear array, uses a multi-channel speech separation network to further learn the energy differences of audio signals in different channels and different sound zones, and more accurately separate the speech signals generated by each sound zone.
[0183] In this embodiment, the audio data of the data set is further processed using the room impulse response so that the processed audio data better simulates the actual environment inside the car, so that the trained multi-channel speech separation model can produce audio with less distortion in actual applications.
[0184] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0185] S71: Acquire a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0186] Among them, at least a first microphone and a second microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.
[0187] The first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range.
[0188] S72 , in response to the target distance between the first microphone and the second microphone being less than a preset threshold, obtaining a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range.
[0189] Specifically, in response to a target distance between the first microphone and the second microphone being less than a preset threshold, the speech processing device determines that the first microphone and the second microphone constitute a linear array microphone. Furthermore, the speech processing device determines a first suppression angle θ2 of the first microphone relative to the second sound range, and a second suppression angle θ1 of the second microphone relative to the first sound range, based on a large amount of pre-acquired sound pickup data from the first and second microphones collected from real vehicles, as well as theoretical modeling results.
[0190] S73 , performing differential beam processing on the first audio signal using the first suppression angle to obtain a first processed signal; and performing differential beam processing on the second audio signal using the second suppression angle to obtain a second processed signal.
[0191] Specifically, the speech processing device performs differential beam processing on the first audio signal using the first suppression angle θ2 to obtain a first processed signal.
[0192] Furthermore, the speech processing device performs differential beam processing on the second audio signal using the second suppression angle θ1 to obtain a second processed signal.
[0193] S74: Input the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone.
[0194] The first separated signal does not include audio signals generated by a sound zone not corresponding to the first microphone, and the second separated signal does not include audio signals generated by a sound zone not corresponding to the second microphone.
[0195] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0196] In addition, the in-car speech processing method provided in this application uses an on-board microphone array to obtain in-car audio signals, determines the sound source angles on both sides of the array, performs spatial filtering through a differential beam algorithm, suppresses the audio signals generated by the sound zones corresponding to non-microphones themselves, and increases the energy difference of the output speech in the corresponding sound zones.
[0197] Another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0198] S81: Acquire a first audio signal collected by a first microphone and a second audio signal collected by a second microphone.
[0199] Among them, at least a first microphone and a second microphone are arranged inside the vehicle, and each microphone corresponds to a sound zone.
[0200] The first microphone corresponds to the first sound zone, and the second microphone corresponds to the second sound zone.
[0201] S82 , in response to the target distance between the first microphone and the second microphone being less than a preset threshold, obtaining a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range.
[0202] S83: Determine a first filter coefficient according to the first suppression angle, the target distance, and the signal angular frequency.
[0203] In some embodiments, the speech processing device calculates a ratio of the target distance to the speed of sound to obtain a target ratio.
[0204] Furthermore, the speech processing device determines a first filter coefficient according to the first suppression angle, the target ratio, and the signal angular frequency.
[0205] Among them, the first filter coefficient satisfies the following relationship:
[0206]
[0207] Wherein, C1 is the first filter coefficient.
[0208] S84: Perform differential beam processing on the first audio signal using the first filter coefficient to obtain a first processed signal.
[0209] Specifically, the speech processing device performs differential beam processing on the first audio signal using the first filter coefficient to obtain a first processed signal.
[0210] S85: Determine a second filter coefficient according to the second suppression angle, the target distance, and the signal angular frequency.
[0211] In some embodiments, the speech processing device determines the second filter coefficient based on the second suppression angle, the target coefficient, and the signal angular frequency.
[0212] The second filter coefficient satisfies the following relationship:
[0213]
[0214] Wherein, C2 is the first filter coefficient.
[0215] S86: Perform differential beam processing on the second audio signal using the second filter coefficient to obtain a second processed signal.
[0216] Specifically, the speech processing device performs differential beam processing on the second audio signal using the second filter coefficient to obtain a second processed signal.
[0217] S87, input the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone.
[0218] The first separated signal does not include audio signals generated by a sound zone not corresponding to the first microphone, and the second separated signal does not include audio signals generated by a sound zone not corresponding to the second microphone.
[0219] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0220] In addition, the in-car speech processing method provided in this application uses an on-board microphone array to obtain in-car audio signals, determines the sound source angles on both sides of the array, performs spatial filtering through a differential beam algorithm, suppresses the audio signals generated by the sound zone where the microphone itself is not located, and increases the energy difference of the output speech in the corresponding sound zone.
[0221] See also Figure 3 , Figure 3 This is a flow chart of another embodiment of the method for processing in-vehicle voice provided by the present application.
[0222] like Figure 3As shown, another embodiment of the in-vehicle speech processing method provided by the present application may specifically include the following steps:
[0223] S101, obtaining microphone input signals corresponding to different sound zones.
[0224] Specifically, the voice processing device obtains a microphone input signal corresponding to each vehicle-mounted microphone, wherein each vehicle-mounted microphone corresponds to a sound zone.
[0225] S102: Is there a microphone linear array?
[0226] Specifically, the voice processing device determines whether there is a microphone linear array according to the model of the vehicle.
[0227] If the judgment result is yes, jump to S103. If the judgment result is no, jump to S104.
[0228] In this embodiment, the microphone layout inside the vehicle can be as follows: Figure 2A-2E As shown. Among them, Figure 2A and Figure 2B is the microphone layout of the linear microphone array, Figure 2C The microphone layout for distributed microphones, Figure 2D and Figure 2E It is a hybrid microphone layout, that is, it contains both linear array microphones and distributed microphones.
[0229] S103 , performing differential beam processing on microphone input signals corresponding to the microphone linear array to obtain multiple processed signals.
[0230] Specifically, if a linear microphone array is present, the speech processing device applies a differential beamforming algorithm to the speech signal of the corresponding channel to enhance the target direction and suppress interference from other directions.
[0231] Among them, the differential beamforming algorithm is a sound source localization and sound separation technology based on a microphone array, which reflects the differences in spatial sound pressure.
[0232] Differential microphone arrays are different from additive microphone arrays, such as conventional beamforming. Additive microphone arrays determine the shape of their main lobe by setting the main lobe itself. For vehicle-mounted microphone arrays, the microphone array spacing is small and the beam main lobe width is large, resulting in poor directivity, insignificant array gain, weak enhancement effect in the target direction, and poor suppression effect in the interference direction. Differential microphone arrays, on the other hand, determine the shape of their main lobe by setting the null point direction. By placing a null point in the interference direction, the speech energy ratio in the target direction and the interference direction is increased. The sensitivity of the array to different incident directions is represented by the beam pattern, and the specific formula is as follows:
[0233] B[h(ω),θ]=dG (ω,cosθ)h(ω)
[0234] Where d(ω, cosθ) is the array's steering vector, ω is the signal angular frequency, θ is the incident direction, h(ω) is the output weight of each microphone, and H is the conjugate transpose. For example, with dual microphones with an end-fire direction of 0° and a suppression direction of 180°, the beam pattern calculation formula for differential beamforming is:
[0235]
[0236] h′(ω) is τ is the ratio of the microphone array spacing to the speed of sound.
[0237] In this embodiment, the differential beam algorithm is applied to the speech enhancement of the linear array microphone array of a real vehicle. For example, when the microphone linear array is arranged above the central control screen, the zero point is set to the angle of the driver's seat relative to the microphone array, thereby effectively suppressing the speech energy in the driver's direction, thereby highlighting the passenger's voice signal and enhancing the difference in speech energy between the driver and passenger.
[0238] Based on the above differential beamforming algorithm, S103 may specifically include the following steps:
[0239] S1031, collect a large amount of real-car linear microphone array sound pickup data.
[0240] S1032: Based on theoretical modeling and actual vehicle data, analyze the angles of the sound zones on both sides of the microphone relative to the linear array, that is, determine the suppression angles θ2 and θ1 on each side.
[0241] S1033, using the suppression angle, determine the microphone array filter coefficient -e jστcosθ , perform differential beam spatial filtering on multi-channel speech signals.
[0242] Distributed microphone arrays capture distinct energy differences in speech information across different vocal ranges. However, due to the smaller array spacing, linear array microphones pick up smaller energy and phase differences in speech signals. Using a differential beamforming algorithm to spatially filter the multi-channel speech signals from a linear array microphone array creates distinct energy differences in speech output across different vocal ranges. This facilitates the subsequent multi-channel speech separation model's accurate recognition of voices at different locations, improving the quality and accuracy of voice interaction pickup.
[0243] S104: Input the multiple processed signals and / or other microphone input signals into a multi-channel speech separation model to separate and obtain speech signals corresponding to different sound ranges.
[0244] Specifically, the speech processing device inputs the multi-channel signals (processed signals and / or other microphone input signals) obtained through the above steps into a multi-channel speech separation model, and the result output by the model is the speech information corresponding to different sound zones.
[0245] The multi-channel speech separation model uses speech signals captured by multiple microphones to distinguish and extract independent sound sources in mixed speech. In noisy environments, it can accurately separate and recognize speech and distinguish different speakers. It has good adaptability to different acoustic environments and noise conditions.
[0246] The structural diagram of an embodiment of the multi-channel speech separation model provided in this application is as follows Figure 4 shown.
[0247] like Figure 4 As shown, the multi-channel speech separation model consists of a sequentially connected encoder, a temporal convolutional network (TCN), and a decoder. The TCN can capture long-range dependencies in sequences, making it very useful in speech-related tasks. For linear microphone arrays and hybrid distributions, the differential beamforming algorithm increases the energy difference between signals in different sound zones of the corresponding channels, and the signal energy distribution is similar to the sound pickup results of distributed microphones. Therefore, for microphone layouts of different formations, combined with the differential beamforming algorithm, the multi-channel speech separation model can learn the energy differences between passenger speech signals in different channels and sound zones, and separate the speech signals in each sound zone.
[0248] In some embodiments, the model training and lightweighting process is as follows:
[0249] S1041: Calculate the room impulse response (RIR) based on the actual vehicle size, microphone positions, passenger positions, etc.
[0250] S1042 randomly selects DNS-Challenge2 single-channel clean speech, noise, and RIR for convolution to obtain multi-channel speech signals based on real-car scenarios as training sets, validation sets, and test sets.
[0251] S1043, if the microphone layout includes a linear microphone array, apply differential beamforming algorithm processing to the corresponding channel signal; if not, jump to S1044.
[0252] S1044, convert the multi-channel speech signal into the time-frequency domain, Figure 4The multi-channel separation model shown is trained for speech separation, and the model outputs masks corresponding to the sound zones. The microphone input is multiplied by the mask for each channel to obtain the speech signal for each sound zone. If a passenger is speaking, the corresponding passenger's speech is output; if no one is speaking, the sound zone is output as silence.
[0253] S1045 adjusts network parameters and quantizes the model to obtain a lightweight in-vehicle multi-channel speech separation model, enabling accurate recognition of human voices at different locations.
[0254] Based on the above differential beamforming and lightweight multi-channel speech separation model, the specific process of vehicle-mounted speech multi-zone separation is as follows:
[0255] S1111, each car microphone captures the voice signal in the car, including human voice, noise, echo, etc., and obtains multi-channel voice output s1...s n , n is the total number of channels.
[0256] S1112: Determine whether there is a microphone linear array in the vehicle microphone layout. If so, perform differential beam filtering on the corresponding channels. For example, if the current two channels are signals picked up by the microphone linear array, apply the differential beam algorithm to s1 and s2 to obtain x1 and x2. If there is no linear array, proceed directly to S1113.
[0257] S1113, convert x1x2…s n or s1s2…s n The multi-channel data is converted to the time-frequency domain and input into a lightweight multi-channel speech separation model.
[0258] S1114, obtain the separated voice outputs y1…y n The output data can be used for subsequent voice interaction tasks such as speech recognition and voice wake-up.
[0259] In the above scheme, in response to the target distance between the first microphone and the second microphone being less than a preset threshold, the first audio signal and the second audio signal are respectively subjected to beamforming processing, and the first processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the first microphone and suppressing the audio signals generated from other sound zones; and the second processed signal is obtained by enhancing the audio signal generated by the sound zone corresponding to the second microphone and suppressing the sound generated from other sound zones. The first processed signal and the second processed signal are respectively input into the multi-channel speech separation model to obtain the first separated signal and the second separated signal, so as to more effectively filter out the audio signals that do not correspond to the own sound zone. Furthermore, the separated signals obtained by the in-vehicle speech processing method provided in the present application can more accurately perform voice interaction tasks such as speech recognition and voice wake-up, thereby improving the user experience.
[0260] Furthermore, the in-vehicle speech processing method provided in this application utilizes an onboard microphone array to capture in-vehicle audio signals, determine the sound source angles on either side of the array, and perform spatial filtering using a differential beamforming algorithm. This method suppresses audio signals generated by non-native microphones in the corresponding sound zones, thereby increasing the energy variance of the output speech in those zones. Based on preprocessing with the microphone array, a multi-channel speech separation network is used to further learn the energy variance of audio signals in different channels and sound zones, more accurately separating the speech signals generated by each sound zone.
[0261] Please continue to see Figure 5 , Figure 5 The terminal device 500 of the embodiment of the present application includes a processor 51 and a memory 52 .
[0262] The processor 51 and the memory 52 are connected to the bus. The memory 52 stores program data. The processor 51 is used to execute the program data to implement the in-vehicle speech processing method described in the above embodiment.
[0263] In the embodiment of the present application, the processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip having signal processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 51 may be any conventional processor.
[0264] This application also provides a computer storage medium, please continue to refer to Figure 6 , Figure 6 It is a structural diagram of an embodiment of a computer storage medium provided in the present application. The computer storage medium 600 stores program data 61. When the program data 61 is executed by the processor, it is used to implement the in-vehicle voice processing method of the above embodiment.
[0265] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0266] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Equivalent structures or equivalent process changes made by utilizing the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for processing in-vehicle speech, characterized in that: At least a first microphone and a second microphone are provided inside the vehicle, each microphone corresponding to a sound zone, and the method includes: Acquire a first audio signal collected by the first microphone and a second audio signal collected by the second microphone; In response to a target distance between the first microphone and the second microphone being less than a preset threshold, performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal; Inputting the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone; wherein the first separated signal does not include audio signals generated in a sound range other than that corresponding to the first microphone, and the second separated signal does not include audio signals generated in a sound range other than that corresponding to the second microphone; The step of inputting the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone includes: Performing a short-time Fourier transform on the first processed signal to obtain a first frequency domain signal; and performing a short-time Fourier transform on the second processed signal to obtain a second frequency domain signal; Inputting the first frequency domain signal and the second frequency domain signal into the multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal; Calculating the product of the first frequency domain signal and the first frequency domain mask to obtain a separated first frequency domain signal; and calculating the product of the second frequency domain signal and the second frequency domain mask to obtain a separated second frequency domain signal; Performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain the first separated signal; and performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain the second separated signal.
2. The method according to claim 1, characterized in that A third microphone is further provided inside the vehicle, and the distance between the third microphone and the first microphone and the second microphone is respectively greater than the preset threshold; The step of inputting the first processed signal and the second processed signal into a multi-channel speech separation model to obtain a first separated signal corresponding to the first microphone and a second separated signal corresponding to the second microphone includes: The first processed signal, the second processed signal, and the third audio signal collected by the third microphone are input into the multi-channel speech separation model to obtain the speech processing signal generated by each sound zone, and the first separated signal, the second separated signal and the third separated signal are obtained, wherein the third separated signal does not include the audio signal generated by the sound zone not corresponding to the third microphone.
3. The method according to claim 1, characterized in that The multi-channel speech separation model includes an encoder, a time domain convolutional network and a decoder connected in sequence. Inputting the first frequency domain signal and the second frequency domain signal into the multi-channel speech separation model to obtain a first frequency domain mask corresponding to the first frequency domain signal and a second frequency domain mask corresponding to the second frequency domain signal includes: Inputting the first frequency domain signal to the encoder to obtain a first encoding result, and inputting the second frequency domain signal to the encoder to obtain a second encoding result; Inputting the first encoding result into the time-domain convolutional network to obtain a first convolution result; and inputting the second encoding result into the time-domain convolutional network to obtain a second convolution result; The first convolution result is input to the decoder to obtain the first frequency domain mask; and the second convolution result is input to the decoder to obtain the second frequency domain mask.
4. The method according to claim 1, wherein The first audio signal is obtained by superimposing a first initial audio signal and a first noise signal, and the second audio signal is obtained by superimposing a second initial audio signal and a second noise signal; performing an inverse short-time Fourier transform on the separated first frequency domain signal to obtain the first separated signal; And, after the step of performing an inverse short-time Fourier transform on the separated second frequency domain signal to obtain the second separated signal, the method further comprises: determining a first loss using a first difference between the first separated signal and the first original audio signal; and determining a second loss using a difference between the second separated signal and the second original audio signal; determining a total loss using the first loss and the second loss; The total loss is used to adjust parameters of the multi-channel speech separation model.
5. The method according to claim 4, characterized in that The acquiring of the first audio signal collected by the first microphone and the second audio signal collected by the second microphone includes: Acquire a first reference audio signal, a second reference audio signal, a first reference noise signal, and a second reference noise signal; determining a room impulse response based on the dimensions of the vehicle, the positions of the microphones, and the positions of the passengers; Calculating a convolution result of the first reference audio signal and the room impulse response to obtain the first initial audio signal; and calculating a convolution result of the second reference audio signal and the room impulse response to obtain the second initial audio signal; Calculating a convolution result of the first reference noise signal and the room impulse response to obtain the first noise signal; and calculating a convolution result of the second reference noise signal and the room impulse response to obtain the second noise signal; The first initial audio signal and the first noise signal are superimposed to obtain the first audio signal; and the second initial audio signal and the second noise signal are superimposed to obtain the second audio signal.
6. The method according to claim 1, characterized in that The first microphone corresponds to the first sound range, and the second microphone corresponds to the second sound range; The performing differential beam processing on the first audio signal to obtain a first processed signal, and performing differential beam processing on the second audio signal to obtain a second processed signal, includes: Acquire a first suppression angle of the first microphone relative to the second sound range, and a second suppression angle of the second microphone relative to the first sound range; The first audio signal is differentially beam processed using the first suppression angle to obtain the first processed signal; and the second audio signal is differentially beam processed using the second suppression angle to obtain the second processed signal.
7. The method according to claim 6, characterized in that The performing differential beam processing on the first audio signal by using the first suppression angle to obtain the first processed signal includes: determining a first filter coefficient according to the first suppression angle, the target distance, and the signal angular frequency; performing differential beam processing on the first audio signal using the first filter coefficient to obtain the first processed signal; The performing differential beam processing on the second audio signal by using the second suppression angle to obtain the second processed signal includes: determining a second filter coefficient according to the second suppression angle, the target distance, and the signal angular frequency; Perform differential beam processing on the second audio signal using the second filter coefficient to obtain the second processed signal.
8. A terminal device, characterized in that: The terminal device includes a processor and a memory connected to the processor, wherein: The memory stores program instructions; The processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The storage medium stores program instructions, and when the program instructions are executed, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Audio data processing method and device
CN113096679A
Vehicle-mounted multi-sound-area voice interaction method and device, electronic equipment and storage medium
CN115881125A