Speech recognition methods, devices, storage media and electronic devices
By applying echo cancellation, beamforming, and signal enhancement technologies to the microphone extension arrays of in-vehicle and mobile terminals, combined with a speech recognition model, the problem of decreased voice interaction accuracy caused by in-vehicle noise interference was solved, achieving higher speech recognition accuracy and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUBEI XINGJI MEIZU TECH CO LTD
- Filing Date
- 2022-11-22
- Publication Date
- 2026-07-31
AI Technical Summary
In-vehicle noise interference reduces the accuracy of voice interaction and user experience, and existing technologies struggle to effectively eliminate noise and improve voice recognition accuracy.
By employing a microphone extension array and combining echo cancellation, beamforming, and signal enhancement technologies with a speech recognition model, the system utilizes the microphone resources of both in-vehicle and mobile terminals for speech acquisition and processing. It identifies the target microphone, enhances the signal, and reduces noise, ultimately achieving accurate speech recognition.
Without increasing vehicle hardware costs, it improves the accuracy of voice wake-up and recognition, enhances the user's voice interaction experience, reduces noise interference, and improves the response speed of the in-vehicle terminal.
Smart Images

Figure CN115862598B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and more specifically, to a speech recognition method, device, storage medium, and electronic device. Background Technology
[0002] With the advancement of voice interaction technology, users are increasingly using voice commands to operate in-vehicle electronic systems.
[0003] While a car is in motion, the noise inside the vehicle can include engine noise, air conditioning noise, tire rubbing noise, air friction noise, passengers talking, and sounds from electronic devices used for entertainment. These sounds interfere with the user's voice interaction with the vehicle's electronic systems. Summary of the Invention
[0004] This application provides a voice recognition method for use in an in-vehicle terminal, including:
[0005] First voice was acquired using a microphone extension array;
[0006] It is determined that the speech recognition result of the first speech contains a preset wake word;
[0007] Based on the sound propagation parameters of the first speech collected by each microphone in the microphone extension array, the target microphone is determined;
[0008] Signal acquisition is performed based on the microphone extension array, and the signal acquired by the target microphone is enhanced to determine the second speech.
[0009] Determine the speech recognition result of the second speech;
[0010] The microphone extension array includes a first microphone array in the vehicle terminal and one or more microphones in a mobile terminal connected to the vehicle terminal.
[0011] According to the speech recognition method provided in this application, the microphone extension array is determined based on the following steps:
[0012] It is confirmed that the mobile terminal and the vehicle terminal are successfully connected;
[0013] Send a microphone function query request to the mobile terminal;
[0014] Receive the microphone function configuration response sent by the mobile terminal based on the microphone function query request;
[0015] Based on the microphone function configuration response, it is determined whether the mobile terminal supports microphone expansion;
[0016] The first microphone array in the vehicle terminal is extended based on one or more microphones in the mobile terminal to obtain the extended microphone array.
[0017] According to the speech recognition method provided in this application, the step of acquiring the first speech based on a microphone extended array includes:
[0018] Echo cancellation is performed on the first audio signal acquired by each microphone in the microphone extension array;
[0019] Beamforming is applied to the first audio signal after echo cancellation to obtain the first speech collected by each microphone.
[0020] According to the speech recognition method provided in this application, the echo cancellation of the first audio signal collected by each microphone in the microphone extension array includes:
[0021] Determine the acquisition terminal corresponding to any microphone in the microphone extension array;
[0022] When the acquisition terminal is a vehicle-mounted terminal, echo cancellation is performed on the first audio signal acquired by any of the microphones;
[0023] When the acquisition terminal is a mobile terminal, an echo cancellation command is sent to the mobile terminal; the echo cancellation command is used to control the mobile terminal to cancel the echo of the first audio signal acquired by any of the microphones.
[0024] According to the speech recognition method provided in this application, the step of determining the target microphone by collecting sound propagation parameters of the first speech from each microphone in the microphone extension array includes:
[0025] Based on the sound propagation parameters of the first speech collected by each microphone, the arrival time of the first speech from the sound source to each microphone is determined.
[0026] Based on the arrival time of each microphone, the distance between each microphone and the sound source is determined;
[0027] The microphone that is closest to the sound source is identified as the target microphone.
[0028] According to the speech recognition method provided in this application, the step of acquiring signals based on the microphone extension array and enhancing the signals acquired by the target microphone to determine the second speech includes:
[0029] Echo cancellation is performed on the second audio signal acquired by the target microphone;
[0030] Based on the position of the target microphone in the microphone extension array, the second audio signal collected by the target microphone is subjected to beamforming processing to obtain the second speech collected by the target microphone.
[0031] According to the speech recognition method provided in this application, after obtaining the second speech collected by the target microphone, the process includes:
[0032] The second voice recorded by the target microphone is subjected to noise reduction processing.
[0033] According to the speech recognition method provided in this application, determining the speech recognition result of the second speech includes:
[0034] The second speech is input into the speech recognition model to obtain the speech recognition result output by the speech recognition model;
[0035] The speech recognition model includes a feature extraction layer, a silence detection layer, and a speech recognition layer; the silence detection layer and the speech recognition layer are respectively connected to the feature extraction layer.
[0036] The feature extraction layer is used to divide the second speech into multiple speech frames and extract the acoustic recognition features of each speech frame; the silence detection layer is used to determine the speech frame to be recognized in the second speech based on the acoustic recognition features of each speech frame; the speech recognition layer is used to determine the speech recognition result of the second speech based on the acoustic recognition features of the speech frame to be recognized.
[0037] According to the speech recognition method provided in this application, the speech recognition model is deployed on the vehicle terminal or the cloud server corresponding to the vehicle terminal.
[0038] This application provides a voice recognition device, including:
[0039] The first acquisition module is used to acquire the first voice based on the microphone extension array;
[0040] A wake-up module is used to determine that the speech recognition result of the first speech contains a preset wake-up word;
[0041] The determination module is used to determine the target microphone based on the sound propagation parameters corresponding to the first speech collected by each microphone in the microphone extension array;
[0042] The second acquisition module is used to acquire signals based on the microphone extension array and enhance the signals acquired by the target microphone to determine the second speech.
[0043] A recognition module is used to determine the speech recognition result of the second speech;
[0044] The microphone extension array includes a first microphone array in the vehicle terminal and one or more microphones in a mobile terminal connected to the vehicle terminal.
[0045] This application provides a computer-readable storage medium comprising a stored program, wherein the program executes the speech recognition method when it runs.
[0046] This application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the speech recognition method through the computer program. Attached Figure Description
[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0048] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is one of the flowcharts illustrating a speech recognition method provided in an embodiment of this application;
[0050] Figure 2 This is a second schematic flowchart of a speech recognition method provided in one embodiment of this application;
[0051] Figure 3 This is a timing diagram of a speech recognition method provided in one embodiment of this application;
[0052] Figure 4 This is a software module diagram provided in one embodiment of this application;
[0053] Figure 5 This is a schematic diagram of the structure of a speech recognition device provided in one embodiment of this application;
[0054] Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0055] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0056] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0057] The speech recognition method provided in this application is applicable to terminal devices equipped with a human-computer voice interaction system. A human-computer voice interaction system is a system that uses voice as a medium to interact with users.
[0058] Terminal equipment includes various handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem, such as mobile phones, tablets, desktop laptops, and smart devices that can run applications, including the central control console of a smart car. Specifically, it can refer to User Equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device.
[0059] Terminal devices can also be satellite phones, cellular phones, smartphones, wireless data cards, wireless modems, machine-type communication devices, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, Personal Digital Assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices or wearable devices, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, terminal devices in 5G networks or future communication networks, etc.
[0060] The terminal device can be powered by a battery or attached to and powered by the power system of a vehicle or vessel. The vehicle or vessel's power system can also charge the terminal device's battery to extend its communication time.
[0061] Figure 1 This is one of the flowcharts illustrating a speech recognition method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes steps 110, 120, 130, 140, and 150. These method steps are merely one possible implementation of this application.
[0062] Step 110: Acquire first speech based on a microphone extension array. The microphone extension array includes a first microphone array in the vehicle-mounted terminal and one or more microphones in a mobile terminal connected to the vehicle-mounted terminal.
[0063] Specifically, the speech recognition method in this application embodiment is implemented by a speech recognition device. The speech recognition device can be a separate hardware module in the vehicle terminal, or it can be a software program running in the vehicle terminal.
[0064] A microphone array is an arrangement of multiple microphones, that is, a system composed of a certain number of microphones, used to sample and process the spatial characteristics of a sound field.
[0065] The in-vehicle terminal is a terminal device installed inside a vehicle, such as an intelligent vehicle control system. The first microphone array is a microphone array set within the in-vehicle terminal, which can contain 2, 4, 6, or 8 microphones, etc. One microphone can be installed near each seat inside the vehicle to form the first microphone array.
[0066] A mobile terminal is a smart device carried by the user, including smartphones, tablets, and smart wearable devices such as wireless headphones, smartwatches, and AR glasses. Mobile terminals typically contain one or more microphones. For example, smartphones generally have one microphone located at the bottom.
[0067] When a mobile terminal is connected to an in-vehicle terminal, one or more microphones from the mobile terminal and the first microphone array from the in-vehicle terminal can be combined to form a microphone extension array. For example, the first microphone array in the in-vehicle terminal includes four microphones, and the mobile terminal includes two microphones. After a user brings the mobile terminal into the vehicle, the mobile terminal establishes a communication connection with the in-vehicle terminal, and the mobile terminal transmits the audio signals acquired by the two microphones to the in-vehicle terminal in real time via communication. For the in-vehicle terminal, the acquired audio signals come not only from the first microphone array but also from the microphones in the mobile terminal, thus extending the first microphone array in the in-vehicle terminal by one or more microphones from the mobile terminal, resulting in a microphone extension array.
[0068] When a user interacts with the in-vehicle terminal via voice, they can emit an initial voice message. The in-vehicle terminal then captures this initial voice message using a microphone array. For example, it captures the initial voice messages from each microphone individually and then fuses these messages.
[0069] Step 120: Determine that the speech recognition result of the first speech contains a preset wake word.
[0070] Specifically, the wake word is used to trigger the vehicle terminal to end the low power state or sleep state, so that the vehicle terminal can continue to acquire the voice issued by the user and respond to the voice to execute the corresponding control operation.
[0071] A preset wake-up word can be set in the vehicle terminal to determine whether the first voice message reflects the user's intention to wake up the vehicle terminal.
[0072] Preset wake-up words can be set as combinations of multiple words to closely match the user's expression habits and improve the wake-up accuracy of the in-vehicle terminal. For example, if the wake-up word for the voice assistant on a mobile phone is "Xiao Meng," multiple preset wake-up words can be set such as "Xiao Meng Xiao Meng," "Xiao Meng Tongxue," and "Hello Xiao Meng." The in-vehicle terminal can call a speech recognition model to perform speech recognition on the user's first speech, obtaining the speech recognition result. The in-vehicle terminal can then perform semantic similarity matching between the speech recognition result of the first speech and the preset wake-up words. If they match, it means the first speech contains the preset wake-up word; if they do not match, it means the first speech does not contain the preset wake-up word.
[0073] Sound propagation parameters are parameters that describe the characteristics of sound during propagation, such as propagation speed, time of arrival, direction of arrival, loudness, and phase. Because the distribution of microphones in a microphone array varies, the sound propagation parameters determined by each microphone when acquiring the first speech sound also differ.
[0074] Step 130: Collect the sound propagation parameters of the first speech based on each microphone in the microphone extension array to determine the target microphone.
[0075] Specifically, if the speech recognition result of the first speech contains a preset wake-up word, the sound propagation parameters of the first speech can be collected by each microphone in the microphone extension array to determine the target microphone. The target microphone is the microphone that can achieve the best sound acquisition effect. For example, the microphone with the loudest sound acquisition of the first speech can be used as the target microphone, or the microphone that first acquired the first speech can be used as the target microphone, or the microphone with the higher signal-to-noise ratio of the acquired first speech can be used as the target microphone, etc.
[0076] Step 140: Acquire signals based on the microphone extension array and enhance the signals acquired by the target microphone to determine the second speech.
[0077] Specifically, the first and second voice messages were sent by the same user.
[0078] After identifying the target microphone, signal enhancement can be performed on the target microphone to obtain more accurate second speech.
[0079] Signal enhancement can be achieved by calculating the signal delay time between the target microphone and other microphones based on the sound propagation parameters of the first speech sample collected from each microphone. When collecting the second speech sample, the signals from each microphone are aligned and then fused based on the signal delay time.
[0080] The simplest way to enhance speech using a microphone array is to first locate the target direction (the microphone at 0 degrees to the front), calculate the delays of the other microphones relative to the microphone at the front, and then sum the delays of all microphone signals to achieve the enhancement. In practical applications, the sound signal source and any microphone are not necessarily at 0 degrees to the front, and the timing and phase of the signals received by all microphones are inconsistent. Therefore, beamforming can be used to enhance the signals of a single microphone or all microphones.
[0081] Step 150: Determine the speech recognition result of the second speech.
[0082] Specifically, the collected second speech is subjected to speech recognition to obtain the speech recognition result of the second speech.
[0083] The second voice recognition result can be "Start navigation" or "Start recording," etc. Based on the voice recognition result, the in-vehicle terminal automatically calls the application installed on the in-vehicle terminal to execute the corresponding navigation or recording function.
[0084] The speech recognition method provided in this application involves: acquiring first speech using a microphone extension array; determining that the speech recognition result of the first speech includes a preset wake-up word; identifying a target microphone based on the sound propagation parameters of the first speech acquired from each microphone in the microphone extension array; enhancing the signal of the target microphone and acquiring second speech; and determining the speech recognition result of the second speech. Since the microphone extension array includes not only the first microphone array in the vehicle terminal but also one or more microphones in the mobile terminal, it increases the number of microphones used for speech acquisition without increasing vehicle hardware costs. Simultaneously, by enhancing the signal of the microphone array based on the acquisition result of the wake-up voice, the vehicle terminal can more accurately perceive the location of the user's voice, acquire more accurate user speech, improve the accuracy of voice wake-up and speech recognition, and enhance the user's voice interaction experience.
[0085] It should be noted that each implementation method of this application can be freely combined, rearranged, or executed individually, and does not need to rely on or depend on a fixed execution order.
[0086] In some embodiments, the microphone extension array is determined based on the following steps:
[0087] Confirm that the mobile terminal and the vehicle terminal are successfully connected;
[0088] Send a microphone function query request to the mobile terminal;
[0089] Receive the microphone function configuration response sent by the mobile terminal based on the microphone function query request;
[0090] The response based on microphone function configuration determines whether the mobile terminal supports microphone expansion;
[0091] The first microphone array in the vehicle terminal is extended based on one or more microphones in the mobile terminal to obtain a microphone extension array.
[0092] Specifically, after a user brings a mobile terminal into the vehicle, the mobile terminal and the vehicle terminal can establish a communication connection, for example, through Bluetooth or WiFi.
[0093] The vehicle-mounted terminal can send a connection request to the mobile terminal. The fields in the connection request may include a request to establish a connection and a session identifier. The request to establish a connection is used to indicate a request to establish a connection to the mobile terminal, and the session identifier is used to mark the current session between the vehicle-mounted terminal and the mobile terminal. For example, the format of the connection request is shown in Table 1.
[0094] Table 1 Connection Request Format
[0095]
[0096]
[0097] The mobile terminal can send a connection response to the vehicle terminal based on the connection request. Fields in the connection response may include a session identifier, a terminal identifier, a terminal device type, a terminal operating system type, a terminal registration username, and a terminal registration user identifier. The terminal identifier is used to identify the mobile terminal; the terminal device type indicates the type of the mobile terminal; the terminal operating system type indicates the type of operating system running on the mobile terminal; the terminal registration username indicates the name of the user using the mobile terminal; and the terminal registration user identifier identifies the user.
[0098] For example, the format of the connection response is shown in Table 2.
[0099] Table 2 Connection Response Format
[0100]
[0101] Once the mobile terminal and the vehicle-mounted terminal are successfully connected, the vehicle-mounted terminal sends a microphone function query request to the mobile terminal. This request is used to retrieve the functions supported by the microphone on the mobile terminal. Fields in the microphone function query request may include a session identifier and a protocol type check. The protocol type check is specifically used to retrieve the functions supported by the microphone on the mobile terminal.
[0102] For example, the format of a microphone function query request is shown in Table 3.
[0103] Table 3 Microphone Function Query Request Format
[0104]
[0105]
[0106] The mobile terminal sends a microphone function configuration response to the vehicle terminal based on the microphone function query request. The microphone function configuration response indicates the functions supported by the mobile terminal. Fields in the response may include a session identifier, a function list, and the number of microphone channels. The function list represents all functions supported by the mobile terminal, and the number of microphone channels indicates the number of extended microphones supported by the mobile terminal.
[0107] For example, the format of the microphone function configuration response is shown in Table 4.
[0108] Table 4 Microphone Function Configuration Response Format
[0109]
[0110] After determining whether the mobile terminal supports microphone extension, the vehicle-mounted terminal can also send a confirmation message to the mobile terminal. The confirmation message includes fields such as a session identifier and an extension result confirmation. The extension result confirmation indicates whether the microphone extension was successful.
[0111] For example, the confirmation message is shown in Table 5.
[0112] Table 5 Confirmation Message Format
[0113] Fields type describe session_id String Session ID result boolean 0---fail, 1---success
[0114] If the vehicle terminal determines that the mobile terminal supports microphone expansion, one or more microphones in the mobile terminal are used to expand the first microphone array in the vehicle terminal to obtain a microphone expansion array.
[0115] The voice recognition method provided in this application determines the microphone expansion array by automatically querying the mobile terminal from the vehicle terminal, without requiring user operation, thus improving the user's voice interaction experience.
[0116] In some embodiments, step 110 includes:
[0117] Echo cancellation is performed on the first audio signal acquired by each microphone in the microphone extension array;
[0118] Beamforming is applied to the first audio signal after echo cancellation to obtain the first speech collected by each microphone.
[0119] Specifically, during the voice acquisition process, the signal formed after the sound played by the speaker of the vehicle terminal is picked up by the microphone is called the echo. The purpose of the microphone is to pick up the user's voice, and the echo needs to be eliminated, otherwise it will interfere with the acquisition of human voice and speech recognition by each microphone.
[0120] A common echo cancellation method involves estimating the echo delay time based on the signal played from the speaker and the signal captured by the microphone. A Kalman filter algorithm can be used for this delay estimation. Based on the obtained delay time, the echo signal is aligned with the recorded signal from the microphone, and the echo signal is subtracted from the recording data. For example, a filter can be designed whose output signal has the same waveform as the echo but opposite phase; this output signal is then added to the recording data to cancel out the echo signal. The NLMS (Normalized Least Mean Square) algorithm can be used as a filter algorithm.
[0121] The first audio signal is the raw signal captured by each microphone when it picks up the first speech. Echo cancellation can be performed on the first audio signal captured by each microphone in the microphone array. Then, beamforming processing is applied to the echo-cancelled first audio signal to obtain the first speech captured by each microphone.
[0122] Beamforming primarily utilizes the spatial characteristics of signals acquired by multiple microphones to enhance signals in the target direction and suppress signals in non-target directions, thereby improving the signal-to-noise ratio. Beamforming processing can employ algorithms such as the General Side-lobe Canceller (GSC).
[0123] The speech recognition method provided in this application performs echo cancellation and beamforming on the first audio signals collected by each microphone, thereby improving the accuracy of voice wake-up.
[0124] In some embodiments, echo cancellation is performed on the first audio signal acquired by each microphone in the microphone extension array, including:
[0125] Determine the acquisition terminal corresponding to any microphone in the microphone extension array;
[0126] When the acquisition terminal is a vehicle-mounted terminal, echo cancellation is performed on the first audio signal acquired by any microphone.
[0127] When the acquisition terminal is a mobile terminal, an echo cancellation command is sent to the mobile terminal; the echo cancellation command is used to control the mobile terminal to cancel the echo of the first audio signal acquired by any microphone.
[0128] Specifically, when performing echo cancellation on the first audio signals collected by each microphone, either an in-vehicle terminal or a mobile terminal can be selected for echo cancellation. An echo cancellation algorithm or an echo cancellation module can be set up in either the in-vehicle terminal or the mobile terminal to perform echo cancellation on the first audio signal.
[0129] For any given microphone, the corresponding acquisition terminal can be determined first. If the acquisition terminal is an in-vehicle terminal, the in-vehicle terminal can perform echo cancellation on the first audio signal acquired by the microphone; if the acquisition terminal is a mobile terminal with one or more microphones, the in-vehicle terminal can send an echo cancellation command to the mobile terminal, and the mobile terminal can perform echo cancellation on the first audio signal acquired by the microphone.
[0130] The speech recognition method provided in this application reduces the computational load of the vehicle terminal, improves the response speed of the vehicle terminal, and enhances the user's voice interaction experience by selecting an appropriate terminal for echo cancellation of the first audio signals collected by microphones at different locations.
[0131] In some embodiments, step 130 includes:
[0132] Based on the sound propagation parameters of the first speech collected by each microphone, the arrival time of the first speech from the sound source to each microphone is determined.
[0133] Based on the arrival time of each microphone, the distance between each microphone and the sound source is determined;
[0134] The microphone closest to the sound source is identified as the target microphone.
[0135] Specifically, the sound source is the source that emits the first speech sound. Based on the sound propagation parameters of the first speech sound collected by each microphone, the arrival time of the first speech sound from the sound source to each microphone can be calculated. Based on the speed of sound propagation and the corresponding arrival time at each microphone, the distance between each microphone and the sound source can be determined. The microphone with the smallest distance from the sound source can be identified as the target microphone.
[0136] In this embodiment, the location coordinates of the sound source can also be calculated by using the TDOA (Time Difference of Arrival) algorithm, which calculates the time difference between the arrival of the signal at the microphone.
[0137] The speech recognition method provided in this application collects the sound propagation parameters of the first speech through each microphone to determine the target microphone. The calculation is simple and easy to execute.
[0138] In some embodiments, step 130 includes:
[0139] The spatial spectrum of the first speech is determined by collecting sound propagation parameters of the first speech from each microphone.
[0140] The target microphone is determined based on the spatial spectrum.
[0141] Specifically, the spatial spectrum of the first speech can be determined by collecting sound propagation parameters from each microphone. The spatial spectrum can be used to represent the energy distribution of the first speech signal in various spatial directions.
[0142] High-resolution spectral estimation techniques can be used to determine the direction angle and location of the sound source by solving the correlation matrix between microphones using the spatial spectrum of the first speech. This allows for the identification of the target microphone closest to the sound source. The location of the sound source can be determined using methods such as minimum variance spectral estimation and eigenvalue decomposition.
[0143] In some embodiments, step 140 includes:
[0144] Echo cancellation is performed on the second audio signal captured by the target microphone;
[0145] Based on the position of the target microphone in the microphone extension array, beamforming processing is performed on the second audio signal acquired by the target microphone to obtain the second speech acquired by the target microphone.
[0146] Specifically, the second audio signal is the original signal captured by each microphone when collecting the second speech. The echo cancellation of the second audio signal collected by the target microphone is similar to the method for echo cancellation of the first audio signal in the above embodiments, and will not be described again here.
[0147] Beamforming can be used to enhance the second audio signal, for example, by employing a generalized side-lobe canceller (GSC) algorithm or adding post-filtering.
[0148] Unlike the above embodiment where beamforming is performed on each microphone when acquiring the first audio signal, when acquiring the second audio signal, beamforming is performed only on the second audio signal acquired by the target microphone based on the position of the target microphone in the microphone extension array. That is, only the target microphone is signal-enhanced to obtain the second speech acquired by the target microphone.
[0149] In some embodiments, after obtaining the second speech captured by the target microphone, the process includes:
[0150] The second speech captured by the target microphone is subjected to noise reduction processing.
[0151] Specifically, the purpose of noise reduction is to reduce noise components, improve the signal-to-noise ratio, and facilitate subsequent speech recognition.
[0152] The noise reduction algorithm can be based on neural networks, such as the RNNoise algorithm, which can achieve fast noise reduction, improve the signal-to-noise ratio, and improve the recognition accuracy of the second speech.
[0153] In some embodiments, step 150 includes:
[0154] The second voice is input into the speech recognition model to obtain the speech recognition result output by the speech recognition model;
[0155] The speech recognition model includes a feature extraction layer, a silence detection layer, and a speech recognition layer; the silence detection layer and the speech recognition layer are connected to the feature extraction layer, respectively.
[0156] The feature extraction layer is used to divide the second speech into multiple speech frames and extract the acoustic recognition features of each speech frame; the silence detection layer is used to determine the speech frame to be recognized in the second speech based on the acoustic recognition features of each speech frame; and the speech recognition layer is used to determine the speech recognition result of the second speech based on the acoustic recognition features of the speech frame to be recognized.
[0157] Specifically, in actual voice interaction, the second voice emitted by the user can include both spoken and non-spoken parts. The non-spoken part can be silence or ambient sound. For example, if the user is not speaking for more than half the time, more than half of the collected second voice will be silent. Recognizing and processing the second voice containing silence wastes the computing resources of the in-vehicle terminal.
[0158] A speech recognition model can be established using a neural network model as the initial model to process the second speech and obtain the speech recognition result.
[0159] Considering that both silence detection and speech recognition, when implemented using neural network models, can be based on the analysis of the acoustic features of the second speech, the speech recognition model established in this embodiment can structurally include a feature extraction layer, a silence detection layer, and a speech recognition layer. The silence detection layer and the speech recognition layer are respectively connected to the feature extraction layer.
[0160] The feature extraction layer is used to divide the second speech into multiple speech frames and extract the acoustic recognition features of each speech frame. First, the feature extraction layer can divide the second speech into multiple speech frames. For example, the second speech can be divided into frames of 10ms each. Second, the feature extraction layer extracts the acoustic recognition features of each speech frame. These acoustic recognition features describe the physical quantities of the speech frame in terms of acoustic characteristics. For example, acoustic recognition features can be prosodic features, timbre features, and loudness features; they can also be time-domain features and frequency-domain features. Frequency-domain features can further include Mel-frequency cepstral coefficients (MFCC) and filter-bank features (FBANK).
[0161] The silence detection layer is used to determine the speech frames to be recognized in the second speech based on the acoustic recognition features output by the feature extraction layer. The speech frames to be recognized are those determined to contain the user's speech after silence detection is performed on each speech frame. By extracting the speech frames to be recognized, the useful parts (speech components) of the second speech can be extracted, reducing the processing of useless parts (non-speech components), thereby reducing the computational load of the system.
[0162] The speech recognition layer is used to determine the speech recognition result of the second speech based on the acoustic recognition features of the speech frame to be recognized.
[0163] The feature extraction layer, silence detection layer, and speech recognition layer can be implemented using different initial neural network models. The types of initial neural network models used in each layer can be the same or different; this application does not specifically limit this. Initial neural network models may include Convolutional Neural Networks (CNNs), Deep Feedforward Sequential Memory Networks (DFSMNs), Long Short-Term Memory Networks (LSTMs), and Transformers, etc.
[0164] To reduce the size of the speech recognition model, the silence detection layer and speech recognition layer can be implemented using parts of a neural network structure, such as fully connected layers. Since each layer performs a different task, although all are implemented using fully connected layers, the number of neurons and weight parameters differ between layers.
[0165] The speech recognition method provided in this application embodiment, since the silence detection layer and speech recognition layer in the speech recognition model share a feature extraction layer, enables the speech recognition model to implement silence detection and speech recognition functions separately through model fusion. This reduces the network size and computational parameters of the speech recognition model, improves the operation speed and response speed of the speech recognition model, and reduces the computational resource requirements of the speech recognition model. This allows the speech recognition model to be deployed on platforms with limited hardware resources, improving the convenience of users using the voice interaction system and enhancing the user experience of the in-vehicle terminal.
[0166] In some embodiments, the speech recognition model is deployed on an in-vehicle terminal or a cloud server corresponding to the in-vehicle terminal.
[0167] Specifically, the computational complexity of the speech recognition model can be calculated based on its structure and parameters. If the computational complexity is less than a preset data threshold, the speech recognition model can be deployed in the vehicle-mounted terminal. If the computational complexity is greater than or equal to the preset data threshold, the speech recognition model can be deployed on the cloud server corresponding to the vehicle-mounted terminal. The vehicle-mounted terminal then communicates with the cloud server to perform speech recognition on the second voice message. In other words, the vehicle-mounted terminal sends the second voice message to the cloud server, where the speech recognition model recognizes the message and sends the recognition result back to the vehicle-mounted terminal.
[0168] The preset data volume threshold is determined based on the computing power of the vehicle terminal.
[0169] In some embodiments, Figure 2 This is a second schematic flowchart of a speech recognition method provided in one embodiment of this application. Figure 3 This is a timing diagram of a speech recognition method provided in one embodiment of this application, as shown below. Figure 2 and Figure 3 As shown, this method is applied to an in-vehicle terminal and includes:
[0170] Step 210: Detect the mobile terminal
[0171] When a user enters the car with a mobile device, Bluetooth automatically connects, and the vehicle's onboard terminal and the mobile device establish a communication connection. Messages to detect the mobile device are initiated by the onboard terminal, and the mobile device responds.
[0172] Step 220: Voice function matching
[0173] After the communication connection is established, it is necessary to check whether the voice functions of the mobile terminal and the vehicle terminal are compatible, mainly checking whether the mobile terminal supports microphone extension. If the functions and software match, a success confirmation message is returned; otherwise, a failure message is returned.
[0174] Step 230, Microphone Extension
[0175] The microphone array in the mobile terminal is combined with the microphone array in the vehicle terminal to obtain a microphone extension array.
[0176] The mobile terminal first performs echo cancellation on the local microphone signal, and then transmits the collected audio to the vehicle terminal. The vehicle terminal combines the audio transmitted from the mobile terminal with the audio from the vehicle to form a voice signal combination.
[0177] Step 240, Voice Wake-up
[0178] In-vehicle terminal voice interaction does not recognize and respond to all user speech; it only recognizes speech in limited areas, such as navigation, music, and weather. Generally, a wake-up word is used to activate the system before further interaction can occur. Using a wake-up word not only filters out irrelevant speech but also saves computing resources. Only after activation will the in-vehicle terminal recognize subsequent recordings.
[0179] Step 250, Speech Recognition
[0180] After recognizing the user's voice, the in-vehicle terminal executes the corresponding operation.
[0181] Accordingly, when the above method is implemented in vehicle terminals and mobile terminals, corresponding software modules need to be set up in each terminal. Figure 4 This is a software module diagram provided in one embodiment of this application, such as... Figure 4 As shown, the software modules in the vehicle terminal include an array merging module, an echo cancellation module, a sound source localization module, a beamforming module, a noise reduction module, a silence detection module, a voice wake-up module, and a voice recognition module; the software modules in the mobile terminal include an audio transmission module and an echo cancellation module.
[0182] The speech recognition device provided in this application is described below. The speech recognition device described below can be referred to in correspondence with the speech recognition method described above.
[0183] Figure 5 This is a schematic diagram of the structure of a speech recognition device provided in one embodiment of this application, as shown below. Figure 5 As shown, the device includes:
[0184] The first acquisition module 510 is used to acquire the first voice based on a microphone extension array;
[0185] The wake-up module 520 is used to determine that the speech recognition result of the first speech contains a preset wake-up word;
[0186] The determination module 530 is used to determine the target microphone based on the sound propagation parameters of the first speech collected from each microphone in the microphone extension array.
[0187] The second acquisition module 540 is used to acquire signals based on the microphone extension array and enhance the signals acquired by the target microphone to determine the second speech.
[0188] The recognition module 550 is used to determine the speech recognition result of the second speech.
[0189] The microphone extension array includes a first microphone array in the vehicle terminal and one or more microphones in a mobile terminal connected to the vehicle terminal.
[0190] The speech recognition device provided in this application embodiment acquires first speech based on a microphone extension array; determines that the speech recognition result of the first speech contains a preset wake-up word; identifies a target microphone based on the sound propagation parameters of the first speech acquired by each microphone in the microphone extension array; enhances the signal of the target microphone and acquires second speech; and determines the speech recognition result of the second speech. Since the microphone extension array includes not only the first microphone array in the vehicle terminal but also one or more microphones in the mobile terminal, it increases the number of microphones used for speech acquisition without increasing the vehicle's hardware costs. Simultaneously, by enhancing the signal of the microphone array based on the acquisition result of the wake-up speech, the vehicle terminal can more accurately perceive the location of the user's voice, acquire more accurate user speech, improve the accuracy of voice wake-up and speech recognition, and enhance the user's voice interaction experience.
[0191] In some embodiments, the device further includes:
[0192] The microphone expansion module is used to confirm a successful connection between the mobile terminal and the vehicle terminal; and to send a microphone function query request to the mobile terminal.
[0193] Receive the microphone function configuration response sent by the mobile terminal based on the microphone function query request;
[0194] The response based on microphone function configuration determines whether the mobile terminal supports microphone expansion;
[0195] The first microphone array in the vehicle terminal is extended based on one or more microphones in the mobile terminal to obtain a microphone extension array.
[0196] In some embodiments, the first acquisition module is specifically used for:
[0197] Echo cancellation is performed on the first audio signal acquired by each microphone in the microphone extension array;
[0198] Beamforming is applied to the first audio signal after echo cancellation to obtain the first speech collected by each microphone.
[0199] In some embodiments, the first acquisition module is further specifically used for:
[0200] Determine the acquisition terminal corresponding to any microphone in the microphone extension array;
[0201] When the acquisition terminal is a vehicle-mounted terminal, echo cancellation is performed on the first audio signal acquired by any microphone.
[0202] When the acquisition terminal is a mobile terminal, an echo cancellation command is sent to the mobile terminal; the echo cancellation command is used to control the mobile terminal to cancel the echo of the first audio signal acquired by any microphone.
[0203] In some embodiments, the determining module is specifically used for:
[0204] Based on the sound propagation parameters of the first speech collected by each microphone, the arrival time of the first speech from the sound source to each microphone is determined.
[0205] Based on the arrival time of each microphone, the distance between each microphone and the sound source is determined;
[0206] The microphone closest to the sound source is identified as the target microphone.
[0207] In some embodiments, the second acquisition module is specifically used for:
[0208] Echo cancellation is performed on the second audio signal captured by the target microphone;
[0209] Based on the position of the target microphone in the microphone extension array, beamforming processing is performed on the second audio signal acquired by the target microphone to obtain the second speech acquired by the target microphone.
[0210] In some embodiments, the device further includes:
[0211] The noise reduction module is used to perform noise reduction processing on the second speech captured by the target microphone.
[0212] In some embodiments, the identification module is specifically used for:
[0213] The second voice is input into the speech recognition model to obtain the speech recognition result output by the speech recognition model;
[0214] The speech recognition model includes a feature extraction layer, a silence detection layer, and a speech recognition layer; the silence detection layer and the speech recognition layer are connected to the feature extraction layer, respectively.
[0215] The feature extraction layer is used to divide the second speech into multiple speech frames and extract the acoustic recognition features of each speech frame; the silence detection layer is used to determine the speech frame to be recognized in the second speech based on the acoustic recognition features of each speech frame; and the speech recognition layer is used to determine the speech recognition result of the second speech based on the acoustic recognition features of the speech frame to be recognized.
[0216] In some embodiments, the speech recognition model is deployed on an in-vehicle terminal or a cloud server corresponding to the in-vehicle terminal.
[0217] In some embodiments, Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application, as shown below. Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communications bus 640. The processor 610 can call logical commands stored in the memory 630 to execute the following methods:
[0218] The process involves: acquiring a first speech based on a microphone extension array; determining that the speech recognition result of the first speech includes a preset wake-up word; acquiring sound propagation parameters of the first speech based on each microphone in the microphone extension array to determine a target microphone; acquiring signals based on the microphone extension array and enhancing the signals acquired by the target microphone to determine a second speech; and determining the speech recognition result of the second speech. The microphone extension array includes a first microphone array in the vehicle terminal and one or more microphones in a mobile terminal connected to the vehicle terminal.
[0219] Furthermore, the logical commands in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0220] The processor in the electronic device provided in this application embodiment can call logical instructions in the memory to implement the above method. Its specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effect, which will not be repeated here.
[0221] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.
[0222] The specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, so it will not be repeated here.
[0223] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0224] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0225] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A speech recognition method, characterized in that, Applications in vehicle-mounted terminals include: First voice was acquired using a microphone extension array; It is determined that the speech recognition result of the first speech contains a preset wake word; Based on the sound propagation parameters of the first speech collected by each microphone in the microphone extension array, the target microphone is determined; Signal acquisition is performed based on the microphone extension array, and the signal acquired by the target microphone is enhanced to determine the second speech. Determine the speech recognition result of the second speech; The microphone extension array includes a first microphone array in the vehicle terminal and one or more microphones in a mobile terminal connected to the vehicle terminal. The acquisition of the first voice based on the microphone extended array includes: Echo cancellation is performed on the first audio signal acquired by each microphone in the microphone extension array; Beamforming is applied to the first audio signal after echo cancellation to obtain the first speech collected by each microphone. The step of acquiring signals based on the microphone extension array and enhancing the signals acquired by the target microphone to determine the second speech includes: Echo cancellation is performed on the second audio signal acquired by the target microphone; Based on the position of the target microphone in the microphone extension array, the second audio signal collected by the target microphone is subjected to beamforming processing to obtain the second speech collected by the target microphone.
2. The speech recognition method according to claim 1, characterized in that, The microphone extension array was determined based on the following steps: It is confirmed that the mobile terminal and the vehicle terminal are successfully connected; Send a microphone function query request to the mobile terminal; Receive the microphone function configuration response sent by the mobile terminal based on the microphone function query request; Based on the microphone function configuration response, it is determined that the mobile terminal supports microphone expansion; The first microphone array in the vehicle terminal is extended based on one or more microphones in the mobile terminal to obtain the extended microphone array.
3. The speech recognition method according to claim 1, characterized in that, The echo cancellation of the first audio signal acquired by each microphone in the microphone extension array includes: Determine the acquisition terminal corresponding to any microphone in the microphone extension array; When the acquisition terminal is a vehicle-mounted terminal, echo cancellation is performed on the first audio signal acquired by any of the microphones; When the acquisition terminal is a mobile terminal, an echo cancellation command is sent to the mobile terminal; the echo cancellation command is used to control the mobile terminal to cancel the echo of the first audio signal acquired by any of the microphones.
4. The speech recognition method according to claim 1, characterized in that, The step of determining the target microphone by acquiring the sound propagation parameters corresponding to the first speech from each microphone in the microphone extension array includes: Based on the sound propagation parameters of the first speech collected by each microphone, the arrival time of the first speech from the sound source to each microphone is determined. Based on the arrival time of each microphone, the distance between each microphone and the sound source is determined; The microphone that is closest to the sound source is identified as the target microphone.
5. The speech recognition method according to claim 1, characterized in that, After obtaining the second voice recorded by the target microphone, the process includes: The second voice recorded by the target microphone is subjected to noise reduction processing.
6. The speech recognition method according to any one of claims 1 to 5, characterized in that, The determination of the speech recognition result of the second speech includes: The second speech is input into the speech recognition model to obtain the speech recognition result output by the speech recognition model; The speech recognition model includes a feature extraction layer, a silence detection layer, and a speech recognition layer; the silence detection layer and the speech recognition layer are respectively connected to the feature extraction layer. The feature extraction layer is used to divide the second speech into multiple speech frames and extract the acoustic recognition features of each speech frame; the silence detection layer is used to determine the speech frame to be recognized in the second speech based on the acoustic recognition features of each speech frame; the speech recognition layer is used to determine the speech recognition result of the second speech based on the acoustic recognition features of the speech frame to be recognized.
7. The speech recognition method according to claim 6, characterized in that, The speech recognition model is deployed on the vehicle terminal or the cloud server corresponding to the vehicle terminal.
8. A voice recognition device, characterized in that, include: The first acquisition module is used to acquire the first voice based on the microphone extension array; A wake-up module is used to determine that the speech recognition result of the first speech contains a preset wake-up word; The determination module is used to determine the target microphone based on the sound propagation parameters corresponding to the first speech collected by each microphone in the microphone extension array; The second acquisition module is used to acquire signals based on the microphone extension array and enhance the signals acquired by the target microphone to determine the second speech. A recognition module is used to determine the speech recognition result of the second speech; The microphone extension array includes a first microphone array in the vehicle terminal and one or more microphones in a mobile terminal connected to the vehicle terminal. The acquisition of the first voice based on the microphone extended array includes: Echo cancellation is performed on the first audio signal acquired by each microphone in the microphone extension array; Beamforming is applied to the first audio signal after echo cancellation to obtain the first speech collected by each microphone. The step of acquiring signals based on the microphone extension array and enhancing the signals acquired by the target microphone to determine the second speech includes: Echo cancellation is performed on the second audio signal acquired by the target microphone; Based on the position of the target microphone in the microphone extension array, the second audio signal collected by the target microphone is subjected to beamforming processing to obtain the second speech collected by the target microphone.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the speech recognition method according to any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the speech recognition method according to any one of claims 1 to 7 through the computer program.