Vehicle exterior voice control method and device, storage medium and electronic equipment
Audio signals are collected through the microphone array around the vehicle, preprocessing and sound source separation, which realizes high accuracy and personalized interaction of voice control outside the vehicle, solves the voice recognition problem under noise interference in the environment outside the vehicle, and adapts to the needs of various scenarios.
Patent Information
- Application Number
- CN202510489304.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
AI Technical Summary
The accuracy of the voice control outside the vehicle is low in complex environments, making it difficult to effectively block noise interference and identify multiple pronunciators or different locations to communicate accurately.
The microphone array around the vehicle is used to collect audio signals, and through preprocessing, sound source positioning and separation, multi-channel audio signals are generated, wake-up words are identified and corresponding vehicle voice control commands are executed.
The accuracy of voice control is improved in complex environments, the commands of the target pronunciation person can be accurately identified, and the various sound zone division modes are supported, which is adapted to different models and scenarios, and is highly versatile and scalable.
Smart Images

Figure CN120375822A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to an out-of-vehicle voice control method, apparatus, storage medium, and electronic device. Background Art
[0002] Out-of-vehicle voice control refers to a technology in which a user interacts with a vehicle by voice in an external environment of the vehicle, and this technology allows the user to control the vehicle outside the vehicle. With the continuous improvement of the intelligence level of vehicles, the demand for out-of-vehicle voice control is growing rapidly.
[0003] Currently, due to the complex out-of-vehicle environment, it is easy to interfere with the acquisition and recognition of voice signals, resulting in a low accuracy rate of out-of-vehicle voice control. Summary of the Invention
[0004] The present disclosure provides an out-of-vehicle voice control method, apparatus, storage medium, and electronic device, mainly aiming to improve the technical problem that the current out-of-vehicle environment interferes with the acquisition and recognition of voice signals, resulting in a low accuracy rate of out-of-vehicle voice control.
[0005] According to a first aspect of an embodiment of the present disclosure, an out-of-vehicle voice control method is provided, including:
[0006] Obtaining audio signals collected by a microphone array on the vehicle for different regions outside the vehicle;
[0007] Preprocessing the audio signals;
[0008] Performing sound source localization and sound source separation according to the preprocessed audio signals to obtain multi-channel audio signals;
[0009] Performing wake word recognition based on the multi-channel audio signals to determine a target channel where the wake word is located;
[0010] Executing corresponding vehicle voice control commands according to the audio signals corresponding to the target channel.
[0011] Optionally, preprocessing the audio signals includes:
[0012] Performing preprocessing of echo cancellation and noise suppression on the audio signals.
[0013] Optionally, before preprocessing the audio signals, the method further includes:
[0014] Using the original audio played by the in-vehicle speaker as a reference signal to establish a transfer function model between the signal emitted by the in-vehicle speaker and the signal received by the microphone array;
[0015] Preprocessing the audio signals includes:
[0016] Using the transfer function model, identify and eliminate the signal part emitted by the vehicle-mounted speaker in the audio signal.
[0017] Optionally, the microphone array includes microphone modules respectively arranged for different areas outside the vehicle, and each microphone module has at least two microphones;
[0018] Performing sound source localization and sound source separation on the preprocessed audio signal to obtain a multi-channel audio signal, including:
[0019] Performing sound source localization on the preprocessed audio signal of at least one microphone module, and removing the sound source signal part that does not belong to the corresponding area of the microphone module according to the sound source localization result;
[0020] Performing sound source separation on the audio signal of at least one microphone module after removing the sound source signal part to obtain the multi-channel audio signal.
[0021] Optionally, performing sound source localization on the preprocessed audio signal of at least one microphone module, and removing the sound source signal part that does not belong to the corresponding area of the microphone module according to the sound source localization result includes:
[0022] For at least one target microphone module in the microphone module, using the audio signal of other microphone modules except the target microphone module as an interference signal, and removing the interference signal part of the sound source that does not belong to the corresponding area of the target microphone module in the target audio signal of the target microphone module.
[0023] Optionally, performing sound source separation on the audio signal of at least one microphone module after removing the sound source signal part to obtain the multi-channel audio signal includes:
[0024] Using a blind source separation algorithm to perform sound source separation on the audio signal of at least one microphone module after removing the sound source signal part according to the number of microphones in the microphone array to obtain the multi-channel audio signal.
[0025] Optionally, based on the multi-channel audio signal, performing wake word recognition to determine the target channel where the wake word is located, including:
[0026] Detecting whether each channel of the audio signal contains a wake word;
[0027] Determining the channel corresponding to the audio signal detected to contain the wake word as the target channel.
[0028] Optionally, before executing the corresponding vehicle voice control command according to the audio signal corresponding to the target channel, the method further includes:
[0029] Determine a target area where the audio signal corresponding to the target channel is located in different areas outside the vehicle;
[0030] Execute corresponding vehicle voice control commands according to the audio signal corresponding to the target channel, including:
[0031] Use the target area as a voice interaction area, and execute the vehicle voice control commands corresponding to the voice interaction area according to the audio signal corresponding to the target channel.
[0032] According to a second aspect of the embodiments of the present disclosure, there is provided an out-of-vehicle voice control device, including: an acquisition module, a processing module, a positioning module, an identification module, and a control module; the out-of-vehicle voice control device is configured to execute the method described in the first aspect.
[0033] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the out-of-vehicle voice control method described in the first aspect is implemented.
[0034] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the out-of-vehicle voice control method described in the first aspect is implemented.
[0035] According to a fifth aspect of the embodiments of the present disclosure, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the out-of-vehicle voice control method described in the first aspect is implemented..
[0036] With the above technical solution, the present disclosure provides an external vehicle voice control method, device, storage medium and electronic device. Compared with the current related technologies, the present disclosure first obtains the audio signals collected by the microphone array on the vehicle for different regions outside the vehicle; then preprocesses the audio signals; performs sound source localization and sound source separation on the preprocessed audio signals to obtain multi-channel audio signals; then performs wake word recognition based on the multi-channel audio signals to determine the target channel where the wake word is located; and executes the corresponding vehicle voice control command according to the audio signal corresponding to the target channel. By applying the technical solution of the present disclosure, audio signals are collected by the microphone arrays in different regions around the vehicle, and through preprocessing, sound source localization and separation, multi-channel audio signals are obtained, thereby realizing accurate wake word recognition and execution of voice control commands. It can not only effectively shield the complex outdoor environmental noise, but also accurately recognize the commands of the target speaker, greatly improving the accuracy of external vehicle voice control; at the same time, it supports multiple sound zone division modes and the operation mode without sound zone division, enabling the embodiments of the present disclosure to meet the requirements of different vehicle models and usage scenarios, and having strong versatility and scalability.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0039] Figure 1 The flowchart showing a method for external vehicle voice control provided by an embodiment of the present disclosure;
[0040] Figure 2 The flowchart showing another method for external vehicle voice control provided by an embodiment of the present disclosure;
[0041] Figure 3 The schematic diagram showing an example provided by an embodiment of the present disclosure;
[0042] Figure 4 The schematic diagram showing another example provided by an embodiment of the present disclosure;
[0043] Figure 5 The schematic diagram showing yet another example provided by an embodiment of the present disclosure;
[0044] Figure 6 The schematic diagram showing the structure of an external vehicle voice control device provided by an embodiment of the present disclosure;
[0045] Figure 7 The schematic diagram showing the structure of a vehicle provided by an embodiment of the present disclosure. Detailed implementation manners
[0046] Some embodiments of the present disclosure will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those set forth herein, but may be changed as will be apparent after understanding the present disclosure, except for operations that must be performed in a specific order. Additionally, descriptions of features known in the art may be omitted for the sake of clarity and conciseness. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.
[0047] The implementation manners described in some embodiments of the present disclosure below do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0048] For vehicle exterior voice control technology, as an example, in a parking assist scenario, a user can control the vehicle's automatic parking system through voice commands and obtain real-time information about the surrounding environment. This convenient interaction method not only improves the user experience but also enhances driving safety to a certain extent. In specific scenarios such as camping, the demand for vehicle exterior voice interaction is also significant. A user can control in-vehicle devices through voice commands, such as adjusting lighting, playing music, or even controlling camping equipment outside the vehicle. This interaction method not only improves the convenience of the user but also enhances the interactivity between the vehicle owner and the vehicle.
[0049] Generally speaking, the development of vehicle exterior intelligent voice interaction technology is moving towards a more intelligent, personalized, and multifunctional direction. Related research and technological progress are also continuously driving the advancement of this field. However, currently, vehicle exterior voice interaction still mainly focuses on single functions (such as trunk opening) and lacks a systematic solution for multi-person and multi-zone collaboration in complex application environments.
[0050] Due to the complex interactive acoustic environment outside the vehicle, such as in basements, on roads, etc., there are characteristics such as high noise and large reverberation, which easily interfere with the acquisition and recognition of voice signals, resulting in a decrease in the accuracy of interaction; traditional voice interaction systems lack effective discrimination and targeted interaction capabilities when dealing with the situation where multiple speakers speak simultaneously or at different positions, and it is difficult to meet the diverse usage needs of users; when performing voice interaction, the vehicle head unit usually plays stereo sources such as TTS reply sounds, music, and audiobooks through external vehicle speakers, and the speaker signals collected by the signal acquisition system will affect the extraction of the sound source of the target speaker, reducing the voice wake-up and recognition performance.
[0051] In order to improve the technical problem that the current external vehicle environment interferes with the acquisition and recognition of voice signals, resulting in a low accuracy of external vehicle voice control, the embodiments of the present disclosure provide an external vehicle voice control method, as Figure 1 shown, this method can be applied to be executed by an external vehicle voice control device or equipment, etc., and this method includes the following steps.
[0052] Step 101: Obtain the audio signals collected by the microphone array on the vehicle for different regions outside the vehicle.
[0053] Exemplarily, a plurality of microphones are arranged at different positions of the vehicle (such as around the vehicle body, on the roof, on the door, etc.) to form an array. Through the vehicle control system, audio signals from different directions and distances can be captured, including voice commands of target people inside and outside the vehicle and the sounds of the external vehicle environment. Such a multi-region microphone layout can provide a rich data basis for subsequent voice processing, achieve omnidirectional and multi-angle voice acquisition, cover the sound sources in all directions around the vehicle, and provide support for operations such as sound source localization, noise suppression, and voice recognition, thereby improving the performance and reliability of the vehicle control system.
[0054] Step 102: Perform preprocessing on the audio signals.
[0055] In some examples, the audio quality can be optimized by performing preprocessing on the audio signals to lay a foundation for subsequent processing. The preprocessing usually includes operations such as denoising, filtering, and gain adjustment, aiming to remove environmental noise, suppress interference signals, enhance the amplitude and clarity of voice signals, and at the same time correct the signal amplitude differences, so as to improve the signal-to-noise ratio and recognizability of the audio signals, and provide more accurate and reliable input data for subsequent steps such as sound source localization, sound source separation, and voice recognition.
[0056] Step 103: Perform sound source localization and sound source separation according to the preprocessed audio signals to obtain multi-channel audio signals.
[0057] For example, the spatial information of a microphone array can be utilized to optimize audio processing. Sound source localization determines the direction and position of the sound source by analyzing the time delay differences of the signals received by different microphones. Sound source separation decomposes the mixed audio signals into independent sound source signals through algorithms, and can effectively separate the target speech. The multi-channel audio signals generated after sound source localization and sound source separation processing can provide clearer and more accurate signal inputs for subsequent speech recognition and interaction.
[0058] Step 104: Perform wake word recognition based on the multi-channel audio signals to determine the target channel where the wake word is located.
[0059] In some examples, the multi-channel audio signals can provide rich information for wake word recognition, and can effectively improve the accuracy and robustness of recognition. The vehicle control system uses efficient algorithms to search for preset wake words in each channel. Once the wake word is detected, it determines the target channel where it is located, thereby locking the position and sound area of the target speaker. This not only enhances the accuracy of voice interaction, but also enables directional interaction. Even when multiple speakers are present simultaneously or in a complex noise environment, it can accurately identify the source of the instruction, providing a more fluent and personalized voice interaction experience for users, effectively improving the response efficiency and interaction accuracy of the system, and is applicable to various scenarios that require high-precision speech recognition.
[0060] Exemplarily, the method of this embodiment can not only improve the recognition rate of the wake word, but also reduce the false wake-up rate, thereby providing a more reliable trigger signal for subsequent voice interaction.
[0061] Step 105: Execute corresponding vehicle voice control commands according to the audio signals corresponding to the target channel.
[0062] Exemplarily, when the target speaker expresses an interaction intention through a wake word instruction outside the vehicle, the in-vehicle processor performs processing such as echo cancellation, noise suppression, and sound source separation on the sound signals received by the microphone, and transmits the processed sound signals to the wake-up decision module to identify the current target speaker's sound track, and then transmits them to the voice recognition module of the intelligent cockpit system or other in-vehicle devices for further processing or response. This system can quickly and accurately parse the instructions of the target speaker, ensuring the fluency and accuracy of interaction in complex outdoor environments, and providing a convenient and efficient out-of-vehicle voice interaction service for users.
[0063] In specific application scenarios, the method of the embodiments of the present disclosure can be applied to various carriers or scenarios, such as automotive intelligent cockpits, etc., to provide users with a more intelligent and convenient interaction experience.
[0064] Compared with the current related technologies, in the embodiments of the present disclosure, audio signals are collected by microphone arrays in different areas around the vehicle, and after preprocessing, sound source localization and separation, multi-channel audio signals are obtained, thereby realizing accurate wake-word recognition and execution of voice control commands. It can not only effectively shield the complex outdoor environmental noise, but also accurately recognize the commands of the target speaker, greatly improving the accuracy of out-of-vehicle voice control. At the same time, it supports multiple sound zone division modes and the operation mode without sound zone division, enabling the embodiments of the present disclosure to meet the requirements of different vehicle models and usage scenarios, and having strong versatility and scalability.
[0065] To further illustrate the specific implementation process of the method as Figure 1 shown, as an optional way, the embodiments of the present disclosure provide the specific method as Figure 2 shown, and the method includes:
[0066] Step 201: Obtain the audio signals collected by the microphone array on the vehicle for different areas outside the vehicle.
[0067] Exemplarily, as Figure 3 shown, the vehicle control system in this embodiment includes multiple microphone modules and speakers, which are used to receive and play sound signals. The specific functions include: echo cancellation to ensure clear voice reception; noise shielding to reduce external noise interference; interference source shielding to identify and shield interference sources in specific directions; wake-up providing azimuth information to provide azimuth information according to the position of the target speaker. At the same time, multiple sound zones are set around the vehicle, and these sound zones can be fixed or dynamically adjusted according to the movement of the vehicle and the surrounding environment. The system includes one or more microphone modules, which are used to receive sound signals from each sound zone, support different configurations such as 2 modules corresponding to 2 sound zones, 4 modules corresponding to 4 sound zones, etc., and can also operate stably without sound zone division. At the same time, the system is equipped with speakers, which can play sounds to each sound zone.
[0068] For example, as Figure 4 shown, four modules each containing 2 microphones are deployed around the vehicle, and the space outside the vehicle is divided into "front-of-vehicle sound zone", "rear-of-vehicle sound zone", "left-side-of-vehicle sound zone" and "right-side-of-vehicle sound zone". The sound source signals in each sound zone will be received by the corresponding microphone array, and these signals are mixed with multiple sound sources, echoes and noises. In order to accurately capture audio information, all microphones need to collect data synchronously and send the data to the central processing unit for processing through an appropriate transmission method. Using signal processing technologies such as beamforming and sound source localization, the system can identify and separate the sound sources of interest, thereby improving the audio quality in a complex acoustic environment and applying it to multiple aspects such as enhancing driving safety and user experience.
[0069] Step 202: Preprocess the audio signals.
[0070] Exemplarily, the main purpose of the signal processing stage for preprocessing the audio signal is to accurately extract the voice track of the target speaker from the collected audio containing mixed echoes, noise, interference, and the target speaker. This process involves using advanced signal processing techniques, such as echo cancellation, noise suppression, and source separation methods, to purify the audio signal, remove unnecessary noise and echoes, and thus accurately capture the voice information of the target speaker. The processed clear voice signal will be sent to the wake-up decision module for further identification and response to the user's instructions, ensuring an efficient and accurate voice interaction experience in a complex acoustic environment.
[0071] Optionally, step 202 may specifically include: preprocessing the audio signal by performing echo cancellation and noise suppression.
[0072] For example, to perform noise suppression on four groups of microphone signals respectively, traditional signal processing methods such as spectral subtraction, filtering method, based on statistical models and subspace can be used. At the same time, neural network methods can also be used to achieve a more accurate noise reduction effect. These processing means aim to effectively remove background noise and echoes, ensure the extraction of high-quality voice signals of the target speaker from the mixed audio, and provide clear data support for further speech recognition and processing.
[0073] Optionally, the method of this embodiment may specifically further include: using the original audio played by the in-vehicle speaker as a reference signal to establish a transfer function model between the signal emitted by the in-vehicle speaker and the signal received by the microphone array; correspondingly, the above-mentioned echo cancellation of the audio signal may specifically further include: using the transfer function model to identify and eliminate the signal part emitted by the in-vehicle speaker in the audio signal.
[0074] In some examples, using the original audio played by the in-vehicle speaker (such as voice prompts, media playback, etc.) as a reference signal to establish a transfer function model between the speaker signal and the echo received by the microphone, reducing the interference of echoes caused by sound reflections inside and outside the vehicle to the signal. Usually, an adaptive filter or a model-based method is used to achieve echo cancellation. The adaptive filter will automatically adjust the filter parameters according to the characteristics of the sound inside and outside the vehicle to minimize the impact of the echo signal.
[0075] Step 203: Perform sound source localization on the preprocessed audio signal of at least one microphone module, and remove the sound source signal part that does not belong to the corresponding area of the microphone module according to the sound source localization result.
[0076] Exemplarily, the direction and position of a sound source can be accurately determined through sound source localization. On this basis, removing the sound source signals that do not belong to the target area can significantly reduce the background noise and the influence of other interfering sound sources, thereby improving the quality and clarity of the target voice signal and providing a better input for subsequent speech recognition and interaction.
[0077] In some embodiments, the microphone array includes microphone modules respectively arranged for different areas outside the vehicle, and each microphone module has at least two microphones.
[0078] Optionally, step 203 may specifically include: for at least one target microphone module in the microphone modules, using the audio signals of other microphone modules except the target microphone module as interference signals, and removing the interference signal part of the sound source that does not belong to the area corresponding to the target microphone module in the target audio signal of the target microphone module through an adaptive interference canceller.
[0079] In some examples, an adaptive interference canceller can be used to update the interference reference signal in real time according to the energy magnitude at each time-frequency point by using the normalized least mean squares algorithm, and compare the updated interference reference signal with the target audio signal of the target microphone module to remove the interference signal part of the sound source that does not belong to the area corresponding to the target microphone module in the target audio signal.
[0080] Exemplarily, based on the orthogonality hypothesis of speech, there is only the frequency-domain information of one sound source at each time-frequency point. Taking the first group of microphone signals X1 as an example, when using an adaptive interference canceller for sound zone cancellation, using the other three groups of microphone signals as interference signals, and updating the interference reference signal in real time according to the energy magnitude at each time-frequency point by using the Normalized Least Mean Squares (NLMS) algorithm to suppress interference and remove the interference of sound sources in other sound zones to the current sound zone.
[0081] Step 204: Perform sound source separation on the audio signals of at least one microphone module after removing the sound source signal part to obtain multi-channel audio signals.
[0082] Exemplarily, through sound source separation techniques such as beamforming or blind source separation algorithms, the mixed audio signals can be decomposed into multiple independent sound source signals, thereby achieving the precise extraction and enhancement of the target voice signal.
[0083] Optionally, step 204 may specifically include: using a blind source separation algorithm to perform sound source separation on the audio signals of at least one microphone module after removing the sound source signal part according to the number of microphones in the microphone array to obtain multi-channel audio signals.
[0084] Exemplarily, the blind source separation algorithm is used to separately separate the signals of two microphones from four groups of microphone signals, and two separated signals can be obtained for each sound zone, namely the target person signal and the interfering person signal. This method can achieve sound source separation only relying on the mixed signals received by the microphone array without explicitly knowing the specific information of the sound source, thereby effectively extracting the speech signal of the target speaker and suppressing other interfering signals. It can also improve the signal-to-noise ratio and clarity of the speech signal, provide clearer and more accurate audio data for subsequent speech processing tasks, and significantly enhance the performance of the system and the user experience.
[0085] Step 205: Perform wake word recognition based on the multi-channel audio signal to determine the target channel where the wake word is located.
[0086] Exemplarily, the multi-channel audio signal provides rich spatial information, enabling the vehicle control system to more accurately recognize the wake word while effectively reducing the false wake-up rate. By determining the specific channel where the wake word is located, the system can further focus on the audio signal of this channel for subsequent processing, thereby improving the response speed of voice interaction and the user experience.
[0087] Optionally, step 205 may specifically include: detecting whether the audio signal of each channel contains the wake word; determining the channel corresponding to the detected audio signal containing the wake word as the target channel.
[0088] Exemplarily, the two separated signals obtained from the four sound zones respectively pass through the wake-up module. At this time, once the wake-up module detects the wake word, it will give a wake-up event. The path where the wake word is located is the sound track where the target person is located, the sound zone where the wake-up audio is located is the interaction sound zone, and the sound track where the target person in the wake-up sound zone is located is further handed over to the recognition system for speech recognition and corresponding control commands are executed.
[0089] Step 206: Execute the corresponding vehicle voice control command according to the audio signal corresponding to the target channel.
[0090] In some examples, through the processing of multi-channel audio signals and wake word recognition, the vehicle control system can accurately locate the source of the voice command and focus on the audio signal of the target channel, thereby effectively avoiding misoperations.
[0091] Optionally, the method of this embodiment may specifically further include: determining the target area where the audio signal corresponding to the target channel is located in different areas outside the vehicle; correspondingly, step 206 may specifically include: using the target area as the voice interaction area and executing the corresponding vehicle voice control command of the voice interaction area according to the audio signal corresponding to the target channel.
[0092] In some examples, the pure voice signal after front - end signal processing is transmitted to the speech recognition and interaction module. This module uses an automatic speech recognition model to accurately recognize speech commands and make a quick and accurate response to the recognized speech commands. When the speech recognition module completes the command recognition, it transmits the command information to the system response module. This module, according to the command type and combined with the information of multiple sound zones, realizes personalized interaction for speakers in different sound zones. And it executes corresponding operations through the vehicle control system. For example, when receiving the command "open the door", the system response module controls the door unlocking motor to work to unlock the door; at the same time, the system response module will feedback the operation status to the speaker outside the vehicle through the vehicle's lights, horn or in - vehicle display screen, etc., so that the user can timely know whether the interaction is successful.
[0093] Exemplarily, the method of this embodiment can also support multi - sound - zone interaction. For a target person in different regions or a target person moving between different regions, the vehicle control system can independently perform speech recognition and command parsing according to the region where the target person is located, and switch the target region as the speech interaction region. For example, multiple microphone arrays deployed around the vehicle can be used to collect audio signals in real - time, and advanced signal processing algorithms can be used for echo cancellation and noise suppression to ensure high - quality audio input, identify and track the specific position of the target person, automatically switch the target region for speech interaction according to the region where the target person is located, realize accurate speech recognition and command parsing for the specified user, ensure accurate response to the commands of specific users even in complex environments, meet the speech control requirements of target people in different positions outside the vehicle, and at the same time allocate different permissions according to the sound zone to further optimize the flexibility and security of the interaction.
[0094] In some embodiments, when the target speaker issues an interaction intention through a wake - up word command outside the vehicle, it is often affected by complex environments, such as various noises in the surrounding environment, interference from other voices, and music played by the vehicle head unit itself. This patent uses front - end signals such as echo cancellation, dynamic noise suppression, multi - source signal separation, and sound - zone locking to solve problems such as poor noise robustness / multi - source interference, extracts multiple audio signals including the target speaker, and transmits them to the wake - up decision module to lock the interaction sound zone of the target speaker, and selects the sound track where the "target speaker" is located through the wake - up word decision module. The voice signal after front - end signal processing is transmitted to the speech recognition and interaction module to realize personalized interaction for speakers in different sound zones without being interfered by irrelevant voices in other sound zones, and realize the function of multi - interference - person and multi - sound - zone collaborative interaction in complex application environments.
[0095] For example, such as Figure 5As shown, first, audio signals in different sound zones are collected by multiple microphone arrays deployed around the vehicle. Then, during the signal processing stage, these audio signals are preprocessed for echo cancellation and noise suppression, using traditional methods such as spectral subtraction and filtering methods or neural network methods. Then, a blind source separation algorithm is used to process the microphone signals in each sound zone to separate the target speaker signal and the interference signal. The processed signals are sent to the wake-up decision module to identify and confirm the instructions of the target speaker, and then the processing results are transmitted to the Automatic Speech Recognition (ASR) system for further processing, so as to achieve the efficient response and smooth interaction experience of the vehicle control system, ensuring that the user's voice instructions can be accurately captured and parsed even in a complex acoustic environment.
[0096] The out-of-vehicle voice control solution provided by the embodiments of the present disclosure can be applied to complex outdoor scenarios. Especially when there are various noises and human voice interferences in the surrounding environment, it can effectively perform echo cancellation, noise suppression, distinguish the target speaker and the interfering speaker, and lock the sound zone where the target speaker is located, thus providing a more accurate and customized voice interaction experience for the target speaker. By applying the technical solution of the embodiments of the present disclosure, audio signals are collected by microphone arrays in different areas around the vehicle, and after preprocessing, sound source localization and separation, multi-channel audio signals are obtained, and then accurate wake-up word recognition and voice control command execution are realized. Through multi-sound zone division and accurate front-end signal processing technology, it can effectively shield the noise of the complex outdoor environment, accurately identify the instructions of the target speaker, and greatly improve the accuracy of voice interaction; at the same time, it can independently interact with speakers in different sound zones, avoiding command confusion when multiple speakers speak simultaneously, and providing a more personalized and efficient interaction experience for users; it also supports multiple sound zone division modes and the operation mode without sound zone division, and can adapt to the needs of different vehicle models and different usage scenarios, with strong versatility and scalability.
[0097] Figure 6 is a block diagram of an out-of-vehicle voice control device shown according to some embodiments of the present disclosure. Referring to Figure 6 , the device includes: an acquisition module 31, a processing module 32, a positioning module 33, an identification module 34, and a control module 35. The device is configured to execute Figures 1 to 2 the out-of-vehicle voice control method shown.
[0098] The acquisition module 31 is configured to acquire audio signals collected by the microphone array on the vehicle for different areas outside the vehicle;
[0099] The processing module 32 is configured to preprocess the audio signals;
[0100] A positioning module 33, configured to perform sound source positioning and sound source separation according to the preprocessed audio signal to obtain a multi-channel audio signal;
[0101] An identification module 34, configured to perform wake-up word identification based on the multi-channel audio signal to determine the target channel where the wake-up word is located;
[0102] A control module 35, configured to execute corresponding vehicle voice control commands according to the audio signal corresponding to the target channel.
[0103] In some embodiments, the processing module 32 is specifically configured to perform preprocessing of echo cancellation and noise suppression on the audio signal.
[0104] In some embodiments, using the original audio played by the in-vehicle speaker as a reference signal, a transfer function model between the signal emitted by the in-vehicle speaker and the signal received by the microphone array is established; correspondingly, the processing module 32 is specifically further configured to use the transfer function model to identify and eliminate the signal part emitted by the in-vehicle speaker in the audio signal.
[0105] In some embodiments, each area around the vehicle is correspondingly provided with a microphone module, and the microphone module includes the microphone array having at least two microphones; correspondingly, the positioning module 33 is specifically configured to perform sound source positioning on the preprocessed audio signal of at least one microphone module, and remove the sound source signal part that does not belong to the area corresponding to the microphone module according to the sound source positioning result; perform sound source separation on the audio signal of at least one microphone module after removing the sound source signal part to obtain the multi-channel audio signal.
[0106] In some embodiments, the positioning module 33 is specifically further configured to, for at least one target microphone module in the microphone module, use the audio signals of other microphone modules except the target microphone module as interference signals, and an adaptive interference canceller removes the interference signal part of the sound source that does not belong to the area corresponding to the target microphone module in the target audio signal of the target microphone module.
[0107] In some embodiments, the positioning module 33 is specifically further configured to use a blind source separation algorithm to perform sound source separation on the audio signal of at least one microphone module after removing the sound source signal part according to the number of microphones in the microphone array to obtain the multi-channel audio signal.
[0108] In some embodiments, the identification module 34 is specifically configured to detect whether the wake-up word is included in the audio signal of each channel; determine the channel corresponding to the audio signal detected to include the wake-up word as the target channel.
[0109] In some embodiments, a target area where an audio signal corresponding to the target channel is located is determined in different areas outside the vehicle; correspondingly, the control module 35 is specifically configured to use the target area as a voice interaction area and execute vehicle voice control commands corresponding to the voice interaction area according to the audio signal corresponding to the target channel.
[0110] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0111] It should be noted that for other corresponding descriptions of each functional unit involved in the vehicle external voice control device provided in the embodiments of the present disclosure, reference can be made to Figures 1 to 2 the corresponding descriptions therein, which will not be repeated here.
[0112] Based on the above method as Figures 1 to 2 shown, correspondingly, the embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as Figures 1 to 2 shown is implemented.
[0113] Based on such an understanding, the technical solution of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various implementation scenarios of the present disclosure.
[0114] Based on the above method as Figures 1 to 2 shown, and Figure 6 the virtual device embodiments shown, for the purpose of achieving the above object, the embodiments of the present disclosure also provide an electronic device, such as a high-definition television, a computer, a projector, a vehicle, etc. with a liquid crystal display, a plasma display, a digital light processing display, etc., and the device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as Figures 1 to 2 shown.
[0115] Optionally, the above physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc., and optionally the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0116] Those skilled in the art can understand that the above-mentioned physical device structure provided by the embodiments of the present disclosure does not limit the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0117] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication between other hardware and software in the information processing physical device.
[0118] Figure 7 It is a block diagram of a vehicle 600 shown according to an exemplary embodiment. For example, the vehicle 600 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0119] Refer to Figure 7 , the vehicle 600 may include various subsystems. For example, the infotainment system 610, the perception system 620, the decision control system 630, the drive system 640, and the computing platform 650. Among them, the vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 may be interconnected by wired or wireless means.
[0120] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, etc.
[0121] The perception system 620 may include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.
[0122] The decision control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0123] The drive system 640 may include components that provide motive power for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.
[0124] Some or all functions of the vehicle 600 are controlled by the computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652. The processor 651 may execute instructions 653 stored in the memory 652.
[0125] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0126] The memory 652 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0127] In addition to the instructions 653, the memory 652 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 652 can be used by the computing platform 650.
[0128] In an embodiment of the present disclosure, the processor 651 may execute the instructions 653 to complete all or part of the steps of the above-mentioned audio processing method.
[0129] Based on the method as Figures 1 to 2 shown above, and Figure 6 the virtual device embodiment as Figures 1 to 2 shown above, an embodiment of the present disclosure also provides a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from the memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the method as
[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. Compared with the current related technologies, the embodiments of the present disclosure, through multi-zone division and precise front-end signal processing technology, can not only effectively shield the complex outdoor environmental noise, but also accurately identify the instructions of the target speaker, greatly improving the accuracy of voice interaction; at the same time, it can independently interact with speakers in different zones, avoiding the command confusion caused when multiple speakers speak simultaneously, and providing users with a more personalized and efficient interaction experience. In addition, supporting multiple zone division modes and the operation mode without zone division enables the embodiments of the present disclosure to meet the requirements of different vehicle models and usage scenarios, and has strong versatility and scalability.
[0131] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0132] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An in-vehicle voice control method, characterized in that, Including: Obtaining audio signals collected by a microphone array on a vehicle for different areas outside the vehicle; Preprocessing the audio signals; Performing sound source localization and sound source separation based on the preprocessed audio signals to obtain multi-channel audio signals; Performing wake word recognition based on the multi-channel audio signals to determine the target channel where the wake word is located; Executing corresponding vehicle voice control commands according to the audio signals corresponding to the target channel.
2. The method according to claim 1, characterized in that Before preprocessing the audio signals, the method further includes: Using the original audio played by an in-vehicle speaker as a reference signal to establish a transfer function model between the signal emitted by the in-vehicle speaker and the signal received by the microphone array; Preprocessing the audio signals includes: Using the transfer function model to identify and eliminate the signal part emitted by the in-vehicle speaker in the audio signals.
3. The method according to claim 1, characterized in that, The microphone array includes microphone modules respectively arranged for different areas outside the vehicle, and each microphone module has at least two microphones; The performing sound source localization and sound source separation based on the preprocessed audio signals to obtain multi-channel audio signals includes: Performing sound source localization on the preprocessed audio signals of at least one microphone module, and removing the sound source signal part that does not belong to the area corresponding to the microphone module according to the sound source localization result; Performing sound source separation on the audio signals of at least one microphone module after removing the sound source signal part to obtain the multi-channel audio signals.
4. The method according to claim 3, wherein The performing sound source localization on the preprocessed audio signals of at least one microphone module, and removing the sound source signal part that does not belong to the area corresponding to the microphone module according to the sound source localization result includes: For at least one target microphone module in the microphone module, using the audio signals of other microphone modules except the target microphone module as interference signals, and removing the interference signal part of the sound source that does not belong to the area corresponding to the target microphone module in the target audio signal of the target microphone module through an adaptive interference canceller.
5. The method according to claim 3, characterized in that The performing sound source separation on the audio signals of at least one microphone module after removing the sound source signal part to obtain the multi-channel audio signals includes: Using a blind source separation algorithm to perform sound source separation on the audio signals of at least one microphone module after removing the sound source signal part according to the number of microphones in the microphone array to obtain the multi-channel audio signals.
6. The method according to claim 1, characterized in that, The performing wake word recognition based on the multi-channel audio signals to determine the target channel where the wake word is located includes: Detecting whether the wake word is included in the audio signals of each channel; Determining the channel corresponding to the audio signal detected to include the wake word as the target channel.
7. The method according to any one of claims 1 to 6, characterized in that, Before executing corresponding vehicle voice control commands according to the audio signals corresponding to the target channel, the method further includes: Determining the target area where the audio signal corresponding to the target channel is located among different areas outside the vehicle; Executing corresponding vehicle voice control commands according to the audio signals corresponding to the target channel includes: Taking the target area as a voice interaction area, and executing corresponding vehicle voice control commands for the voice interaction area according to the audio signal corresponding to the target channel.
8. An in-vehicle voice control device, characterized in that, Including: An acquisition module, a processing module, a positioning module, an identification module and a control module; the vehicle exterior voice control device is configured to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Vehicle exterior voice interaction method, system and device, storage medium and program product
CN121708894A
Speech recognition method and device and vehicle
CN122050376A