A vehicle voice instruction response method and device and a storage medium
By combining a microphone array with a mobile terminal microphone, the problem of vehicles being unable to accurately locate sound sources has been solved, thus improving the accuracy of voice command response without increasing hardware costs.
Patent Information
- Application Number
- CN202211482991.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-11-24
AI Technical Summary
In existing technologies, vehicles cannot accurately locate the position of sound sources, resulting in inaccurate responses to vehicle voice commands, and adding rear microphones would increase hardware costs.
By combining a microphone array with a mobile terminal microphone, and through time delay calculation and ultra-wideband signal localization, the location of the sound source is accurately located. This includes voice localization and recognition information from vehicle microphones and mobile terminal microphones, combined with echo cancellation algorithm noise reduction processing to achieve sound source localization.
Without increasing vehicle hardware costs, this method accurately locates sound sources, improves the accuracy of vehicle voice command responses, and enhances the vehicle's voice command execution capabilities.
Smart Images

Figure CN115877756B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to a vehicle voice command response method, device and storage medium. Background Technology
[0002] With the rapid development of intelligent vehicles, the voice function of in-vehicle systems is being used more and more frequently. Currently, most vehicles have microphones installed on both sides of the front row. The position of the sound source relative to the vehicle is determined based on the phase and loudness of the sound source. However, the judgment result can only indicate the left and right position of the sound source relative to the vehicle, but cannot indicate the front and rear position of the sound source relative to the vehicle, so it is impossible to accurately locate the sound source.
[0003] In the existing technology, in order to solve the problem of inaccurate sound source localization, two rear microphones are added to the vehicle to help determine the front and rear positions of the sound source relative to the vehicle. However, adding microphones to the rear of the vehicle will inevitably increase the hardware cost of the vehicle, and it is not convenient to add rear microphones to vehicles that have already been manufactured. Summary of the Invention
[0004] This application provides a vehicle voice command response method, device, and storage medium, which can accurately locate the sound source without increasing the vehicle's hardware cost, thereby improving the accuracy of the vehicle's voice command response.
[0005] On the one hand, this application provides a vehicle voice command response method, including receiving output voice data of a target object based on a microphone array corresponding to the target vehicle, wherein the microphone array includes at least one vehicle-mounted microphone and at least one mobile microphone of the target mobile terminal;
[0006] Acquire the voice positioning and recognition information corresponding to each of the at least one vehicle-mounted microphone and the at least one mobile microphone, wherein the voice positioning and recognition information is the recognition information related to the audio transmission distance generated when the output voice data is transmitted to the vehicle-mounted microphone or to the mobile microphone;
[0007] Based on the voice localization and recognition information corresponding to each of the at least one vehicle-mounted microphone and the at least one mobile microphone, the output voice data is subjected to sound source localization and recognition to obtain the sound source localization result.
[0008] Based on the sound source localization result, execute the voice command corresponding to the output voice data.
[0009] Furthermore, before receiving the output voice data of the target object based on the microphone array corresponding to the target vehicle, the method further includes:
[0010] Obtain the relative position information between the mobile terminal and the target vehicle;
[0011] If the relative position information indicates that the mobile terminal is located within a preset area of the target vehicle, the mobile terminal is determined to be the target mobile terminal.
[0012] Furthermore, before performing sound source localization and recognition on the output speech data based on the speech localization and recognition information corresponding to the at least one vehicle-mounted microphone and the at least one mobile microphone to obtain the sound source localization result, the method further includes:
[0013] Acquire the first location information corresponding to each of the at least one target mobile terminal and the second location information corresponding to each of the at least one vehicle microphone;
[0014] Based on the first location information and the second location information, the microphone array information corresponding to the microphone array is determined.
[0015] Furthermore, the voice location and recognition information includes voice reception time information for the output voice data; the step of performing sound source location and recognition on the output voice data based on the voice location and recognition information corresponding to the at least one vehicle-mounted microphone and the at least one mobile microphone to obtain the sound source location result includes:
[0016] Based on the voice reception time information corresponding to the at least one vehicle-mounted microphone and the at least one mobile microphone, the time delay difference between the sound source corresponding to the output voice data and the at least one vehicle-mounted microphone and the at least one mobile microphone is obtained by calculating the time delay.
[0017] Based on the time delay difference and audio transmission speed information, distance calculation is performed to obtain the distance difference corresponding to the transmission distance between the sound source and the at least one vehicle-mounted microphone and the at least one mobile microphone;
[0018] Based on the distance difference and the microphone array information, the output voice data is used to locate and identify the sound source, and the sound source localization result is obtained.
[0019] Furthermore, before performing sound source localization and identification on the output speech data based on the distance difference and the microphone array information to obtain the sound source localization result, the method further includes:
[0020] Monitor the first location information;
[0021] Upon detecting an update to the first location information, the microphone array information is updated based on the second location information and the updated first location information.
[0022] Furthermore, the target vehicle includes multiple anchor antennas, and acquiring the relative position information between the mobile terminal and the target vehicle includes:
[0023] Obtain the ultra-wideband signal flight time information corresponding to each of the mobile terminal and the plurality of anchor antennas;
[0024] Based on the ultra-wideband signal flight time information and signal flight speed information, the flight distance is calculated to obtain the ultra-wideband signal flight distance information between the mobile terminal and the multiple anchor antennas respectively.
[0025] Based on the ultra-wideband signal flight distance information, the relative position information between the mobile terminal and the target vehicle is calculated.
[0026] Furthermore, the step of receiving the output voice data of the target object based on the microphone array corresponding to the target vehicle includes:
[0027] The initial voice data of the target object is received based on the microphone array corresponding to the target vehicle;
[0028] Based on the echo cancellation algorithm, the initial speech data is subjected to environmental noise reduction processing to obtain the output speech data.
[0029] On the other hand, this application provides a vehicle voice command response device, including:
[0030] Receiving module: used to receive the output voice data of the target object based on the microphone array corresponding to the target vehicle, wherein the microphone array includes at least one vehicle microphone and at least one mobile microphone of the target mobile terminal;
[0031] First acquisition module: used to acquire voice positioning and recognition information corresponding to the at least one vehicle microphone and the at least one mobile microphone respectively, wherein the voice positioning and recognition information is recognition information related to audio transmission distance generated when the output voice data is transmitted to the vehicle microphone or to the mobile microphone;
[0032] Processing module: used to perform sound source localization and recognition on the output voice data based on the voice localization and recognition information corresponding to the at least one vehicle-mounted microphone and the at least one mobile microphone, and obtain the sound source localization result;
[0033] Execution module: used to execute the voice command corresponding to the output voice data based on the sound source localization result.
[0034] On the other hand, this application provides a computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or the at least one program is loaded by a processor and executed as described above in the vehicle voice command response method.
[0035] On the other hand, this application provides an electronic device including a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the vehicle voice command response method as described above.
[0036] This application provides a vehicle voice command response method, device, and storage medium, which have the following characteristics:
[0037] Beneficial effects:
[0038] This application receives output voice data from a target object using a microphone array corresponding to the target vehicle. The microphone array includes at least one in-vehicle microphone and at least one mobile microphone of the target mobile terminal. It acquires voice localization and recognition information corresponding to each of the at least one in-vehicle microphone and the at least one mobile microphone. This voice localization and recognition information is recognition information related to the audio transmission distance generated when the output voice data is transmitted to the in-vehicle microphone or the mobile microphone. Based on the voice localization and recognition information corresponding to each of the at least one in-vehicle microphone and the at least one mobile microphone, it performs sound source localization and recognition on the output voice data to obtain a sound source localization result. Based on the sound source localization result, it executes the voice command corresponding to the output voice data. In this way, the sound source can be accurately located without increasing the vehicle's hardware cost, thereby improving the accuracy of the vehicle's voice command response. Attached Figure Description
[0039] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating a vehicle voice command response method provided in an embodiment of this application;
[0041] Figure 2 A flowchart illustrating a method for acquiring output voice data provided in an embodiment of this application;
[0042] Figure 3 A flowchart illustrating a method for determining a target mobile terminal provided in an embodiment of this application;
[0043] Figure 4 A flowchart illustrating a method for obtaining relative position information between a mobile terminal and a target vehicle, provided in an embodiment of this application;
[0044] Figure 5 A flowchart illustrating a method for determining microphone array information corresponding to a microphone array, provided in an embodiment of this application;
[0045] Figure 6 A flowchart illustrating a method for obtaining sound source localization results provided in an embodiment of this application;
[0046] Figure 7 A flowchart illustrating a method for updating microphone array information provided in an embodiment of this application;
[0047] Figure 8 This is a flowchart of a vehicle voice command response method provided in an embodiment of this application;
[0048] Figure 9 A schematic diagram of the structure of a vehicle voice command response device provided in an embodiment of this application;
[0049] Figure 10 This is a hardware structure block diagram of an electronic device for implementing a vehicle voice command response method, provided as an embodiment of this application. Detailed Implementation
[0050] To enable those skilled in the art to better understand the solutions of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0051] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0052] The following combination Figure 1 This application introduces a vehicle voice command response method provided by an embodiment of the present application. The vehicle voice command response method provided by the present application includes:
[0053] S101. Receive the output voice data of the target object based on the microphone array corresponding to the target vehicle. The microphone array includes at least one vehicle-mounted microphone and at least one mobile microphone of the target mobile terminal.
[0054] In some embodiments, the microphone array includes at least two vehicle-mounted microphones and at least one mobile microphone of the target mobile terminal.
[0055] For example, mobile terminals may include, but are not limited to, smartphones, smartwatches, or smart bracelets.
[0056] For example, the target can be the driver, the front passenger, or a passenger in the back seat of the vehicle.
[0057] In this embodiment of the application, please refer to Figure 2 S101 includes:
[0058] S201. Receive the initial voice data of the target object based on the microphone array corresponding to the target vehicle.
[0059] In some embodiments, initial voice data of the target object is collected in real time based on the microphone array corresponding to the target vehicle, wherein the initial voice data represents the audio of the target object in the vehicle environment.
[0060] S202. Based on the echo cancellation algorithm, the initial speech data is subjected to environmental noise reduction processing to obtain the output speech data.
[0061] In some embodiments, the vehicle environment noise in the initial speech data is suppressed based on the echo cancellation algorithm, thereby achieving environmental noise reduction processing of the initial speech data.
[0062] In some embodiments, adaptive filtering is used to simulate the corresponding echo signals transmitted to the vehicle-mounted microphone or mobile microphone in the microphone array from the dynamically acquired initial voice data in real time. The echo signals are then subtracted from the initial voice data to eliminate the self-noise of the vehicle environment, thereby obtaining the output voice data.
[0063] In some embodiments, the signal-to-noise ratio of the output speech data is greater than that of the initial speech data.
[0064] Specifically, echo cancellation is officially called Acoustic Echo Cancellation (AEC). Acoustic echo refers to the collection of echoes produced when sound emitted by a device's own speakers is reflected once or multiple times through different paths before entering the microphone; it can also be called device self-noise. Therefore, when users interact with devices via voice, the echo signal mixes with the clean speech signal, which degrades the signal-to-noise ratio (SNR) of the acquired speech signal and severely interferes with the performance of subsequent signal processing algorithms and wake-up recognition modules. Therefore, an echo cancellation algorithm module is used to process the initial speech data to eliminate environmental self-noise in order to improve the SNR. The main principle of echo cancellation is to use adaptive filtering technology to dynamically track the acoustic channel inside the vehicle in real time. The reference tone is filtered through this channel to simulate the echo reaching the microphone. Finally, this echo signal is subtracted from the initial speech data to eliminate environmental self-noise.
[0065] In this embodiment of the application, please refer to Figure 3 Prior to S101, vehicle voice command response methods also included:
[0066] S301. Obtain the relative position information between the mobile terminal and the target vehicle.
[0067] In some embodiments, the target vehicle includes an ultra-wideband (UWB) based vehicle positioning system, which includes multiple anchor antennas and at least one UWB module.
[0068] Among them, the UWB module is an ultra-wideband module. Ultra-wideband (UWB) technology is a carrier-free communication technology. It does not use a sinusoidal carrier, but instead uses nanosecond-level non-sinusoidal narrow pulses to transmit data. Its signal peaks are steep and narrow, making it easy to identify even in noisy multi-channel environments. Therefore, it can meet the needs of various short-range wireless communications, and is especially suitable for precise positioning in dense multipath environments, such as vehicle unlocking, automatic vehicle start, in-vehicle passenger detection, vehicle-mounted drone operation, automatic valet parking, automatic parking, parking lot entry, and drive-through payment.
[0069] In some embodiments, multiple anchor antennas are grouped and installed in preset installation areas on the target vehicle body, so that the signal areas of the multiple anchor antennas cover preset areas around and / or inside the vehicle; the output of at least one UWB module is time-divisionally connected to two or more anchor antennas located at different positions and / or with different orientations.
[0070] The preset installation area may include, but is not limited to, the front left, front right, rear left, rear right, left side, right side, and roof areas of the vehicle. The specific location and size of the preset installation area can be determined according to different vehicle models or different signal detection requirements.
[0071] Preferably, the signal network areas covered by two adjacent anchor antennas overlap.
[0072] In some embodiments, the mobile terminal can be a mobile terminal capable of communicating with the aforementioned ultra-wideband-based vehicle positioning system via ultra-wideband signals and meeting the communication protocol requirements. For example, the mobile terminal may include, but is not limited to, smartphones, smartwatches, or smart bracelets.
[0073] Furthermore, the anchor antenna can be used to transmit signals to mobile terminals around or inside the target vehicle, thereby locating the mobile terminals around or inside the target vehicle and obtaining the relative position information between the mobile terminals and the target vehicle.
[0074] For details, please see Figure 4 S301 includes:
[0075] S401. Obtain the flight time information of the ultra-wideband signals between the mobile terminal and multiple anchor antennas.
[0076] In some embodiments, the ultra-wideband signal time-of-flight information indicates the corresponding ultra-wideband signal time of flight between the mobile terminal and each of the multiple anchor antennas.
[0077] S402. Calculate the flight distance based on the ultra-wideband signal flight time information and signal flight speed information to obtain the corresponding ultra-wideband signal flight distance information between the mobile terminal and multiple anchor antennas.
[0078] In some embodiments, the signal flight speed information indicates the flight speed of the ultra-wideband signal, and the ultra-wideband signal flight distance information indicates the corresponding ultra-wideband signal flight distance between the mobile terminal and each of the multiple anchor antennas.
[0079] Specifically, the flight time of each ultra-wideband signal is multiplied by its flight speed to obtain the corresponding flight distance.
[0080] S403. Based on the flight distance information of the ultra-wideband signal, the relative position information between the mobile terminal and the target vehicle is calculated.
[0081] Specifically, the relative position information between the mobile terminal and the target vehicle is calculated based on the corresponding ultra-wideband signal flight distance. The algorithm can be the same as the existing ultra-wideband positioning algorithm, and this application does not impose any restrictions on it.
[0082] In some embodiments, the relative position information between the mobile terminal and the target vehicle indicates the relative position between the mobile terminal and the target vehicle. The relative position between the mobile terminal and the target vehicle may be the position of the mobile terminal in the coordinate system pre-stored in the on-board positioning system of the target vehicle.
[0083] S302. If the relative position information indicates that the mobile terminal is located within a preset area of the target vehicle, the mobile terminal is determined to be the target mobile terminal.
[0084] In some embodiments, based on the relative position information between the mobile terminal and the target vehicle, it is determined whether the mobile terminal is within a preset area of the target vehicle. If the relative position information between the mobile terminal and the target vehicle indicates that the mobile terminal is within the preset area of the target vehicle, the mobile terminal is determined to be the target mobile terminal. That is, if the relative position information indicates that the mobile terminal is within the preset area of the target vehicle, the mobile terminal is added to the microphone array as the target mobile terminal to receive the output voice data of the target object.
[0085] In some embodiments, if the relative position information between the mobile terminal and the target vehicle indicates that the mobile terminal is located outside the preset area of the target vehicle, the mobile terminal cannot be identified as the target mobile terminal; that is, if the relative position information indicates that the mobile terminal is located outside the preset area of the target vehicle, the mobile terminal cannot be added to the microphone array as the target mobile terminal, and therefore cannot receive the output voice data of the target object.
[0086] In some embodiments, before S301, it is necessary to pair the mobile terminal with the target vehicle for the first time. Specifically, the pairing of the mobile terminal with the target vehicle can be achieved by the vehicle sending a pairing code and the mobile terminal user entering the pairing code.
[0087] S102. Obtain voice positioning and recognition information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, wherein the voice positioning and recognition information is recognition information related to the audio transmission distance generated when the output voice data is transmitted to the vehicle-mounted microphone or the mobile microphone.
[0088] In some embodiments, the voice location and recognition information may be the data arrival time information corresponding to the output voice data being transmitted to the vehicle-mounted microphone or the mobile microphone, respectively, and the data arrival time information indicates the arrival time corresponding to the output voice data being transmitted to the vehicle-mounted microphone or the mobile microphone.
[0089] S103. Based on the voice localization and recognition information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, perform sound source localization and recognition on the output voice data to obtain the sound source localization result.
[0090] In this embodiment of the application, please refer to Figure 5 Prior to S103, vehicle voice command response methods also included:
[0091] S501. Obtain first location information corresponding to at least one target mobile terminal and second location information corresponding to at least one vehicle microphone.
[0092] In some embodiments, the target mobile terminal is located based on the ultra-wideband vehicle positioning system as described above, to obtain first location information corresponding to at least one target mobile terminal. The first location information corresponding to at least one target mobile terminal can indicate the relative position between at least one target mobile terminal and the target vehicle. The first location information can be the position of at least one target mobile terminal in the coordinate system pre-stored by the vehicle positioning system of the target vehicle.
[0093] In some embodiments, the second position information corresponding to each of the at least one vehicle microphone can indicate the relative position between the at least one vehicle microphone and the target vehicle. The second position information can be the position of each of the at least one vehicle microphone in the coordinate system pre-stored in the vehicle positioning system of the target vehicle.
[0094] S502 determines the microphone array information corresponding to the microphone array based on the first position information and the second position information.
[0095] In some embodiments, microphone array information indicates the relative position between at least one vehicle-mounted microphone and at least one target mobile terminal.
[0096] In this embodiment, a sound source localization method based on time delay difference is used to locate and identify the sound source in the output speech data, thereby obtaining the sound source localization result. For details, please refer to [link to relevant documentation]. Figure 6 S103 includes:
[0097] S601. Based on the voice reception time information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, perform time delay calculation to obtain the time delay difference between the sound source corresponding to the output voice data arriving at at least one vehicle-mounted microphone and at least one mobile microphone.
[0098] In some embodiments, the arrival times of the output voice data transmitted to the vehicle-mounted microphone or the mobile microphone are calculated by differentiating the arrival times to obtain the time delay difference between the sound source and at least one vehicle-mounted microphone and at least one mobile microphone.
[0099] S602. Based on the time delay difference and audio transmission speed information, perform distance calculation to obtain the distance difference corresponding to the transmission distance between the sound source and at least one vehicle-mounted microphone and at least one mobile microphone.
[0100] In some embodiments, audio transmission speed information indicates the speed at which audio travels through the air.
[0101] In some embodiments, the delay difference and transmission speed are multiplied to obtain the distance difference corresponding to the transmission distance between the sound source and at least one vehicle-mounted microphone and at least one mobile microphone.
[0102] S603. Based on the distance difference and microphone array information, perform sound source localization and identification on the output speech data to obtain the sound source localization result.
[0103] In some embodiments, there is a unique coordinate point, i.e., the location of the sound source, in the coordinate system pre-stored in the target vehicle's on-board positioning system, which can satisfy the distance difference corresponding to the transmission distance between the sound source and at least one on-board microphone and at least one mobile microphone.
[0104] Specifically, the location of the sound source is calculated based on the distance difference and microphone array information. The algorithm can be the same as existing sound source localization algorithms, and this application does not impose any restrictions on it.
[0105] In some embodiments, the sound source localization result can indicate the position of the sound source in the coordinate system pre-stored in the target vehicle's on-board positioning system, thereby determining the specific location of the sound source inside the target vehicle. For example, it can be determined that the sound source is located in the passenger seat of the target vehicle.
[0106] In this embodiment of the application, please refer to Figure 7 Prior to S603, vehicle voice command response methods also included:
[0107] S701. Monitor the first location information.
[0108] In some embodiments, the target vehicle may be equipped with a mobile terminal bracket, which is fixedly connected to the fixed components of the target vehicle. The target mobile terminal is fixed by the mobile terminal bracket to achieve a fixed relative position between the target mobile terminal and the target vehicle, thereby obtaining accurate and stable first position information.
[0109] In some embodiments, when the target mobile terminal moves within a preset area of the target vehicle, the relative position between the target mobile terminal and the target vehicle changes, thereby causing a change in the first position information.
[0110] To improve the accuracy of sound source localization results, the target mobile terminal can be repositioned at preset intervals to obtain the latest first location information. The latest first location information is then compared with the previous first location information to determine whether the location of the target mobile terminal has changed.
[0111] S702. Upon detecting an update to the first position information, update the microphone array information based on the second position information and the updated first position information.
[0112] If the latest first position information is different from the previous first position information, update the first position information, and update the microphone array information based on the second position information and the updated first position information.
[0113] S104. Based on the sound source localization result, execute the voice command corresponding to the output voice data.
[0114] In some embodiments, the sound source localization result can indicate the specific location of the sound source within the target vehicle, for example, the driver's seat, front passenger seat, rear right passenger seat, and rear left passenger seat of the target vehicle.
[0115] Based on the specific location of the sound source within the target vehicle, the voice commands corresponding to the output voice data are executed precisely.
[0116] For example, in response to a voice command from a passenger in the front passenger seat of the target vehicle to lower the window, the target vehicle precisely controls the lowering of the front passenger side window.
[0117] For example, if a passenger in the rear right passenger seat issues a voice command to cancel autopilot, for the sake of the vehicle's driving safety, the vehicle may ignore the voice command or ask the passenger to confirm whether to cancel autopilot, thus improving the vehicle's driving safety.
[0118] For details, please see Figure 8 The following is a flowchart of the vehicle voice command response method provided in this application embodiment, combined with specific application scenarios:
[0119] S1. Obtain the flight time information of the ultra-wideband signals between the mobile terminal and multiple anchor antennas.
[0120] S2. Based on the flight time information and flight speed information of the ultra-wideband signal, the flight distance is calculated to obtain the corresponding ultra-wideband signal flight distance between the mobile terminal and multiple anchor antennas.
[0121] S3. Based on the flight distance of the ultra-wideband signal, the relative position information between the mobile terminal and the target vehicle is calculated.
[0122] S4. Determine whether the relative position information indicates that the mobile terminal is located within the preset area of the target vehicle. If the determination result is yes, proceed to step S5; if the determination result is no, proceed to step S1.
[0123] S5. Identify the mobile terminal as the target mobile terminal.
[0124] S6. Receive the output voice data of the target object based on the microphone array corresponding to the target vehicle. The microphone array includes at least one vehicle-mounted microphone and at least one mobile microphone of the target mobile terminal.
[0125] S7. Obtain voice positioning and recognition information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone. The voice positioning and recognition information is recognition information related to the audio transmission distance generated when the output voice data is transmitted to the vehicle-mounted microphone or the mobile microphone. The voice positioning and recognition information includes voice reception time information for the output voice data.
[0126] S8. Obtain first location information corresponding to at least one target mobile terminal and second location information corresponding to at least one vehicle microphone.
[0127] S9. Based on the first position information and the second position information, determine the microphone array information corresponding to the microphone array.
[0128] S10. Based on the voice reception time information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, perform time delay calculation to obtain the time delay difference between the sound source corresponding to the output voice data arriving at at least one vehicle-mounted microphone and at least one mobile microphone.
[0129] S11. Based on the time delay difference and audio transmission speed information, perform distance calculation to obtain the distance difference corresponding to the transmission distance between the sound source and at least one vehicle-mounted microphone and at least one mobile microphone.
[0130] S12. Determine whether the first position information has been updated. If the result is no, proceed to step S13. If the result is yes, proceed to step S8.
[0131] S13. Based on the distance difference and microphone array information, perform sound source localization and identification on the output speech data to obtain the sound source localization result.
[0132] S14. Based on the sound source localization result, execute the voice command corresponding to the output voice data.
[0133] This application receives output voice data from a target object using a microphone array corresponding to the target vehicle. The microphone array includes at least one in-vehicle microphone and at least one mobile microphone of the target mobile terminal. It acquires voice localization and recognition information corresponding to each of the at least one in-vehicle microphone and the at least one mobile microphone. This voice localization and recognition information is recognition information related to the audio transmission distance generated when the output voice data is transmitted to the in-vehicle microphone or the mobile microphone. Based on the voice localization and recognition information corresponding to each of the at least one in-vehicle microphone and the at least one mobile microphone, it performs sound source localization and recognition on the output voice data to obtain a sound source localization result. Based on the sound source localization result, it executes the voice command corresponding to the output voice data. In this way, the sound source can be accurately located without increasing the vehicle's hardware cost, thereby improving the accuracy of the vehicle's voice command response.
[0134] This application also provides a vehicle voice command response device. Please refer to [link to relevant documentation]. Figure 9 The vehicle voice command response device provided in this application includes:
[0135] Receiving module 910: used to receive the output voice data of the target object based on the microphone array corresponding to the target vehicle, the microphone array including at least one vehicle microphone and at least one mobile microphone of the target mobile terminal.
[0136] The first acquisition module 920 is used to acquire voice positioning and recognition information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone. The voice positioning and recognition information is recognition information related to the audio transmission distance generated when the output voice data is transmitted to the vehicle-mounted microphone or the mobile microphone.
[0137] Processing module 930: Used to perform sound source localization and recognition on the output voice data based on the voice localization and recognition information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, and obtain the sound source localization result.
[0138] Execution module 940: Used to execute voice commands corresponding to the output voice data based on the sound source localization results.
[0139] The vehicle voice command response device provided in this application embodiment also includes:
[0140] The second acquisition module is used to acquire the relative position information between the mobile terminal and the target vehicle before receiving the output voice data of the target object based on the microphone array corresponding to the target vehicle.
[0141] First determining module: used to determine the mobile terminal as the target mobile terminal when the relative position information indicates that the mobile terminal is located within a preset area of the target vehicle.
[0142] The vehicle voice command response device provided in this application embodiment also includes:
[0143] The third acquisition module is used to acquire first location information corresponding to at least one target mobile terminal and second location information corresponding to at least one vehicle-mounted microphone before performing sound source localization and recognition on the output voice data based on the voice localization and recognition information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, and obtaining the sound source localization result.
[0144] The second determining module is used to determine the microphone array information corresponding to the microphone array based on the first position information and the second position information.
[0145] In this embodiment of the application, the processing module 930 includes:
[0146] First processing unit: used to perform time delay calculation based on the voice reception time information corresponding to at least one vehicle-mounted microphone and at least one mobile microphone, to obtain the time delay difference between the sound source corresponding to the output voice data arriving at at least one vehicle-mounted microphone and at least one mobile microphone.
[0147] The second processing unit is used to calculate the distance based on the time delay difference and audio transmission speed information, and obtain the distance difference corresponding to the transmission distance between the sound source and at least one vehicle-mounted microphone and at least one mobile microphone.
[0148] The third processing unit is used to perform sound source localization and identification on the output speech data based on the distance difference and microphone array information, and obtain the sound source localization result.
[0149] The vehicle voice command response device provided in this application embodiment also includes:
[0150] Monitoring module: Used to monitor the first position information before obtaining the sound source localization result by performing sound source localization and identification on the output voice data based on distance difference and microphone array information.
[0151] Update module: Used to update microphone array information based on second location information and updated first location information when an update to the first location information is detected.
[0152] In this embodiment of the application, the second acquisition module includes:
[0153] The first acquisition unit is used to acquire the flight time information of the ultra-wideband signals between the mobile terminal and multiple anchor antennas.
[0154] The fourth processing unit is used to calculate the flight distance based on the ultra-wideband signal flight time information and signal flight speed information, so as to obtain the ultra-wideband signal flight distance information between the mobile terminal and each of the multiple anchor antennas.
[0155] The fifth processing unit is used to calculate the relative position information between the mobile terminal and the target vehicle based on the ultra-wideband signal flight distance information.
[0156] In this embodiment of the application, the receiving module 910 includes:
[0157] Receiving unit: Used to receive the initial voice data of the target object based on the microphone array corresponding to the target vehicle.
[0158] The sixth processing unit is used to perform environmental noise reduction processing on the initial speech data based on the echo cancellation algorithm to obtain the output speech data.
[0159] The apparatus and method embodiments described above are based on the same application concept.
[0160] Please refer to Figure 10 This application provides an electronic device for implementing the above-described vehicle voice command response method. The electronic device includes a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or at least one program is loaded and executed by the processor to implement the vehicle voice command response method provided in the above-described method embodiments.
[0161] Memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. Memory can primarily include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for the functions, etc.; the data storage area can store data created based on the use of the device, etc. Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.
[0162] The methods and embodiments provided in this application can be executed in mobile terminals, computer terminals, servers, or similar computing devices. That is, the aforementioned electronic devices may include mobile terminals, computer terminals, servers, or similar computing devices. The server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these.
[0163] Figure 10 This is a hardware structure block diagram of an electronic device for implementing the above-described vehicle voice command response method, provided in an embodiment of this application. For example... Figure 10 As shown, the electronic device 1000 can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) 1010 (CPUs 1010 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 1030 for storing data, and one or more storage media 1020 (e.g., one or more mass storage devices) for storing application programs 1023 or data 1022. The memory 1030 and storage media 1020 may be temporary or persistent storage. The program stored in the storage media 1020 may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the CPU 1010 may be configured to communicate with the storage media 1020 and execute the series of instruction operations in the storage media 1020 on the electronic device 1000. Electronic device 1000 may also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1040, and / or one or more operating systems 1021, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0164] The processor 1010 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0165] The input / output interface 1040 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 1000. In one example, the input / output interface 1040 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 1040 may be a radio frequency (RF) module for wireless communication with the Internet.
[0166] The operating system 1021 may include system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks.
[0167] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 1000 may also include... Figure 10 The more or fewer components shown, or having the same Figure 10 The different configurations shown.
[0168] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a vehicle voice command response method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the vehicle voice command response method provided in the above method embodiment.
[0169] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0170] Embodiments of this application also provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.
[0171] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0172] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0173] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0174] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for responding to vehicle voice commands, characterized in that, include: The target vehicle receives output voice data based on a microphone array corresponding to the target vehicle. The microphone array includes at least one vehicle-mounted microphone and at least one mobile microphone of the target mobile terminal. Acquire voice positioning and recognition information corresponding to each of the at least one vehicle-mounted microphone and the at least one mobile microphone. The voice positioning and recognition information is recognition information related to audio transmission distance generated when the output voice data is transmitted to the vehicle-mounted microphone or the mobile microphone. The voice positioning and recognition information includes voice reception time information for the output voice data. Acquire the first location information corresponding to each of the at least one target mobile terminal and the second location information corresponding to each of the at least one vehicle microphone; Based on the first location information and the second location information, determine the microphone array information corresponding to the microphone array; Based on the voice reception time information corresponding to each of the at least one vehicle-mounted microphone and the at least one mobile microphone, a time delay calculation is performed to obtain the time delay difference between the sound source corresponding to the output voice data and the at least one vehicle-mounted microphone and the at least one mobile microphone; based on the time delay difference and audio transmission speed information, a distance calculation is performed to obtain the distance difference corresponding to the transmission distance between the sound source and the at least one vehicle-mounted microphone and the at least one mobile microphone; based on the distance difference and the microphone array information, the output voice data is used for sound source localization and identification to obtain the sound source localization result; Based on the sound source localization result, execute the voice command corresponding to the output voice data.
2. The vehicle voice command response method according to claim 1, characterized in that, Before receiving the output voice data of the target object based on the microphone array corresponding to the target vehicle, the method further includes: Obtain the relative position information between the mobile terminal and the target vehicle; If the relative position information indicates that the mobile terminal is located within a preset area of the target vehicle, the mobile terminal is determined to be the target mobile terminal.
3. The vehicle voice command response method according to claim 1, characterized in that, Before obtaining the sound source localization result by performing sound source localization and identification on the output speech data based on the distance difference and the microphone array information, the method further includes: Monitor the first location information; Upon detecting an update to the first location information, the microphone array information is updated based on the second location information and the updated first location information.
4. The vehicle voice command response method according to claim 2, characterized in that, The target vehicle includes multiple anchor antennas, and acquiring the relative position information between the mobile terminal and the target vehicle includes: Obtain the ultra-wideband signal flight time information corresponding to each of the mobile terminal and the plurality of anchor antennas; Based on the ultra-wideband signal flight time information and signal flight speed information, the flight distance is calculated to obtain the ultra-wideband signal flight distance information between the mobile terminal and the multiple anchor antennas respectively. Based on the ultra-wideband signal flight distance information, the relative position information between the mobile terminal and the target vehicle is calculated.
5. The vehicle voice command response method according to claim 1, characterized in that, The method of receiving the output voice data of the target object based on the microphone array corresponding to the target vehicle includes: The initial voice data of the target object is received based on the microphone array corresponding to the target vehicle; Based on the echo cancellation algorithm, the initial speech data is subjected to environmental noise reduction processing to obtain the output speech data.
6. A vehicle voice command response device, characterized in that, include: Receiving module: used to receive the output voice data of the target object based on the microphone array corresponding to the target vehicle, wherein the microphone array includes at least one vehicle microphone and at least one mobile microphone of the target mobile terminal; First acquisition module: used to acquire voice positioning and recognition information corresponding to the at least one vehicle microphone and the at least one mobile microphone respectively. The voice positioning and recognition information is recognition information related to audio transmission distance generated when the output voice data is transmitted to the vehicle microphone or to the mobile microphone. The voice positioning and recognition information includes voice reception time information for the output voice data. Processing module: used to acquire the first location information corresponding to each of the at least one target mobile terminal and the second location information corresponding to each of the at least one vehicle microphone; Based on the first location information and the second location information, determine the microphone array information corresponding to the microphone array; Based on the voice reception time information corresponding to each of the at least one vehicle-mounted microphone and the at least one mobile microphone, a time delay calculation is performed to obtain the time delay difference between the sound source corresponding to the output voice data and the at least one vehicle-mounted microphone and the at least one mobile microphone; based on the time delay difference and audio transmission speed information, a distance calculation is performed to obtain the distance difference corresponding to the transmission distance between the sound source and the at least one vehicle-mounted microphone and the at least one mobile microphone; based on the distance difference and the microphone array information, the output voice data is used for sound source localization and identification to obtain the sound source localization result; Execution module: used to execute the voice command corresponding to the output voice data based on the sound source localization result.
7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor according to any one of claims 1 to 5.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the vehicle voice command response method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for realizing sound source localization by utilizing mobile terminal
CN104422922A
A mobile terminal and a voice input method and apparatus thereof
CN106782589A