Vehicle voice interaction wake-up method, device, equipment and storage medium
By using microphone arrays inside and outside the vehicle to analyze voiceprint features and sound source location in new energy vehicles, the problem that voice wake-up systems cannot distinguish between sound sources inside and outside the vehicle has been solved, achieving higher accuracy and security.
Patent Information
- Application Number
- CN202411071640.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2044-08-06
AI Technical Summary
Existing voice wake-up systems for new energy vehicles struggle to accurately distinguish between sound sources inside and outside the vehicle, potentially leading to erroneous execution of vehicle control commands and impacting safety and normal operation.
The system uses two microphone arrays, one inside and one outside the vehicle, to collect sound signals. By analyzing voiceprint characteristics and the location of the sound source, it ensures that voice control commands are only executed when the sound source is inside the vehicle; otherwise, the interaction is terminated.
It improves the accuracy and response speed of the voice interaction system, ensures the safety and accuracy of vehicle control, and avoids interference from external noise.
Smart Images

Figure CN118865973B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction technology, and in particular to a method, apparatus, device, and storage medium for waking up a vehicle's voice interaction system. Background Technology
[0002] In the intelligent configuration of new energy vehicles, voice wake-up functionality has received widespread attention due to its ability to reduce driver reliance on physical controls while driving. Through simple voice commands, drivers can safely control various vehicle functions, such as navigation, music playback, and answering phone calls. However, the widespread adoption of voice wake-up functionality has also brought some new challenges. The complexity of the in-vehicle environment and the uncertainty of the external environment place higher demands on the accuracy and safety of voice recognition.
[0003] Existing voice wake-up functions have limitations in distinguishing between sound sources inside and outside the vehicle. If a voiceprint similar to that of a passenger inside the vehicle is played outside, the existing system may not be able to effectively identify and distinguish the true location of this sound source, potentially leading to the incorrect execution of vehicle control commands. When this happens, it may not only affect the normal use of the vehicle but also pose a threat to passenger safety.
[0004] Therefore, improving the accuracy of sound source localization for voice wake-up functions, ensuring that the system only responds to commands from passengers inside the vehicle, and avoiding interference from external noise or similar voiceprints have become urgent technical challenges in the intelligent development of new energy vehicles. Summary of the Invention
[0005] The main objective of this invention is to provide a vehicle voice interaction wake-up method, device, equipment, and storage medium, aiming to solve the technical problem of how to ensure that the voice wake-up function of new energy vehicles only responds to the commands of passengers inside the vehicle and avoids interference from similar voiceprints outside the vehicle.
[0006] To achieve the above objectives, the present invention provides a vehicle voice interaction wake-up method, the vehicle voice interaction wake-up method comprising:
[0007] Acquire voice information from the in-vehicle ambient sound that conforms to preset acoustic characteristics;
[0008] The location of the sound source of the voice information is determined based on the microphone array at preset positions;
[0009] When the sound source is located inside the vehicle, the vehicle control command in the voice information is executed.
[0010] Optionally, acquiring speech information in the sound signal that conforms to preset acoustic features includes:
[0011] Collect sound signals from the in-vehicle environment;
[0012] When the loudness of the sound signal is greater than the voice control threshold, the voice information in the sound signal is extracted;
[0013] Obtain the voiceprint features of the speech information;
[0014] When the voiceprint feature belongs to the preset wake-up voiceprint feature, it is determined that the current voice information conforms to the preset acoustic feature.
[0015] Optionally, determining the sound source location of the voice information based on a microphone array at preset positions includes:
[0016] The estimated location of the voice information is obtained based on the first microphone array;
[0017] Based on the second microphone array, the sound source pointing matrix of the voice information is obtained;
[0018] The location of the sound source is determined based on the estimated location of the speech information and the sound source pointing matrix.
[0019] Optionally, obtaining the sound source direction and sound source intensity of the voice information based on the first microphone array includes:
[0020] The sound source direction and sound source intensity of the first microphone array relative to the voice information are obtained through the first microphone array.
[0021] Based on the direction and intensity of the sound source, the estimated location of the speech information in space is obtained.
[0022] Optionally, obtaining the sound source pointing matrix of the voice information based on the second microphone array includes:
[0023] The second microphone array is used to obtain the sound source intensity of the voice information at each array unit, wherein all array units of the second microphone array are located outside the vehicle.
[0024] The sound source pointing matrix of the speech information is obtained based on the spatial position of the array unit and the sound source intensity.
[0025] Optionally, determining the sound source location of the speech information based on the estimated location of the speech information and the sound source pointing matrix includes:
[0026] Based on the sound source pointing matrix, the spatial distribution area of the sound source of the speech information is determined;
[0027] If the estimated location falls within the spatial distribution area of the sound source, then it is determined that the sound source of the voice information is located inside the vehicle.
[0028] Optionally, after determining the spatial distribution region of the speech information based on the sound source pointing matrix, the method further includes:
[0029] If the estimated location does not fall within the spatial distribution area of the sound source, the sound source intensity of each array unit in the second microphone array is compared with the sound source intensity of the first microphone array to obtain the sound source intensity comparison result.
[0030] When the sound source intensity of any array unit in the second microphone array is greater than the sound source intensity of the first microphone array, it is determined that the sound source of the voice information is located in the space outside the vehicle.
[0031] If the sound source intensity of all array units in the second microphone array is less than the sound source intensity of the first microphone array, then it is determined that the sound source of the voice information is located in the vehicle interior space.
[0032] Furthermore, to achieve the above objectives, the present invention provides a vehicle voice interaction wake-up device, the vehicle voice interaction wake-up device comprising:
[0033] The data acquisition module is used to acquire voice information that conforms to preset acoustic characteristics in the in-vehicle ambient sound.
[0034] The data processing module is used to determine the sound source location of the voice information based on the microphone array at a preset position;
[0035] The control module is used to execute vehicle control commands in the voice information when the sound source is located inside the vehicle.
[0036] Furthermore, to achieve the above objectives, the present invention provides a vehicle voice interaction wake-up device, the vehicle voice interaction wake-up device comprising: a memory, a processor, and a vehicle voice interaction wake-up program stored in the memory and executable on the processor, the vehicle voice interaction wake-up program being configured to implement the steps of the vehicle voice interaction wake-up method.
[0037] Furthermore, to achieve the above objectives, the present invention provides a storage medium storing a vehicle voice interaction wake-up program, wherein the vehicle voice interaction wake-up program, when executed by a processor, implements the steps of the vehicle voice interaction wake-up method.
[0038] This invention first collects sound signals using two microphone arrays, one inside and one outside the vehicle. When ambient noise exceeds a preset wake-up threshold, the voiceprint features of the speech information are extracted and analyzed, and matched with a preset wake-up voiceprint. If a match is successful, the spoken voice is extracted. Then, the two microphone arrays, located at different spatial positions, collect speech information and analyze the precise location of the acoustic source. If the sound source is confirmed to be inside the vehicle, the corresponding control command within the voice message is executed; if the sound source is determined not to be inside the vehicle, the voice interaction is terminated. This method effectively improves the accuracy and response speed of the vehicle voice interaction system while ensuring the safety of vehicle voice interaction control. Attached Figure Description
[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the first embodiment of the vehicle voice interaction wake-up method of this application;
[0042] Figure 2 This is a microphone array distribution diagram of the first embodiment of the vehicle voice interaction wake-up method of this application;
[0043] Figure 3 This is a flowchart illustrating the second embodiment of the vehicle voice interaction wake-up method of this application;
[0044] Figure 4 This is a schematic diagram of the sound source pointing in the second embodiment of the vehicle voice interaction wake-up method of this application;
[0045] Figure 5 This is a schematic diagram of the functional modules of the vehicle voice interaction wake-up device of this application;
[0046] Figure 6 This is a schematic diagram of the structure of the terminal device in the hardware operating environment involved in the embodiments of this application.
[0047] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0048] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0049] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0050] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a vehicle voice interaction wake-up device. The following description uses a vehicle voice interaction wake-up device as an example to illustrate this embodiment and the subsequent embodiments.
[0051] This application provides a method for waking up a vehicle's voice interaction, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of this application.
[0052] In this embodiment, the vehicle voice interaction wake-up method includes:
[0053] Step S10: Acquire voice information in the in-vehicle ambient sound that conforms to preset acoustic characteristics.
[0054] It should be noted that certain conditions must be met before voice control can be performed in the car. On the one hand, the voice source of the voice control command must be inside the car, and on the other hand, the voiceprint of the voice command must match the voiceprint information of the wake-up person reserved in the vehicle system. Only when both of these conditions are met will the system consider the received voice information to be valid.
[0055] In one embodiment, the sound signal in the in-vehicle environment is collected; when the loudness of the sound signal is greater than the voice control threshold, the voice information in the sound signal is extracted; the voiceprint features of the voice information are obtained; when the voiceprint features belong to the preset wake-up voiceprint features, it is determined that the current voice information conforms to the preset acoustic features.
[0056] It should be noted that voice control requires a certain amount of received voice information to be activated. Generally speaking, if the noise level inside and outside the car is high, the system may need to set a high loudness threshold to ensure that only sufficiently loud voices can activate the system. Secondly, considering the different speaking habits and volume levels of different users, the system may need to adaptively adjust the threshold for different people who need to be activated. That is, each voiceprint feature corresponds to a specific voice control threshold, which can be set manually or automatically generated by the system.
[0057] Understandably, because everyone's voiceprint features are unique, including parameters such as voice characteristics, pitch, and timbre, just like a fingerprint, the in-vehicle voice interaction system authenticates the identity of the voice source by analyzing these acoustic features. It can extract key parameters from the voice signal and compare them with known voiceprint templates to identify the speaker.
[0058] It should be understood that when multiple voice signals with different voiceprints are present simultaneously, it is necessary to determine the sound source location and voiceprint characteristics of each corresponding voice signal one by one, check whether they meet the prerequisites for voice control, and then execute the corresponding vehicle control commands. To provide a better user experience, the system provides clear feedback information through the user interface, such as the status of voice commands being received, recognized, or executed, and requests for user confirmation or repetition of commands when they cannot be recognized or there is doubt.
[0059] Step S20: Determine the sound source location of the voice information based on the microphone array at the preset position.
[0060] It should be noted that if the same or similar voiceprint is played outside the vehicle as the voiceprint of the passenger inside, the existing system may not be able to effectively identify and distinguish the true location of the voice source, which may lead to the incorrect execution of vehicle control commands and potentially have adverse consequences.
[0061] It is understood that the microphone array in this embodiment is configured as follows: Figure 2 As shown, the microphone array consists of two parts: one part is located at the four corners outside the vehicle near the driver's seat, and the other part is located near the gear shift console inside the vehicle. By placing microphones at the four corners outside the vehicle near the driver's seat, the system can capture sound information from outside the vehicle. This helps to identify and locate sound sources outside the vehicle, thereby distinguishing between voice commands from outside and inside the vehicle. Placing the other part of the microphones near the gear shift console inside the vehicle can more accurately capture voice commands from passengers inside the vehicle, especially near the driver. This is crucial for improving the responsiveness and accuracy of voice control during driving.
[0062] It should be understood that the closed state of car windows directly affects the sound pickup performance of each microphone array. For example, when the windows are closed, the glass provides an additional physical barrier for sound propagation. By comparing the actual received voice volume, it's possible to determine whether the sound source is inside or outside the vehicle. Furthermore, in most cases, even if some or all windows are open, the sound pickup performance from external sound sources will still have a specific impact. For instance, when a window in a certain direction is open, if the sound source is inside the vehicle, the microphone at that location will pick up a different volume of the voice signal. The volume of the voice signal collected by the microphone inside the vehicle should be increased, while the volume of the sound source outside the vehicle should increase slightly but not significantly. When the sound source is outside the vehicle, the situation should be exactly the opposite. On the one hand, the volume of the various collection units outside the vehicle should not change much, while the volume of the sound source collected inside the vehicle towards the open window should be greater than that of other directions. Therefore, when the vehicle is driving in different window states and different environments, the voice interaction system has a high degree of dynamic adaptability, can monitor changes in the window state in real time, and quickly adjust its algorithm parameters to adapt to the new acoustic environment.
[0063] Understandably, assuming no physical obstructions, when the sound source location is fixed, the intensity (or loudness) of the sound wave is inversely proportional to the square of the distance between the source and the receiver. This means that when the distance doubles, the intensity of the sound wave decreases to one-quarter of its original value. The specific formula is: Where I is the intensity of the sound wave at a distance r, and P is the total power of the sound source. For a sound source at a fixed location, each acquisition unit can determine the approximate distance based on the loudness of the acquired sound. Considering that the power of human speech will change at different times depending on the content of the speech, but at the same moment, the distance ratio obtained by each acquisition unit should remain constant. Considering the existence of errors, with continuous distance correction, the position of the sound source on the plane should form a small planar range.
[0064] Step S30: When the sound source is located inside the vehicle, execute the vehicle control command in the voice information.
[0065] It should be noted that once the system has confirmed the presence of a sound source in the vehicle's interior through some means and trusts the instructions issued by that sound source, it will execute the vehicle control instructions contained in the voice message.
[0066] Understandably, if the person waking up is having a normal conversation with other passengers in the car, which may meet the prerequisite for executing a control command, but there is no specific control command in the voice message, the system may not perform any operation, but will only remain in a wake-up state. The system may continue to listen and wait for the user to issue a clear control command. If no specific command is received within a certain period of time, the system may automatically return to a sleep or standby state.
[0067] It should be understood that, to avoid the inconvenience or potential risks caused by the above situations, the voice recognition component of the vehicle control system should possess high accuracy and intelligence, capable of distinguishing between normal conversation and actual control commands, and taking appropriate response measures. At the same time, the system design should consider user experience, ensuring simple and intuitive operation.
[0068] This embodiment provides a vehicle voice interaction wake-up method. This embodiment acquires voice information that conforms to preset acoustic characteristics in the in-vehicle ambient sound; determines the sound source location of the voice information according to a microphone array at a preset location; and executes the vehicle control command in the voice information when the sound source location is in the in-vehicle space.
[0069] In summary, this application first collects sound signals using two microphone arrays, one inside and one outside the vehicle. When ambient noise exceeds a preset wake-up threshold, the voiceprint features of the speech information are extracted and analyzed, and matched with a preset wake-up voiceprint. If a match is successful, the spoken voice is extracted. Then, the two microphone arrays, located at different spatial positions, collect speech information and analyze the precise location of the acoustic source. If the sound source is confirmed to be inside the vehicle, the corresponding control command within the voice message is executed; if the sound source is determined not to be inside the vehicle, the voice interaction is terminated. This method effectively improves the accuracy and response speed of the vehicle voice interaction system while ensuring the safety of vehicle voice interaction control.
[0070] Reference Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the vehicle voice interaction wake-up method of this application. Based on the first embodiment described above, a second embodiment of the vehicle voice interaction wake-up method of this application is proposed.
[0071] In this embodiment, step S20 includes:
[0072] Step S201: Obtain the estimated location of the voice information based on the first microphone array.
[0073] It should be noted that the first microphone array refers to the microphone array inside the vehicle. Since the inside of the vehicle is a relatively closed system, echoes will be generated inside the vehicle. In order to accurately detect the location of the sound source, the first microphone array in this embodiment is installed in the center of the roof. Another common choice is to install it below the windshield. Both can reduce the direct reflection of sound and reduce the impact of echoes, so as to accurately locate the sound source.
[0074] In some embodiments, obtaining the sound source direction and sound source intensity of the speech information based on the first microphone array includes: obtaining the sound source direction and sound source intensity of the first microphone array relative to the speech information through the first microphone array; and obtaining the estimated position of the speech information in space based on the sound source direction and sound source intensity.
[0075] Understandably, microphone arrays estimate the direction of a sound source by detecting the time and phase differences in the arrival of the speech signal at different pickup units within the array. By analyzing the signal strength received by different microphone units, the intensity of the sound source can be assessed; generally, microphones with stronger signals are closer to the sound source. Combining the direction and intensity information of the sound source, the approximate location of the sound source inside the vehicle can be predicted. Assume a car has an array containing four microphones (M1, M2, M3, M4) mounted on the ceiling. M1 and M2 are horizontally arranged at the front of the vehicle, while M3 and M4 are horizontally arranged at the rear. The driver issues a voice command from the driver's seat. Since the driver is at the front, the driver's voice reaches M1 first, and then M2. The microphone array's signal processing system detects that M1 received the sound earlier than M2, indicating that the sound source is in the direction of M1. The system calculates the time difference between the arrival of the sound at M1 and M2 and uses triangulation to determine the approximate direction of the sound source towards the front of the vehicle, i.e., the driver's seat. Meanwhile, the system analysis showed that the signal strength received by M1 was stronger than that of M2, M3, and M4, which further confirmed the close proximity of the sound source.
[0076] Step S202: Obtain the sound source pointing matrix of the voice information based on the second microphone array.
[0077] It should be noted that, as Figure 4 As shown, Figure 4 This is a schematic diagram of the sound source direction in this embodiment, where the white squares represent microphone arrays or microphone units. It should be noted that since the various units of the microphone array inside the vehicle are relatively close together, they are considered as a whole for now. Figure 4The white box in the center is the second microphone array. The distances between the sound source and each acquisition unit are a, b, c, d, and E, respectively. As shown in the figure, the location of the sound source is not a specific point in the plane space, but a small area. This is because there are some factors in the process of acquiring and processing the signal that prevent us from accurately locating the sound source at a single point in the plane space. Instead, we need to treat it as a small area.
[0078] In some embodiments, obtaining the sound source pointing matrix of the voice information based on the second microphone array includes: acquiring the sound source intensity of the voice information at each array unit through the second microphone array, wherein all array units of the second microphone array are located outside the vehicle; and obtaining the sound source pointing matrix of the voice information based on the spatial position and sound source intensity of the array units.
[0079] It should be understood that by analyzing the spatial position and corresponding sound source intensity of each unit in the second microphone array, a multi-dimensional sound source directivity vector can be constructed for each acquisition unit outside the vehicle, which is a combination of the sound source direction vector, the sound source intensity vector, and the sound source arrival time vector. The sound source directivity vectors of each unit are then combined into a sound source directivity matrix, which is related to the direction and intensity of the sound source and the geometric layout of the microphone array.
[0080] Step S203: Determine the sound source location of the speech information based on the estimated location of the speech information and the sound source pointing matrix.
[0081] In some embodiments, determining the sound source location of the voice information based on the estimated location of the voice information and the sound source pointing matrix includes: determining the sound source spatial distribution area of the voice information based on the sound source pointing matrix; if the estimated location falls within the sound source spatial distribution area, then it is determined that the sound source location of the voice information is located in the vehicle interior space.
[0082] In some embodiments, determining the sound source location of the voice information based on the estimated location of the voice information and the sound source pointing matrix further includes: if the estimated location does not fall within the sound source spatial distribution area, comparing the sound source intensity of each array unit in the second microphone array with the sound source intensity of the first microphone array to obtain a sound source intensity comparison result; when the sound source intensity of any array unit in the second microphone array is greater than the sound source intensity of the first microphone array, determining that the sound source location of the voice information is located in the external space of the vehicle; when the sound source intensity of all array units in the second microphone array is less than the sound source intensity of the first microphone array, determining that the sound source location of the voice information is located in the internal space of the vehicle.
[0083] It should be noted that the sampling environments of the microphone arrays outside and inside the vehicle differ. Specifically, external sounds are more affected by external factors such as buildings and trees, which may reflect or absorb sound, while sound propagation inside the vehicle is more direct and controllable. Therefore, in the actual data processing, each uses its own acoustic processing algorithm. Consequently, the microphone arrays inside and outside the vehicle will calculate the range of the sound source separately. In other words, there will be two results simultaneously: the predicted location of the sound source by the first microphone array and the spatial distribution area of the sound source by the second microphone array.
[0084] It should be understood that when two location estimation systems point to the same result, it means that the current sound source localization is accurate. At the same time, due to the special nature of the in-vehicle microphone unit localization, if the in-vehicle microphone unit localization result is that the sound source is outside the vehicle, it has the power to veto and directly terminate the current voice service. Therefore, when the estimated location falls within the sound source spatial distribution area, it means that the sound source is located in the in-vehicle space.
[0085] It should be noted that if the sound source intensity of at least one unit in the second microphone array is greater than that of the first microphone array, the system will determine that the sound source of the voice information is located outside the vehicle. If the sound source intensity of all array units in the second microphone array is less than or equal to that of the first microphone array, the system will determine that the sound source of the voice information is located inside the vehicle. This means that sound sources inside the vehicle have a more significant impact on the first microphone array.
[0086] This embodiment provides a vehicle voice interaction wake-up method. In this embodiment, the estimated position of the voice information is obtained based on a first microphone array; the sound source pointing matrix of the voice information is obtained based on a second microphone array; and the sound source position of the voice information is determined based on the estimated position of the voice information and the sound source pointing matrix.
[0087] In summary, by working collaboratively with two microphone arrays, one inside and one outside the vehicle, the location of the voice source can be accurately determined. The in-vehicle microphone array utilizes the sound propagation characteristics within a closed environment to predict the sound source location using sound source direction and intensity information, while the outside microphone array constructs a sound source pointing matrix, combining spatial location and sound source intensity for sound source localization. When the results from the two systems are consistent, the sound source location is confirmed; if there are discrepancies, the final judgment is made by comparing sound source intensities. This scheme fully considers the differences between the in-vehicle and out-of-vehicle environments and employs adaptive acoustic processing algorithms, improving the accuracy of sound source localization and the robustness of the system.
[0088] Reference Figure 5 This application also provides a vehicle voice interaction wake-up device, the vehicle voice interaction wake-up device comprising:
[0089] Data acquisition module 10 is used to acquire voice information that conforms to preset acoustic characteristics in the in-vehicle ambient sound.
[0090] Data processing module 20 is used to determine the sound source location of the voice information based on the microphone array at a preset position;
[0091] The control module 30 is used to execute the vehicle control commands in the voice information when the sound source is located in the vehicle interior space.
[0092] In one embodiment, the data acquisition module 10 is further configured to acquire sound signals in the in-vehicle environment; when the loudness of the sound signal is greater than the voice control threshold, extract the voice information in the sound signal; obtain the voiceprint features of the voice information; and when the voiceprint features belong to the preset wake-up voiceprint features, determine that the current voice information conforms to the preset acoustic features.
[0093] In one embodiment, the data processing module 20 is further configured to obtain the estimated position of the voice information based on the first microphone array; obtain the sound source pointing matrix of the voice information based on the second microphone array; and determine the sound source position of the voice information based on the estimated position of the voice information and the sound source pointing matrix.
[0094] In one embodiment, the data processing module 20 is further configured to obtain the sound source direction and sound source intensity of the first microphone array relative to the voice information through the first microphone array; and obtain the estimated position of the voice information in space based on the sound source direction and sound source intensity.
[0095] In one embodiment, the data processing module 20 is further configured to acquire the sound source intensity of the voice information at each array unit through the second microphone array, wherein all array units of the second microphone array are located outside the vehicle; and to obtain the sound source pointing matrix of the voice information based on the spatial position and sound source intensity of the array units.
[0096] In one embodiment, the data processing module 20 is further configured to determine the spatial distribution area of the sound source of the voice information based on the sound source pointing matrix; if the estimated position falls within the spatial distribution area of the sound source, then it is determined that the sound source position of the voice information is located in the vehicle interior space.
[0097] In one embodiment, the data processing module 20 is further configured to, if the estimated location does not fall within the sound source spatial distribution area, compare the sound source intensity of each array unit in the second microphone array with the sound source intensity of the first microphone array to obtain a sound source intensity comparison result; when the sound source intensity of any array unit in the second microphone array is greater than the sound source intensity of the first microphone array, determine that the sound source location of the voice information is located in the external space of the vehicle; when the sound source intensity of all array units in the second microphone array is less than the sound source intensity of the first microphone array, determine that the sound source location of the voice information is located in the internal space of the vehicle.
[0098] This application also provides a vehicle voice interaction wake-up device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the vehicle voice interaction wake-up method in the first embodiment described above.
[0099] The following is for reference. Figure 6 The diagram illustrates a structural schematic suitable for implementing a vehicle voice interaction wake-up device according to embodiments of this application. The vehicle voice interaction wake-up device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The vehicle voice interaction wake-up device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0100] like Figure 6As shown, the vehicle voice interaction wake-up device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the vehicle voice interaction wake-up device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the vehicle voice interaction wake-up device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show vehicle voice interaction wake-up devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0101] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0102] The vehicle voice interaction wake-up device provided in this application, employing the vehicle voice interaction wake-up method in the above embodiments, can solve the technical problem in the art of how to ensure that the voice wake-up function of new energy vehicles only responds to the commands of passengers inside the vehicle and avoids interference from similar voiceprints outside the vehicle. Compared with the prior art, the beneficial effects of the vehicle voice interaction wake-up device provided in this application are the same as the beneficial effects of the vehicle voice interaction wake-up method provided in the above embodiments, and other technical features in this vehicle voice interaction wake-up device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0103] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0105] This application also provides a storage medium storing a vehicle voice interaction wake-up program, which, when executed by a processor, implements the steps of the vehicle voice interaction wake-up method described in any one of the above descriptions.
[0106] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0107] The aforementioned storage medium may be included in the vehicle's voice interaction wake-up device; or it may exist independently and not be installed in the vehicle's voice interaction wake-up device.
[0108] The aforementioned computer-readable storage medium carries one or more programs. When the one or more programs are executed by the vehicle voice interaction wake-up device, the vehicle voice interaction wake-up device: acquires voice information in the in-vehicle ambient sound that conforms to preset acoustic characteristics; determines the sound source location of the voice information according to a microphone array at a preset location; and executes the vehicle control commands in the voice information when the sound source location is in the in-vehicle space.
[0109] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0111] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0112] The storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described vehicle voice interaction wake-up method. This solves the technical problem of ensuring that the voice wake-up function of a new energy vehicle only responds to commands from passengers inside the vehicle, avoiding interference from similar voiceprints outside the vehicle. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the vehicle voice interaction wake-up method provided in the above embodiments, and will not be repeated here.
[0113] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for waking up a vehicle's voice interaction, characterized in that, The vehicle voice interaction wake-up method includes: Acquire voice information from the in-vehicle ambient sound that conforms to preset acoustic characteristics; The location of the sound source of the voice information is determined based on the microphone array at preset positions; Determining the sound source location of the voice information based on a microphone array at preset positions includes: The estimated location of the voice information is obtained based on the first microphone array; According to the second microphone array, the sound source pointing matrix of the speech information is obtained. The sound source pointing matrix is a multi-dimensional sound source pointing vector set constructed based on the sound source signals collected by each pickup unit in the second microphone array and the spatial position of each pickup unit. The multi-dimensional sound source pointing vector set is generated by combining the sound source direction vector, sound source intensity vector and sound source arrival time vector of each pickup unit relative to the speech information. The location of the sound source of the speech information is determined based on the estimated location of the speech information and the sound source pointing matrix. Determining the sound source location of the speech information based on the estimated location of the speech information and the sound source pointing matrix includes: Based on the sound source pointing matrix, the spatial distribution area of the sound source of the speech information is determined; If the estimated location falls within the spatial distribution area of the sound source, then it is determined that the sound source of the voice information is located in the vehicle interior space. When the sound source is located inside the vehicle, the vehicle control command in the voice information is executed.
2. The vehicle voice interaction wake-up method according to claim 1, characterized in that, The acquisition of voice information conforming to preset acoustic characteristics in the in-vehicle ambient sound includes: Collect sound signals from the in-vehicle environment; When the loudness of the sound signal is greater than the voice control threshold, the voice information in the sound signal is extracted; Obtain the voiceprint features of the speech information; When the voiceprint feature belongs to the preset wake-up voiceprint feature, it is determined that the current voice information conforms to the preset acoustic feature.
3. The vehicle voice interaction wake-up method according to claim 1, characterized in that, The step of obtaining the estimated location of the voice information based on the first microphone array includes: The sound source direction and sound source intensity of the first microphone array relative to the voice information are obtained through the first microphone array. Based on the direction and intensity of the sound source, the estimated location of the speech information in space is obtained.
4. The vehicle voice interaction wake-up method according to claim 1, characterized in that, The step of obtaining the sound source pointing matrix of the voice information based on the second microphone array includes: The second microphone array is used to obtain the sound source intensity of the voice information at each array unit, wherein all array units of the second microphone array are located outside the vehicle. The sound source pointing matrix of the speech information is obtained based on the spatial position of the array unit and the sound source intensity.
5. The vehicle voice interaction wake-up method according to claim 1, characterized in that, After determining the spatial distribution region of the speech information based on the sound source pointing matrix, the method further includes: If the estimated location does not fall within the spatial distribution area of the sound source, the sound source intensity of each array unit in the second microphone array is compared with the sound source intensity of the first microphone array to obtain the sound source intensity comparison result. When the sound source intensity of any array unit in the second microphone array is greater than the sound source intensity of the first microphone array, it is determined that the sound source of the voice information is located in the space outside the vehicle. If the sound source intensity of all array units in the second microphone array is less than the sound source intensity of the first microphone array, then it is determined that the sound source of the voice information is located in the vehicle interior space.
6. A vehicle voice interaction wake-up device, characterized in that, The vehicle voice interaction wake-up device includes: The data acquisition module is used to acquire voice information that conforms to preset acoustic characteristics in the in-vehicle ambient sound. The data processing module is used to determine the sound source location of the voice information based on the microphone array at a preset position; The data processing module is further configured to: obtain the estimated location of the speech information based on the first microphone array; obtain the sound source pointing matrix of the speech information based on the second microphone array, wherein the sound source pointing matrix is a multi-dimensional sound source pointing vector set constructed based on the sound source signals collected by each pickup unit in the second microphone array and the spatial position of each pickup unit, wherein the multi-dimensional sound source pointing vector set is generated by combining the sound source direction vector, sound source intensity vector, and sound source arrival time vector of each pickup unit relative to the speech information; and determine the sound source location of the speech information based on the estimated location of the speech information and the sound source pointing matrix. The data processing module is further configured to determine the spatial distribution area of the sound source of the voice information based on the sound source pointing matrix; if the estimated position falls within the spatial distribution area of the sound source, then it is determined that the sound source position of the voice information is located in the vehicle interior space. The control module is used to execute vehicle control commands in the voice information when the sound source is located inside the vehicle.
7. A vehicle voice interaction wake-up device, characterized in that, The vehicle voice interaction wake-up device includes: a memory, a processor, and a vehicle voice interaction wake-up program stored in the memory and executable on the processor, wherein the vehicle voice interaction wake-up program is configured to implement the steps of the vehicle voice interaction wake-up method as described in any one of claims 1 to 5.
8. A storage medium, characterized in that, The storage medium stores a vehicle voice interaction wake-up program, which, when executed by a processor, implements the steps of the vehicle voice interaction wake-up method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method for starting voice assistant and electronic device with voice assistant
CN110175016A
Voice interaction method and device and mobile carrier
CN115482816A
Vehicle control method and control device
CN117711394A
Vehicle awakening method and device, electronic equipment and storage medium
CN118366460A