Voice interaction method and device, vehicle, storage medium and program product

By tracking the user's location in a smart car and performing voice separation technology, the problem that users find it difficult to interact with vehicles outside the car is solved, and high accuracy and efficiency voice interaction is achieved.

CN120220681APending Publication Date: 2025-06-27XIAOMI EV TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510436131.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In smart cars, it is difficult for users to interact with the vehicle through the voice system when they are outside the car, which is limited by the complex environment outside the car and noise interference.

Method used

The vehicle is awakened through the user's voice, tracked the user's location, and obtained the audio signal collected by the voice acquisition device in the corresponding location. Use the user's voiceprint and audio signal to separate voiceprints, and select signals with high matching soundprints for interaction.

Benefits of technology

It realizes accurate collection of user voice commands in an off-vehicle environment, improves the accuracy and efficiency of voice interaction, and ensures that users can interact with vehicles effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220681A_ABST
    Figure CN120220681A_ABST
Patent Text Reader

Abstract

The invention relates to a voice interaction method and device, a vehicle, a storage medium and a program product, and belongs to the technical field of artificial intelligence, and the method comprises the steps: responding to a user voice wake-up vehicle, and tracking the position of the user; acquiring an audio signal acquired by a voice acquisition device corresponding to the position to obtain a first audio signal; and performing voice interaction with the user according to the voiceprint of the user and the first audio signal. Therefore, after the user wakes up the vehicle by voice, the corresponding voice acquisition device can be selected to acquire the audio signal by tracking the position of the user. In this way, the voice instruction of the user can be collected more accurately. In addition, voice interaction with the user may be performed through a voiceprint of the user and the first audio signal. By combining the voiceprint information, the user can be recognized in the voice interaction, and the accuracy of the voice interaction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a voice interaction method, apparatus, vehicle, storage medium, and program product. Background Art

[0002] In intelligent vehicles, in-vehicle voice interaction systems have been widely applied, which can provide many conveniences for drivers and passengers. However, when the user is outside the vehicle, it may be difficult to interact with the vehicle through the voice system. Summary of the Invention

[0003] To overcome the problems in the related art, the present disclosure provides a voice interaction method, apparatus, vehicle, storage medium, and program product.

[0004] According to a first aspect of the embodiments of the present disclosure, a voice interaction method is provided, including: In response to a user's voice waking up the vehicle, tracking the position of the user; Obtaining an audio signal collected by a voice collection device corresponding to the position to obtain a first audio signal; Performing voice interaction with the user according to the user's voiceprint and the first audio signal.

[0005] Optionally, the performing voice interaction with the user according to the user's voiceprint and the first audio signal includes: Performing voice separation on the first audio signal to obtain a main signal and a reference signal; Taking the one with a higher voiceprint matching degree with the user's voiceprint among the main signal and the reference signal as a second audio signal; Performing voice interaction with the user through the second audio signal.

[0006] Optionally, the performing voice interaction with the user through the second audio signal includes: Transmitting the second audio signal to a processing device; Obtaining a voice recognition result of the processing device for the second audio signal; Performing a response action for voice interaction according to the voice recognition result.

[0007] Optionally, it includes: In response to a user's voice waking up the vehicle, extracting the user's voiceprint according to the user's voice; or, Obtaining the voiceprint pre-recorded by the user.

[0008] Optionally, the vehicle includes a plurality of collection devices, and the method includes: Determine the first acquisition device closest to the position, where the voice acquisition device includes the first acquisition device.

[0009] Optionally, it includes: Obtain a second audio signal collected by the acquisition device of the vehicle; Perform at least one of the following processes on the second audio signal to obtain a third audio signal corresponding to the acquisition device: echo cancellation; noise reduction; dereverberation; blind source separation; In response to the acoustic similarity between the third audio signal and a preset wake-up word signal being greater than a threshold, wake up the vehicle.

[0010] Optionally, the vehicle includes a plurality of acquisition devices, and the method includes: For the third audio signal of each acquisition device, determine the acoustic similarity between the third audio signal and a preset wake-up word signal, and the sound energy of the wake-up word signal in the third audio signal; Determine a second acquisition device from the plurality of acquisition devices according to the acoustic similarity and the sound energy; Identify the user within the audio acquisition range of the second acquisition device.

[0011] According to a second aspect of the embodiments of the present disclosure, there is provided a voice interaction device, including: A first module configured to track the position of the user in response to the vehicle being woken up by the user's voice; A second module configured to obtain an audio signal collected by a voice acquisition device corresponding to the position to obtain a first audio signal; A third module configured to perform voice interaction with the user according to the voiceprint of the user and the first audio signal.

[0012] Optionally, the third module is configured to: Perform voice separation on the first audio signal to obtain a main signal and a reference signal; Use the one with a higher voiceprint matching degree between the main signal and the reference signal and the voiceprint of the user as a second audio signal; Perform voice interaction with the user through the second audio signal.

[0013] Optionally, the third module is configured to: Transmit the second audio signal to a processing device; Obtain a voice recognition result of the second audio signal by the processing device; Execute a response action for voice interaction according to the voice recognition result.

[0014] Optionally, it includes: The fourth module is configured to extract the voiceprint of the user according to the user's voice in response to the user's voice waking up the vehicle; or, obtain the voiceprint pre-recorded by the user.

[0015] Optionally, the vehicle includes a plurality of collection devices, and the devices include: The fifth module is configured to determine a first collection device closest to the position, and the voice collection device includes the first collection device.

[0016] Optionally, it includes: The sixth module is configured to obtain a second audio signal collected by the collection device of the vehicle; The seventh module is configured to perform at least one of the following processes on the second audio signal to obtain a third audio signal corresponding to the collection device: echo cancellation; noise reduction; dereverberation; blind source separation; The eighth module is configured to wake up the vehicle in response to the acoustic similarity between the third audio signal and a preset wake-up word signal being greater than a threshold.

[0017] Optionally, the vehicle includes a plurality of collection devices, and the devices include: The ninth module is configured to determine the acoustic similarity between the third audio signal of each collection device and a preset wake-up word signal, and the sound energy of the wake-up word signal in the third audio signal; The tenth module is configured to determine a second collection device from the plurality of collection devices according to the acoustic similarity and the sound energy; The eleventh module is configured to identify the user within the audio collection range of the second collection device.

[0018] According to a third aspect of the embodiments of the present disclosure, there is provided a vehicle, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the steps of the method described in any one of the first aspect.

[0019] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in any one of the first aspect.

[0020] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in any one of the first aspect.

[0021] In the above solution, the vehicle can be woken up in response to the user's voice, and the position of the user can be tracked. And the first audio signal collected by the voice collection device corresponding to the position can be obtained. In this way, voice interaction with the user can be performed according to the user's voiceprint and the first audio signal.

[0022] In this way, after the vehicle is woken up by the user's voice, the corresponding voice collection device can be selected to collect the audio signal by tracking the user's position. In this way, the user's voice command can be collected more accurately. In addition, voice interaction with the user can be performed through the user's voiceprint and the first audio signal. By combining the voiceprint information, it helps to identify the user in the voice interaction, and further helps to improve the accuracy of the voice interaction.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0024] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0025] Figure 1 is a schematic diagram of a vehicle-human interaction shown according to an exemplary embodiment.

[0026] Figure 2 is a flowchart of a voice interaction method shown according to an exemplary embodiment.

[0027] Figure 3 is an implementation flowchart of a step S23 shown according to an exemplary embodiment.

[0028] Figure 4 is a block diagram of a voice interaction device shown according to an exemplary embodiment.

[0029] Figure 5 is a block diagram of a vehicle shown according to an exemplary embodiment. Detailed Description of the Embodiments

[0030] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0031] It should be noted that all actions of obtaining signals, information, or data in this disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and with the authorization given by the owner of the corresponding device.

[0032] Before introducing the voice interaction method, device, vehicle, storage medium, and program product of this disclosure, first, an exemplary description of the relevant scenarios of the embodiments of this disclosure is given.

[0033] In intelligent vehicles, in-vehicle voice interaction systems have been widely used, which can provide many conveniences for drivers and passengers. However, when the user is outside the vehicle, it may be difficult to interact with the vehicle through the voice system.

[0034] For example, in some scenarios, the external environment of the vehicle may be relatively complex, with high noise and reverberation, and there may also be interference from other voices, resulting in difficult voice interaction. In addition, the user outside the vehicle may be in a moving state when using the voice function, resulting in poor interactivity. Figure 1 is a schematic diagram of vehicle-human interaction shown in an exemplary embodiment of this disclosure. Refer to Figure 1 , the user may wake up the vehicle by voice at position A and then move to position B. In this case, since the user moves to position B, which is far from position A, and the external environment of the vehicle is complex, the vehicle may not be able to accurately obtain the user's voice signal, resulting in poor voice interaction effect.

[0035] Therefore, the embodiments of this disclosure provide a voice interaction method. The method can be executed by a vehicle or a related in-vehicle device, for example. Figure 2 is a flowchart of a voice interaction method shown in an exemplary embodiment. As Figure 2 shown, the method includes the following steps.

[0036] In step S21, in response to the user's voice waking up the vehicle, track the position of the user.

[0037] In step S22, obtain the audio signal collected by the voice collection device corresponding to the position to obtain a first audio signal.

[0038] In step S23, perform voice interaction with the user according to the user's voiceprint and the first audio signal.

[0039] The following gives an exemplary description of the implementation manners of the above steps S21 to S23.

[0040] In step S21, in response to the user's voice waking up the vehicle, track the position of the user.

[0041] In one embodiment, a user can wake up the vehicle by voice containing a wake word outside the vehicle. The vehicle can be provided with one or more voice collection devices to collect the user's voice, and trigger a wake-up event when the preset wake word is included in the voice.

[0042] In one embodiment, the voice collection device can include a microphone array. Exemplarily, multiple groups of microphone arrays can be arranged outside the vehicle to continuously collect ambient sounds. The layout of the microphone arrays can be arranged based on requirements. As an example, the microphone array can include four groups, which are respectively arranged at the front of the vehicle, the rear of the vehicle, and the left and right rearview mirrors. One group of microphone arrays can include two or more microphones. In this way, the ambient sounds can be collected through the microphone array, and when the wake-up condition is met, the vehicle wake-up event can be triggered.

[0043] For example, in one possible embodiment, before step S21, the method further includes: Obtaining a second audio signal collected by the collection device of the vehicle; Performing at least one of the following processes on the second audio signal to obtain a third audio signal corresponding to the collection device: echo cancellation; noise reduction; reverberation removal; blind source separation; Waking up the vehicle in response to the acoustic similarity between the third audio signal and the preset wake word signal being greater than a threshold.

[0044] Exemplarily, the second audio signal can be collected by some or all of the collection devices of the vehicle. In one embodiment, the collected second audio signal can be directly processed to determine whether the acoustic similarity between the second audio signal and the preset wake word signal is greater than the threshold. When the acoustic similarity is greater than the threshold, the voice collected by the collection device may include the wake word used by the user to wake up the vehicle. Therefore, the vehicle wake-up event can be triggered to wake up the vehicle.

[0045] As an example, the wake word can be "Little A classmate". When the user utters a wake-up voice including "Little A classmate", the collection device can collect the second audio signal and identify whether the acoustic similarity between the second audio signal and the preset wake word signal (Little A classmate) is greater than the threshold.

[0046] In a possible implementation, the acoustic similarity between the second audio signal and the preset wake-up word signal can also be determined by a wake-up model. The wake-up model may include a speech recognition model pre-trained for detecting specific wake-up words. Exemplarily, the wake-up model can be deployed in a vehicle. The vehicle can extract features from the second audio signal and input the extracted features into the wake-up model to obtain the output result of the wake-up model. The output result may include, for example, the acoustic similarity (acoustic score) between the second audio signal and the preset wake-up word signal, the sound energy (e.g., loudness, power, etc.) of the wake-up word signal in the second audio signal, etc. In this way, the wake-up model can be used to determine whether the acoustic similarity between the second audio signal and the preset wake-up word signal (classmate Xiao A) is greater than a threshold, and trigger a vehicle wake-up event when the acoustic similarity between the second audio signal and the preset wake-up word signal is greater than the threshold. In addition, when the acoustic similarity between the second audio signal and the preset wake-up word signal is less than or equal to the threshold, the second audio signal collected this time can be ignored.

[0047] In the above embodiments, the present disclosure is exemplarily described by taking the direct processing of the second audio signal as an example. In some implementations, considering that the second audio signal collected by the acquisition device may include interference signals, the second signal can also be processed.

[0048] For example, in one implementation, the second audio signal can be subjected to one or more of the following processes to obtain a third audio signal corresponding to the acquisition device: echo cancellation, noise reduction, dereverberation, blind source separation.

[0049] Echo cancellation processing can be used to remove the echo signal collected by the acquisition device to avoid sound repetition and interference. Noise reduction processing can be used to weaken or eliminate background noise in the speech signal and improve the clarity and intelligibility of the speech. Dereverberation processing can be used to reduce the reverberation effect caused by sound reflection in an enclosed space and make the speech signal cleaner and clearer. Blind source separation processing can be used to separate multiple independent sound sources from the mixed signal without prior knowledge of the number or location of the sound sources.

[0050] Through the above processing, the interference in the second audio signal can be removed to obtain a third audio signal. In this way, the vehicle can be awakened in response to the acoustic similarity between the third audio signal and the preset wake-up word signal being greater than the threshold.

[0051] Exemplarily, feature extraction can be performed on the third audio signal, and the extracted features can be input into the wake-up model to obtain the output result of the wake-up model. The output result can, for example, include the acoustic similarity between the third audio signal and the preset wake-up word signal, the sound energy of the wake-up word signal in the third audio signal, etc. Thus, it is possible to determine whether the acoustic similarity between the third audio signal and the preset wake-up word signal is greater than a threshold through the wake-up model, and trigger a vehicle wake-up event when the acoustic similarity between the third audio signal and the preset wake-up word signal is greater than the threshold. In addition, when the acoustic similarity between the third audio signal and the preset wake-up word signal is less than or equal to the threshold, the third audio signal collected this time can be ignored.

[0052] After the vehicle is woken up, the position of the user can be tracked. For example, in one implementation, images around the vehicle can be collected by an image acquisition device of the vehicle, so as to identify the user based on the images. Exemplarily, multiple consecutive images can be captured by one or more cameras of the vehicle, and the moving trajectory, real-time position, predicted arrival position, etc. of the user outside the vehicle can be determined through image matching and target tracking. In one implementation, the position of the user can also be determined through Bluetooth positioning. Exemplarily, the vehicle's Bluetooth anchor points can be used to perform Bluetooth positioning on the user's car key to determine the position of the user.

[0053] In one implementation, the vehicle includes a plurality of acquisition devices, and the method includes: For the third audio signal of each acquisition device, determine the acoustic similarity between the third audio signal and the preset wake-up word signal, and the sound energy of the wake-up word signal in the third audio signal; Determine a second acquisition device from the plurality of acquisition devices according to the acoustic similarity and the sound energy; Identify the user within the audio acquisition range of the second acquisition device.

[0054] Continuing with the above example, the vehicle can include 4 sets of acquisition devices. When the user issues a wake-up voice, the four sets of acquisition devices can respectively collect the second audio signal, and after processing, four sets of third audio signals are obtained.

[0055] For the third audio signal of each acquisition device, the acoustic similarity between the third audio signal and the preset wake-up word signal, and the sound energy of the wake-up word signal in the third audio signal can be determined. Thus, a second acquisition device can be determined from the plurality of acquisition devices according to the acoustic similarity and the sound energy.

[0056] As an example, with the goal of relatively high acoustic similarity and relatively high sound energy, the second acquisition device can be comprehensively determined from the multiple acquisition devices. Relatively high acoustic similarity and relatively high sound energy can indicate that the audio signals collected by the acquisition device are more likely to include a preset wake-up word.

[0057] Exemplarily, weights and scores can be set for acoustic similarity and sound energy respectively, and the comprehensive score of each acquisition device can be obtained by the method of weighted summation. In this way, the acquisition device with the highest score can be used as the second acquisition device.

[0058] Since the second acquisition device is comprehensively determined with the goal of relatively high acoustic similarity and relatively high sound energy, it is very likely that the user waking up the vehicle is within the audio acquisition range (sound zone) of the second acquisition device. Thus, after determining the second acquisition device, it is possible to be within the audio acquisition range of the second acquisition device, so as to identify the user. This helps to quickly track the position of the user.

[0059] In one implementation, the user position tracking can be combined with the audio acquisition range of the second acquisition device. Taking tracking via Bluetooth as an example, the position of the user can be determined by Bluetooth positioning. If the position where the user is located is outside the audio acquisition range of the second acquisition device, an abnormality can be determined and the current vehicle wake-up event is invalid. If the position where the user is located is within the audio acquisition range of the second acquisition device, Bluetooth positioning can be continued in real time, so as to track the position of the user. In one implementation, the position of the Bluetooth signal can be tracked in real time, and by removing abnormal points and smoothing the tracked position, the moving direction of the user outside the vehicle and the predicted final arrival position can be finally determined.

[0060] In one possible implementation, it is also possible to combine the above Bluetooth positioning and image recognition methods to track the user position at the same time.

[0061] Refer to Figure 2 , in step S22, the audio signal collected by the voice acquisition device corresponding to the position is obtained to get the first audio signal.

[0062] Exemplarily, the vehicle can include multiple acquisition devices, and different acquisition devices can be configured to collect audio information in different (or not completely the same) areas. For example, the acquisition device at the front of the vehicle can mainly collect the audio information at the front of the vehicle, the acquisition device at the rear of the vehicle can mainly collect the audio information at the rear of the vehicle. The acquisition device at the position of the left rearview mirror can mainly collect the audio information on the left side of the vehicle. The acquisition device at the position of the right rearview mirror can mainly collect the audio information on the right side of the vehicle.

[0063] In this way, according to the user's location, the audio signal collected by the voice collection device corresponding to the location can be selected to obtain the first audio signal.

[0064] For example, in one implementation, the first collection device closest to the location can be determined, and the voice collection device includes the first collection device.

[0065] As an example, if it is traced that the user is located at the rear of the vehicle, the voice collection device at the nearest rear of the vehicle can be obtained to collect the first audio signal. If it is traced that the user is located at the front of the vehicle, the voice collection device at the front of the vehicle can be obtained to collect the first audio signal. In this way, the accuracy of voice collection can be improved.

[0066] In step S23, voice interaction is performed with the user according to the user's voiceprint and the first audio signal.

[0067] Figure 3 FIG. is a flowchart of an implementation of step S23 shown in an exemplary embodiment of the present disclosure. Refer to Figure 3 The voice interaction with the user according to the user's voiceprint and the first audio signal includes: In step S31, the first audio signal is subjected to voice separation to obtain a main signal and a reference signal.

[0068] Exemplarily, the first audio signal collected by the voice collection device can be subjected to voice separation processing to obtain a main signal and a reference signal. The main signal may include the user's voice, and the reference signal may include noise interference, such as ambient noise, the voice of other speakers, the echo of the device itself, and so on.

[0069] In step S32, the one with a higher voiceprint matching degree between the user's voiceprint in the main signal and the reference signal is used as the second audio signal.

[0070] In one implementation, the vehicle can be woken up in response to the user's voice, and the user's voiceprint can be extracted from the user's voice. Exemplarily, the user's voiceprint can be extracted from the user's voice to obtain the user's voiceprint. In one implementation, the voice of the wake-up word part in the wake-up voice can also be obtained, and the user's voiceprint can be extracted therefrom.

[0071] In one possible implementation, the user's pre-recorded voiceprint can also be obtained. Exemplarily, the user can pre-record voiceprint information in the vehicle, so that the pre-recorded voiceprint can be obtained when needed, thereby improving the processing speed.

[0072] In a possible implementation, the voiceprint of the main signal can also be obtained to get the first voiceprint. In addition, the voiceprint of the reference signal can be obtained to get the second voiceprint. In this way, the user's voiceprint can be matched with the first voiceprint and the second voiceprint. When the user's voiceprint is consistent with the first voiceprint, it can be determined that the user's speech is located in the main signal, and at this time the main signal is the second audio signal. When the user's voiceprint is consistent with the second voiceprint, it can be determined that the user's speech is located in the reference signal, and at this time the reference signal is the second audio signal.

[0073] In a possible implementation, the user's voiceprint does not match either the first voiceprint or the second voiceprint. In this case, the current interaction can be terminated.

[0074] In step S33, a voice interaction with the user is performed through the second audio signal.

[0075] In one implementation, the vehicle can process the second audio signal by itself. For example, the vehicle can perform speech recognition on the second audio signal to obtain a speech recognition result. In this way, the vehicle can respond to the speech recognition result to achieve a voice interaction with the user. As an example, the speech recognition result may include an instruction from the user to "play music", so the vehicle can respond to the speech recognition result to play music.

[0076] In a possible implementation, the performing the voice interaction with the user through the second audio signal includes: Transmitting the second audio signal to a processing device; Obtaining the speech recognition result of the processing device for the second audio signal; Performing a response action for the voice interaction according to the speech recognition result.

[0077] Exemplarily, the processing device can be, for example, a cloud server or other relevant processing devices, such as a computer, a tablet device, a mobile phone, etc. Taking the cloud server as an example, the vehicle can send the second audio signal to the cloud server. The cloud server can perform speech recognition on the second audio signal to obtain a speech recognition result. In this way, the vehicle obtains the speech recognition result from the cloud server and responds to the speech recognition result to achieve a voice interaction with the user. As an example, the speech recognition result may include an instruction from the user to "turn on the air conditioner", so the vehicle can respond to the speech recognition result to turn on the air conditioner.

[0078] In this way, since the vehicle separates the voice from the first audio signal and selects the second audio signal representing the user's voice for processing, the above solution can reduce the transmission volume of the voice signal, help improve the efficiency of voice interaction, reduce the power consumption of voice recognition, and reduce the latency of voice interaction.

[0079] In the above solution, the vehicle can be woken up in response to the user's voice, the position of the user can be tracked, and the first audio signal collected by the voice collection device corresponding to the position can be obtained. In this way, voice interaction with the user can be performed according to the user's voiceprint and the first audio signal.

[0080] In this way, after the vehicle is woken up by the user's voice, the position of the user can be tracked to select the corresponding voice collection device to collect the audio signal. In this way, the user's voice command can be collected more accurately. In addition, voice interaction with the user can be performed through the user's voiceprint and the first audio signal. By combining the voiceprint information, it helps to identify the user in voice interaction, and further helps to improve the accuracy of voice interaction.

[0081] Based on the same inventive concept, an embodiment of the present disclosure provides a voice interaction device. Figure 4 is a block diagram of a voice interaction device shown in an exemplary embodiment of the present disclosure. Refer to Figure 4 , the voice interaction device includes: A first module 401, configured to track the position of the user in response to the user's voice waking up the vehicle; A second module 402, configured to obtain the audio signal collected by the voice collection device corresponding to the position to obtain a first audio signal; A third module 403, configured to perform voice interaction with the user according to the user's voiceprint and the first audio signal.

[0082] In the above solution, the vehicle can be woken up in response to the user's voice, the position of the user can be tracked, and the first audio signal collected by the voice collection device corresponding to the position can be obtained. In this way, voice interaction with the user can be performed according to the user's voiceprint and the first audio signal.

[0083] In this way, after the vehicle is woken up by the user's voice, the position of the user can be tracked to select the corresponding voice collection device to collect the audio signal. In this way, the user's voice command can be collected more accurately. In addition, voice interaction with the user can be performed through the user's voiceprint and the first audio signal. By combining the voiceprint information, it helps to identify the user in voice interaction, and further helps to improve the accuracy of voice interaction.

[0084] Optionally, the third module 403 is configured to: Perform voice separation on the first audio signal to obtain a main signal and a reference signal; Among the main signal and the reference signal, use the one with a higher voiceprint matching degree with the user's voiceprint as the second audio signal; Perform voice interaction with the user through the second audio signal.

[0085] Optionally, the third module 403 is configured to: Transmit the second audio signal to a processing device; Obtain the speech recognition result of the processing device for the second audio signal; Perform a response action for voice interaction according to the speech recognition result.

[0086] Optionally, it includes: A fourth module, configured to extract the user's voiceprint according to the user's voice in response to waking up the vehicle by the user's voice; or, obtain the voiceprint pre-recorded by the user.

[0087] Optionally, the vehicle includes a plurality of collection devices, and the devices include: A fifth module, configured to determine the first collection device closest to the position, and the voice collection device includes the first collection device.

[0088] Optionally, it includes: A sixth module, configured to obtain a second audio signal collected by the collection device of the vehicle; A seventh module, configured to perform at least one of the following processes on the second audio signal to obtain a third audio signal corresponding to the collection device: echo cancellation; noise reduction; reverberation removal; blind source separation; An eighth module, configured to wake up the vehicle in response to the acoustic similarity between the third audio signal and a preset wake-up word signal being greater than a threshold.

[0089] Optionally, the vehicle includes a plurality of collection devices, and the devices include: A ninth module, configured to determine the acoustic similarity between the third audio signal of each collection device and a preset wake-up word signal, and the sound energy of the wake-up word signal in the third audio signal; A tenth module, configured to determine a second collection device from the plurality of collection devices according to the acoustic similarity and the sound energy; An eleventh module, configured to identify the user within the audio collection range of the second collection device.

[0090] An embodiment of the present disclosure provides a vehicle, including: A processor; A memory for storing processor-executable instructions; wherein, the processor is configured to execute the steps of the voice interaction method provided in at least one embodiment of the present disclosure.

[0091] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the voice interaction method provided in at least one embodiment of the present disclosure.

[0092] An embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the voice interaction method provided in at least one embodiment of the present disclosure.

[0093] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0094] Figure 5 It is a block diagram of a vehicle 600 shown according to an exemplary embodiment. For example, the vehicle 600 can be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 can be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0095] Referring to Figure 5 , the vehicle 600 may include various subsystems. For example, the infotainment system 610, the perception system 620, the decision control system 630, the drive system 640, and the computing platform 650. Among them, the vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 can be interconnected in a wired or wireless manner.

[0096] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, a navigation system, etc.

[0097] The perception system 620 may include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system can be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.

[0098] The decision control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0099] The drive system 640 may include components that provide motive power for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.

[0100] Some or all of the functions of the vehicle 600 are controlled by the computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652, and the processor 651 may execute instructions 653 stored in the memory 652.

[0101] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0102] The memory 652 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0103] In addition to the instructions 653, the memory 652 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 652 can be used by the computing platform 650.

[0104] In an embodiment of the present disclosure, the processor 651 may execute the instructions 653 to complete all or part of the steps of the above-described voice interaction method.

[0105] In addition, as used herein, the word "exemplary" is used to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the word exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless specified otherwise, or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied under any of the foregoing instances. Additionally, unless specified otherwise or clear from the context that it refers to the singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0106] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such feature may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Further, with respect to the use of "comprises", "comprising", "has", "having", "includes", or variants thereof in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "including".

[0107] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known art or conventional techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

[0108] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A voice interaction method, characterized in that: include: In response to a user's voice waking up the vehicle, tracking the user's location; Acquire an audio signal collected by a voice collection device corresponding to the position to obtain a first audio signal; Perform voice interaction with the user according to the voiceprint of the user and the first audio signal.

2. The method according to claim 1, characterized in that The performing voice interaction with the user according to the voiceprint of the user and the first audio signal includes: Performing speech separation on the first audio signal to obtain a main signal and a reference signal; Using the one of the main signal and the reference signal whose voiceprint matches the user's voiceprint more highly as the second audio signal; Perform voice interaction with the user through the second audio signal.

3. The method according to claim 2, characterized in that The performing voice interaction with the user through the second audio signal includes: transmitting the second audio signal to a processing device; Acquiring a speech recognition result of the processing device on the second audio signal; A response action of the voice interaction is executed according to the voice recognition result.

4. The method according to any one of claims 1 to 3, characterized in that include: In response to the user's voice waking up the vehicle, extracting the user's voiceprint according to the user's voice; or, Obtain the voiceprint pre-recorded by the user.

5. The method according to any one of claims 1 to 3, characterized in that The vehicle includes a plurality of collection devices, and the method includes: A first collection device that is closest to the location is determined, and the voice collection device includes the first collection device.

6. The method according to any one of claims 1 to 3, characterized in that include: Acquiring a second audio signal collected by a collection device of the vehicle; Perform at least one of the following processing on the second audio signal to obtain a third audio signal corresponding to the acquisition device: echo cancellation, noise reduction, dereverberation, blind source separation; In response to the acoustic similarity between the third audio signal and a preset wake-up word signal being greater than a threshold, waking up the vehicle.

7. The method according to claim 6, characterized in that The vehicle includes a plurality of collection devices, and the method includes: For the third audio signal of each acquisition device, determine the acoustic similarity between the third audio signal and the preset wake-up word signal, and the sound energy of the wake-up word signal in the third audio signal; Determining a second collection device from the multiple collection devices according to the acoustic similarity and the sound energy; The user is identified in the audio collection range of the second collection device.

8. A voice interaction device, characterized in that: include: A first module is configured to wake up the vehicle in response to a user's voice and track the user's location; The second module is configured to obtain an audio signal collected by a voice collection device corresponding to the position to obtain a first audio signal; The third module is configured to perform voice interaction with the user according to the voiceprint of the user and the first audio signal.

9. A vehicle, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

11. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.