Vehicle voice interaction methods, devices, vehicles and storage media

By using voiceprint recognition and natural language understanding, the target sound zone inside the vehicle cabin is determined, which solves the problems of low signal separation and control accuracy in multi-sound zone audio processing technology, and enables the vehicle to perform precise actions in the correct sound zone and position.

CN116564310BActive Publication Date: 2026-04-17GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU XIAOPENG MOTORS TECH CO LTD
Filing Date
2023-05-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In intelligent vehicles, due to the complex and ever-changing environment and the variable location of sound sources when the vehicle is in motion, multi-zone audio processing technology leads to the separation of effective voice signals and a reduction in the accuracy of vehicle motion control.

Method used

By using voiceprint recognition and natural language understanding, the target sound zone in the vehicle cabin is determined, and vehicle control commands are generated based on the target sound zone to achieve precise voice interaction.

Benefits of technology

It improves the accuracy of voice signal separation and vehicle motion control, ensuring that the vehicle performs the correct actions in the correct voice range and position.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116564310B_ABST
    Figure CN116564310B_ABST
Patent Text Reader

Abstract

This application discloses a vehicle voice interaction method, a vehicle voice interaction device, a vehicle, and a computer-readable storage medium, comprising: acquiring a voice request; performing voiceprint recognition and natural language understanding on the voice request; determining the target sound region in the vehicle cabin corresponding to the voice request based on a first result of voiceprint recognition; and generating a vehicle control command corresponding to the voice request based on the target sound region and a second result of natural language understanding, so that the vehicle can complete the voice interaction according to the vehicle control command. This application can, after acquiring a voice request, perform voiceprint recognition and natural language understanding processing on the voice request separately, and distinguish and filter the sound region of the voice request based on the result of voiceprint recognition to determine the sound region in the vehicle cabin corresponding to the generated control command. Furthermore, it determines the content and direction of the control command based on the result of natural language understanding, so that the vehicle can perform the correct action in the correct sound region and at the correct location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle voice interaction technology, specifically to a vehicle voice interaction method, a vehicle voice interaction device, a vehicle, and a computer-readable storage medium. Background Technology

[0002] In current intelligent vehicle technology, multi-zone audio processing technology is used in the vehicle cabin to simultaneously process audio signals from multiple zones and separate the effective voice signals to form multiple independent audio channels for controlling vehicle movements. However, due to the complex and changing environment and the variable location of sound sources during vehicle operation, the processing performance of multi-zone technology is reduced, and signal residue is unavoidable, thus affecting the separation of effective voice signals and the accuracy of vehicle movement control. Summary of the Invention

[0003] This application provides a vehicle voice interaction method, a vehicle voice interaction device, a vehicle, and a computer-readable storage medium.

[0004] The vehicle voice interaction method according to the embodiments of this application includes:

[0005] Get the voice request;

[0006] The voice request is then subjected to voiceprint recognition and natural language understanding.

[0007] Based on the first result of the voiceprint recognition, the target sound region in the vehicle cabin corresponding to the voice request is determined;

[0008] Based on the target voice region and the second result of natural language understanding, a vehicle control command corresponding to the voice request is generated so that the vehicle can complete the voice interaction according to the vehicle control command.

[0009] Thus, after obtaining a voice request, this application can perform voiceprint recognition and natural language understanding processing on the voice request separately, and distinguish and filter the voice region of the voice request based on the voiceprint recognition result to determine the corresponding voice region in the vehicle cabin for the generated control command. In addition, the content and direction of the control command are determined based on the natural language understanding result, so as to enable the vehicle to perform the correct action in the correct voice region and the correct position.

[0010] In some implementations, the process of obtaining a voice request further includes:

[0011] In response to the voice request, an initial audio signal including the voice request is acquired by multiple audio collection devices installed in the vehicle cabin;

[0012] The initial audio signal is preprocessed to determine multiple audio signals corresponding to multiple sound zones in the vehicle cabin.

[0013] Thus, this application can preprocess the acquired voice requests so that subsequent audio processing methods can be easily implemented.

[0014] In some implementations, performing voiceprint recognition and natural language understanding on the voice request includes:

[0015] Perform voiceprint recognition on the audio signal to determine the voiceprint information included in the audio signal;

[0016] Perform speech recognition on the audio signal to determine the speech recognition information included in the audio signal;

[0017] Natural language understanding is performed on the speech recognition information to determine the semantic information corresponding to the speech recognition information;

[0018] The step of determining the target voice region in the vehicle cabin corresponding to the voice request based on the first result of the voiceprint recognition includes:

[0019] The target sound region is determined based on the voiceprint information.

[0020] Thus, this application is able to perform voiceprint recognition and natural language understanding processing on voice requests, and respectively determine the voiceprint recognition result used for subsequent matching and comparison and the natural language understanding result used to finally determine the content of vehicle control commands.

[0021] In some implementations, determining the target voice region based on the voiceprint information includes:

[0022] Based on the voiceprint information and the preset speech synthesis voiceprint database, determine the comparison result between the voiceprint information and the speech synthesis voiceprint database;

[0023] If, after comparison, it is determined that the voiceprint information does not belong to the speech synthesis voiceprint database, the target voice region corresponding to the speech request is determined based on the voiceprint information that does not belong to the voiceprint database.

[0024] Thus, this application can avoid interference from synthesized speech emitted by the vehicle to the voice request by comparing the determined voiceprint information with the speech synthesis voiceprint database.

[0025] In some embodiments, the method further includes:

[0026] If, after comparison, it is determined that the voiceprint information belongs to the speech synthesis voiceprint database, the audio signal corresponding to the voiceprint information is discarded.

[0027] Thus, this application can also discard the voice request corresponding to the voiceprint information when there is a matching relationship between the voiceprint information and the speech synthesis voiceprint database, thereby eliminating the interference of the synthesized speech information on the voice request.

[0028] In some embodiments, the

[0029] Thus, this application can also filter multiple audio signals by determining whether multiple voiceprint information comes from the same object, so as to determine the unique target sound region corresponding to the same object.

[0030] In some embodiments, determining the target voice region based on the voiceprint information further includes:

[0031] If it is determined that multiple voiceprint information sources are from different objects other than the first object, the target voice region is determined based on the multiple voiceprint information sources.

[0032] Thus, this application can also directly determine the corresponding audio region based on the multi-channel voiceprint information when the multi-channel voiceprint signals do not originate from the same object, without going through a screening process, thereby achieving the goal of maintaining the information integrity of the multi-channel audio signals as realistically as possible.

[0033] In some implementations, determining the target voice region based on the voiceprint information includes:

[0034] The voiceprint information is compared with a preset historical voiceprint database;

[0035] Based on the comparison results, the voiceprint information corresponding to the voice region is corrected to determine the target voice region.

[0036] Thus, this application can also improve the accuracy of target voice region judgment by comparing voiceprint information with information in the historical voiceprint database and correcting the corresponding voice region information.

[0037] In some implementations, generating vehicle control commands corresponding to the voice request based on the target voice region and the second result of natural language understanding includes:

[0038] Based on the target voice region and the semantic information, a vehicle control command corresponding to the voice request is generated;

[0039] According to the vehicle control command, execute the action corresponding to the vehicle control command to complete the voice interaction.

[0040] Thus, after the voiceprint recognition and matching process is completed, this application can combine the information obtained through natural language understanding processing with the target voice region that has just been determined as a basis to determine the vehicle control command for controlling the vehicle to perform the correct actions, and correctly control the vehicle according to the vehicle control command.

[0041] The vehicle voice interaction device according to the embodiments of this application includes:

[0042] The information acquisition module is used to acquire voice requests;

[0043] The audio processing module is used to perform voiceprint recognition and natural language understanding on the voice request;

[0044] The voice region determination module is used to determine the target voice region in the vehicle cabin corresponding to the voice request based on the first result of the voiceprint recognition.

[0045] The instruction control module is used to generate vehicle control instructions corresponding to the voice request based on the target voice region and the second result of natural language understanding, so that the vehicle can complete the voice interaction according to the vehicle control instructions.

[0046] The vehicle according to the embodiments of this application includes a memory and a processor. The memory stores a computer program that, when executed by the processor, causes the vehicle to perform the vehicle voice interaction method as described in any of the above embodiments.

[0047] The computer-readable storage medium of this application embodiment stores a computer program that, when executed by one or more processors, implements the vehicle voice interaction method as described in any of the above embodiments.

[0048] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description

[0049] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0050] Figure 1 This is a flowchart illustrating the vehicle voice interaction method in the embodiments of this application;

[0051] Figure 2 This is a schematic diagram illustrating the application scenario of the vehicle voice interaction method in the embodiments of this application.

[0052] Figure 3 This is a flowchart illustrating the vehicle voice interaction method in the embodiments of this application;

[0053] Figure 4 This is a flowchart illustrating the vehicle voice interaction method in the embodiments of this application;

[0054] Figure 5 This is a flowchart illustrating the vehicle voice interaction method in the embodiments of this application;

[0055] Figure 6 This is a flowchart illustrating the vehicle voice interaction method in the embodiments of this application. Detailed Implementation

[0056] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.

[0057] like Figure 1 As shown, the vehicle voice interaction method in this application includes the following steps:

[0058] 01: Obtain voice request;

[0059] 02: Perform voiceprint recognition and natural language understanding on voice requests;

[0060] 03: Based on the first result of voiceprint recognition, determine the target sound region in the vehicle cabin corresponding to the voice request;

[0061] 04: Based on the target sound region and the second result of natural language understanding, generate vehicle control commands corresponding to the voice request so that the vehicle can complete the voice interaction according to the vehicle control commands.

[0062] The vehicle voice interaction device in this application embodiment can realize the vehicle voice interaction method in this application embodiment. Specifically, the vehicle voice interaction device includes an information acquisition module, an audio processing module, a voice region determination module, and an instruction control module. The information acquisition module is used to acquire voice requests. The audio processing module is used to perform voiceprint recognition and natural language understanding on the voice requests. The voice region determination module is used to determine the target voice region in the vehicle cabin corresponding to the voice request based on the first result of the voiceprint recognition. The instruction control module is used to generate a vehicle control instruction corresponding to the voice request based on the target voice region and the second result of natural language understanding, so that the vehicle can complete the voice interaction according to the vehicle control instruction.

[0063] The vehicle in this application includes a memory and a processor. The memory stores a computer program, and the processor is used to acquire a voice request, and to perform voiceprint recognition and natural language understanding on the voice request, and to determine the target sound zone in the vehicle cabin corresponding to the voice request based on the first result of voiceprint recognition, and to generate vehicle control commands corresponding to the voice request based on the target sound zone and the second result of natural language understanding, so that the vehicle can complete voice interaction according to the vehicle control commands.

[0064] Specifically, the voice signals in the voice interaction function inside the cabin of current intelligent vehicles generally include voice requests issued by the user and voice feedback generated by the in-vehicle system through voice synthesis.

[0065] To facilitate accurate identification of the source of voice signals and ensure the precise execution of actions corresponding to voice requests, the interior space of a vehicle cabin is typically divided into multiple sound zones. The method of division can be adjusted according to the actual situation; in some examples, such as... Figure 2 The application scenario shown divides the vehicle cabin into four sound zones based on seating: the driver's seat zone, the passenger's seat zone, the driver's seat rear zone, and the passenger's seat rear zone, corresponding to the driver's seat, passenger's seat, driver's seat rear zone, and passenger's seat rear zone, respectively.

[0066] The vehicle voice interaction method in this application involves processing a voice request obtained from the vehicle cabin using voiceprint recognition and natural language understanding in the vehicle system. This process involves breaking down the information contained in the voice request to obtain voiceprint and semantic information. Then, by matching the voiceprint information, the target sound region corresponding to the voice request is determined, thereby relatively accurately identifying the source of the voice request within the vehicle cabin. Simultaneously, the obtained semantic information is combined with the determined target sound region to generate a highly accurate vehicle control signal, thereby controlling the vehicle to perform the correct actions in the correct area and position within the vehicle cabin.

[0067] Thus, after obtaining a voice request, this application can perform voiceprint recognition and natural language understanding processing on the voice request separately, and distinguish and filter the voice region of the voice request based on the voiceprint recognition result to determine the corresponding voice region in the vehicle cabin for the generated control command. In addition, the content and direction of the control command are determined based on the natural language understanding result, so as to enable the vehicle to perform the correct action in the correct voice region and the correct position.

[0068] In some implementations, step 01 is followed by:

[0069] 011: In response to a voice request, the initial audio signal, including the voice request, is acquired through multiple audio collection devices installed in the vehicle cabin;

[0070] 012: Preprocess the initial audio signal to determine multiple audio signals corresponding to multiple sound zones in the vehicle cabin.

[0071] In some implementations, the information acquisition module is also configured to, in response to a voice request, acquire an initial audio signal including the voice request through multiple audio collection devices installed in the vehicle cabin, and preprocess the initial audio signal to determine multiple audio signals corresponding to multiple sound zones in the vehicle cabin.

[0072] In some implementations, the processor is also configured to, in response to a voice request, acquire an initial audio signal including the voice request via a plurality of audio collection devices disposed within the vehicle cabin, and preprocess the initial audio signal to determine a plurality of audio signals corresponding to a plurality of sound zones within the vehicle cabin.

[0073] Specifically, the process from acquiring a voice request to converting the voice request into an audio signal suitable for voiceprint recognition and natural language understanding can be carried out in the following manner:

[0074] In some examples, the user makes a voice request "open the window" in the 1st audio zone corresponding to the driver's seat. At this time, the six sets of microphones m1 to m6 in the vehicle cabin acquire the sound inside the vehicle cabin, acquiring a total of 6 initial audio signals.

[0075] Next, based on the sound zone distribution in the vehicle cabin, the acquired 6 initial audio signals are processed by multi-zone audio front-end calculation according to the sound zone. By separating and recombining, 4 audio signals corresponding to the 4 sound zones are determined, namely, driver's signal, passenger's signal, driver's rear position signal, and passenger's rear position signal.

[0076] Finally, the four audio signals are processed by VAD (Voice Activity Detection / Silence Suppression) to eliminate the parts of each audio signal that are identified as non-speech or residual speech, so as to retain as much of the effective audio as possible, which is convenient for subsequent voiceprint recognition and natural language understanding processing and improves the accuracy of processing.

[0077] Thus, this application can preprocess the acquired voice requests so that subsequent audio processing methods can be easily implemented.

[0078] In some implementations, step 02 includes:

[0079] 021: Perform voiceprint recognition on the audio signal to determine the voiceprint information included in the audio signal;

[0080] 022: Perform speech recognition on the audio signal to determine the speech recognition information included in the audio signal;

[0081] 023: Perform natural language understanding on the speech recognition information to determine the semantic information corresponding to the speech recognition information;

[0082] Based on this, step 03 includes:

[0083] The target sound region is determined based on the voiceprint information.

[0084] In some implementations, the audio processing module is used to perform voiceprint recognition on the audio signal to determine the voiceprint information included in the audio signal, and to perform speech recognition on the audio signal to determine the speech recognition information included in the audio signal, and to perform natural language understanding on the speech recognition information to determine the semantic information corresponding to the speech recognition information. The voice region determination module is used to determine the target voice region based on the voiceprint information.

[0085] In some embodiments, the processor is further configured to perform voiceprint recognition on the audio signal to determine the voiceprint information included in the audio signal, and to perform speech recognition on the audio signal to determine the speech recognition information included in the audio signal, and to perform natural language understanding on the speech recognition information to determine the semantic information corresponding to the speech recognition information, and to determine the target voice region based on the voiceprint information.

[0086] Specifically, after acquiring the four pre-processed audio signals, it is necessary to perform audio processing on the four audio signals to convert them into a format that is easy to match or recognize.

[0087] In some examples, after the in-vehicle system obtains four audio signals through multi-zone calculation at the audio front end and VAD processing, it uses a voiceprint recognition module (VPR) to perform voiceprint recognition on each of these four audio signals to determine the voiceprint information corresponding to each audio signal. This allows the system to match and compare the corresponding audio signal with other voiceprint information to determine the target zone. Simultaneously, an ASR module (ASR) performs speech recognition on each audio signal to determine the speech recognition information corresponding to each audio signal. This allows the Natural Language Understanding (NLU) module to identify the semantic information corresponding to the audio signal based on the speech recognition information, thereby deriving the correct vehicle control commands.

[0088] To ensure the correct vehicle control commands are obtained while reducing the complexity of the audio processing, in some examples, the process of identifying semantic information from speech recognition information via the Natural Language Understanding (NLU) module can be placed after the process of matching voiceprint information to determine the target voice region. The advantage of doing so is that speech recognition information corresponding to the target voice region can be selected for NLU processing based on the determined target voice region, while speech recognition information corresponding to non-target voice regions can be discarded without further processing, effectively reducing the number of processing items and improving processing efficiency.

[0089] Thus, this application is able to perform voiceprint recognition and natural language understanding processing on voice requests, and respectively determine the voiceprint recognition result used for subsequent matching and comparison and the natural language understanding result used to finally determine the content of vehicle control commands.

[0090] In some implementations, determining the target sound region based on voiceprint information includes:

[0091] 031: Based on the voiceprint information and the preset speech synthesis voiceprint database, determine the comparison result between the voiceprint information and the speech synthesis voiceprint database;

[0092] 032: If it is determined through comparison that the voiceprint information does not belong to the speech synthesis voiceprint database, the target voice region corresponding to the voice request is determined based on the voiceprint information that does not belong to the voiceprint database.

[0093] In some implementations, the voice region determination module is further configured to determine the comparison result between the voiceprint information and the speech synthesis voiceprint database based on the voiceprint information and the preset speech synthesis voiceprint database, and to determine the target voice region corresponding to the voice request based on the voiceprint information that does not belong to the speech synthesis voiceprint database if the comparison determines that the voiceprint information does not belong to the voiceprint database.

[0094] In some embodiments, the processor is further configured to determine the comparison result between the voiceprint information and the speech synthesis voiceprint database based on the voiceprint information and the preset speech synthesis voiceprint database, and to determine the target voice region corresponding to the voice request based on the voiceprint information that does not belong to the speech synthesis voiceprint database if the comparison determines that the voiceprint information does not belong to the voiceprint database.

[0095] Specifically, in the process of determining the target audio region through voiceprint matching, in some examples, to exclude the case where the voice request corresponding to the determined audio signal is synthesized speech issued by the in-vehicle system, the voiceprint database used by the in-vehicle system to store synthesized speech voiceprints is called. The voiceprint information of each audio signal is compared with each voiceprint information in the voiceprint database. The voiceprint database stores the voiceprint information corresponding to synthesized speech issued by the in-vehicle system within a preset time period before the current moment, such as the voiceprint information corresponding to synthesized speech issued by the in-vehicle system within the last 5 days. If, after comparison, no voiceprint information in the voiceprint database is the same as the voiceprint information corresponding to the audio signal being matched, it means that the voiceprint information compared with the voiceprint database is not synthesized speech issued by the vehicle. In this case, the above voiceprint information is retained, and then, after all the voiceprint information of all 4 audio signals has been matched, the target audio region is determined based on the retained voiceprint information.

[0096] Thus, this application can avoid interference from synthesized speech emitted by vehicles on voice request recognition by comparing the determined voiceprint information with a speech synthesis voiceprint database.

[0097] In some implementations, the vehicle voice interaction method further includes:

[0098] 033: If the voiceprint information is determined to belong to the speech synthesis voiceprint database after comparison, the audio signal corresponding to the voiceprint information is discarded.

[0099] In some implementations, the voice region determination module is also used to discard the audio signal corresponding to the voiceprint information if it is determined by comparison that the voiceprint information belongs to the speech synthesis voiceprint database.

[0100] In some implementations, the processor is also configured to discard the audio signal corresponding to the voiceprint information if the comparison determines that the voiceprint information belongs to the speech synthesis voiceprint database.

[0101] Specifically, based on the above implementation method, if, after comparison, at least one voiceprint information in the voice synthesis voiceprint database is identical to the voiceprint information corresponding to the audio signal being matched, it indicates that the voiceprint information being compared with the voice synthesis voiceprint database is synthesized speech emitted by the vehicle. In order to avoid interference from the synthesized speech to the recognition of the voice request, the vehicle system will discard the audio signal corresponding to the voiceprint information.

[0102] Thus, this application can also discard the voice request corresponding to the voiceprint information when there is a matching relationship between the voiceprint information and the speech synthesis voiceprint database, thereby eliminating the interference of the synthesized speech information on the voice request.

[0103] In some implementations, determining the target sound region based on voiceprint information further includes:

[0104] 034: Compare multiple voiceprint information to determine whether the multiple voiceprint information comes from the first object;

[0105] 035: When it is determined that multiple voiceprint information comes from the first object, the audio region corresponding to the audio signal that includes voiceprint information that meets the preset conditions of the silence suppression processing is determined as the target audio region based on the silence suppression processing result.

[0106] In some implementations, the voice region determination module is also used to compare multiple voiceprint information to determine whether the multiple voiceprint information comes from the first object, and to determine the voice region corresponding to the audio signal including voiceprint information that meets the preset conditions of the silence suppression processing as the target voice region when it is determined that the multiple voiceprint information comes from the first object, based on the silence suppression processing result.

[0107] In some implementations, the processor is further configured to compare multiple voiceprint information to determine whether the multiple voiceprint information comes from the first object, and to determine the audio region corresponding to the audio signal including voiceprint information that meets the preset conditions of the silence suppression processing as the target audio region when it is determined that the multiple voiceprint information comes from the first object, based on the silence suppression processing result.

[0108] Specifically, generally speaking, in the vehicle cabin described in the above embodiments, voice requests from the same object within the target sound zone are typically recorded by multiple audio signals. For example, if a user in the driver's seat makes a voice request, each of the four pre-processed audio signals may still contain the user's voice. Therefore, to more accurately determine the target sound zone, it is necessary to compare the voiceprint information of multiple audio signals. When the comparison reveals that the voiceprint information of two or more audio signals from different channels originates from the same object, it indicates that only one set of these audio signals corresponds to the target sound zone, and the audio signals other than this set are not the target sound zone.

[0109] Therefore, in some examples, when the voiceprint information of two or more audio signals from different sources originates from the same object, it is necessary to filter based on the VAD processing results of the audio signals. When obtaining voiceprint information from audio signals originating from the same object through VAD processing, the VAD model first divides the audio signal to be processed into multiple segments. For example, an audio signal of approximately 1 second in length is truncated into 50 segments, each 20ms long. Then, the VAD model determines the score of each segment, where the score reflects the probability that each segment is valid information, ranging from 0 to 1. Next, the score of each segment is compared with a preset score threshold. Segments with scores greater than or equal to the threshold are considered valid information, while those less than the threshold are considered residual invalid information. Finally, the scores of all the segments considered valid information are averaged to obtain the overall VAD model score for the corresponding audio signal. The audio region corresponding to the highest-scoring audio signal is determined as the target audio region, and the other audio signals are discarded. This allows for the filtering of audio signals when multiple audio signals contain the voiceprint of the same object, narrowing down the range of target audio regions and improving the efficiency of determining target audio regions.

[0110] Thus, this application can also filter multiple audio signals by determining whether multiple voiceprint information comes from the same object, so as to determine the unique target sound region corresponding to the same object.

[0111] In some implementations, determining the target sound region based on voiceprint information further includes:

[0112] 036: When it is determined that multiple voiceprint information comes from different objects other than the first object, the target sound region is determined based on the multiple voiceprint information.

[0113] In some implementations, the voice region determination module is further configured to determine the target voice region based on the voiceprint information from the non-first object when it is determined that multiple voiceprint information sources are from a non-first object.

[0114] In some implementations, the processor is also configured to determine the target sound region based on the voiceprint information from the non-first object when it is determined that multiple voiceprint information sources are from a non-first object.

[0115] Specifically, based on the above implementation method, if multiple voiceprint information comes from different objects, it means that the audio signals corresponding to these voiceprint information may come from different sound regions. In order to ensure the integrity of the information in the voice request, all of the above audio signals should be retained, and the sound regions corresponding to the above audio signals should also be determined as the target sound region.

[0116] Specifically, in some examples, among the four audio signals used for voiceprint information comparison, the following situation may exist: the voiceprints of two of the four audio signals come from the first object, while the voiceprints of the other two are different and both come from objects other than the first object. For such cases, the solution provided in the above embodiments can be applied. For the two audio signals from the first object, based on the VAD model's assignment of scores according to the VAD model for each audio signal during VAD processing, the audio signal with the higher VAD model score is determined as the target audio region. Simultaneously, for the other two audio signals that do not come from the first object, both audio regions corresponding to these two audio signals are determined as target audio regions.

[0117] Thus, this application can also directly determine the corresponding audio region based on the multi-channel voiceprint information when the multi-channel voiceprint signals do not originate from the same object, without going through a screening process, thereby achieving the goal of maintaining the information integrity of the multi-channel audio signals as realistically as possible.

[0118] In some implementations, determining the target sound region based on voiceprint information includes:

[0119] 037: Compare the voiceprint information with the preset historical voiceprint database;

[0120] 038: Based on the comparison results, the corresponding voiceprint information is corrected to determine the target voice region.

[0121] In some implementations, the voice region determination module is also used to compare the voiceprint information with a preset historical voiceprint database, and to correct the voice region corresponding to the voiceprint information based on the comparison results, so as to determine the target voice region.

[0122] In some implementations, the processor is also used to compare the voiceprint information with a preset historical voiceprint database, and to correct the voice region corresponding to the voiceprint information based on the comparison results, thereby determining the target voice region.

[0123] Specifically, in long-term vehicle applications, a user's position inside the vehicle is relatively fixed. Therefore, during repeated use of the voice interaction function, the content of the voice request is correlated with the user's position inside the vehicle. Thus, after each use of the voice interaction function, voiceprint information corresponding to multiple voice requests with clear vocal range information and high confidence in voice recognition results can be selected and stored in a database to establish a preset historical voiceprint database. This database can then provide a reference and basis for subsequent target vocal range determination.

[0124] In order to improve the efficiency and accuracy of the process of determining the target voice region, before comparing the voiceprint information with the preset historical voiceprint database, based on the implementation method described above, the number of voiceprint information for determining the target voice region can be compressed and the range of the target voice region can be narrowed by prioritizing the matching of the speech synthesis voiceprint database and the filtering process of audio signals corresponding to voiceprints from the same object, thereby effectively improving the range of the target voice region.

[0125] Then, in some examples, the remaining voiceprint information is compared with each voiceprint in the preset historical voiceprint database. If there is at least one voiceprint in the preset historical voiceprint database that is the same as the voiceprint used for matching, the sound region corresponding to the voiceprint in the preset historical voiceprint database is redefined as the target sound region corresponding to the voiceprint used for matching.

[0126] For example, after matching with the voiceprint database and filtering out audio signals corresponding to voiceprints from the same object, two of the four audio signals are discarded. The voiceprint signals corresponding to the remaining two audio signals are then compared with each voiceprint information in the preset historical voiceprint database. If one of the two sets of voiceprint signals corresponding to these two audio signals is the same as one voiceprint information in the preset historical voiceprint database, then the voice region corresponding to the voiceprint information in the preset historical voiceprint database is determined as the target voice region.

[0127] In addition, if there are multiple voiceprints in the preset historical voiceprint database that are the same as the voiceprint information used for matching, then one of the voiceprints that are successfully matched in the preset historical voiceprint database can be randomly selected as the target voice region.

[0128] Thus, this application can also improve the accuracy of target voice region judgment by comparing voiceprint information with information in the historical voiceprint database and correcting the corresponding voice region information.

[0129] In some implementations, step 04 includes:

[0130] 041: Generate vehicle control commands corresponding to the voice request based on the target voice region and semantic information;

[0131] 042: Execute the actions corresponding to the vehicle control commands according to the vehicle control commands to complete the voice interaction.

[0132] In some implementations, the command control module is also used to generate vehicle control commands corresponding to the voice request based on the target voice region and semantic information, and to execute actions corresponding to the vehicle control commands to complete the voice interaction.

[0133] In some implementations, the processor is also configured to generate vehicle control commands corresponding to the voice request based on the target voice region and semantic information, and to execute actions corresponding to the vehicle control commands to complete the voice interaction.

[0134] Specifically, after the target sound zone is determined, the semantic information corresponding to the voiceprint information used to determine the target sound zone is combined with the target sound zone information and input into the vehicle system. The vehicle system then generates vehicle control commands based on the target sound zone and the corresponding semantic information, and controls the vehicle to perform the correct actions in the correct position or the correct area through the vehicle control commands.

[0135] For example, in some implementations, based on the above implementation, the user sends a voice request to the driver's seat to "open the window". After the target sound zone determination process described above, the driver's sound zone is determined as the target sound zone. Then, the vehicle system generates a vehicle control command "open the driver's side window" based on the semantic information of "open the window" and the target sound zone of the driver's sound zone, and controls the driver's side window to open according to the vehicle control command.

[0136] Thus, after the voiceprint recognition and matching process is completed, this application can combine the information obtained through natural language understanding processing with the target voice region that has just been determined as a basis to determine the vehicle control command for controlling the vehicle to perform the correct actions, and correctly control the vehicle according to the vehicle control command.

[0137] The computer-readable storage medium in the embodiments of this application stores a computer program that, when executed by one or more processors, implements the vehicle voice interaction method as described in any of the above embodiments.

[0138] In the description of this specification, the references to terms such as "some embodiments," "in one example," "exemplarily," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0139] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.

[0140] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A vehicle voice interaction method, characterized in that, The method includes: Get the voice request; In response to the voice request, an initial audio signal including the voice request is acquired by multiple audio collection devices installed in the vehicle cabin; The initial audio signal is preprocessed to determine multiple audio signals corresponding to multiple sound zones in the vehicle cabin. Performing voiceprint recognition and natural language understanding on the voice request includes: Perform voiceprint recognition on the audio signal to determine the voiceprint information included in the audio signal; Based on the first result of the voiceprint recognition, the target voice region corresponding to the voice request within the vehicle cabin is determined, including: The target sound region is determined based on the voiceprint information; Based on the target voice region and the second result of natural language understanding, a vehicle control command corresponding to the voice request is generated so that the vehicle can complete the voice interaction according to the vehicle control command. Determining the target voice region based on the voiceprint information includes: Based on the voiceprint information and the preset speech synthesis voiceprint database, determine the comparison result between the voiceprint information and the speech synthesis voiceprint database; If it is determined through comparison that the voiceprint information does not belong to the speech synthesis voiceprint database, the target voice region corresponding to the speech request is determined based on the voiceprint information that does not belong to the voiceprint database.

2. The method of claim 1, wherein, The process of performing voiceprint recognition and natural language understanding on the voice request also includes: Perform speech recognition on the audio signal to determine the speech recognition information included in the audio signal; Natural language understanding is performed on the speech recognition information to determine the semantic information corresponding to the speech recognition information.

3. The method of claim 1, wherein, The method further includes: If, after comparison, it is determined that the voiceprint information belongs to the speech synthesis voiceprint database, the audio signal corresponding to the voiceprint information is discarded.

4. The method of claim 2, wherein, Determining the target voice region based on the voiceprint information includes: The multiple voiceprint information pieces are compared to determine whether the multiple voiceprint information pieces come from the first object; If it is determined that multiple voiceprint information sources originate from the first object, the audio region corresponding to the audio signal that includes voiceprint information that meets the preset conditions for silence suppression processing is determined as the target audio region based on the silence suppression processing result.

5. The method according to claim 4, characterized in that, The step of determining the target voice region based on the voiceprint information further includes: If it is determined that multiple voiceprint information sources are from different objects other than the first object, the target voice region is determined based on the multiple voiceprint information sources.

6. The method of claim 2, wherein, Determining the target voice region based on the voiceprint information includes: The voiceprint information is compared with a preset historical voiceprint database; Based on the comparison results, the voiceprint information corresponding to the voice region is corrected to determine the target voice region.

7. The method of claim 2, wherein, The step of generating vehicle control commands corresponding to the voice request based on the target voice region and the second result of natural language understanding includes: Based on the target voice region and the semantic information, a vehicle control command corresponding to the voice request is generated; According to the vehicle control command, execute the action corresponding to the vehicle control command to complete the voice interaction.

8. A vehicle voice interaction apparatus characterized by comprising: The device includes: An information acquisition module is used to acquire a voice request; in response to the voice request, it acquires an initial audio signal including the voice request through multiple audio collection devices installed in the vehicle cabin; and preprocesses the initial audio signal to determine multiple audio signals corresponding to multiple sound zones in the vehicle cabin. The audio processing module is used to perform voiceprint recognition and natural language understanding on the voice request; The process of performing voiceprint recognition and natural language understanding on the voice request includes: Perform voiceprint recognition on the audio signal to determine the voiceprint information included in the audio signal; The voice region determination module is used to determine the target voice region in the vehicle cabin corresponding to the voice request based on the first result of the voiceprint recognition. The step of determining the target voice region in the vehicle cabin corresponding to the voice request based on the first result of the voiceprint recognition includes: The target sound region is determined based on the voiceprint information; Based on the target voice region and the second result of natural language understanding, a vehicle control command corresponding to the voice request is generated so that the vehicle can complete the voice interaction according to the vehicle control command. Determining the target voice region based on the voiceprint information includes: Based on the voiceprint information and the preset speech synthesis voiceprint database, determine the comparison result between the voiceprint information and the speech synthesis voiceprint database; If it is determined through comparison that the voiceprint information does not belong to the speech synthesis voiceprint database, the target voice region corresponding to the voice request is determined based on the voiceprint information that does not belong to the voiceprint database. The instruction control module is used to generate vehicle control instructions corresponding to the voice request based on the target voice region and the second result of natural language understanding, so that the vehicle can complete the voice interaction according to the vehicle control instructions.

9. A vehicle characterized by comprising: The vehicle includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the vehicle to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Vehicle control method and system and vehicle

    CN111754994A

  • In-vehicle user positioning method, vehicle-mounted interaction method, vehicle-mounted device and vehicle

    CN112655000A