Voice interaction method, vehicle and computer readable storage medium

By using a pre-trained voice request processing model in the vehicle for noise reduction and voice activity detection, the problem that vehicles are difficult to distinguish human voices under noise interference is solved, and a robust voice interaction function and user experience improvement is achieved.

CN120071908APending Publication Date: 2025-05-30GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510229359.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In vehicles, noise interference makes it difficult for the vehicle to distinguish human voices from mixed sounds, and thus cannot wake up the voice interaction function, affecting the user experience.

Method used

The pre-trained voice request processing model is used to reduce noise and voice activity detection processing on the current voice request, and determine the noise reduction voice request and voice activity detection results, thereby controlling the vehicle for voice interaction.

Benefits of technology

Through noise reduction and voice activity detection processing, the vehicle can accurately recognize the vocal part, steadily wake up the voice interaction function, and respond to the user's voice commands to improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071908A_ABST
    Figure CN120071908A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method, a vehicle and a computer readable storage medium, and the method comprises the steps: obtaining a current voice request, carrying out the noise reduction and voice activity detection processing of the current voice request according to a voice request processing model which is trained in advance, determining a noise reduction voice request and a voice activity detection result, and controlling the vehicle according to the noise reduction voice request and the voice activity detection result, and performing voice interaction. Therefore, the noise reduction and voice activity detection processing of the current voice request can be jointly performed by the voice request processing model which is trained in advance, so that the noise reduction processing of the current voice request can be associated with the voice activity detection processing of the current voice request, and the effectiveness of the noise reduction processing and the voice activity detection is ensured; therefore, the use experience of the user on the vehicle voice interaction function can be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of vehicles, and particularly to a voice interaction method, a vehicle, and a computer-readable storage medium. Background Art

[0002] In related vehicles, the vehicle can be equipped with a voice interaction function for the user to control the vehicle by voice. However, during driving, various noises such as wind noise and road noise may exist in the vehicle cockpit space. After the user utters a voice command, the human voice is mixed with the noise, making it difficult for the vehicle to distinguish the human voice from the mixed sound, thus unable to wake up the vehicle voice interaction function and failing to respond to the user's voice command, which affects the user's experience of using the vehicle voice interaction function. Summary of the Invention

[0003] This application provides a voice interaction method, a vehicle, and a computer-readable storage medium.

[0004] An embodiment of this application provides a voice interaction method for a vehicle. The method includes:

[0005] Obtain a current voice request;

[0006] According to a pre-trained voice request processing model, perform noise reduction and voice activity detection processing on the current voice request to determine a noise-reduced voice request and a voice activity detection result;

[0007] Control the vehicle according to the noise-reduced voice request and the voice activity detection result to perform the voice interaction.

[0008] In this way, in the embodiment of this application, the noise reduction and voice activity detection processing of the current voice request can be jointly performed by a pre-trained voice request processing model, so that the noise reduction processing of the current voice request can be associated with the voice activity detection processing of the current voice request. Compared with the situation where the noise reduction processing and the voice activity detection processing are independently set, the effectiveness of the noise reduction processing of the current voice request and the voice activity detection of the current voice request can be guaranteed to a certain extent. Moreover, the vehicle can be controlled according to the noise-reduced voice request and the voice activity detection result, so that the vehicle can accurately identify the human voice part in the current voice request through the noise-reduced voice request and the voice activity detection result, and then can robustly wake up the vehicle voice interaction function and respond to the voice command pointed to by the current voice request, thereby guaranteeing the user's experience of using the vehicle voice interaction function.

[0009] In some embodiments of this application, the voice request processing model includes an encoder module, a voice activity detection module, and a decoder module connected in sequence;

[0010] The encoder module is configured to encode the current voice request and determine the voice request encoding result;

[0011] The voice activity detection module is configured to determine the voice activity detection result according to the voice request encoding result;

[0012] The decoder module is configured to perform decoding processing on the voice request encoding result to determine the noise-reduced voice request.

[0013] In this way, in the embodiment of the present application, the voice request processing model can be implemented by sequentially connecting an encoder module, a voice activity detection module, and a decoder module.

[0014] In some embodiments of the present application, the voice activity detection module is configured to perform temporal information modeling processing on the voice request encoding result to determine the temporal information modeling result of the voice request encoding result, and perform voice activity detection on the temporal information modeling result to determine the voice activity detection result;

[0015] The decoder module is configured to perform decoding processing on the temporal information modeling result of the voice request encoding result to determine the noise-reduced voice request.

[0016] In this way, in the embodiment of the present application, the voice activity detection module can perform temporal information modeling processing on the voice request encoding result to determine the temporal information modeling result of the voice request encoding result, and perform voice activity detection on the temporal information modeling result to determine the voice activity detection result. Correspondingly, the decoder module can perform decoding processing on the temporal information modeling result of the voice request encoding result to determine the noise-reduced voice request, thereby realizing noise reduction and voice activity detection of the current voice request.

[0017] In some embodiments of the present application, the encoder module is connected to the decoder module;

[0018] The decoder module is configured to perform decoding processing on the temporal information modeling result and the voice request encoding result to determine the noise-reduced voice request.

[0019] In this way, in the embodiment of the present application, the decoder module can jointly perform decoding processing on the temporal information modeling result and the voice request encoding result to determine the noise-reduced voice request, thereby ensuring the effectiveness and reliability of the noise-reduced voice request to a certain extent.

[0020] In some embodiments of the present application, the encoder module includes a plurality of downsampling units connected in sequence, the decoder module includes a plurality of upsampling units connected in sequence, and one downsampling unit is connected to one upsampling unit;

[0021] The downsampling unit is configured to perform downsampling processing on the first input data to determine the downsampling processing result of the first input data. The first input data includes the voice request sample and / or the downsampling processing result of the previous downsampling unit, and the voice request coding result is the downsampling processing result of the last downsampling unit;

[0022] The upsampling unit is configured to perform upsampling processing on the second input data to determine the upsampling result of the second input data. The noise-reduced voice request is the upsampling result of the last upsampling unit. The second input data includes the timing information modeling information, the first input data of the downsampling unit connected to the upsampling unit, or includes the upsampling processing result of the previous upsampling unit and the downsampling processing result of the downsampling unit connected to the upsampling unit.

[0023] In this way, in the embodiment of the present application, the encoder module can be implemented based on a plurality of sequentially connected downsampling units, and the decoder module can be implemented based on a plurality of sequentially connected upsampling units, thereby ensuring the stable operation of the encoder module and the decoder module.

[0024] In some embodiments of the present application, the noise reduction and voice activity detection processing of the current voice request according to the pre-trained voice request processing model to determine the noise-reduced voice request and the voice activity detection result includes:

[0025] Preprocess the current voice request to determine the processed voice request;

[0026] According to the voice request processing model, perform noise reduction and voice activity detection processing on the processed voice request to determine the noise-reduced voice request and the voice activity detection result.

[0027] In this way, in the embodiment of the present application, the current voice request can be preprocessed to determine the processed voice request, and the processed voice request can be subjected to noise reduction and voice activity detection processing according to the voice request processing model to determine the noise-reduced voice request and the voice activity detection result, so that the current voice request can be preprocessed before being input into the voice request processing model to improve the quality of the current voice request, thereby improving the voice activity detection result and the effectiveness of the noise-reduced voice request to a certain extent.

[0028] In some embodiments of the present application, controlling the vehicle according to the noise-reduced voice request and the voice activity detection result to perform the voice interaction includes:

[0029] Perform processing on the noise-reduced voice request for gain to determine the voice request after gain processing;

[0030] Control the vehicle according to the voice request after gain processing and the voice activity detection result.

[0031] In this way, in the embodiment of the present application, the noise-reduced voice request can be processed for gain to determine the voice request after gain processing, so that the quality of the noise-reduced voice request can be guaranteed to a certain extent, and further the robustness of controlling the vehicle according to the voice request after gain processing and the voice activity detection result can be guaranteed.

[0032] In some embodiments of the present application, the method further includes:

[0033] Perform model format conversion processing on the pre-trained voice request processing model to determine the voice request processing model in a preset format;

[0034] Configure a runtime library corresponding to the voice request processing model in the preset format, and perform deployment of the voice request processing model in the vehicle.

[0035] In this way, in the embodiment of the present application, the pre-trained voice request processing model can be subjected to model format conversion processing to determine the voice request processing model in a preset format, and a runtime library corresponding to the voice request processing model in the preset format can be configured, so as to perform deployment of the voice request processing model in the vehicle, so that the vehicle can process the current voice request through the voice request processing model deployed locally.

[0036] An embodiment of the present application provides a vehicle, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.

[0037] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by one or more processors, the above-mentioned voice interaction method is implemented.

[0038] The vehicle and computer-readable storage medium provided by the embodiments of the present application enable the noise reduction and voice activity detection processing of the current voice request to be jointly performed by a pre-trained voice request processing model, enabling the noise reduction processing of the current voice request to be associated with the voice activity detection processing of the current voice request. Compared with the situation where the noise reduction processing and the voice activity detection processing are independently set, it can ensure the effectiveness of the noise reduction processing of the current voice request and the voice activity detection of the current voice request to a certain extent. Moreover, the vehicle can be controlled according to the noise-reduced voice request and the voice activity detection result, enabling the vehicle to accurately identify the human voice part in the current voice request through the noise-reduced voice request and the voice activity detection result, and then robustly wake up the vehicle voice interaction function and respond to the voice command pointed to by the current voice request. Thus, the user experience of using the vehicle voice interaction function can be guaranteed.

[0039] Additional aspects and advantages of the embodiments of the present application will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0041] Figure 1 is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0042] Figure 2 is a schematic diagram of a voice request processing model in some embodiments of the present application;

[0043] Figure 3 is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0044] Figure 4 is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0045] Figure 5 is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0046] Figure 6 is a schematic flow chart of a voice interaction method in some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and should not be construed as a limitation on the embodiments of the present application.

[0048] In in-vehicle high-noise scenarios, such as high-speed driving, high air volume of the air conditioner, opening windows, etc., the performance of the vehicle voice wake-up system is often severely interfered by environmental noise and deteriorates. Moreover, when the user speaks the wake-up word in a relatively low voice, the wake-up rate of the voice wake-up system will further decrease. It can be understood that this situation not only affects the user experience but may also cause the driver's attention to be distracted, posing a safety hazard.

[0049] Therefore, in related technologies, solutions such as noise reduction enhancement and joint optimization of the front-end signal and the wake-up model are usually adopted to improve the wake-up rate of the vehicle voice wake-up system. However, these solutions still have limitations in in-vehicle high-noise scenarios. For example: 1) The low-volume wake-up word in a high-noise scenario belongs to a very low signal-to-noise ratio situation. Therefore, when performing noise reduction processing on such data, it is easy to accidentally damage the human voice part where the harmonic information is not clear originally; 2) Since the voice wake-up model usually has the characteristics of a simple structure and a small number of network parameters, in the scenarios of high noise and low-volume voice, the generalization performance of the back-end wake-up model is insufficient; 3) Affected by the extremely low signal-to-noise ratio factor, the VAD (Voice Activity Detection) module before the front-end signal is processed and enters the wake-up engine is also prone to misjudging small-volume wake-up words. For example, because the VAD module usually has a smoothing strategy, it is more difficult to judge the beginning part of the wake-up word.

[0050] Based on the above possible problems, please refer to Figure 1 , an embodiment of the present application provides a voice interaction method for a vehicle, and the method includes:

[0051] 01: Obtain a current voice request;

[0052] 02: According to a pre-trained voice request processing model, perform noise reduction and voice activity detection processing on the current voice request to determine a noise-reduced voice request and a voice activity detection result;

[0053] 03: Control the vehicle according to the noise-reduced voice request and the voice activity detection result to perform voice interaction.

[0054] Embodiments of the present application provide a voice interaction device. The voice interaction method according to the embodiments of the present application can be implemented by the voice interaction device according to the embodiments of the present application. Specifically, the voice interaction device includes an acquisition module, a processing module, and an interaction module. Among them, the acquisition module is used to acquire the current voice request. The processing module is used to perform noise reduction and voice activity detection processing on the current voice request according to the pre-trained voice request processing model, and determine the noise-reduced voice request and the voice activity detection result. The interaction module is used to control the vehicle according to the noise-reduced voice request and the voice activity detection result to perform voice interaction.

[0055] Embodiments of the present application further provide a vehicle, which includes a memory and a processor. The voice interaction method according to the embodiments of the present application can be implemented by the vehicle according to the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is used to acquire the current voice request, and is used to perform noise reduction and voice activity detection (Voice Activity Detection, VAD) processing on the current voice request according to the pre-trained voice request processing model, determine the noise-reduced voice request and the voice activity detection result, and is used to control the vehicle according to the noise-reduced voice request and the voice activity detection result to perform voice interaction.

[0056] Specifically, in the embodiments of the present application, when a user in the vehicle cockpit wants to control the vehicle to perform actions such as turning off the air conditioner and opening the window at the current moment, and thus says statements such as "turn off the air conditioner" and "Xiaop, open the window", the vehicle can capture the corresponding sound through a sound collection component such as a microphone to obtain the current voice request. Then, the vehicle can call the pre-trained voice request processing model to perform noise reduction and voice activity detection (Voice Activity Detection, VAD) processing on the current voice request, determine the current voice request after noise reduction, that is, the noise-reduced voice request, and determine the voice activity detection result of the current voice request. Finally, the vehicle can perform corresponding control operations according to the noise-reduced voice request and the voice activity detection result, such as controlling the vehicle air conditioner to turn off and controlling the window to open.

[0057] Thus, in the embodiments of the present application, the noise reduction and voice activity detection processing of the current voice request can be jointly performed by a pre-trained voice request processing model, so that the noise reduction processing of the current voice request can be associated with the voice activity detection processing of the current voice request. Compared with the situation where the noise reduction processing and the voice activity detection processing are independently set, to a certain extent, the effectiveness of the noise reduction processing of the current voice request and the voice activity detection of the current voice request can be ensured. Moreover, the vehicle can be controlled according to the noise-reduced voice request and the voice activity detection result, so that the vehicle can accurately identify the human voice part in the current voice request through the noise-reduced voice request and the voice activity detection result, and then can robustly wake up the vehicle voice interaction function and respond to the voice command pointed to by the current voice request. Thereby, the user experience of using the vehicle voice interaction function can be guaranteed.

[0058] In addition, since the voice request processing model is responsible for both the noise reduction and voice activity detection of the current voice request, in order to ensure the effective noise reduction of the current voice request and the accuracy of the voice activity detection of the current voice request, the parameters such as the number of parameters and the network depth of the voice request processing model can be set to a relatively large value. Further, when the parameters such as the number of parameters and the network depth of the voice request processing model are relatively large, the noise reduction of the voice request and the voice activity detection of the voice request can be implemented based on the voice request processing model with relatively large parameters such as the number of parameters and / or the network depth. Furthermore, compared with the VAD (Voice Activity Detection) module (or voice wake-up module) with a simple structure and fewer parameters in the related art, in the embodiments of the present application, the vehicle can relatively reliably perform the voice activity detection of the current voice request through the voice request processing model with relatively large parameters such as the number of parameters and / or the network depth, so as to accurately identify whether the user has the intention to wake up the vehicle voice interaction function (or in-vehicle computer), and then can ensure the effective wake-up of the vehicle voice interaction function (or in-vehicle computer) and the reliable response to the user's voice command. Therefore, in the scenarios of high noise and low-volume voice, the noise reduction and voice activity detection of the current voice request can be performed through the voice request processing model, thereby accurately responding to the user's voice command.

[0059] It can be understood that in the embodiments of the present application, the current voice request can be understood as the voice information obtained by the vehicle through the sound collection components such as microphones set in the cockpit space. It can also be understood that in addition to the user command voice (such as "turn on the air conditioner", "Xiaop, close the window") issued by the user, the sound in the cockpit space may also include environmental noise, such as road noise, tire noise, wind noise and other sounds made by users.

[0060] It can be understood that in the embodiments of the present application, the structure, framework, etc. of the voice request processing model can be set according to the actual situation. For example, in one example, the voice request processing model is a CRN (Convolutional Recurrent Network) model with a CNN (Convolutional Neural Network) as the main structure.

[0061] It can also be understood that in the embodiments of the present application, noise reduction for voice requests can be used to suppress sounds other than human voices in voice requests, such as road noise, tire noise, wind noise, and conversations of other users.

[0062] In addition, in the embodiments of the present application, voice activity detection for voice requests can be used to determine whether there is voice in the voice request, such as whether there is voice when the user says a sentence like "turn on the air conditioner".

[0063] Furthermore, in the embodiments of the present application, the voice activity detection result can be used to represent whether there is a sound (or human voice) emitted by the user in the current voice request. Furthermore, when the voice activity detection result represents that there is a sound (or human voice) emitted by the user in the current voice request, the vehicle can perform natural language processing on the noise-reduced voice request, such as recognizing the text corresponding to the noise-reduced voice request, and controlling vehicle components to perform corresponding operations according to the text recognition result. For example, when the text recognized corresponding to the noise-reduced voice request is "turn on the air conditioner", the vehicle air conditioner is controlled to turn on.

[0064] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the voice request processing model in some embodiments of the present application. That is, as Figure 2 shown, in some embodiments of the present application, the voice request processing model includes an encoder module 101, a voice activity detection module 102, and a decoder module 103 connected in sequence. The encoder module 101 is configured to encode the current voice request 104 to determine a voice request encoding result. The voice activity detection module 102 is configured to determine a voice activity detection result according to the voice request encoding result. The decoder module 103 is configured to perform decoding processing on the voice request encoding result to determine a noise-reduced voice request 105.

[0065] Specifically, as Figure 2 shown, in the embodiments of the present application, the voice request processing model is composed of an encoder module 101, a voice activity detection module 102, and a decoder module 103 connected in sequence.

[0066] In one example, the Encoder module 101 can downsample the current voice request 104 to encode the current voice request 104, and then output the encoded result of the current voice request, that is, the voice request encoding result.

[0067] In one example, the voice activity detection module 102 can detect whether there is voice in the current voice request 104 according to the voice request encoding result output by the Encoder module 101, such as whether there is voice when the user says statements like "turn on the air conditioner", and then output the voice activity detection result.

[0068] In one example, the Decoder module 103 can upsample the voice request encoding result output by the Encoder module 101 to decode the voice request encoding result, and then output the current voice request after noise reduction, that is, the noise-reduced voice request 105.

[0069] In this way, in the embodiment of the present application, the voice request processing model can be implemented by the sequentially connected Encoder module 101, voice activity detection module 102, and Decoder module 103.

[0070] In some embodiments of the present application, the voice activity detection module 102 is configured to perform temporal information modeling processing on the voice request encoding result to determine the temporal information modeling result of the voice request encoding result, and perform voice activity detection on the temporal information modeling result to determine the voice activity detection result. The Decoder module 103 is configured to perform decoding processing on the temporal information modeling result of the voice request encoding result to determine the noise-reduced voice request.

[0071] Specifically, in the embodiment of the present application, the voice activity detection module 102 can perform temporal information modeling processing on the voice request encoding result, and thus obtain the temporal information modeling result of the voice request encoding result.

[0072] As Figure 2 shown, in one example, the voice activity detection module 102 includes two layers of LSTM (Long Short-Term Memory) networks connected in sequence. Correspondingly, the voice activity detection module 102 can model temporal information through these two layers of LSTM networks, that is, perform temporal information modeling processing on the voice request encoding result, and output the temporal information modeling result of the voice request encoding result.

[0073] In addition, in the embodiment of the present application, the voice activity detection module 102 can also perform voice activity detection on the temporal information modeling result of the voice request encoding result, and output the corresponding voice activity detection result.

[0074] In addition, in the embodiment of the present application, the decoder module 103 connected to the voice activity detection module 102 can also receive or obtain the timing information modeling result when the voice activity detection module 102 outputs the timing information modeling result of the voice request coding result, and decode the timing information modeling result to restore the timing information modeling result to a "voice request with the same dimension as the current voice request", that is, the noise-reduced voice request 105.

[0075] In this way, in the embodiment of the present application, the voice activity detection module 102 can perform timing information modeling processing on the voice request coding result to determine the timing information modeling result of the voice request coding result, and perform voice activity detection on the timing information modeling result to determine the voice activity detection result. Correspondingly, the decoder module 103 can perform decoding processing on the timing information modeling result of the voice request coding result to determine the noise-reduced voice request, thereby realizing noise reduction and voice activity detection of the current voice request.

[0076] In some embodiments of the present application, the encoder module is connected to the decoder module, and the decoder module is configured to perform decoding processing on the timing information modeling result and the voice request coding result to determine the noise-reduced voice request.

[0077] Specifically, in order to avoid problems such as gradient disappearance or loss of detailed information, in the embodiment of the present application, the decoder module 103 is also connected to the encoder module 101. Furthermore, the decoder module 103 can receive the timing information modeling result output by the voice activity detection module 102, and can obtain the voice request coding result output by the encoder module 101, and jointly perform decoding processing on the timing information modeling result and the voice request coding result to obtain the noise-reduced voice request 105.

[0078] In this way, in the embodiment of the present application, the decoder module 103 can jointly perform decoding processing on the timing information modeling result and the voice request coding result to determine the noise-reduced voice request 105, thereby ensuring the effectiveness and reliability of the noise-reduced voice request 105 to a certain extent.

[0079] In some embodiments of the present application, the encoder module 101 includes a plurality of downsampling units connected in sequence, and the decoder module 103 includes a plurality of upsampling units connected in sequence. A downsampling unit is connected to an upsampling unit. The downsampling unit is configured to perform downsampling processing on the first input data to determine the downsampling processing result of the first input data. The first input data includes a voice request sample and / or the downsampling processing result of a previous downsampling unit. The voice request encoding result is the downsampling processing result of the last downsampling unit. The upsampling unit is configured to perform upsampling processing on the second input data to determine the upsampling result of the second input data. The denoised voice request is the upsampling result of the last upsampling unit. The second input data includes temporal information modeling information, the first input data of the downsampling unit connected to the upsampling unit, or includes the upsampling processing result of a previous upsampling unit and the downsampling processing result of the downsampling unit connected to the upsampling unit.

[0080] Specifically, in the embodiments of the present application, the encoder module 101 is composed of a plurality of downsampling units connected in sequence. Further, after receiving the input of the current voice request 104, the encoder module 101 can perform successive downsampling processing on the current voice request 104 through the plurality of downsampling units connected in sequence, thereby outputting the voice request encoding result of the current voice request 104.

[0081] In an example as Figure 2 shown, the encoder module 101 includes 6 downsampling units connected in sequence, and each downsampling unit is composed of a Conv layer, a BN layer, and an ELU layer. Among them, the Conv layer refers to the Convolutional layer, the BN layer refers to the Batch Normalization layer, and the ELU layer refers to the Exponential Linera Unit layer.

[0082] In an example, the convolutional kernel size of the Conv layer in these 6 downsampling units increases layer by layer to gradually capture higher-dimensional feature information.

[0083] In one example, the convolutional kernel sizes of the Conv layers in these 6 downsampling units are 16, 32, 64, 128, 256, and 512 in sequence. Correspondingly, let the downsampling unit be Conv, then the encoder module 101 can be expressed as: Input-->Conv1[Convolutional layer 1, number of kernels 16]-->Conv2[Convolutional layer 2, number of kernels 32]-->Conv3[Convolutional layer 3, number of kernels 64]-->Conv4[Convolutional layer 4, number of kernels 128]-->Conv5[Convolutional layer 5, number of kernels 256]-->Conv6[Convolutional layer 6, number of kernels 512]. Here, Input represents the voice request input to the encoder module 101, that is, the current voice request 104.

[0084] In addition, in the embodiment of the present application, the decoder module 103 is composed of a plurality of upsampling units connected in sequence. Furthermore, the decoder module 103 can perform successive upsampling processing on the voice request encoding result and the timing information modeling result through the plurality of upsampling units connected in sequence, so as to output the current voice request 104 after noise reduction, that is, the noise-reduced voice request 105.

[0085] In an example as Figure 2 shown, the decoder module 103 includes 6 upsampling units connected in sequence, and each upsampling unit is composed of a Deconv layer, a BN layer, and an ELU layer. Among them, the Deconv layer refers to the Deconvolution layer, the BN layer refers to the Batch Normalization layer, and the ELU layer refers to the Exponential Linera Unit layer.

[0086] In one example, the convolutional kernel sizes of the Deconv layers in these 6 downsampling units decrease layer by layer to gradually restore the high-dimensional information, so as to output the current voice request 104 after noise reduction, that is, the noise-reduced voice request 105.

[0087] In one example, the convolutional kernel sizes of the Deconv layers in these 6 downsampling units are 16, 32, 64, 128, 256, and 512 in sequence. Correspondingly, let the upsampling unit be Deconv, then the decoder module 103 can be expressed as: Input-->Deconv6[Transposed convolution layer 6, number of kernels 512]-->Deconv5[Transposed convolution layer 5, number of kernels 256]-->Deconv4[Transposed convolution layer 4, number of kernels 128]-->Deconv3[Transposed convolution layer 3, number of kernels 64]-->Deconv2[Transposed convolution layer 2, number of kernels 32]-->Deconv1[Transposed convolution layer 1, number of kernels 16]. Here, Input represents the voice request input to the encoder module 101, that is, the current voice request 104.

[0088] In addition, it should be noted that in the embodiment of the present application, the downsampling units in the encoder module 101 are connected to the upsampling units in the decoder module 103 in a one-to-one correspondence. Or rather, the downsampling units in the encoder module 101 are skip-connected to the upsampling units in the decoder module 103 (Skip Connections). It can be understood that the skip connections (Skip Connections) enable the upsampling units in the decoder module 103 to directly access the output features of the downsampling units in the encoder module 101, thus facilitating the upsampling units to recover the details of the features.

[0089] For example, taking the above-mentioned Deconv layer and Conv layer as an example, the connection relationship between the downsampling units in the encoder module 101 and the upsampling units in the decoder module 103 can be expressed as: Conv1--Skip-->Deconv1 / Conv2--Skip-->Deconv2 / …… / Conv6--Skip-->Deconv6. Where Skip represents skip connection. "Conv1--Skip-->Deconv1 / Conv2" means that the downsampling result output by the Conv1 layer can be provided to the Deconv1 layer or the Conv1 layer.

[0090] In addition, when the voice activity detection module 102 includes two layers of LSTM (Long Short-Term Memory) networks connected in sequence, the connection relationship between the downsampling units in the encoder module 101, the LSTM networks in the voice activity detection module 102, and the upsampling units in the decoder module 103 can be expressed as: Conv1-->Conv2-->Conv3-->Conv4-->Conv5-->Conv6-->LSTM1-->LSTM2-->Deconv6-->Deconv5-->Deconv4-->Deconv3-->Deconv2-->Deconv1.

[0091] Among them, it can be understood that in the encoder module 101, the first downsampling unit (such as Conv1) can receive the input of the current voice request 104 and process the current voice request 104, so as to output the downsampling processing result of this unit. For any one of the second to sixth downsampling units, the input of the downsampling unit is the downsampling processing result output by the previous downsampling unit.

[0092] It can also be understood that in the decoder module 103, the first upsampling unit (such as Deconv1) can receive the downsampling processing result output by the last downsampling unit (such as Conv6) and the timing information modeling result output by the voice activity detection module 102, and perform upsampling processing on the downsampling processing result and the timing information modeling result together, thereby outputting the upsampling processing result of this unit. For any one of the second to sixth upsampling units, the input of the upsampling unit is the upsampling processing result output by the previous upsampling unit and the downsampling processing result output by the downsampling unit with the same convolutional kernel size as this unit, and perform upsampling processing on the upsampling processing result and the downsampling processing result together, so as to output the upsampling processing result of this unit.

[0093] In this way, in the embodiment of the present application, the encoder module can be implemented based on a plurality of sequentially connected downsampling units, and the decoder module 103 can be implemented based on a plurality of sequentially connected upsampling units, thereby ensuring the stable operation of the encoder module and the decoder module 103.

[0094] To more clearly illustrate the voice request processing model in the embodiment of the present application, please refer to Figure 2 again. Specifically, in the embodiment of the present application, the chip carried by the vehicle has strong parallel computing support capabilities, so it is suitable for the deployment and application of CNN (Convolutional Neural Network). Furthermore, the model deployed to this chip selects a model with a CNN (Convolutional Neural Network) as the main structure, and as much as possible deepens the number of convolutional layers of the model and expands the number of convolutional kernels of the convolutional layer to improve the model's ability to abstract voice features, thereby obtaining a model with the structure shown in Figure 2 (that is, the voice request processing model).

[0095] Among them, from the perspective of model output, the model includes two branches, one is the noise reduction branch, and the other is the VAD branch. The noise reduction branch refers to restoring the high-dimensional features to the time-frequency features after noise reduction through the transposed convolution layer, that is, the noise-reduced voice request. The VAD branch refers to accessing the fully connected layer according to the output of the timing network after multi-layer convolution operations to predict the voice activity information of each frame.

[0096] Furthermore, as shown in Figure 2As shown, the model includes an encoder module 101, a voice activity detection module 102, and a decoder module 103. Among them, the encoder module 101 includes six convolutional layers, extracting features layer by layer, and the number of kernels gradually increases to capture more high-dimensional features, that is: Input-->Conv1 [Number of kernels in convolutional layer 1: 16]-->Conv2 [Number of kernels in convolutional layer 2: 32]-->Conv3 [Number of kernels in convolutional layer 3: 64]-->Conv4 [Number of kernels in convolutional layer 4: 128]-->Conv5 [Number of kernels in convolutional layer 5: 256]-->Conv6 [Number of kernels in convolutional layer 6: 512].

[0097] The voice activity detection module 102 includes two layers of LSTM, which are used to model temporal information. At the same time, these two layers of LSTM can connect the encoder module 101 and the decoder module 103, thus forming a connection relationship: Conv6-->LSTM1-->LSTM2-->Deconv6. Among them, Conv6 refers to the sixth convolutional layer in the encoder module 101, LSTM1 and LSTM2 represent the first layer of LSTM and the second layer of LSTM in the voice activity detection module 102 in sequence, and Deconv6 refers to the first convolutional layer in the decoder module 103.

[0098] It can be understood that the voice activity detection module 102 is used for voice activity detection. In other words, there is an output relationship: LSTM2-->VADBranch [fully connected layer]-->VADOutput [Output: VAD information].

[0099] The decoder module 103 includes six transposed convolutional layers, restoring features layer by layer to generate the final output, and the number of transposed convolutional kernels gradually decreases, that is: Deconv6 [Number of kernels in transposed convolutional layer 6: 512]-->Deconv5 [Number of kernels in transposed convolutional layer 5: 256]-->Deconv4 [Number of kernels in transposed convolutional layer 4: 128]-->Deconv3 [Number of kernels in transposed convolutional layer 3: 64]-->Deconv2 [Number of kernels in transposed convolutional layer 2: 32]-->Deconv1 [Number of kernels in transposed convolutional layer 1: 16].

[0100] In addition, the decoder module 103 is also responsible for the output of the noise reduction branch, that is: Deconv1-->EnhancedSpectrum [Output: enhanced spectrum].

[0101] In addition, there can be skip connections between the convolutional layers in the encoder module 101 and the encoder layers in the decoder module, so that the decoder module 103 can directly access the features of the encoder module 101, which helps to restore detailed information. Therefore, there is a connection relationship: Conv1--Skip-->Deconv1 / Conv2--Skip-->Deconv2 / …… / Conv6--Skip-->Deconv6

[0102] Please refer to Figure 3 , in the embodiment of the present application, step 02 includes:

[0103] 020: Preprocess the current voice request to determine the processed voice request;

[0104] 021: According to the voice request processing model, perform noise reduction and voice activity detection processing on the processed voice request to determine the noise-reduced voice request and the voice activity detection result.

[0105] The processing module in the embodiment of the present application is also used to preprocess the current voice request to determine the processed voice request, and according to the voice request processing model, perform noise reduction and voice activity detection processing on the processed voice request to determine the noise-reduced voice request and the voice activity detection result.

[0106] The processor in the embodiment of the present application is also used to preprocess the current voice request to determine the processed voice request, and according to the voice request processing model, perform noise reduction and voice activity detection processing on the processed voice request to determine the noise-reduced voice request and the voice activity detection result.

[0107] Specifically, in the audio obtained by the vehicle through a receiving unit such as a microphone (i.e., the current voice request), the intensity of the low-frequency noise is positively correlated with the speed of the vehicle. Therefore, when a user in the vehicle cockpit wants to wake up the vehicle machine or the vehicle voice interaction function by voice, the wake-up rate is low due to the low signal-to-noise ratio in the high-speed driving scenario. Therefore, in order to suppress the influence of the low-frequency noise on the vehicle voice interaction function, the current voice request can be subjected to high-pass filtering.

[0108] In one example, since the vehicle voice interaction function mainly depends on sounds in the frequency band of 300 Hz to 4000 Hz, and the sounds within 200 Hz usually contain a large amount of background noise, high-pass filtering can be performed on the components above 200 Hz.

[0109] In one example, a third-order FIR filter designed based on the window function method can be used to perform high-pass filtering processing on the current voice request to suppress the sounds within 200 Hz.

[0110] In addition, it can also be understood that when the user speaks a voice command such as "turn on the air conditioner" in a relatively low voice, the harmonic information of the human voice in the current voice request obtained by the vehicle is weak. Therefore, to prevent the vehicle from truncating the voice command such as "turn on the air conditioner" spoken by the user in a relatively low voice during the VAD processing, or accidentally suppressing the harmonic information of the human voice in the current voice request during the noise reduction processing, preprocessing for gain enhancement can be performed on the current voice request. Furthermore, after the passenger speaks a voice command such as "turn on the air conditioner" in a relatively low voice, the vehicle can perform gain enhancement processing on the current voice request, thereby enhancing the harmonic information of the human voice in the current voice request, and ensuring the effectiveness of subsequent noise reduction processing and VAD processing.

[0111] In one example, the vehicle can first perform high-pass filtering on the current voice request, and then perform gain enhancement processing on the current voice request after high-pass filtering, thereby completing the preprocessing of the current voice request. Furthermore, the vehicle can input the preprocessed current voice request into a pre-trained voice request processing model to obtain the VAD result of the current voice request and the current voice request after noise reduction (i.e., the noise-reduced voice request).

[0112] Thus, in the embodiments of the present application, the current voice request can be preprocessed to determine the processed voice request, and the processed voice request can be subjected to noise reduction and voice activity detection processing according to the voice request processing model to determine the noise-reduced voice request and the voice activity detection result, so that the current voice request can be preprocessed before being input into the voice request processing model to improve the quality of the current voice request, thereby improving the effectiveness of the voice activity detection result and the noise-reduced voice request to a certain extent.

[0113] Please refer to Figure 4 , in some embodiments of the present application, step 03 includes:

[0114] 030: Perform gain processing on the noise-reduced voice request to determine the voice request after gain processing;

[0115] 031: Control the vehicle according to the voice request after gain processing and the voice activity detection result.

[0116] The voice interaction device further includes an interaction module, which is also used to perform gain processing on the noise-reduced voice request to determine the voice request after gain processing, and control the vehicle according to the voice request after gain processing and the voice activity detection result.

[0117] The processor in the embodiments of the present application is also used to perform gain processing on the noise-reduced voice request to determine the voice request after gain processing, and control the vehicle according to the voice request after gain processing and the voice activity detection result.

[0118] Specifically, in the embodiments of the present application, after the voice request processing model performs noise reduction processing on the current voice request of the float type, it can output a voice request with noise reduction of the short type. Furthermore, to avoid clipping when converting the float type signal into a short type audio stream due to an overly large value, the vehicle can perform processing to reduce the gain of the voice request with noise reduction, thereby ensuring the robustness of subsequent vehicle control.

[0119] In this way, in the embodiments of the present application, it is possible to perform processing on the voice request with noise reduction for gain to determine the voice request after gain processing, so that the quality of the voice request with noise reduction can be guaranteed to a certain extent, and further the robustness of controlling the vehicle based on the voice request after gain processing and the voice activity detection result can be ensured.

[0120] Please refer to Figure 5 , in some embodiments of the present application, the voice interaction method further includes:

[0121] 04: Perform model format conversion processing on the pre-trained voice request processing model to determine a voice request processing model in a preset format;

[0122] 05: Configure a runtime library corresponding to the voice request processing model in the preset format and perform the deployment of the voice request processing model in the vehicle.

[0123] The voice interaction device in the embodiments of the present application further includes a conversion module and a configuration module. Among them, the conversion module is used to perform model format conversion processing on the pre-trained voice request processing model to determine a voice request processing model in a preset format. The configuration module is used to configure a runtime library corresponding to the voice request processing model in the preset format and perform the deployment of the voice request processing model in the vehicle.

[0124] The processor in the embodiments of the present application is further used to perform model format conversion processing on the pre-trained voice request processing model to determine a voice request processing model in a preset format, and configure a runtime library corresponding to the voice request processing model in the preset format and perform the deployment of the voice request processing model in the vehicle.

[0125] Specifically, in the embodiments of the present application, the voice request processing model can be pre-deployed locally in the vehicle for the vehicle to call the locally deployed voice request processing model to achieve efficient processing of the current voice request.

[0126] Specifically, in one example, after the voice request processing model is trained, it will be exported in ONNX format. Among them, the torch.onnx.export() function can be called to export the voice request processing model. For example, "torch.onnx.export(model, dummy_input, \"model.onnx\", opset_version = 11)".

[0127] Next, a format conversion tool (such as the qnn-convert tool) can be used to convert the ONNX format voice request processing model into a format supported by the QNN (Quantum Neural Network) framework, such as the ".qnn" format. In one example, this process can be described as: qnn-convert --input_model.onnx --output_model.qnn.

[0128] Finally, the QNN runtime library can be integrated into the application to load and run the QNN format voice request processing model, thus completing the deployment of the voice request processing model in the vehicle. In one example, the integration of the runtime library of the voice request processing model can be expressed as "#include <qnn_runtime.h>".

[0129] In this way, in the embodiment of the present application, the model format conversion process can be performed on the pre-trained voice request processing model to determine the voice request processing model in the preset format, and configure the runtime library corresponding to the voice request processing model in the preset format, so as to deploy the voice request processing model in the vehicle, enabling the vehicle to process the current voice request through the locally deployed voice request processing model.

[0130] To more clearly illustrate the processing process of the current voice request in the embodiment of the present application, please refer to Figure 6 , Figure 6 which is the schematic flow chart of the voice interaction method in some embodiments of the present application, that is, as Figure 6 shown, in the embodiment of the present application, after the vehicle obtains the current voice request, first, the audio with a frequency band below 200 hz in the current voice request is filtered to suppress the noise in the current voice request.

[0131] Then, after the filtering process, the current voice request can be processed to increase the signal gain (Gain) to enhance the harmonic information corresponding to the human voice part in the current voice request.

[0132] Then, for the current voice request after signal gain improvement, the vehicle can call the voice request processing model pre-deployed in the vehicle chip to perform noise reduction and VAD processing on the current voice request, so as to obtain the current voice request after noise reduction and the VAD result of the current voice request.

[0133] Next, for the current voice request after noise reduction, or rather, for the noise-reduced voice request, the signal gain (Gain) of the noise-reduced voice request can be reduced to avoid the occurrence of clipping.

[0134] Finally, based on the VAD result of the current voice request and the noise-reduced voice request after signal gain reduction, the wake-up engine in the vehicle can perform corresponding natural language processing to control the vehicle to perform corresponding operations.

[0135] For example, in one example, both the VAD result and the noise-reduced voice request are frame-level outputs. Furthermore, when the VAD result indicates the start of speech, the wake-up engine starts to decode the noise-reduced voice request; when the VAD result indicates the end of speech, the wake-up engine ends the decoding of the noise-reduced voice request and outputs a wake-up message at the same time. Among them, if the wake-up occurs before the speech activity ends, the wake-up engine can output the wake-up message in advance.

[0136] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by one or more processors, implements the above-mentioned voice interaction method.

[0137] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0138] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including substantially concurrently or in reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0139] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A voice interaction method for a vehicle, characterized in that: The method comprises: Get the current voice request; According to the pre-trained voice request processing model, the current voice request is subjected to noise reduction and voice activity detection processing to determine the noise reduction voice request and the voice activity detection result; The vehicle is controlled according to the noise reduction voice request and the voice activity detection result to perform the voice interaction.

2. The method according to claim 1, characterized in that The voice request processing model includes an encoder module, a voice activity detection module and a decoder module connected in sequence; The encoder module is configured to encode the current voice request and determine a voice request encoding result; The voice activity detection module is configured to determine the voice activity detection result according to the voice request encoding result; The decoder module is configured to decode the voice request encoding result to determine the noise reduction voice request.

3. The method according to claim 2, characterized in that The voice activity detection module is configured to perform timing information modeling processing on the voice request encoding result to determine the timing information modeling result of the voice request encoding result, and perform voice activity detection on the timing information modeling result to determine the voice activity detection result; The decoder module is configured to decode the timing information modeling result of the speech request encoding result to determine the noise reduction speech request.

4. The method according to claim 3, characterized in that The encoder module is connected to the decoder module; The decoder module is configured to decode the timing information modeling result and the voice request encoding result to determine the noise reduction voice request.

5. The method according to claim 4, characterized in that The encoder module comprises a plurality of down-sampling units connected in sequence, and the decoder module comprises a plurality of up-sampling units connected in sequence, wherein one down-sampling unit is connected to one up-sampling unit; The downsampling unit is configured to perform downsampling processing on first input data to determine a downsampling processing result of the first input data, wherein the first input data includes the voice request sample and / or a downsampling processing result of a previous downsampling unit, and the voice request encoding result is a downsampling processing result of the last downsampling unit; The upsampling unit is configured to perform upsampling processing on the second input data and determine the upsampling result of the second input data, the noise reduction speech request is the upsampling result of the last upsampling unit, and the second input data includes the timing information modeling information, the first input data of the downsampling unit connected to the upsampling unit, or includes the upsampling processing result of the previous upsampling unit and the downsampling processing result of the downsampling unit connected to the upsampling unit.

6. The method according to claim 1, characterized in that The performing noise reduction and voice activity detection processing on the current voice request according to the pre-trained voice request processing model to determine the noise reduction voice request and the voice activity detection result includes: Preprocessing the current voice request to determine a processed voice request; According to the voice request processing model, noise reduction and voice activity detection processing are performed on the processed voice request to determine the noise reduction voice request and the voice activity detection result.

7. The method according to claim 1, characterized in that The controlling the vehicle and performing the voice interaction according to the noise reduction voice request and the voice activity detection result includes: Performing gain processing on the noise reduction voice request to determine a voice request after gain processing; The vehicle is controlled based on the gain processed voice request and the voice activity detection result.

8. The method according to claim 1, characterized in that The method further comprises: Performing model format conversion processing on the pre-trained voice request processing model to determine the voice request processing model in a preset format; A runtime library corresponding to the voice request processing model in the preset format is configured to deploy the voice request processing model in the vehicle.

9. A vehicle, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 8 is implemented.