Audio interaction method and device

By combining built-in and external voice activity detection units, the one-shot wake-up timing was optimized, solving the problems of success rate and stability of voice wake-up in complex environments, and improving the accuracy and experience of user voice interaction.

CN122090836APending Publication Date: 2026-05-26BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DIDI INFINITY TECH & DEV CO LTD
Filing Date
2024-11-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing one-shot voice wake-up technology performs poorly in complex environments, with insufficient success rate and stability, resulting in a poor user voice interaction experience.

Method used

A combination of built-in and external voice activity detection units is adopted. The built-in VAD unit traces back the wake word audio and detects the end point. If the end point is not detected, the external VAD unit is activated to perform more accurate end point detection and provide an end signal to the speech recognition unit to optimize the wake-up timing.

Benefits of technology

It improves wake-up success rate and robustness, optimizes user voice interaction experience, and enhances voice interaction accuracy in vehicle cabin environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090836A_ABST
    Figure CN122090836A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio interaction method and device, electronic equipment, a computer storage medium and a computer program product. The method described herein includes: in response to detecting that received audio data includes a wake-up word, activating a first voice activity detection unit; determining, using a first voice activity detection unit, a first tail point of the audio data; starting a second voice activity detection unit in response to the situation that a first tail point of the audio data is not detected within a preset time period after the starting moment of the first voice activity detection unit; determining, using a second voice activity detection unit, a second end point of the audio data; and providing an end signal corresponding to the second tail point to the voice recognition unit to provide a response based on a recognition result of the voice recognition unit on the audio data. Based on the above mode, the embodiment of the invention can improve the success rate of identifying the one-shot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, electronic devices, computer storage media, and computer program products for audio interaction. Background Technology

[0002] With advancements in technology, voice recognition is being widely applied in various scenarios. Taking smart cockpits as an example, people can control vehicles to perform various actions through voice interaction. For instance, users can use voice commands to control the vehicle to turn on the air conditioning, close the windows, and so on. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for audio interaction is provided. The method includes: activating a first voice activity detection unit in response to detecting that received audio data includes a wake word; determining a first end point of the audio data using the first voice activity detection unit; activating a second voice activity detection unit in response to no detection of the first end point of the audio data within a preset time period after the activation time of the first voice activity detection unit; determining a second end point of the audio data using the second voice activity detection unit; and providing an end signal corresponding to the second end point to a speech recognition unit to provide a response based on the recognition result of the audio data by the speech recognition unit.

[0004] In a second aspect of this disclosure, an apparatus for audio interaction is provided. The apparatus includes: a first activation module configured to activate a first endpoint detection module in response to detecting received audio data including a wake word; a first determination module configured to determine a first end point of the audio data using the first endpoint detection module; a second activation module configured to activate a second endpoint detection module in response to no detection of the first end point of the audio data within a preset time period after the activation time of the first endpoint detection module; a second determination module configured to determine a second end point of the audio data using the second endpoint detection module; and a signal providing module configured to provide a termination signal corresponding to the second end point to a speech recognition module, thereby providing a response based on the speech recognition module's recognition result of the audio data.

[0005] In a third aspect of this disclosure, an electronic device is provided, comprising: a memory and a processor; wherein the memory is configured to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to a first aspect of this disclosure.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having one or more computer instructions stored thereon, wherein the one or more computer instructions are executed by a processor to implement the method according to a first aspect of this disclosure.

[0007] In a fifth aspect of this disclosure, a computer program product is provided, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to a first aspect of this disclosure.

[0008] According to various embodiments of this disclosure, embodiments of this disclosure can transmit audio signals from the vehicle's underlying system back to the vehicle's operating system, thereby enabling wireless playback devices (e.g., Bluetooth headsets, etc.) to play audio from the vehicle's underlying system. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0011] Figure 2 A flowchart illustrating an audio interaction process according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic diagram of audio interaction according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic structural block diagram of an apparatus for audio interaction according to some embodiments of the present disclosure is shown; and

[0014] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] As discussed above, in the intelligent cockpit of modern cars, intelligent sensing technology is key to enhancing the driving experience, allowing drivers to interact with vehicle systems via voice. This interaction comes in two forms: conventional voice wake-up, where the driver must wait for the vehicle system to respond after uttering the wake-up phrase before issuing further commands; and the more efficient one-shot voice wake-up, where the driver can issue commands directly after uttering the wake-up phrase without waiting for a system response. However, existing one-shot technology performs poorly in complex environments, exhibiting insufficient success rate and stability.

[0018] In view of this, embodiments of the present disclosure provide a scheme for audio interaction. According to this scheme, in response to detecting that received audio data includes a wake word, a first voice activity detection unit can be activated; the first voice activity detection unit can be used to determine a first end point of the audio data; in response to no detection of the first end point of the audio data within a preset time period after the activation time of the first voice activity detection unit, a second voice activity detection unit can be activated; the second voice activity detection unit can be used to determine a second end point of the audio data; and an end signal corresponding to the second end point can be provided to a speech recognition unit to provide a response based on the speech recognition unit's recognition result of the audio data.

[0019] In this way, by setting up both a built-in voice activity detection unit and an external voice activity detection unit, the embodiments of this disclosure optimize the determination of one-shot wake-up timing. Therefore, the embodiments of this disclosure can improve wake-up success rate and robustness, and optimize the user's voice interaction experience.

[0020] Example Environment

[0021] First see Figure 1 The diagram illustrates an example environment 100 in which embodiments of the present disclosure may be implemented.

[0022] like Figure 1 As shown, example environment 100 may include vehicle 110. Exemplarily, vehicle 110 can be any type of vehicle capable of carrying people and / or goods and moving via a power system such as an engine, including but not limited to cars, trucks, buses, electric vehicles, motorcycles, motorhomes, trains, etc. Such a vehicle 110 may be a vehicle with or without intelligent driving capabilities.

[0023] like Figure 1 As shown, vehicle 110 can receive user audio data 115 and provide corresponding responses. As an example, such audio data 115 may include the user's wake-up word and voice commands. Vehicle 110 may, for example, utilize a speech recognition model to recognize the audio data 115 and execute actions corresponding to the voice commands.

[0024] For example, the wake word could include "Hello, XX," and the voice command could include "Open the window." Vehicle 110, for example, could respond to detecting the wake word and recognize the corresponding voice command to perform the action of opening the window.

[0025] The following section will describe the specific process of voice interaction.

[0026] Example process

[0027] Figure 2 A flowchart of a process 200 for audio interaction according to several embodiments of the present disclosure is shown. Process 200 can be implemented by a suitable speech recognition device. As an example, the speech recognition device can be deployed, for example, in a… Figure 1 In the vehicle 110 shown. The process 200 will be described below using vehicle 110 as an example. It should be understood that such a process can also be implemented by other devices.

[0028] like Figure 2 As shown in box 210, vehicle 110 activates the first voice activity detection unit (also known as VAD unit) in response to detecting received audio data including a wake word.

[0029] The following will be referenced Figure 3 To describe process 200. Figure 3 A schematic diagram of audio interaction according to some embodiments of the present disclosure is shown. For example... Figure 3 As shown, vehicle 110 may, for example, use an audio acquisition unit to receive audio stream 330. Audio stream 330 may, for example, include the voice content of the driver or passenger of vehicle 110.

[0030] As shown in the figure, the vehicle 110 may be equipped with a voice module 310, which may include, for example, a cascaded voice wake-up unit 305 and a built-in VAD unit 308. The basic principle of the VAD unit is to automatically identify and distinguish between the speech and non-speech parts (such as silence, background noise, etc.) in the audio signal.

[0031] As an example, the voice wake-up unit 305 can be activated after detecting a voice signal (i.e., time point A 350) and can be used to determine whether the received audio data includes a wake-up word. As an example, the wake-up word may include a preset word. Alternatively, the wake-up word may also be a word configured by the user.

[0032] As shown in the figure, the built-in VAD unit 308 can be woken up when the voice wake-up unit 305 detects that the audio data (e.g., audio stream 330) includes the wake word speech 335.

[0033] Continue to refer to Figure 2 In frame 220, vehicle 110 uses the first voice activity detection unit to determine the first tail point of the audio data.

[0034] In traditional voice wake-up schemes, the voice module typically has only one voice interaction detection unit. Traditionally, after detecting a wake-up event, the voice module can use the VAD unit to determine whether human voice is detected within a predetermined timeframe. If human voice is detected, the voice module can, for example, enter a one-shot wake-up mode and stream the audio to text from the beginning of the human voice until the end of the human voice is detected.

[0035] However, this wake-up scheme has several drawbacks. First, in one-shot wake-up, the interval between the wake-up module's start time and the user's start time when speaking the command is relatively short. This results in a high probability that the starting point of the user's speech will not be detected, leading to the failure of the one-shot judgment.

[0036] Furthermore, VAD (Visual Analog Detection) relies on certain speech frames to determine the starting point of speech. If ASR (Automatic Speech Retrieval) is performed after VAD has detected the starting point of speech, word loss will occur, which significantly impacts semantic integrity.

[0037] Conversely, according to the scheme of this disclosure, the built-in VAD unit 308 can recall the received wake word audio 335 and continuously received audio frames (e.g., Figure 3 Tail point detection is performed from time point B (355) to time point C (360).

[0038] In frame 230, vehicle 110 activates the second voice activity detection unit in response to the first tail point of no audio data being detected within a preset time period after the activation time of the first voice activity detection unit.

[0039] Furthermore, such as Figure 3 As shown, the built-in VAD unit 308 can determine in Figure 3 The one-shot determination time period 340 (i.e., from time point B 355 to time point C 360) is shown to determine whether an end point is detected. If an end point is detected, the built-in VAD unit 308 can determine that it is in non-one-shot wake-up mode. Conversely, if no end point is detected, the built-in VAD unit 308 can determine that it is in one-shot wake-up mode.

[0040] by Figure 3As an example, since no voice signal is received from time point B 355 to time point C 360, the built-in VAD unit 308 can, for example, determine time point B 355 as the end point of the audio stream 330 and determine to start the non-one-shot mode.

[0041] Specifically, in one-shot wake-up mode, the voice broadcast unit 312 of the voice module 310 will not be invoked. For example, the vehicle 110 will not immediately respond with a preset reply voice, such as "Hey".

[0042] Furthermore, the voice module 310 can activate the external VAD unit 320. For example... Figure 3 As shown, unlike the built-in VAD unit 308, the external VAD unit 320 can operate independently of the voice wake-up unit 305.

[0043] Furthermore, the external VAD unit 320 may, for example, have a larger model structure than the built-in VAD unit 308. In some examples, the built-in VAD unit 308 may include a 9-state acoustic model plus a state machine. In contrast, the external VAD unit 320 only includes two states: human voice and silence.

[0044] Using this modeling approach, the built-in VAD unit 308 can better align with Chinese pronunciation habits, thereby improving accuracy and recall. Furthermore, the built-in VAD unit 308 and the voice wake-up unit 305 can be configured as a combined model, and the built-in VAD unit 308 can be specifically optimized based on business needs or changes in the wake-up word.

[0045] In some embodiments, the built-in VAD unit 308 can also be trained for a specific wake word. Specifically, the built-in VAD unit 308 can be trained based on training data associated with a preset wake word, thereby improving the accuracy of VAD detection after voice wake-up.

[0046] In frame 240, vehicle 110 uses the second voice activity detection unit to determine the second tail point of the audio data.

[0047] For example, the external VAD unit 320 can detect the end point of the audio stream 330. If the audio stream 330 is "Hello XX, please turn on the air conditioner." as an example, the external VAD unit 320 can detect that no human voice signal is detected after the user finishes speaking the sentence, and determine that the moment after the word "turn on" is the end point of the audio stream 330.

[0048] In frame 250, vehicle 110 provides a termination signal corresponding to the second tail point to the voice recognition unit in order to provide a response based on the recognition result of the voice recognition unit on the audio data.

[0049] In some embodiments, the external VAD can send an end signal to the speech recognition unit (ASR) in response to detecting the end of the user's speech, to indicate that the user's voice input for this round has ended.

[0050] In some embodiments, the vehicle 110 may provide audio data from the start point of the wake-up word to the second end point to the voice recognition unit to determine the recognition result of the audio data. Accordingly, the voice module 310 may determine that the recognition result of the voice content input by the user in this round is "Hello XX, please turn on the air conditioner".

[0051] Accordingly, vehicle 110 can, for example, perform a corresponding action based on the recognition result of the voice content. For example, vehicle 110 can turn on the vehicle's air conditioning system. In other examples, such a response may also include a voice response or other appropriate action response.

[0052] In some embodiments, if a first tail point of audio data is detected within a preset time period after the activation time of the first voice activity detection unit, the vehicle 110 may provide a preset response audio.

[0053] by Figure 3 As an example, the built-in VAD unit can detect the tail point at the one-shot determination time 340 and determine that it has entered a non-one-shot wake-up mode. Accordingly, the vehicle 110 can use the voice broadcast unit 312 to provide a preset response audio, such as "Hey".

[0054] Furthermore, vehicle 110 can activate a second voice activity detection unit to detect the starting point of the received additional audio data.

[0055] For example, the external VAD unit 320 can be activated to detect the start point of received additional audio data (e.g., interactive voice 345). Figure 3 (C time point 360).

[0056] Furthermore, the speech recognition unit can determine the corresponding recognition result based on the start and end points detected by the external VAD unit 320. For example, after the user says "Hello XX" and pauses for a while, and continues to say "Please turn on the air conditioner" after pausing to hear the reply audio "Yes", the external VAD unit 320 can determine that the content of the interactive voice 345 is "Please turn on the air conditioner" and can accordingly perform the action of turning on the air conditioner.

[0057] In this way, by setting up both a built-in voice activity detection unit and an external voice activity detection unit, the embodiments of this disclosure optimize the determination of one-shot wake-up timing. Therefore, the embodiments of this disclosure can improve wake-up success rate and robustness, and optimize the user's voice interaction experience.

[0058] Furthermore, since the vehicle cabin environment is relatively noisy, the embodiments of this disclosure can further improve the accuracy of in-vehicle voice interaction through the voice interaction method described above.

[0059] Example devices and equipment

[0060] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an apparatus 400 for audio interaction according to some embodiments of the present disclosure is shown.

[0061] like Figure 4 As shown, the device 400 includes a first activation module 410 configured to activate a first endpoint detection module in response to detecting received audio data including a wake word; a first determination module 420 configured to determine a first end point of the audio data using the first endpoint detection module; a second activation module 430 configured to activate a second endpoint detection module in response to no detection of the first end point of the audio data within a preset time period after the activation time of the first endpoint detection module; a second determination module 440 configured to determine a second end point of the audio data using the second endpoint detection module; and a signal providing module 450 configured to provide an end signal corresponding to the second end point to a speech recognition module, so as to provide a response based on the recognition result of the speech recognition module on the audio data.

[0062] In some embodiments, the apparatus 400 further includes a first providing module configured to provide a preset response audio in response to detecting a first end point of audio data within a preset time period after the activation time of the first voice activity detection unit; and to activate a second voice activity detection unit to detect the start point of received additional audio data.

[0063] In some embodiments, the first voice activity detection unit is cascaded with the voice wake-up unit, and the second voice activity detection unit is independent of the voice wake-up unit.

[0064] In some embodiments, the model size of the first speech activity detection unit is smaller than that of the second speech activity detection unit.

[0065] In some embodiments, the first voice activity detection unit is trained based on training data associated with a preset wake word.

[0066] In some embodiments, the device 400 further includes a second providing module configured to provide audio data from the start point of the wake word to the second end point to the speech recognition unit in order to determine the recognition result of the audio data.

[0067] In some embodiments, the voice recognition system is integrated into the target vehicle, and the target vehicle is configured to perform a response action based on the recognition result of the audio data.

[0068] The units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0069] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0070] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0071] Electronic device 500 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 500.

[0072] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0073] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0074] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 570 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0075] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the methods described above.

[0076] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0077] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0078] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0080] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

1. A method for audio interaction, comprising: In response to the detection that the received audio data includes a wake word, the first voice activity detection unit is activated; The first end point of the audio data is determined using the first voice activity detection unit; In response to the fact that no first tail point of the audio data is detected within a preset time period after the start time of the first voice activity detection unit, the second voice activity detection unit is started; The second voice activity detection unit is used to determine the second tail point of the audio data; as well as An end signal corresponding to the second tail point is provided to the speech recognition unit to provide a response based on the speech recognition unit's recognition result of the audio data.

2. The method according to claim 1, further comprising: In response to detecting the first tail point of the audio data within a preset time period after the start time of the first voice activity detection unit, a preset response audio is provided; as well as The second voice activity detection unit is activated to detect the starting point of the received additional audio data.

3. The method according to claim 1, wherein the first voice activity detection unit is cascaded with the voice wake-up unit, and the second voice activity detection unit is independent of the voice wake-up unit.

4. The method according to claim 3, wherein the model size of the first speech activity detection unit is smaller than that of the second speech activity detection unit.

5. The method of claim 3, wherein the first voice activity detection unit is trained based on training data associated with a preset wake word.

6. The method according to claim 1, further comprising: The audio data from the start point of the wake word to the second end point is provided to the speech recognition unit to determine the recognition result of the audio data.

7. The method of claim 1, wherein the speech recognition system is integrated in the target vehicle, and the target vehicle is configured to perform a response action based on the recognition result of the audio data.

8. A device for audio interaction, comprising: The first startup module is configured to start the first endpoint detection module in response to detecting received audio data including a wake word; The first determining module is configured to determine the first tail point of the audio data using the first endpoint detection module; The second startup module is configured to start the second endpoint detection module in response to the first tail point of the audio data not being detected within a preset time period after the startup time of the first endpoint detection module. The second determining module is configured to use the second endpoint detection module to determine the second tail point of the audio data; as well as The signal providing module is configured to provide a termination signal corresponding to the second tail point to the speech recognition module in order to provide a response based on the recognition result of the speech recognition module on the audio data.

9. An electronic device, comprising: Memory and processor; The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to any one of claims 1 to 7.

11. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 7.