Voice interaction method, device and system, electronic equipment and storage medium

By detecting the time information of the wake-up command in real time, determining the target position of the wake-up command in the audio, and extracting the complete recognition voice segment, the problem of recognition language truncation caused by the wake-up word and user recognition language in the same sentence is solved, thereby improving the accuracy of voice interaction and user experience.

CN120636401APending Publication Date: 2025-09-12XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511092946.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In complex voice interaction scenarios, the wake-up word and user identification language in the same sentence will cause the user identification language to be truncated, reducing the accuracy of voice recognition and user experience.

Method used

By detecting the time information of the wake-up command in real time and determining the target position of the wake-up command in the audio, the complete recognition voice segment can be extracted for voice interaction, avoiding the truncation of the recognition voice segment caused by collecting it after the wake-up state.

Benefits of technology

It ensures the integrity of the voice segments sent by users during voice interaction, and improves the accuracy of voice interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636401A_ABST
    Figure CN120636401A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice interaction method, device and system, electronic equipment and a storage medium. And performing real-time detection on a wake-up instruction in the voice input signal, determining a first audio based on time information corresponding to the detected wake-up instruction, determining a target position according to the time information corresponding to the wake-up instruction, determining a first recognition voice segment from the first audio according to the target position, and performing voice interaction based on the first recognition voice segment. Therefore, the first recognition voice segment is determined according to the target position information of the wake-up instruction, the integrity of the first recognition voice segment can be effectively ensured, the problem that the audio is cut off due to the fact that the first recognition voice segment is collected after the wake-up state is avoided, it is ensured that the first recognition voice segment can be accurately recognized, and the user experience is improved. And thus, the accuracy of voice interaction and the use experience of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of human-computer interaction technology and voice interaction technology, and specifically to voice interaction methods, devices, systems, electronic devices, and storage media. Background Art

[0002] Voice interaction is a form of human-computer communication that allows users to communicate with devices or applications through voice commands. During voice interaction, the device / application must first be woken up, and then the user recognition language used for voice interaction is determined based on the device / application's wakeup time.

[0003] However, with the continuous increase in voice interaction application scenarios, especially in complex application scenarios such as continuous conversations, multi-round conversations, manual conversations, and wake-up continuous speaking mode (one-word answer), since the wake-up word and the user identification word are in the same sentence, if the user identification word is determined based on the time when the device / application is awakened, it may cause a delay in collecting the user identification word, causing the user identification word to be truncated, resulting in incomplete user identification word, thereby reducing the accuracy of voice recognition through user identification word and affecting the user voice interaction experience. Summary of the Invention

[0004] In order to solve the above technical problems, the embodiments of the present disclosure provide a voice interaction method, device, system, electronic device and storage medium.

[0005] One aspect of an embodiment of the present disclosure provides a voice interaction method, including: real-time detection of wake-up instructions in a voice input signal; determining a first audio within a target time period where the wake-up instruction is located based on time information corresponding to the detected wake-up instruction, the time information corresponding to the wake-up instruction including a start time and an end time of the wake-up instruction; determining a target position of the wake-up instruction in the first audio based on the time information corresponding to the wake-up instruction; determining a first recognized voice segment from the first audio based on the target position, and performing voice interaction based on the first recognized voice segment.

[0006] Another aspect of the embodiments of the present disclosure provides a voice interaction device, including: an instruction detection module, for performing real-time detection of wake-up instructions in a voice input signal; an audio determination module, for determining, based on the time information corresponding to the detected wake-up instruction, a first audio within a target time period in which the wake-up instruction is located, wherein the time information corresponding to the wake-up instruction includes the start time and the end time of the wake-up instruction; a positioning module, for determining, based on the time information corresponding to the wake-up instruction, a target position of the wake-up instruction in the first audio; and a voice interaction module, for determining, based on the target position, a first recognized voice segment from the first audio, and performing voice interaction based on the first recognized voice segment.

[0007] Another aspect of the embodiments of the present disclosure provides a voice interaction system, including: an audio acquisition device, an audio playback device and an interaction device; the audio acquisition device is used to acquire voice input signals, and the audio playback device is used to play the device response voice segment in voice interaction; the interaction device is used to: perform real-time detection of wake-up instructions in the voice input signal; based on the time information corresponding to the detected wake-up instruction, determine the first audio within the target time period where the wake-up instruction is located, the time information corresponding to the wake-up instruction includes the start time and end time of the wake-up instruction; according to the time information corresponding to the wake-up instruction, determine the target position of the wake-up instruction in the first audio; according to the target position, determine a first recognized voice segment from the first audio, and perform voice interaction based on the first recognized voice segment.

[0008] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the above-mentioned voice interaction method.

[0009] Another aspect of the embodiments of the present disclosure provides an electronic device, comprising: a processor, a memory communicatively connected to the processor, and the above-mentioned voice interaction device;

[0010] The memory is used to store the processor executable instructions; the processor is used to read the executable instructions from the memory to control the voice interaction device to implement the above-mentioned voice interaction method.

[0011] Based on the embodiment of the present disclosure, the first audio is determined by obtaining the time information corresponding to the wake-up instruction, and the first audio includes audio whose timing is before and after the wake-up instruction, and the target position of the wake-up instruction in the first audio is determined according to the time information corresponding to the wake-up instruction. Thereafter, the first recognized voice segment is determined from the first audio according to the target position for voice interaction, which can effectively ensure the integrity of the voice segment sent by the user during the voice interaction process, and avoid the problem of audio being truncation caused by collecting the voice segment after the wake-up state. Therefore, voice interaction based on the determined first recognized voice segment can ensure complete and accurate recognition of the user's recognition language, which helps to improve the accuracy of voice interaction and the user's experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 An exemplary application scenario of the voice interaction method provided by the present disclosure is shown.

[0013] Figure 2 It is a flowchart of a voice interaction method provided by an exemplary embodiment of the present disclosure.

[0014] Figure 3 It is a flowchart of step S230 provided by an exemplary embodiment of the present disclosure.

[0015] Figure 4 It is a flowchart of step S233 provided by an exemplary embodiment of the present disclosure.

[0016] Figure 5 It is a flowchart of a voice interaction method provided by another exemplary embodiment of the present disclosure.

[0017] Figure 6 It is a flowchart of step S240 provided by an exemplary embodiment of the present disclosure.

[0018] Figure 7 This is a flowchart of a voice interaction method provided by an application example of the present disclosure.

[0019] Figure 8 It is a structural block diagram of a voice interaction device provided by an exemplary embodiment of the present disclosure.

[0020] Figure 9 It is a structural block diagram of a voice interaction device provided by another exemplary embodiment of the present disclosure.

[0021] Figure 10 It is a structural block diagram of a voice interaction system provided by an exemplary embodiment of the present disclosure.

[0022] Figure 11 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.

[0024] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0025] Application Overview

[0026] In the process of realizing the present disclosure, the inventor discovered through research that in voice interaction, the conventional process is: first wake up the application / device, wait for it to enter the awakened state, then collect the user identification language after the awakening moment, and perform voice recognition on the user identification language. However, in actual interaction scenarios, the wake-up word and the user identification language used by the user to wake up the application / device are usually in the same sentence. In this case, since the application / device starts the voice collection process after awakening, this will cause a collection delay, resulting in some voice information adjacent to the wake-up word not being effectively captured, resulting in truncation of the user identification language and missing content. This problem directly reduces the accuracy of voice recognition, which in turn has a negative impact on the user's voice interaction experience.

[0027] Example Applications

[0028] The technical solution disclosed herein can be applied to voice interaction in any field and any application scenario, and can operate devices or applications through voice interaction. For example, it can be applied to driving scenarios, and can control devices or applications in the vehicle through voice interaction. Alternatively, it can also be applied to the control of any human-computer interaction device such as mobile terminals, smart home appliances, smart phones, tablets, notebooks, computers, and robots through voice interaction. Mobile terminals can be, for example, aircraft. Smart home appliances can be, for example, refrigerators, color TVs, or washing machines. The following uses the application in a driving scenario to illustrate the control of devices or applications in a vehicle through voice interaction as an example, but the application scenarios of the technical solution disclosed herein are not limited to this.

[0029] Figure 1 An exemplary application scenario of the voice interaction method provided by the present disclosure is shown. Figure 1As shown, vehicle 1 is equipped with a voice interaction device 2, which may include: a voice acquisition device 21, a voice storage device 22, and an interaction device 23. The voice interaction device 2 is communicatively connected to the speech interaction model (Speech Interaction Model) and the control system 11 of vehicle 1. The voice acquisition device 21 is used to acquire voice signals in real time. For example, the voice acquisition device 21 may be a microphone array, a pickup, or a sound source acquisition device. The voice storage device 22 is used to store voice signals of a preset duration. For example, the voice storage device 22 may be a static random access memory (SRAM) or a dynamic random access memory (DRAM). The interaction device 23 is used to perform noise reduction processing such as denoising and echo cancellation on the audio signal, perform wake-up detection, determine recognized voice segments, and perform voice interaction based on the recognized voice segments. The interaction device 23 may be, for example, a single-chip microcomputer (MCU), a microprocessor (MCU), or an AI chip (AI chip). The voice interaction model can be set in the cloud, the voice interaction device 2 or the vehicle 1.

[0030] Exemplarily, the voice collection device 21 collects the voice signal inside the vehicle, stores the collected voice signal in the voice storage device 22, and transmits the voice signal to the interaction device 23. The interaction device 23 performs noise reduction processing on the voice signal, and then performs wake-up instruction detection on the voice signal. For example, when the voice signal is "Xiao A opens the window", it can be detected that it has a wake-up instruction (Xiao A). Based on the time information of the wake-up instruction, the audio including the time period of the wake-up instruction is obtained from the voice storage device 22, and the recognized voice segment (open the window) is determined in the audio based on the time information of the wake-up instruction, and the recognized voice segment is sent to the voice interaction model for voice interaction, and the operation instruction (open the window) is obtained according to the voice interaction, and the window of the vehicle 1 is opened according to the operation instruction.

[0031] In the embodiment of the present disclosure, the recognition voice segment is determined by the target position information of the wake-up instruction, which can effectively ensure the integrity of the recognition voice segment sent by the user, avoid the problem of the recognition voice segment being truncated due to the collection of the recognition voice segment after the wake-up state, thereby ensuring that the recognition voice segment can be accurately recognized, and improving the accuracy of voice interaction and the user experience.

[0032] Exemplary Methods

[0033] Figure 2This is a flow chart of a voice interaction method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to voice interaction devices, such as Figure 2 As shown, the following steps are included:

[0034] Step S200: Real-time detection of the wake-up instruction in the voice input signal.

[0035] The wake-up instruction may be a specific voice signal for waking up a device or application for voice interaction. The wake-up instruction may include a wake-up word (Wake Word), which may be pre-set according to actual conditions.

[0036] In one embodiment, a voice collection device, such as a microphone array, can be used to collect ambient audio in real time, and the collected audio can be used as a voice input signal. Specifically, the voice collection device can include an ambient sound collection channel for collecting ambient audio and an echo cancellation reference channel for collecting reference audio. The echo cancellation reference channel can collect audio from a playback device (e.g., a speaker) as reference audio, which can serve as a reference for subsequent echo cancellation of the voice input signal.

[0037] Exemplarily, before performing wake-up command detection on the voice input signal, anti-interference processing such as noise reduction and echo cancellation may be performed on the voice input signal. Wake-up command detection may be performed on the voice input signal using automatic speech recognition (ASR) technology. For example, a pre-set wake-up engine may be used to perform wake-up command detection on the voice input signal in real time to determine whether there is a wake-up command in the voice input signal. The wake-up engine may, for example, employ a Porcupine voice wake-up engine or a pre-trained deep neural network (DNN) for wake-up command recognition.

[0038] In one implementation, the wake-up instruction may further include a wake-up signal. For example, when a wake-up signal is detected by a manual wake-up method such as a control button or a virtual switch, it is determined that the wake-up instruction is detected.

[0039] Exemplarily, when the wake-up engine detects the presence of a wake-up word (wake-up instruction) in the voice input signal, the wake-up engine outputs a state identifier 1; when it detects that there is no wake-up word in the voice input signal, the wake-up engine outputs a state identifier 0, so that it is possible to determine whether to enter the wake-up state based on the state identifier, and the wake-up engine also outputs the start time and end time of the wake-up word, and determines the start time and the end time as the time information of the wake-up instruction; when it is detected that the user wakes up by sending a wake-up signal through manual wake-up, a manual wake-up flag is generated. If the manual wake-up flag is detected, the state identifier is set to 2; if the manual wake-up flag is not detected, the state identifier is set to 0, so that it is possible to determine whether to enter the wake-up state based on the state identifier, and determine the time information of the wake-up instruction based on the triggering time of the wake-up signal (the triggering time of manual wake-up).

[0040] Step S210 : determining a first audio within a target time period of the wake-up instruction based on the time information corresponding to the detected wake-up instruction.

[0041] The time information corresponding to the wake-up instruction includes: the start time and the end time of the wake-up instruction. The first audio includes the audio corresponding to the wake-up instruction, and the time period corresponding to the first audio is the target time period.

[0042] In one embodiment, when the wake-up instruction includes a wake-up signal, that is, when the wake-up is performed manually, the triggering time of the wake-up signal (the triggering time of manual wake-up) is determined as the start time and the end time of the wake-up instruction.

[0043] For example, a buffer area can be pre-set to store the collected audio in the buffer area, and then the audio corresponding to the target time period is obtained from the buffer area based on the start time or end time of the wake-up instruction.

[0044] Step S220: Determine a target position of the wake-up instruction in the first audio according to the time information corresponding to the wake-up instruction.

[0045] The target position may be determined based on the start time or the end time of the wake-up instruction. For example, the position corresponding to the start time of the wake-up instruction in the first audio may be determined as the target position, or the position corresponding to the end time of the wake-up instruction in the first audio may be determined as the target position.

[0046] In one embodiment, when the wake-up instruction includes a wake-up signal, the position of the triggering moment corresponding to the wake-up signal in the first audio may be determined as the target position.

[0047] Step S230: Determine a first recognized voice segment from the first audio according to the target position, and perform voice interaction based on the first recognized voice segment.

[0048] The first audio includes a first recognized voice segment, and the first recognized voice segment is used for voice interaction.

[0049] In one implementation, the first audio can be segmented by the target position to obtain a wake-up segment and an initial recognition voice segment. The wake-up segment is the audio segment located before the target position in the first audio. The wake-up segment includes the audio corresponding to the wake-up instruction (wake-up word). The initial recognition voice segment is the audio segment located after the target position in the first audio. The first recognition voice segment is determined based on the initial recognition voice segment. Exemplarily, the initial voice segment can be determined as the first recognition voice segment, and based on the first recognition voice segment, voice interaction is performed using a voice interaction model.

[0050] In the embodiment of the present disclosure, the first recognized voice segment is determined by the target position information of the wake-up instruction, which can effectively ensure the integrity of the recognized voice segment sent by the user, avoid the problem of the recognized voice segment being truncated due to the collection of the recognized voice segment after the wake-up state, thereby ensuring that the recognized voice segment can be accurately recognized, thereby improving the accuracy of voice interaction and the user experience.

[0051] In some optional implementations, step S210 in the embodiment of the present disclosure may include: based on the start time and / or end time corresponding to the wake-up instruction, obtaining the cached audio within the target time period of the wake-up instruction as the first audio.

[0052] The buffered audio includes: audio whose timing is before the wake-up instruction and audio whose timing is after the wake-up instruction.

[0053] In one embodiment, a preset cache duration corresponding to a cache area can be pre-set, i.e., the cache area can cache audio for the preset cache duration. When newly acquired audio is cached in the cache area, the audio corresponding to the first audio captured in the cache area is deleted, so that the duration of the audio stored in the cache area remains at the preset cache duration. For example, the preset cache duration corresponding to the cache area is 8000ms, i.e., the cache area can store 500 frames of audio, each frame of audio having a duration of 16ms. Assuming that 200 frames of newly acquired audio need to be stored, the 200 frames of audio are stored in the cache area, and the first 200 frames of audio in the corresponding time sequence are deleted.

[0054] When the preset cache duration is equal to the duration corresponding to the target interval, the cached audio stored in the cache area is obtained as the first audio; when the preset cache duration is greater than the duration corresponding to the target time period, the target time interval can be determined by the time interval corresponding to the start time and the end time of the wake-up instruction, or the start time or the end time can be extended forward and backward to obtain the target time interval, and then the target time interval is extended forward and backward to obtain the target time period including the target time interval, and then based on the moment corresponding to the cached audio in the cache area, the cached audio corresponding to the target time period is obtained from the cache area as the first audio.

[0055] In the embodiment of the present disclosure, by caching the audio (voice input signal) collected in real time, the start time and / or end time of the wake-up instruction are used to backtrack and determine the first audio in the cached audio, so that the first audio can include the content of the entire recognition language, avoiding the problem of the recognition language being truncated.

[0056] Figure 3 FIG. 1 is a flow chart of step S230 provided by an exemplary embodiment of the present disclosure. In some optional implementations, such as Figure 3 As shown, step S230 may include the following steps:

[0057] Step S231: Divide the first audio based on the target position to obtain at least two first initial audio segments.

[0058] The first audio may be divided according to the target position and the time corresponding to each audio in the first audio to obtain two first initial audio segments.

[0059] Exemplarily, the target position is the termination moment of the wake-up instruction. Based on the termination moment and the moments corresponding to each frame of audio in the first audio, the first audio is divided into two first initial audio segments, and the two first initial audio segments are respectively located before and after the termination moment in time sequence.

[0060] Step S232: Determine the first audio segment corresponding to the first recognized speech segment based on the time information corresponding to each of the at least two first initial audio segments.

[0061] The first audio segment is the first initial audio segment that is located after the wake-up instruction in time sequence among the at least two first initial audio segments. The time information corresponding to each first initial audio segment may include the start time and the end time of the first initial audio segment.

[0062] Exemplarily, also taking the example in step S231 as an example, the first initial audio segment whose timing is after the end time of the wake-up instruction can be determined as the first segment, that is, the first initial audio segment whose starting time is after the end time of the wake-up instruction among the two first initial audio segments is determined as the first segment.

[0063] Step S233: Determine a first recognized speech segment based on the first sound segment.

[0064] In one embodiment, the first sound segment may be determined as the first recognized speech segment.

[0065] In the embodiment of the present disclosure, the boundary between the wake-up instruction and the first recognized voice segment is determined using the target position corresponding to the wake-up instruction, so that the determined first voice segment can include the entire content of the first recognized voice segment, avoiding the situation where the recognized voice segment is truncated due to the determination of the recognized voice segment after the wake-up moment, thereby ensuring the integrity and accuracy of the recognized voice segment.

[0066] Figure 4 is a flow chart of step S233 provided by an exemplary embodiment of the present disclosure. In some optional implementations, such as Figure 4 As shown, step S233 may include the following steps:

[0067] Step S2331: perform speech endpoint detection on the first sound segment and determine the speech endpoint detection result.

[0068] The speech endpoint detection results include the start and end times of the first recognized speech segment. Voice endpoint detection, also known as Voice Activity Detection (VAD), automatically identifies and distinguishes valid speech segments (including human voices) from invalid speech segments (such as silence, background noise, and non-human vocal interference).

[0069] Exemplarily, a feature and threshold-based method can be used for speech endpoint detection, that is, specific features of the audio are calculated, and then the specific features are compared with a preset threshold to determine the time period corresponding to the valid speech segment (speech endpoint detection result). The specific features may include, for example, the energy, zero crossing rate (Zero Crossing Rate) or spectral features of the audio; a statistical model-based machine learning method can also be used for speech endpoint detection, that is, a model is trained using audio data with valid speech segment labels or invalid speech segment labels, and the model is allowed to learn the statistical distribution differences between valid speech segments and invalid speech segments, so that the trained model can identify valid speech segments (first recognized speech segments) and output the time period corresponding to the valid speech segments. The model can, for example, use a support vector machine (SVM) or a hidden Markov model (HMM).

[0070] Step S2332: Determine a first recognized speech segment from the first speech segment based on the speech endpoint detection result.

[0071] In one embodiment, based on the time corresponding to each frame of audio in the first sound segment, the sound segment corresponding to the start time and the end time of the first recognized speech segment can be determined in the first sound segment as the first recognized speech segment.

[0072] In the embodiment of the present disclosure, by performing speech endpoint detection on the first speech segment, the first recognized speech segment and the invalid speech segment in the first audio can be accurately distinguished, thereby improving the accuracy of the obtained first recognized speech segment.

[0073] In some optional embodiments, in the implementation of the present disclosure, after step S230, it may also include: determining the duration corresponding to the first recognized voice segment based on the voice endpoint detection result, and in response to the duration corresponding to the first recognized voice segment being greater than or equal to a preset duration threshold, performing voice interaction based on the first recognized voice segment.

[0074] Among them, a preset duration threshold can be set in advance, for example, the preset duration threshold can be 500ms. According to the starting time and the ending time of the first recognized voice segment, the duration corresponding to the first recognized voice segment is determined, and the duration corresponding to the first recognized voice segment is compared with the preset duration threshold. When the duration corresponding to the first recognized voice segment is greater than or equal to the preset duration threshold, the first recognized voice segment is determined to be a valid voice segment, that is, it is determined that the first recognized voice segment has practical meaning and the first recognized voice segment can be used for voice interaction; when the duration corresponding to the first recognized voice segment is less than the preset duration threshold, the first recognized voice segment is determined to be an invalid voice segment, that is, it is determined that the first recognized voice segment has no practical meaning. For example, when the first recognized speech segment only includes short voices such as "um" and "ah" that have no practical meaning, the first recognized voice segment is an invalid voice segment. At this time, voice interaction is not performed based on the first recognized voice segment, and the operation of step S200 is performed.

[0075] In the disclosed embodiment, whether the first recognized speech segment is a valid speech segment is determined by the duration of the first recognized speech segment, thereby avoiding the use of invalid speech segments for voice interaction, improving the efficiency of voice interaction, and ensuring the quality of voice interaction.

[0076] In some optional embodiments, in the implementation of the present disclosure, after step S230, the following may be included: in response to the voice interaction being a single-round voice interaction, determining a control instruction corresponding to the first recognized voice segment.

[0077] Single-turn voice interaction refers to a conversational model in which a user fully expresses their intent with a single voice input, and the system directly satisfies the user's needs with a single response. The interaction process is independent of historical context, and the task naturally ends after a single request-response cycle. For example, single-turn voice interaction can be achieved through one-shot communication.

[0078] In one embodiment, a voice interaction model can be pre-set, which can support both single-round voice interaction and multi-round voice interaction. The voice interaction model can adopt, for example, the Mini-Omni voice model or the Step-Audio voice model.

[0079] The voice interaction model determines the mode of voice interaction based on the first recognized voice segment. When the voice interaction mode is determined to be single-round voice interaction, the voice interaction model determines the control instruction based on the first recognized voice segment and outputs the control instruction. The corresponding device or application can be operated based on the control instruction.

[0080] Exemplarily, in this example, the voice interaction method is applied in a smart refrigerator. The smart refrigerator includes a voice interaction device, which is communicatively connected to the control system of the smart refrigerator, and the wake-up instruction is pre-set as "Xiao A". The microphone array in the voice interaction device collects audio from the surrounding environment in real time, and stores the collected audio in a cache area. Suppose the user makes a voice "Xiao A sets the preservation temperature to 5°C", the microphone array in the voice interaction device collects the voice, and transmits the voice to the wake-up engine in the voice interaction device for wake-up instruction detection. The wake-up engine detects that the voice includes the wake-up instruction (Xiao A), and the voice interaction device enters the wake-up state. The wake-up engine outputs the time information corresponding to the wake-up instruction. The interaction device in the voice interaction device determines the target time period based on the start time and / or end time in the time information corresponding to the wake-up instruction, obtains the cached audio in the target time period from the cache area as the first audio, and determines the position corresponding to the end time of the wake-up instruction in the first audio as the target position. According to the target position, the first The audio is divided into two first initial audio segments, and the first initial audio segment whose timing is after the termination moment of the wake-up instruction is determined as the first audio segment. Voice endpoint detection is performed on the first audio segment to determine the voice endpoint detection result. Based on the voice endpoint detection result, a first recognized voice segment (the fresh-keeping temperature is set to 5°C) is determined from the first audio segment. When it is determined that the duration corresponding to the first recognized voice segment is greater than or equal to the preset duration threshold, the first recognized voice segment is sent to the cloud. The voice interaction model in the cloud determines that the voice interaction mode is single-round voice interaction based on the first recognized voice segment, and outputs a control instruction (fresh-keeping temperature 5°C), and transmits the control instruction to the voice interaction device. The voice interaction device sends the control instruction to the control system of the smart refrigerator, and the control system adjusts the fresh-keeping temperature of the smart refrigerator to 5°C based on the control instruction.

[0081] In the embodiment of the present disclosure, when the voice interaction is a single-round voice interaction, the control instruction is determined by the first recognized voice segment, thereby efficiently obtaining the control instruction and improving the user experience.

[0082] Figure 5is a flow chart of a voice interaction method provided by another exemplary embodiment of the present disclosure. Figure 5 As shown, after step S230, the following steps are also included:

[0083] Step S240 , in response to the voice interaction being a multi-round voice interaction, determining a second recognized voice segment in the voice interaction based on the playback status of the audio playback device corresponding to the device response voice segment in the multi-round voice interaction.

[0084] Among them, multi-round voice interaction refers to the process of multiple alternating conversations between the user and the system around one or more related tasks or topics through continuous voice input and output. The device response segment is the voice segment that responds to the first recognized voice segment or the second recognized voice segment in the multi-round voice interaction. The audio playback device may include, for example, a sound or a speaker. The audio playback status may include: playing state (the process state of the audio being played) and playback completion state (the final state of the audio having finished playing).

[0085] Exemplarily, the voice interaction device may also include an audio playback device for playing a device response segment. When the voice interaction model determines that the voice interaction mode is multi-round voice interaction based on the first recognized voice segment, the voice interaction model outputs the corresponding device response voice segment based on the first recognized voice segment, and the device response voice segment is played by the audio playback device, and the playback status of the audio playback device is detected in real time. When it is detected that the audio playback status is completed, the second recognized voice segment issued by the user is collected.

[0086] Step S250: Determine whether to proceed to the next round of voice interaction based on the second recognized voice segment.

[0087] Wherein, whether to conduct the next round of voice interaction can be determined based on the semantics of the second recognized voice segment. Exemplarily, the voice interaction model recognizes and analyzes the second recognized voice segment to determine whether to conduct the next round of voice interaction.

[0088] In an embodiment of the present disclosure, in a scenario where the voice interaction is determined to be a multi-round voice interaction, the second recognized voice segment in the voice interaction is located by monitoring the playback status of the audio playback device, so that the voice information emitted by the user during the multi-round voice interaction can be efficiently and accurately collected, thereby ensuring the overall quality of the multi-round voice interaction.

[0089] Figure 6 is a flow chart of step S240 provided by an exemplary embodiment of the present disclosure. In some optional implementations, such as Figure 6 As shown, step S240 may include the following steps:

[0090] Step S241 : In response to the playback state of the audio playback device being a playback completion state, the second audio corresponding to the voice interaction is divided based on the time corresponding to the playback completion state to obtain at least two second initial audio segments.

[0091] In terms of time sequence, the two second initial audio segments are respectively located before and after the moment corresponding to the playback completion state.

[0092] In one embodiment, the moment corresponding to the playback completion state is extended forward and backward to obtain a time period including the moment, and based on the moment corresponding to the cached audio in the cache area, the cached audio corresponding to the time period is obtained from the cache area as the second audio, and then the second audio is divided based on the moment to obtain two second initial audio segments.

[0093] Step S242: Determine the second sound segment corresponding to the second recognized speech segment based on the time information corresponding to the at least two second initial sound segments.

[0094] The time information corresponding to each second initial sound segment includes the start time and the end time of the second initial sound segment.

[0095] Exemplarily, a second initial sound segment whose starting time is after the time corresponding to the playback completion state may be determined as the second sound segment.

[0096] Step S243: Determine a second recognized speech segment based on the second sound segment.

[0097] The second sound segment can be determined as the second recognized speech segment. Alternatively, voice endpoint detection can be performed on the second sound segment to determine the start time and end time corresponding to the second recognized speech segment, and the duration corresponding to the second recognized speech segment can be determined based on the start time and the end time. When the duration is greater than or equal to a preset duration threshold, the second recognized speech segment is obtained from the second speech segment based on the start time and the end time. When the duration is less than the preset duration threshold, voice interaction is not performed based on the second recognized speech segment, and the operation of collecting user recognition language is performed.

[0098] In the disclosed embodiment, the audio is accurately divided with the help of the moments corresponding to the audio playback status, and the effective sound segments (second recognized voice segments) are filtered through time information, so as to achieve efficient and accurate extraction of the user's subsequent voice commands in multiple rounds of voice interaction, reduce invalid audio interference, ensure the accuracy and reliability of information collection in multiple rounds of voice interaction, and improve the overall interaction quality and efficiency.

[0099] In some optional implementations, in the embodiments of the present disclosure, the playback state of the audio playback device may be determined in the following manner: in response to the target audio signal energy being greater than the first preset audio signal energy, the playback state of the audio playback device is changed to the playback completion state; or, in response to the target audio signal energy being less than the second preset audio signal energy, the playback state of the audio playback device is changed to the playback completion state.

[0100] The target audio signal energy is the audio signal energy corresponding to the device's response voice segment. Audio signal energy represents the actual physical energy carried in the audio signal and is related to the audio's amplitude. Specifically, audio signal energy is proportional to the square of the amplitude. The first preset audio signal energy is greater than the second preset audio signal energy.

[0101] In one embodiment, a target audio signal energy can be obtained and compared with a second preset audio signal energy. If the target audio signal energy is less than the second preset audio signal energy, it indicates that the audio playback device has completed playing the device response voice segment, and the playback state of the audio playback device is switched from a playing state to a playback completed state. If the target audio signal energy is greater than or equal to the second preset audio signal energy, the target audio signal energy is compared with the first preset audio signal energy. If the target audio signal energy is less than or equal to the second preset audio signal energy, it indicates that the audio playback device is normally playing the device response voice segment, and the playback state of the audio playback device is not processed. If the target audio signal energy is greater than the first preset audio signal energy, it indicates that the user has continued to issue an instruction to start the next recognition speech (the second recognition voice segment) before the audio playback device has completed playing the device response voice segment, that is, the second recognition voice segment issued by the user overlaps with the device response voice segment. At this time, the playback state of the audio playback device is switched from a playing state to a playback completed state.

[0102] For example, the target audio signal energy can be obtained by the total energy in the time domain. Specifically, n sampling points can be set in advance in the device response voice segment, and then the target audio signal energy can be obtained by the total energy in the time domain. Determine the target audio signal energy, E signal represents the target audio signal energy, x i represents the amplitude value of the i-th sampling point; alternatively, the target audio signal energy can be obtained by decibel conversion. Specifically, the decibel value dB corresponding to the device response voice segment can be obtained first, and the square value of the maximum amplitude value in the device response voice segment can be used as the reference energy E reference , then based on Determine the target audio signal energy.

[0103] In one embodiment, the first preset audio signal energy may be obtained by acquiring an average audio signal energy of audio played by the audio playback device within a preset time period, and determining the average audio signal energy as the first preset audio signal energy.

[0104] The preset time period can be set according to actual needs. For example, the average audio signal energy of the audio within 2 seconds to 3 seconds (preset duration) after the audio playback device enters the playback state can be obtained as the first preset audio signal energy. The method for obtaining the average audio signal energy is the same as the method for obtaining the target audio signal energy. The method for obtaining the target audio signal energy can be referred to and will not be repeated here.

[0105] In the embodiment of the present disclosure, by real-time detection of the target audio signal energy of the audio playback device when the playback device responds to the voice segment, and combining the first preset audio signal energy and the second preset audio signal energy, it is possible to accurately determine whether the audio playback device has completed the playback and whether the user has issued a new recognition voice segment. This allows the playback status of the audio playback device to be switched instantly, and the user's recognition voice segment can be efficiently collected.

[0106] In some optional implementations, the voice interaction method in the embodiments of the present disclosure may further include: determining control instructions corresponding to the multiple rounds of voice interaction based on audio corresponding to the multiple rounds of voice interaction.

[0107] Among them, the semantics corresponding to the audio corresponding to multiple rounds of voice interaction can be determined through the voice interaction model, and the control instructions can be determined based on the semantics.

[0108] Exemplarily, the voice interaction method in this example is applied in a vehicle. The vehicle includes a voice interaction device, which is communicatively connected to the control system of the vehicle, and the wake-up instruction is set to "Xiao A". The microphone array in the voice interaction device collects audio from the surrounding environment in real time, and stores the collected audio in a cache area. Suppose the user utters a voice "Xiao A, open the window", the microphone array in the voice interaction device collects the voice, and sends the voice to the wake-up engine in the voice interaction device for wake-up instruction detection. The wake-up engine detects that the voice includes the wake-up instruction (Xiao A), and the voice interaction device enters the wake-up state. The wake-up engine outputs the time information corresponding to the wake-up instruction. The interaction device in the voice interaction device determines the target time period based on the start time and / or end time in the time information corresponding to the wake-up instruction, obtains the cached audio in the target time period from the cache area as the first audio, and uses the wake-up instruction as the first audio. The position corresponding to the termination moment of the command in the first audio is determined as the target position, the first audio is divided into two first initial audio segments according to the target position, the first initial audio segment whose timing is after the termination moment of the wake-up command is determined as the first sound segment, voice endpoint detection is performed on the first sound segment, and the start time and end time of the first recognized voice segment are determined (voice endpoint detection result), and the duration corresponding to the first recognized voice segment is determined based on the start time and the end time. When the duration is greater than or equal to the preset duration threshold, the first recognized voice segment (open the car window) is determined from the first sound segment, and the first recognized voice segment is sent to the cloud. The voice interaction model in the cloud determines the voice interaction mode based on the first recognized voice segment. For multi-round voice interaction, the device responds to the voice segment "Please ask which window of the vehicle should be opened", and the device responds to the voice segment. The audio playback device in the voice interaction device plays the device responds to the voice segment, and detects the target audio signal energy of the audio playback device in real time. When it is detected that the target audio signal energy is greater than the first preset audio signal energy, or the target audio signal energy is less than the second preset audio signal energy, the playback state of the audio playback device is changed to the playback completion state, and the second audio is obtained from the buffer area based on the time corresponding to the playback completion state, and the second audio is divided based on the time to obtain two second initial sounds, and the starting time is positioned at The second initial sound segment after this moment is determined as the second sound segment, and voice endpoint detection is performed on the second sound segment to determine the starting time and ending time corresponding to the second recognized voice segment. Based on the starting time and the ending time, the duration corresponding to the second recognized voice segment is determined. When the duration is greater than or equal to the preset duration threshold, the second recognized voice is obtained from the second sound segment based on the starting time and the ending time, and the second recognized voice segment is sent to the voice interaction model. The voice interaction model recognizes the second recognized voice segment and determines whether to proceed to the next round of voice interaction. Assuming that the second recognized voice segment is "open the window in the driving position", the voice interaction model determines that the voice interaction is completed and there is no need to proceed to the next round of voice interaction.Output the operation instruction "open the window at the driver's position" and transmit the control instruction to the voice interaction device. The voice interaction device sends the control instruction to the vehicle control system, and the control system opens the window at the driver's position based on the control instruction.

[0109] In the disclosed embodiment, control instructions can be accurately determined through audio in multiple rounds of voice interaction, thereby improving the user experience.

[0110] For example, Figure 7 This is a flow chart of a voice interaction method provided by an application example of the present disclosure. Figure 7 As shown, the following steps are included:

[0111] S1, real-time detection of wake-up commands in voice input signals;

[0112] S2, when a wake-up instruction is detected, based on the start time and / or end time corresponding to the wake-up instruction, obtaining the cached audio within the target time period of the wake-up instruction as the first audio, and then executing step S4;

[0113] S3, when no wake-up instruction is detected, execute step S1;

[0114] S4, determining the termination time of the wake-up instruction as the target position;

[0115] S5, dividing the first audio based on the target position to obtain two first initial audio segments;

[0116] S6, determining the first initial audio segment whose timing is after the termination moment of the wake-up instruction as the first audio segment;

[0117] S7, performing speech endpoint detection on the first sound segment and determining a speech endpoint detection result;

[0118] S8, determining the duration of the first recognized speech segment based on the start time and end time of the first recognized speech segment in the speech endpoint detection result;

[0119] S9, determining whether the duration is greater than or equal to a preset duration threshold, if yes, executing step S10, if not, executing step S1;

[0120] S10, determining a first recognized speech segment from the first sound segment based on the start time and end time of the first recognized speech segment;

[0121] S11, determining whether it is a single-round voice interaction based on the first recognized voice segment, if yes, executing step 12, if not, executing step 13;

[0122] S12, determining a control instruction corresponding to the first recognized speech segment based on the first recognized speech segment, and then executing step S1;

[0123] S13, generating a device response voice segment, and playing the device response voice segment through an audio playback device;

[0124] S14, detecting the playback state of the audio playback device, and when the target audio signal energy is greater than a first preset audio signal energy, or the target audio signal energy is less than a second preset audio signal energy, changing the playback state of the audio playback device to a playback completion state, where the target audio signal energy is the audio signal energy corresponding to the device's response voice segment;

[0125] S15, when the playback state of the audio playback device is a playback completion state, dividing the second audio based on the time corresponding to the playback completion state to obtain two second initial audio segments;

[0126] S16, determining a second sound segment based on the time information corresponding to the two second initial sound segments;

[0127] S17, determining a second recognized speech segment based on the second sound segment;

[0128] S18, determining whether to perform the next round of voice interaction based on the second recognized voice segment, and if it is determined that the next round of voice interaction is to be performed, executing steps S13 to S18; if it is determined that the next round of voice interaction is not to be performed, executing step S19;

[0129] S19, based on the audio corresponding to the multiple rounds of voice interaction, determine the control instructions corresponding to the multiple rounds of voice interaction, and then execute step S1.

[0130] Exemplary devices

[0131] Figure 8 This is a structural block diagram of a voice interaction device provided by an exemplary embodiment of the present disclosure. Figure 8 As shown, the voice interaction device is applied to an electronic device, and the voice interaction device includes: an instruction detection module 400, an audio determination module 410, a positioning module 420 and a voice interaction module 430.

[0132] The command detection module 400 is used to detect the wake-up command in the voice input signal in real time;

[0133] An audio determination module 410 is configured to determine, based on the detected time information corresponding to the wake-up instruction, a first audio within a target time period of the wake-up instruction, where the time information corresponding to the wake-up instruction includes a start time and an end time of the wake-up instruction;

[0134] a positioning module 420, configured to determine a target position of the wake-up instruction in the first audio according to time information corresponding to the wake-up instruction;

[0135] The voice interaction module 430 is configured to determine a first recognized voice segment from the first audio according to the target location, and perform voice interaction based on the first recognized voice segment.

[0136] In some optional embodiments, the audio determination module 410 in the embodiment of the present disclosure is specifically used to obtain the cached audio within the target time period of the wake-up instruction as the first audio based on the start time and / or end time corresponding to the wake-up instruction, and the cached audio includes audio whose timing is before the wake-up instruction and audio whose timing is after the wake-up instruction.

[0137] Figure 9 This is a structural block diagram of a voice interaction device provided by another exemplary embodiment of the present disclosure. Figure 9 As shown, in some optional implementations, the positioning module 420 in the embodiment of the present disclosure includes:

[0138] A first audio segmentation submodule 421 is configured to segment the first audio based on the target position to obtain at least two first initial audio segments;

[0139] a first audio segment determination submodule 422 configured to determine, based on time information corresponding to each of the at least two first initial audio segments, a first audio segment corresponding to the first recognized speech segment, wherein the first audio segment is the first initial audio segment of the at least two first initial audio segments that is located after the wake-up instruction in terms of time sequence;

[0140] The first recognized language determination submodule 423 is configured to determine the first recognized speech segment according to the first sound segment.

[0141] In some optional implementations, the first recognition term determination submodule 423 in the embodiment of the present disclosure includes:

[0142] An endpoint detection unit 4231 is configured to perform speech endpoint detection on the first speech segment and determine a speech endpoint detection result, wherein the speech endpoint detection result includes a start time and an end time of the first recognized speech segment;

[0143] The recognition language determination unit 4232 is used to determine the first recognition speech segment from the first speech segment based on the speech endpoint detection result.

[0144] In some optional implementations, the voice interaction device in the embodiment of the present disclosure further includes:

[0145] a duration determination module 440 for determining a duration corresponding to the first recognized speech segment based on the speech endpoint detection result;

[0146] The segment validity judgment module 450 is configured to perform the voice interaction based on the first recognized voice segment in response to the duration corresponding to the first recognized voice segment being greater than or equal to a preset duration threshold.

[0147] In some optional implementations, the voice interaction device in the embodiment of the present disclosure further includes:

[0148] The first control instruction determination module 460 is configured to determine a control instruction corresponding to the first recognized voice segment in response to the voice interaction being a single-round voice interaction.

[0149] In some optional implementations, the voice interaction device in the embodiment of the present disclosure further includes:

[0150] a recognition language determination module 470 for determining, in response to the voice interaction being a multi-round voice interaction, a second recognition voice segment in the voice interaction based on a playback state of an audio playback device corresponding to a device response voice segment in the multi-round voice interaction;

[0151] The voice interaction determination module 480 is configured to determine whether to perform the next round of voice interaction based on the second recognized voice segment.

[0152] In some optional implementations, the recognition term determination module 470 in the embodiment of the present disclosure includes:

[0153] A second audio segmentation submodule 471 is configured to, in response to the playback state of the audio playback device being a playback completion state, segment the second audio corresponding to the voice interaction based on a time corresponding to the playback completion state to obtain at least two second initial audio segments;

[0154] A second sound segment determination submodule 472 is configured to determine a second sound segment corresponding to the second recognized speech segment based on time information corresponding to at least two of the second initial sound segments;

[0155] The second recognized language determination submodule 473 is configured to determine the second recognized speech segment according to the second sound segment.

[0156] In some optional implementations, the voice interaction device in the embodiment of the present disclosure further includes:

[0157] The playback status determination module 490 is used to change the playback status of the audio playback device to the playback completion status in response to the target audio signal energy being greater than the first preset audio signal energy, where the target audio signal energy is the audio signal energy corresponding to the device's response voice segment; or, in response to the target audio signal energy being less than the second preset audio signal energy, change the playback status of the audio playback device to the playback completion status.

[0158] In some optional implementations, the voice interaction device in the embodiment of the present disclosure further includes:

[0159] The audio signal energy determination module 500 is configured to obtain an average audio signal energy of the audio played by the audio playback device within a preset time period; and determine the average audio signal energy as the first preset audio signal energy.

[0160] In some optional implementations, the voice interaction device in the embodiment of the present disclosure further includes:

[0161] The second control instruction determination module 510 determines the control instructions corresponding to the multiple rounds of voice interaction based on the audio corresponding to the multiple rounds of voice interaction.

[0162] The voice interaction device of the embodiment of the present disclosure corresponds to the embodiment of the above-mentioned voice interaction method of the present disclosure, and the relevant contents can be referenced to each other and will not be repeated here.

[0163] The beneficial technical effects corresponding to the exemplary embodiments of the voice interaction device of the embodiments of the present disclosure can be found in the corresponding beneficial technical effects of the above-mentioned corresponding exemplary method part, which will not be repeated here.

[0164] Figure 10 This is a structural block diagram of a voice interaction system provided by an exemplary embodiment of the present disclosure. Figure 10 As shown, the voice interaction system 600 includes: an audio collection device 610, an audio playback device 620 and an interaction device 630; the audio collection device 610 is used to collect voice input signals, and the audio playback device 620 is used to play the device response voice segment in the voice interaction;

[0165] The interaction device 630 is used to: perform real-time detection of wake-up instructions in a voice input signal; determine the first audio within the target time period of the wake-up instruction based on the time information corresponding to the detected wake-up instruction, and the time information corresponding to the wake-up instruction includes the start time and end time of the wake-up instruction; determine the target position of the wake-up instruction in the first audio according to the time information corresponding to the wake-up instruction; determine a first recognized voice segment from the first audio according to the target position, and perform voice interaction based on the first recognized voice segment.

[0166] In some optional implementations, the determining, based on the time information corresponding to the wake-up instruction, the first audio within the target time period of the wake-up instruction in the embodiment of the present disclosure is further configured to:

[0167] Based on the start time and / or end time corresponding to the wake-up instruction, the cached audio within the target time period of the wake-up instruction is obtained as the first audio, and the cached audio includes the audio whose timing is before the wake-up instruction and the audio whose timing is after the wake-up instruction.

[0168] In some optional embodiments, the method of determining the first recognized voice segment from the first audio according to the target position in the embodiment of the present disclosure is further used to: divide the first audio based on the target position to obtain at least two first initial sound segments; determine the first sound segment corresponding to the first recognized voice segment according to the time information corresponding to each first initial audio segment in the at least two first initial sound segments, the first sound segment being the first initial sound segment in the at least two first initial sound segments that is located after the wake-up instruction in time sequence; determine the first recognized voice segment based on the first sound segment.

[0169] In some optional embodiments, the method of determining the first recognized speech segment based on the first sound segment in the embodiment of the present disclosure is further used to: perform speech endpoint detection on the first sound segment to determine a speech endpoint detection result, wherein the speech endpoint detection result includes the start time and end time of the first recognized speech segment; and determine the first recognized speech segment from the first sound segment based on the speech endpoint detection result.

[0170] In some optional embodiments, the interaction device 630 in the embodiment of the present disclosure is also used to: determine the duration corresponding to the first recognized voice segment based on the voice endpoint detection result; in response to the duration corresponding to the first recognized voice segment being greater than or equal to a preset duration threshold, perform the voice interaction based on the first recognized voice segment.

[0171] In some optional implementations, the interaction device 630 in the embodiment of the present disclosure is further used to: in response to the voice interaction being a single-round voice interaction, determine the control instruction corresponding to the first recognized voice segment.

[0172] In some optional embodiments, the interaction device 630 in the embodiment of the present disclosure is also used to determine the second recognized voice segment in the voice interaction based on the playback status of the audio playback device corresponding to the device response voice segment in the multiple rounds of voice interaction in response to the voice interaction being a multi-round voice interaction; and determine whether to perform the next round of voice interaction based on the second recognized voice segment.

[0173] In some optional implementations, the determining of the second recognized voice segment in the voice interaction based on the playback status of the audio playback device corresponding to the device response voice segment in the voice interaction in the embodiment of the present disclosure is further used to:

[0174] In response to the playback status of the audio playback device being a playback completion status, the second audio corresponding to the voice interaction is divided based on the time corresponding to the playback completion status to obtain at least two second initial sound segments; based on the time information corresponding to the at least two second initial sound segments, the second sound segment corresponding to the second recognized voice segment is determined; based on the second sound segment, the second recognized voice segment is determined.

[0175] In some optional embodiments, the interactive device 630 in the embodiment of the present disclosure is further used to change the playback state of the audio playback device to the playback completion state in response to the target audio signal energy being greater than the first preset audio signal energy, where the target audio signal energy is the audio signal energy corresponding to the device's response voice segment; or, in response to the target audio signal energy being less than the second preset audio signal energy, change the playback state of the audio playback device to the playback completion state.

[0176] In some optional implementations, the interactive device 630 in the embodiment of the present disclosure is further configured to obtain an average audio signal energy of the audio played by the audio playback device within a preset time period; and determine the average audio signal energy as the first preset audio signal energy.

[0177] In some optional implementations, the interaction device 630 in the embodiment of the present disclosure is further configured to determine control instructions corresponding to the multiple rounds of voice interactions based on audio corresponding to the multiple rounds of voice interactions.

[0178] The voice interaction system of the embodiment of the present disclosure corresponds to the embodiment of the above-mentioned voice interaction method of the present disclosure, and the relevant contents can be referenced to each other and will not be repeated here.

[0179] The beneficial technical effects corresponding to the exemplary embodiments of the voice interaction system of the embodiments of the present disclosure can be found in the corresponding beneficial technical effects of the above-mentioned corresponding exemplary method part, which will not be repeated here.

[0180] Exemplary electronic devices

[0181] Figure 11 This is a structural diagram of an electronic device provided in an embodiment of the present disclosure. The electronic device 700 includes at least one processor 710 and a memory 720.

[0182] The processor 710 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 700 to perform desired functions.

[0183] The memory 720 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 710 may execute one or more computer program instructions to implement the voice interaction method and / or other desired functions of the various embodiments of the present disclosure described above.

[0184] In one example, the electronic device 700 may further include an input device 730 and an output device 740 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0185] The input device 730 may also include, for example, a keyboard, a mouse, and the like.

[0186] The output device 740 can output various information to the outside, and may include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.

[0187] Of course, to simplify, Figure 11 Only some of the components related to the present disclosure in the electronic device 700 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 700 may further include any other appropriate components according to specific application scenarios.

[0188] Exemplary computer program products and computer-readable storage media

[0189] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the voice interaction method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.

[0190] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0191] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps in the voice interaction method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.

[0192] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or component comprising electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0193] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be considered as essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0194] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.

Claims

1. A voice interaction method, comprising: Real-time detection of wake-up commands in voice input signals; Determining, based on the detected time information corresponding to the wake-up instruction, a first audio within a target time period in which the wake-up instruction is located, where the time information corresponding to the wake-up instruction includes a start time and an end time of the wake-up instruction; determining, according to time information corresponding to the wake-up instruction, a target position of the wake-up instruction in the first audio; According to the target position, a first recognized voice segment is determined from the first audio, and voice interaction is performed based on the first recognized voice segment.

2. The method according to claim 1, wherein determining the first audio within the target time period of the wake-up instruction based on the time information corresponding to the wake-up instruction comprises: Based on the start time and / or end time corresponding to the wake-up instruction, the cached audio within the target time period of the wake-up instruction is obtained as the first audio, and the cached audio includes the audio whose timing is before the wake-up instruction and the audio whose timing is after the wake-up instruction.

3. The method according to claim 1 or 2, wherein: The step of determining a first recognized speech segment from the first audio according to the target position includes: Dividing the first audio based on the target position to obtain at least two first initial audio segments; Determining, based on time information corresponding to each of the at least two first initial audio segments, a first audio segment corresponding to the first recognized speech segment, where the first audio segment is a first initial audio segment that is located after the wake-up instruction in time sequence among the at least two first initial audio segments; The first recognized speech segment is determined according to the first sound segment.

4. The method according to claim 3, wherein: The determining the first recognized speech segment according to the first sound segment includes: Performing speech endpoint detection on the first speech segment to determine a speech endpoint detection result, the speech endpoint detection result including a start time and an end time of the first recognized speech segment; Based on the speech endpoint detection result, the first recognized speech segment is determined from the first speech segment.

5. The method according to claim 4, further comprising: after determining the first recognized speech segment from the first audio according to the target location information; Determining a duration corresponding to the first recognized speech segment based on the speech endpoint detection result; In response to the duration corresponding to the first recognized voice segment being greater than or equal to a preset duration threshold, the voice interaction is performed based on the first recognized voice segment.

6. The method according to any one of claims 1 to 5, further comprising: after determining a first recognized speech segment from the first audio according to the target location; In response to the voice interaction being a single-round voice interaction, a control instruction corresponding to the first recognized voice segment is determined.

7. The method according to any one of claims 1 to 5, further comprising: after determining a first recognized speech segment from the first audio according to the target location; In response to the voice interaction being a multi-round voice interaction, determining a second recognized voice segment in the voice interaction based on a playback state of an audio playback device corresponding to a device response voice segment in the multi-round voice interaction; Determine whether to perform the next round of voice interaction based on the second recognized voice segment.

8. The method according to claim 7, wherein: The determining, based on a playback state of an audio playback device corresponding to a device response voice segment in the voice interaction, a second recognized voice segment in the voice interaction includes: In response to the playback state of the audio playback device being a playback completion state, dividing the second audio corresponding to the voice interaction based on a time corresponding to the playback completion state to obtain at least two second initial audio segments; determining, based on time information corresponding to at least two of the second initial sound segments, a second sound segment corresponding to the second recognized speech segment; The second recognized speech segment is determined according to the second sound segment.

9. The method according to claim 8, further comprising: In response to a target audio signal energy being greater than a first preset audio signal energy, changing the playback state of the audio playback device to the playback completion state, wherein the target audio signal energy is an audio signal energy corresponding to the device response voice segment; or, In response to the target audio signal energy being less than the second preset audio signal energy, the playing state of the audio playing device is changed to the playing completion state.

10. The method according to claim 9, further comprising: Obtaining average audio signal energy of the audio played by the audio playback device within a preset time period; The average audio signal energy is determined as the first preset audio signal energy.

11. The method according to any one of claims 8 to 10, further comprising: Based on the audio corresponding to the multiple rounds of voice interaction, control instructions corresponding to the multiple rounds of voice interaction are determined.

12. A voice interaction device, comprising: A command detection module is used to detect the wake-up command in the voice input signal in real time; an audio determination module, configured to determine, based on the detected time information corresponding to the wake-up instruction, a first audio within a target time period in which the wake-up instruction is located, wherein the time information corresponding to the wake-up instruction includes a start time and an end time of the wake-up instruction; a positioning module, configured to determine a target position of the wake-up instruction in the first audio according to time information corresponding to the wake-up instruction; A voice interaction module is used to determine a first recognized voice segment from the first audio according to the target position, and perform voice interaction based on the first recognized voice segment.

13. A voice interaction system, comprising: Audio acquisition device, audio playback device and interactive device; The audio acquisition device is used to collect voice input signals, and the audio playback device is used to play the device response voice segment in voice interaction; The interactive device is used for: Real-time detection of wake-up commands in voice input signals; Determining, based on the detected time information corresponding to the wake-up instruction, a first audio within a target time period in which the wake-up instruction is located, where the time information corresponding to the wake-up instruction includes a start time and an end time of the wake-up instruction; determining, according to time information corresponding to the wake-up instruction, a target position of the wake-up instruction in the first audio; According to the target position, a first recognized voice segment is determined from the first audio, and voice interaction is performed based on the first recognized voice segment.

14. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the voice interaction method according to any one of claims 1 to 11.

15. An electronic device, comprising: A processor, and a memory communicatively connected to the processor, further comprising the voice interaction device according to claim 12; The memory is used to store instructions executable by the processor; The processor is used to read the executable instructions from the memory to control the voice interaction device to implement the voice interaction method described in any one of claims 1-11 above.