Voice interaction method and related device

By comparing the voiceprint features of wake-up voice and voice commands using voiceprint recognition technology, the problem of erroneous responses from smart devices in noisy environments is solved, enabling more efficient and accurate voice interaction.

CN119380728BActive Publication Date: 2026-06-16BEIJING JINGDONG TUOXIAN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JINGDONG TUOXIAN TECH CO LTD
Filing Date
2024-09-24
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In noisy environments, smart devices may fail to respond to voice commands that wake up the user when multiple people are talking, leading to false responses and a reduced user experience.

Method used

By performing voiceprint recognition on the wake-up voice and subsequent voice commands, comparing voiceprint features, and determining processing strategies based on similarity and validity, the corresponding voice interaction operations are executed to avoid false responses.

Benefits of technology

It improves the accuracy and efficiency of voice interaction, ensures consistency between the person waking up the user and the person issuing the voice command, reduces false responses, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380728B_ABST
    Figure CN119380728B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice interaction method and related equipment, and relates to the technical field of voice processing. The method comprises: in response to the target wake-up word recognized by the acquired first voice, performing voiceprint recognition on the first voice to obtain first voiceprint features of the first voice; acquiring second voice, performing voiceprint recognition on the second voice to obtain second voiceprint features of the second voice; determining a processing strategy for the second voice according to the comparison result of the first voiceprint features and the second voiceprint features and the validity of the second voice; and executing the processing strategy to perform voice interaction. The present disclosure can effectively avoid the problem of false response by comparing the voiceprint features in the wake-up stage and the voiceprint features of the person issuing the voice instruction, thereby improving the accuracy and efficiency of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech processing technology, and in particular to a speech interaction method, speech interaction device, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] More and more smart devices are using intelligent voice for human-computer interaction. Typically, when interacting with a smart device via voice, you need to wake it up first, and then input a voice command or engage in a voice conversation. However, in noisy environments, multiple people talking can cause the smart device to fail to respond to the voice command used to wake it up, resulting in false responses and significantly reducing the user experience.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] This disclosure provides a voice interaction method and related device, which at least to some extent overcomes the problem of false responses in voice interaction in related technologies.

[0005] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0006] According to one aspect of this disclosure, a voice interaction method is provided, comprising: in response to acquiring a target wake word of a first voice, performing voiceprint recognition on the first voice to obtain a first voiceprint feature of the first voice; acquiring a second voice, performing voiceprint recognition on the second voice to obtain a second voiceprint feature of the second voice; determining a processing strategy for the second voice based on a comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second voice; and executing the processing strategy to perform voice interaction.

[0007] In one embodiment of this disclosure, determining the processing strategy for the second speech based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, includes: if the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset first voiceprint threshold, and the second speech is valid, then determining the processing strategy for the second speech as responding to the second speech.

[0008] In one embodiment of this disclosure, the method further includes: if the first voiceprint feature and the second voiceprint feature are less than a preset second voiceprint threshold, then determining that the processing strategy for the second speech is to not respond to the second speech, or not perform speech recognition on the second speech, or discard the second speech, wherein the preset second voiceprint threshold is less than the preset first voiceprint threshold.

[0009] In one embodiment of this disclosure, when the second voiceprint feature includes multiple sub-voiceprint features, the method further includes: calculating the similarity between the first voiceprint feature and each sub-voiceprint feature respectively; determining a target sub-voiceprint feature whose similarity to the first voiceprint feature is greater than the preset first voiceprint threshold; extracting a speech segment corresponding to the target sub-voiceprint feature from the second speech; performing speech recognition on the speech segment to obtain the recognition result of the speech segment; and determining a processing strategy for the second speech based on the recognition result of the speech segment.

[0010] In one embodiment of this disclosure, determining the processing strategy for the second speech based on the recognition result of the speech segment includes: if the recognition result of the speech segment is valid speech, then determining the processing strategy for the second speech to be responding to the speech segment; if the recognition result of the speech segment is invalid speech, then determining the processing strategy for the second speech to be not responding to the second speech, or discarding the second speech.

[0011] In one embodiment of this disclosure, the method further includes: if no voice command of the second voice is recognized, then the validity of the second voice is determined to be invalid voice; if a voice command of the second voice is recognized, then the validity of the second voice is determined to be valid voice.

[0012] In one embodiment of this disclosure, the method further includes: in response to the processing strategy for the second voice being not to respond to the second voice, waiting for or prompting the user to input the next voice segment; if the next voice segment is not obtained within a preset time period, then exiting the voice interaction.

[0013] According to another aspect of this disclosure, a voice interaction device is also provided, comprising: a first voiceprint recognition module, configured to perform voiceprint recognition on the first voice in response to acquiring a target wake-up word of the first voice, and obtain a first voiceprint feature of the first voice; a second voiceprint recognition module, configured to acquire a second voice, perform voiceprint recognition on the second voice, and obtain a second voiceprint feature of the second voice; a processing strategy determination module, configured to determine a processing strategy for the second voice based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second voice; and a voice interaction execution module, configured to execute the processing strategy to perform voice interaction.

[0014] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the above-described voice interaction method by executing the executable instructions.

[0015] According to another aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described voice interaction method.

[0016] According to another aspect of this disclosure, a computer program product is provided, the computer program product comprising a computer program or computer instructions, the computer program or computer instructions being loaded and executed by a processor to enable a computer to implement the above-described voice interaction method.

[0017] In this embodiment of the disclosure, in response to obtaining the target wake-up word of the first speech, voiceprint recognition is performed on the first speech to obtain the first voiceprint feature of the first speech; the second speech is obtained, and voiceprint recognition is performed on the second speech to obtain the second voiceprint feature of the second speech; based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, a processing strategy for the second speech is determined; the processing strategy is executed to perform voice interaction. By comparing the voiceprint feature of the wake-up stage with the voiceprint feature of the person issuing the voice command, the problem of false response can be effectively avoided, and the accuracy and efficiency of voice interaction can be improved.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0020] Figure 1 This diagram illustrates a system architecture for a voice interaction method or apparatus provided in an embodiment of this disclosure.

[0021] Figure 2 This diagram illustrates a flowchart of a voice interaction method provided in an embodiment of the present disclosure;

[0022] Figure 3 This diagram illustrates another voice interaction method provided by an embodiment of the present disclosure.

[0023] Figure 4 This diagram illustrates a flowchart of yet another voice interaction method provided in an embodiment of the present disclosure;

[0024] Figure 5 This diagram illustrates a flowchart of yet another voice interaction method provided in an embodiment of the present disclosure;

[0025] Figure 6 This diagram illustrates an example flowchart of a voice interaction method provided in an embodiment of this disclosure.

[0026] Figure 7 This diagram illustrates the structure of a voice interaction device according to an embodiment of the present disclosure.

[0027] Figure 8 A structural block diagram of an electronic device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0028] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0029] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0030] Figure 1 An exemplary system architecture 100 is shown that can be applied to the voice interaction method or voice interaction device of the present disclosure embodiments.

[0031] like Figure 1 As shown, system architecture 100 may include terminal device 110, network 120 and server 130.

[0032] Network 120 is a medium used to provide a communication link between terminal device 110 and server 130, and can be a wired network or a wireless network. Network 120 can include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0033] Users can use terminal device 110 to interact with server 130 via network 120 to receive or send messages.

[0034] Terminal device 110 can be any electronic device equipped with a voice acquisition device and supporting web browsing, including but not limited to smartphones, smart cars, smart speakers, robots, tablets, displays, and desktop computers. The voice acquisition device can capture user-inputted voice.

[0035] It should be noted that the client applications installed on terminal devices 110 are the same, or are clients of the same type of application based on different operating systems. Depending on the terminal platform, the specific form of the application client can also differ; for example, the application client can be a mobile client, a PC client, etc.

[0036] Server 130 can be a server that provides various services, such as a backend management server that supports the voice collected by the user using terminal device 110. The backend server can analyze and process the first and second voices acquired by terminal device 110, and feed back the processing results (such as processing strategies) to terminal device 110.

[0037] Optionally, server 130 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0038] Those skilled in the art will know that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative; any number of terminal devices, networks, and servers can be included depending on actual needs. This disclosure does not limit the scope of the embodiments.

[0039] In related technologies, during voice interaction, after the terminal device is woken up by a wake word, the terminal device will continue to receive the next segment of voice as input for voice commands, and perform speech recognition on the next segment of voice, and the system will make a corresponding response based on the recognized content.

[0040] During this process, the wake-up person and the person issuing the voice command are not identified, which may lead to the terminal device responding as if someone else's voice is valid input after being woken up, resulting in a false response.

[0041] For example, after user A wakes up the terminal device, in the environment where user A is located, there are users B, user C, etc. If user B or user C is speaking, then the voice of user B or user C may be collected as the voice command for this voice interaction, resulting in the inability to effectively respond to user A's voice command.

[0042] To at least partially solve the above-mentioned technical problems, this disclosure provides a voice interaction method that compares the voiceprint features during voice interaction with the voiceprint features during wake-up, thereby ensuring that the person waking up the smart device is consistent with the user issuing the voice command, and avoiding the problem of false responses.

[0043] Specifically, in response to acquiring the target wake-up word of the first speech, voiceprint recognition is performed on the first speech to obtain the first voiceprint feature; the second speech is acquired, and voiceprint recognition is performed on the second speech to obtain the second voiceprint feature; based on the comparison result of the first and second voiceprint features and the validity of the second speech, a processing strategy for the second speech is determined; the processing strategy is executed to perform voice interaction. By comparing the voiceprint features of the wake-up phase with the voiceprint features of the person issuing the voice command, the problem of false responses can be effectively avoided, and the accuracy and efficiency of voice interaction can be improved.

[0044] It should be noted that, unless otherwise specified, the embodiments of the present invention and the technical features thereof can be combined with each other.

[0045] The following detailed description of this exemplary implementation method is provided in conjunction with the accompanying drawings and embodiments.

[0046] First, this disclosure provides a voice interaction method that can be executed by any system with computing power. The voice interaction method provided in this disclosure can be executed by a terminal device, such as a voice interaction device configured on the terminal device.

[0047] Figure 2 This diagram illustrates a flowchart of a voice interaction method provided in an embodiment of the present disclosure, such as... Figure 2 As shown, the voice interaction method provided in this embodiment includes the following steps:

[0048] S202. In response to obtaining the target wake-up word of the first speech, perform voiceprint recognition on the first speech to obtain the first voiceprint feature of the first speech.

[0049] The first voice input method can capture the voice uttered by the person waking up through the voice acquisition device of the terminal device, such as the microphone in a smart vehicle or robot. The microphone converts sound waves into electrical signals. Alternatively, the recording function of a smart vehicle's dashcam can capture and save in-vehicle sounds, such as user conversations and ambient sounds.

[0050] After acquiring the first speech signal, preprocessing steps such as filtering, analog-to-digital (AD) conversion, pre-emphasis, and endpoint detection can be performed. Filtering suppresses all components in the input first speech signal whose frequency exceeds half its range, preventing aliasing interference and also suppressing 50Hz power supply frequency interference. AD conversion converts the analog speech signal into a digital signal, quantizing the first speech signal. The difference between the quantized signal value and the original first speech signal value is the quantization error, or quantization noise. Pre-emphasis boosts the high-frequency components, flattening the spectrum of the first speech signal across the entire frequency band from low to high frequencies, allowing for spectrum analysis with the same signal-to-noise ratio. Endpoint detection is used to determine the start and end points of the first speech signal from a segment containing speech.

[0051] In one embodiment, after the terminal device acquires the first speech, it can perform speech recognition on the first speech to determine the semantic information in the first speech. Speech recognition can realize the conversion of speech to text, and can be performed on the first speech using methods such as template matching, probabilistic modeling, and deep learning. Template matching compares the input speech signal with a pre-stored template, matching the corresponding words or phrases as the semantic information of the first speech. Probabilistic modeling uses probabilistic statistics to model and recognize speech signals. Based on the probability distribution of sound features, it recognizes speech by establishing acoustic and language models. The acoustic model converts the speech signal into a series of feature parameters, and the language model performs grammatical and semantic analysis based on these feature parameters. Deep learning automatically learns speech features by training on a large amount of data. For example, models such as deep neural networks and recurrent neural networks can be used for speech recognition. It should be noted that, in addition to the above-mentioned speech recognition methods, other methods capable of speech recognition are also applicable, and this disclosure does not make specific limitations.

[0052] The target wake-up word is used as a criterion for determining whether the terminal device is woken up. If the target wake-up word is recognized in the first voice, the terminal device is woken up; if the target wake-up word is not recognized in the first voice, the terminal device cannot be woken up and there is no need to enter the voice interaction operation.

[0053] The target wake-up word can include, but is not limited to, words or phrases, and is pre-configured in the terminal device. For example, the target wake-up word can be set to "Hello, AAA", where AAA can be the terminal device model, manufacturer, etc., and this disclosure does not make specific limitations on it.

[0054] The first voiceprint feature is the voice feature contained in the first speech that can characterize and identify the user who made the first speech.

[0055] In one embodiment, when the target wake word is detected in the first speech, a preset voiceprint recognition algorithm can be used to perform voiceprint recognition on the first speech. The preset voiceprint recognition algorithm includes at least one of template matching algorithm, nearest neighbor algorithm, neural network algorithm, deep learning model, VQ clustering algorithm, D-vector, and I-vector. It should be noted that this disclosure does not impose any special limitations on the method of voiceprint recognition.

[0056] It should be noted that after acquiring the first speech, the first speech is cached while performing speech recognition. When the first speech includes the target wake word, voiceprint recognition is performed on the cached first speech. Once the first voiceprint feature is obtained, it can be cached for later use. When the target wake word is not recognized in the first speech, there is no need to extract features from the first speech; simply delete the cached first speech. This reduces the frequency of speech recognition and the computational workload.

[0057] S204. Acquire the second speech, perform voiceprint recognition on the second speech, and obtain the second voiceprint feature of the second speech.

[0058] In one embodiment, when a target wake-up word is recognized in the first voice message, the terminal device is awakened, allowing for human-computer interaction and the acquisition of the second voice message. The second voice message is the voice signal acquired after the terminal device is awakened, serving as the voice command to be executed.

[0059] The acquisition and speech recognition methods for the second speech are similar to those for the first speech, and will not be repeated here.

[0060] The second voice is the voice collected after waking up the terminal device. The second voice can be the voice issued by the person waking up the device, the voice issued by a user other than the person waking up the device, or the voice issued by both the person waking up the device and the user who is not waking up the device at the same time.

[0061] The second voiceprint feature is used to distinguish the user who produces the second voice. The voiceprint recognition method for the second voice is the same as that for the first voice, and will not be described again here.

[0062] S206. Based on the comparison results of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, determine the processing strategy for the second speech.

[0063] In one embodiment, the comparison result of the first voiceprint feature and the second voiceprint feature may include the first voiceprint feature being the same as the second voiceprint feature, the second voiceprint feature being completely different from the first voiceprint feature, or the first voiceprint feature and the second voiceprint feature being partially the same.

[0064] For example, when the similarity between the first voiceprint feature and the second voiceprint feature exceeds a preset first voiceprint threshold, the first voiceprint feature and the second voiceprint feature are determined to be the same; when the similarity between the first voiceprint feature and the second voiceprint feature is less than a preset second voiceprint threshold, the first voiceprint feature and the second voiceprint feature are determined to be completely different; when the similarity between the first voiceprint feature and the second voiceprint feature is between the preset second voiceprint threshold and the preset first voiceprint threshold, the first voiceprint feature and the second voiceprint feature are determined to be partially the same.

[0065] In one embodiment, the similarity between the first voiceprint feature and the second voiceprint feature can be obtained using cosine similarity, or it can be determined by other methods of calculating similarity, and this disclosure does not limit this.

[0066] In one embodiment, the validity of the second speech can be determined by performing speech recognition on the second speech. If no speech command is recognized from the second speech, and the second speech is a meaningless word, noise, or silence, such as "ah," "ee," or "ya," then the validity of the second speech is determined to be invalid speech. If a speech command is recognized from the second speech, and the second speech can clearly convey information without being interfered with by noise, then the validity of the second speech is determined to be valid speech. Valid speech can recognize specific speech commands, such as "open a music player" or "open navigation."

[0067] The processing strategies for the second speech include, but are not limited to, responding to the second speech, not responding to the second speech, and discarding the second speech.

[0068] S208. Execute the processing strategy to perform voice interaction.

[0069] In one embodiment, the execution processing strategy includes executing the voice command corresponding to the valid voice, and also includes not responding to invalid voice, discarding invalid voice, etc.

[0070] In some embodiments, the execution processing strategy may also include issuing a prompt message to re-encode the voice, waiting for the voice to be re-encoded and acquiring the re-encoded voice, etc.

[0071] In this embodiment of the disclosure, in response to obtaining the target wake-up word of the first speech, voiceprint recognition is performed on the first speech to obtain the first voiceprint feature of the first speech; the second speech is obtained, and voiceprint recognition is performed on the second speech to obtain the second voiceprint feature of the second speech; based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, a processing strategy for the second speech is determined; the processing strategy is executed to perform voice interaction. By comparing the voiceprint feature of the wake-up stage with the voiceprint feature of the person issuing the voice command, the problem of false response can be effectively avoided, and the accuracy and efficiency of voice interaction can be improved.

[0072] Figure 3 A flowchart of another voice interaction method provided in an embodiment of this disclosure is shown. Figure 2 Based on the embodiment, S206 is further defined as S2062 to limit the scenario of a processing strategy. For example... Figure 3 As shown, in one embodiment, a voice interaction method includes S202-S204, S2062, and S208. Specifically, S206 determines a processing strategy for the second voice based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second voice, including:

[0073] S2062. If the similarity between the first voiceprint feature and the second voiceprint feature is greater than the preset first voiceprint threshold, and the second speech is valid, then the processing strategy for the second speech is determined to be responding to the second speech.

[0074] It should be noted that the specific implementation methods of S202 to S204 and S208 in this embodiment are the same as those of S202 to S204 and S208 in the previous embodiment, and will not be repeated here.

[0075] In one embodiment, when the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset first voiceprint threshold, it indicates that there is a high probability that the user who made the second voice is the wake-up person. At this time, it is determined that the user who made the second voice is the wake-up person, and speech recognition is performed on the second voice. If the validity of the second voice is valid, the voice command recognized in the second voice is responded to and the operation corresponding to the voice command is executed.

[0076] The preset first voiceprint threshold can be pre-configured in the terminal device. The preset first voiceprint threshold can be determined according to actual needs. For example, the preset first voiceprint threshold can be 90%, 95%, etc. This disclosure does not make any specific limitation in this regard.

[0077] In this embodiment of the disclosure, when the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset first voiceprint threshold and the second voice is valid, it indicates that the person who made the first voice and the person who made the second voice are the same. By responding to the second voice, voice interaction can be realized, the effectiveness of human-computer interaction can be improved, and the influence of environmental factors, interference and other factors can be avoided.

[0078] Figure 4 This diagram illustrates a flowchart of yet another voice interaction method provided in an embodiment of the present disclosure. Figure 2 Based on the embodiment, S206 is further defined as S2064 to limit the case of another processing strategy. For example... Figure 4 As shown, in one embodiment, a voice interaction method includes steps S202-S204, S2064, and S208. Specifically, step S206 determines a processing strategy for the second voice based on the comparison result of the first and second voiceprint features and the validity of the second voice, including:

[0079] S2064. If the first voiceprint feature and the second voiceprint feature are less than the preset second voiceprint threshold, then the processing strategy for the second speech is determined to be not to respond to the second speech, or not to perform speech recognition on the second speech, or to discard the second speech.

[0080] The aforementioned preset second voiceprint threshold is less than the preset first voiceprint threshold. The preset second voiceprint threshold can be pre-configured in the terminal device. The value of the preset second voiceprint threshold can be determined according to actual needs. For example, the preset second voiceprint threshold can be 50%, 60%, etc. This disclosure does not make any specific limitation in this regard.

[0081] In one embodiment, when the first voiceprint feature and the second voiceprint feature are less than a preset second voiceprint threshold, it indicates that the speaker of the second voice is different from the waker of the first voice. In this case, the second voice can be ignored, speech recognition can be skipped, and the second voice can be discarded directly, thereby effectively reducing the computational intensity of the terminal device.

[0082] Figure 5 A flowchart illustrating another voice interaction method provided in an embodiment of this disclosure is shown. Figure 5 As shown, in one embodiment, S206 above determines a processing strategy for the second speech based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, including:

[0083] S502. In response to the recognition that the second voiceprint feature includes multiple sub-voiceprint features, calculate the similarity between the first voiceprint feature and each sub-voiceprint feature respectively.

[0084] S504. Determine the target sub-voiceprint feature whose similarity to the first voiceprint feature is greater than the preset first voiceprint threshold, and extract the speech segment corresponding to the target sub-voiceprint feature from the second speech.

[0085] S506. Perform speech recognition on the speech segment to obtain the recognition result of the speech segment;

[0086] S508. Based on the recognition results of the speech segments, determine the processing strategy for the second speech.

[0087] In one embodiment, when the similarity between the first voiceprint feature and the second voiceprint feature is between a preset second voiceprint threshold and a preset second voiceprint threshold, it indicates that it is impossible to directly determine whether the second voiceprint feature was emitted by the person waking up the user. In this case, the second voiceprint feature can be further processed to determine the multiple sub-voiceprint features included in the second voiceprint feature.

[0088] It should be noted that when there is a sub-voiceprint feature among multiple sub-voiceprint features whose similarity to the first voiceprint feature is greater than the preset first voiceprint threshold, the sub-voiceprint feature is identified as the target sub-voiceprint feature. By determining the start and end times of the target sub-voiceprint feature, the speech segment corresponding to the target sub-voiceprint feature can be obtained.

[0089] In one embodiment, after determining the speech segment corresponding to the target sub-voiceprint features, speech recognition is performed only on the speech segment to obtain the recognition result of the speech segment. The recognition result of the speech segment is similar to the recognition result of the second speech in the aforementioned embodiment, and also includes valid speech, invalid speech, etc., and the similarities will not be repeated.

[0090] If none of the multiple sub-voiceprint features has a similarity greater than the preset first voiceprint threshold with the first voiceprint feature, the second voiceprint is not responded to and can be discarded. There is no need to recognize the second voiceprint, thus reducing the computational intensity of the terminal device.

[0091] In this embodiment of the disclosure, by extracting the speech segments in the second voiceprint features that have a similarity greater than that of the first voiceprint features than a preset first voiceprint threshold, and based on the recognition results of the speech segments, a processing strategy for the second speech is determined, thereby effectively improving the effectiveness of voice interaction.

[0092] In one embodiment, S508 determines a processing strategy for the second speech based on the recognition result of the speech segment, including: if the recognition result of the speech segment is valid speech, then the processing strategy for the second speech is to respond to the speech segment; if the recognition result of the speech segment is invalid speech, then the processing strategy for the second speech is to not respond to the second speech or discard the second speech.

[0093] When the recognition result of a speech segment is valid, the system responds to the speech segment to execute the voice command that wakes the user, thus achieving voice interaction. When the recognition result of a speech segment is invalid, the system does not respond to the speech segment, does not perform speech recognition on the speech segment, and discards the second speech, effectively reducing the computational burden on the terminal device.

[0094] In one embodiment, the voice interaction method provided in this disclosure further includes: waiting for or prompting the user to input the next voice segment in response to the processing strategy for the second voice segment being not to respond to the second voice segment; and exiting the voice interaction if the next voice segment is not obtained within a preset time period.

[0095] When the system does not respond to the second voice command, it waits for the user to input the next voice command, and then performs voiceprint and speech recognition on the next voice command to achieve voice interaction. Alternatively, it can prompt the user to re-enter the next voice command via voice and / or a dialog box to achieve voice interaction. For example, when the first and second voice commands are not issued by the same person, the following prompt can be issued: "The person who woke up the user and the person who issued the command are not the same person." When the second voice command is invalid, the following prompt can be issued: "Unable to recognize voice command, please re-enter."

[0096] The aforementioned preset time period can be pre-configured in the terminal device. The value of the preset time period can be determined according to actual needs, such as 20 seconds, 30 seconds, etc. Alternatively, the voice interaction can exit after repeatedly prompting the user to input the next voice segment without capturing any speech.

[0097] In this embodiment of the disclosure, by prompting or waiting for the user to input voice, and by exiting the voice interaction if no next voice segment is obtained within a preset time period, the user experience can be improved.

[0098] Figure 6 This diagram illustrates an example flowchart of a voice interaction method provided in an embodiment of this disclosure. Figure 6 As shown, the voice interaction method provided in this embodiment includes the following steps:

[0099] S601. When the first voice input terminal device is used, the first voice is recognized.

[0100] S602. Determine that the first voice message includes the target wake-up word, and wake up the terminal device;

[0101] S603, Cache the target wake-up word and the first speech;

[0102] S604. After waking up the terminal device, the voiceprint judgment is triggered, and the voiceprint is recognized on the first voice.

[0103] S605, Obtain the first voiceprint characteristics of the person being awakened;

[0104] S606. After the target wake-up word is recognized in the first speech, the second speech is acquired and speech recognition is performed on the second speech.

[0105] S607, second voice buffer;

[0106] S608. When performing speech recognition on the second speech, a voiceprint judgment is triggered, and voiceprint recognition is performed on the cached second speech.

[0107] S609, Obtain the second voiceprint feature;

[0108] S610. Compare the first voiceprint feature and the second voiceprint feature. If the voiceprint features are consistent, proceed to S611; if the voiceprint features are inconsistent, proceed to S613.

[0109] S611. Determine whether the second speech is valid. If yes, execute S612; otherwise, execute S613.

[0110] S612, responding to the second voice;

[0111] S613: No need to respond to the second voice segment, wait for the next voice segment, return to S606, and re-perform voice recognition and voiceprint feature judgment on the next voice segment.

[0112] This disclosure proposes a method to compare the voiceprint features during voice interaction with the voiceprint features during wake-up, thereby ensuring that the person waking up the terminal device is consistent with the person issuing the voice command, avoiding false responses, improving voice interaction efficiency, and providing a better user experience.

[0113] Based on the same inventive concept, this disclosure also provides a voice interaction device, as shown in the following embodiments. Since the principle by which this device embodiment solves the problem is similar to that of the above-described method embodiments, the implementation of this device embodiment can refer to the implementation of the above-described method embodiments, and repeated details will not be elaborated further.

[0114] Figure 7 A schematic diagram of a voice interaction device provided in an embodiment of this disclosure is shown. Figure 7 As shown, the voice interaction device in this embodiment includes a first voiceprint recognition module 710, a second voiceprint recognition module 720, a processing strategy determination module 730, and a voice interaction execution module 740.

[0115] The first voiceprint recognition module 710 is used to perform voiceprint recognition on the first speech in response to obtaining the target wake-up word of the first speech, and obtain the first voiceprint feature of the first speech.

[0116] The second voiceprint recognition module 720 is used to acquire the second speech, perform voiceprint recognition on the second speech, and obtain the second voiceprint features of the second speech.

[0117] The processing strategy determination module 730 is used to determine the processing strategy for the second speech based on the comparison results of the first voiceprint feature and the second voiceprint feature, as well as the validity of the second speech.

[0118] The voice interaction execution module 740 is used to execute processing strategies for voice interaction.

[0119] In one embodiment, the processing strategy determination module 730 is configured to determine the processing strategy for the second speech as responding to the second speech if the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset first voiceprint threshold and the second speech is valid.

[0120] In one embodiment, the processing strategy determination module 730 is configured to determine, if the first voiceprint feature and the second voiceprint feature are less than a preset second voiceprint threshold, the processing strategy for the second speech is to not respond to the second speech, or not perform speech recognition on the second speech, or discard the second speech.

[0121] In one embodiment, the processing strategy determination module 730 is configured to, in response to the recognition that the second voiceprint feature includes multiple sub-voiceprint features, calculate the similarity between the first voiceprint feature and each sub-voiceprint feature respectively; determine the target sub-voiceprint feature whose similarity to the first voiceprint feature is greater than a preset first voiceprint threshold; extract the speech segment corresponding to the target sub-voiceprint feature from the second speech; perform speech recognition on the speech segment to obtain the recognition result of the speech segment; and determine the processing strategy for the second speech based on the recognition result of the speech segment.

[0122] In one embodiment, the processing strategy determination module 730 is further configured to determine the processing strategy for the second speech as responding to the speech segment if the recognition result of the speech segment is valid speech; and to determine the processing strategy for the second speech as not responding to the second speech or discarding the second speech if the recognition result of the speech segment is invalid speech.

[0123] In one embodiment, the device further includes a semantic recognition module not shown in the figures, which is used to determine the validity of the second speech as invalid speech if no voice command of the second speech is recognized, and to determine the validity of the second speech as valid speech if a voice command of the second speech is recognized.

[0124] In one embodiment, the device further includes a prompting module not shown in the figures, which is used to wait for or prompt the user to input the next voice segment in response to the processing strategy of not responding to the second voice segment; if the next voice segment is not obtained within a preset time period, the voice interaction is exited.

[0125] In this embodiment of the disclosure, in response to the first speech being recognized as a target wake-up word, voiceprint recognition is performed on the first speech to obtain a first voiceprint feature of the first speech, and a second speech is obtained; voiceprint recognition is performed on the second speech to obtain a second voiceprint feature of the second speech; based on the comparison result of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, a processing strategy for the second speech is determined; the processing strategy is executed to perform voice interaction. By comparing the voiceprint features of the wake-up stage with the voiceprint features of the person issuing the voice command, the problem of false responses can be effectively avoided, and the accuracy and efficiency of voice interaction can be improved.

[0126] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuits,” “modules,” or “systems.”

[0127] The following reference Figure 8 To describe an electronic device 800 according to this embodiment of the present invention. Figure 8 The electronic device 800 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0128] like Figure 8 As shown, the electronic device 800 is manifested in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one processing unit 810, at least one storage unit 820, and a bus 830 connecting different system components (including storage unit 820 and processing unit 810).

[0129] The storage unit stores program code, which can be executed by the processing unit 810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 810 can perform actions such as... Figure 2 The process shown in the diagram responds to the first speech being recognized as a target wake word, performs voiceprint recognition on the first speech to obtain the first voiceprint feature of the first speech, and then acquires the second speech; performs voiceprint recognition on the second speech to obtain the second voiceprint feature of the second speech; determines a processing strategy for the second speech based on the comparison result of the first and second voiceprint features and the validity of the second speech; and executes the processing strategy to perform voice interaction.

[0130] Storage unit 820 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 8201 and / or cache memory 8202, and may further include a read-only memory (ROM) 8203.

[0131] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0132] Bus 830 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0133] Electronic device 800 can also communicate with one or more external devices 840 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with the system, and / or with any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed through input / output (I / O) interface 850. Furthermore, the system can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 860. Figure 8 As shown, network adapter 860 communicates with other modules of electronic device 800 via bus 830. It should be understood that, although... Figure 8 As not shown in the diagram, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0134] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0135] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, which may be a readable signal medium or a readable storage medium. In exemplary embodiments of this disclosure, a computer program product is also provided, comprising a computer program or computer instructions, which are loaded and executed by a processor to cause a computer to implement the steps of the methods disclosed in the above embodiments.

[0136] More specific examples of computer-readable storage media in this disclosure may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0137] In this disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device.

[0138] Optionally, the program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0139] In practical implementation, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0140] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0141] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0142] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0143] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

Claims

1. A voice interaction method, characterized in that, include: In response to obtaining the target wake-up word of the first speech, voiceprint recognition is performed on the first speech to obtain the first voiceprint feature of the first speech; Acquire the second speech, perform voiceprint recognition on the second speech, and obtain the second voiceprint feature of the second speech; Based on the comparison results of the first voiceprint feature and the second voiceprint feature, and the validity of the second speech, a processing strategy for the second speech is determined. The processing strategy is executed to perform voice interaction; The step of determining the processing strategy for the second speech based on the comparison results of the first and second voiceprint features and the validity of the second speech includes: If the similarity between the first voiceprint feature and the second voiceprint feature is between a preset first voiceprint threshold and a preset second voiceprint threshold, then the second voiceprint feature is processed, wherein the preset second voiceprint threshold is less than the preset first voiceprint threshold; in response to recognizing that the second voiceprint feature includes multiple sub-voiceprint features, the similarity between the first voiceprint feature and each sub-voiceprint feature is calculated respectively; a target sub-voiceprint feature with a similarity greater than the preset first voiceprint threshold is determined, and a speech segment corresponding to the target sub-voiceprint feature is extracted from the second speech; speech recognition is performed on the speech segment to obtain the recognition result of the speech segment; based on the recognition result of the speech segment, a processing strategy for the second speech is determined.

2. The voice interaction method according to claim 1, characterized in that, The step of determining a processing strategy for the second speech based on the comparison results of the first and second voiceprint features and the validity of the second speech includes: If the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset first voiceprint threshold, and the second speech is valid, then the processing strategy for the second speech is determined to be responding to the second speech.

3. The voice interaction method according to claim 1, characterized in that, The step of determining a processing strategy for the second speech based on the comparison results of the first and second voiceprint features and the validity of the second speech includes: If the first voiceprint feature and the second voiceprint feature are less than a preset second voiceprint threshold, then the processing strategy for the second speech is determined to be not to respond to the second speech, or not to perform speech recognition on the second speech, or to discard the second speech.

4. The voice interaction method according to claim 1, characterized in that, The step of determining the processing strategy for the second speech based on the recognition result of the speech segment includes: If the recognition result of the speech segment is valid speech, then the processing strategy for the second speech is determined to be responding to the speech segment; If the recognition result of the speech segment is invalid speech, then the processing strategy for the second speech is determined to be either not responding to the second speech or discarding the second speech.

5. The voice interaction method according to claim 1, characterized in that, The method further includes: If no voice command is recognized from the second voice, the validity of the second voice is determined to be invalid. If the voice command of the second voice is recognized, the validity of the second voice is determined to be valid.

6. The voice interaction method according to any one of claims 1-5, characterized in that, The method further includes: The processing strategy for the second voice is to not respond to the second voice and wait for or prompt the user to input the next voice segment; If the next segment of voice is not obtained within the preset time period, the voice interaction will be terminated.

7. A voice interaction device, characterized in that, include: The first voiceprint recognition module is used to perform voiceprint recognition on the first speech in response to obtaining the target wake word of the first speech, and obtain the first voiceprint feature of the first speech. The second voiceprint recognition module is used to acquire the second speech, perform voiceprint recognition on the second speech, and obtain the second voiceprint features of the second speech. The processing strategy determination module is used to determine the processing strategy for the second speech based on the comparison results of the first voiceprint feature and the second voiceprint feature, as well as the validity of the second speech. A voice interaction execution module is used to execute the processing strategy to perform voice interaction; The processing strategy determination module is further configured to process the second voiceprint feature if the similarity between the first voiceprint feature and the second voiceprint feature is between a preset first voiceprint threshold and a preset second voiceprint threshold, wherein the preset second voiceprint threshold is less than the preset first voiceprint threshold. In response to the recognition that the second voiceprint feature includes multiple sub-voiceprint features, the similarity between the first voiceprint feature and each sub-voiceprint feature is calculated respectively; a target sub-voiceprint feature with a similarity greater than the preset first voiceprint threshold is determined, and a speech segment corresponding to the target sub-voiceprint feature is extracted from the second speech; speech recognition is performed on the speech segment to obtain the recognition result of the speech segment; based on the recognition result of the speech segment, a processing strategy for the second speech is determined.

8. An electronic device, characterized in that, include: processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the voice interaction method as described in any one of claims 1-6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the voice interaction method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, which are loaded and executed by a processor to enable the computer to implement the voice interaction method as described in any one of claims 1-6.