Voice wake-up method and related apparatus, electronic device, and storage medium

By setting dual wake-up thresholds and voiceprint feature verification, the accuracy and response speed of voice wake-up are improved, solving the problem of insufficient speed and accuracy of voice wake-up in existing technologies.

CN115798468BActive Publication Date: 2026-05-26IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-10-14
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

How to improve the response speed and accuracy of voice wake-up, especially in the fields of smart homes, mobile devices and vehicles.

Method used

By detecting the wake-up confidence of the user's voice and setting dual wake-up thresholds (first wake-up threshold and second wake-up threshold), voice interaction is directly enabled when the wake-up confidence is not less than the second wake-up threshold. When the wake-up confidence is not less than the first wake-up threshold but less than the second wake-up threshold, further verification is performed based on voiceprint features to determine whether to enable voice interaction.

Benefits of technology

It effectively reduces the probability of false wake-up via voice and improves the accuracy and response speed of wake-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798468B_ABST
    Figure CN115798468B_ABST
Patent Text Reader

Abstract

This application discloses a voice wake-up method and related devices, electronic devices, and storage media. The voice wake-up method includes: detecting the wake-up confidence level of a user's voice, and sequentially analyzing the relationship between the wake-up confidence level and a first wake-up threshold and a second wake-up threshold, wherein the first wake-up threshold is less than the second wake-up threshold; initiating voice interaction in response to the wake-up confidence level being not less than the second wake-up threshold; and determining whether to initiate voice interaction based on a first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system, in response to the wake-up confidence level being not less than the first wake-up threshold and less than the second wake-up threshold. This solution can improve wake-up response speed and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a voice wake-up method and related devices, electronic devices, and storage media. Background Technology

[0002] With the rapid development of artificial intelligence technology, intelligent voice technology has become widespread, and voiceprint-based wake-up solutions are widely used in various voice products such as smart homes, mobile devices, and in-vehicle systems.

[0003] Currently, with the widespread adoption of intelligent voice technology, the response speed and accuracy of voice wake-up are becoming increasingly important. Therefore, improving both wake-up response speed and accuracy has become a pressing issue. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a voice wake-up method and related devices, electronic devices, and storage media that can improve wake-up response speed and accuracy.

[0005] To address the aforementioned technical problems, the first aspect of this application provides a voice wake-up method, comprising: detecting the wake-up confidence of a user's voice, and sequentially analyzing the relationship between the wake-up confidence and a first wake-up threshold and a second wake-up threshold, wherein the first wake-up threshold is less than the second wake-up threshold; in response to the wake-up confidence being not less than the second wake-up threshold, initiating voice interaction; and in response to the wake-up confidence being not less than the first wake-up threshold and less than the second wake-up threshold, determining whether to initiate voice interaction based on a first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system.

[0006] To address the aforementioned technical problems, a second aspect of this application provides a voice wake-up device, comprising: a confidence detection module, a numerical analysis module, a first response module, and a second response module; wherein the confidence detection module is used to detect the wake-up confidence level of the user's voice; the numerical analysis module is used to sequentially analyze the relationship between the wake-up confidence level and a first wake-up threshold and a second wake-up threshold, respectively; wherein the first wake-up threshold is less than the second wake-up threshold; the first response module is used to initiate voice interaction in response to a wake-up confidence level not being less than the second wake-up threshold; the second response module is used to determine whether to initiate voice interaction based on a first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system in response to a wake-up confidence level not being less than the first wake-up threshold and less than the second wake-up threshold.

[0007] To address the aforementioned technical problems, a third aspect of this application provides an electronic device, including a memory and a processor coupled to each other. The memory stores program instructions, and the processor executes the program instructions to implement the voice wake-up method of the first aspect.

[0008] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium storing program instructions executable by a processor, the program instructions being used to implement the voice wake-up method of the first aspect described above.

[0009] The above scheme detects the wake-up confidence of the user's voice and analyzes the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold, respectively, where the first wake-up threshold is less than the second wake-up threshold. In response to a wake-up confidence not being less than the second wake-up threshold, voice interaction is initiated. In response to a wake-up confidence not being less than the first wake-up threshold and less than the second wake-up threshold, based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system, it is determined whether to initiate voice interaction. On the one hand, setting dual wake-up thresholds, by using differentiated wake-up thresholds to determine whether to initiate voice interaction, helps reduce the probability of false wake-up. On the other hand, determining whether to initiate voice interaction based on the first and second voiceprint features when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold helps improve wake-up accuracy. Furthermore, during the voice wake-up process, by first judging the wake-up threshold and then, depending on the situation, detecting the voiceprint to determine whether to initiate voice interaction, the wake-up response speed and wake-up accuracy can be improved.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0012] Figure 1 This is a flowchart illustrating an embodiment of the voice wake-up method of this application;

[0013] Figure 2 This is a flowchart illustrating another embodiment of the voice wake-up method of this application;

[0014] Figure 3 This is a schematic diagram of the framework of an embodiment of the voice wake-up device of this application;

[0015] Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of this application;

[0016] Figure 5 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0017] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0018] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0019] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. "Several" means at least one. The terms "first," "second," etc., in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0020] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the voice wake-up method of this application.

[0021] Specifically, this may include the following steps:

[0022] Step S11: Detect the wake-up confidence of the user's voice, and analyze the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold respectively.

[0023] In this embodiment, the first wake-up threshold is less than the second wake-up threshold. It should be noted that the user may be a target user (i.e., a frequent user) or a non-target user (i.e., an infrequent user). Target users have pre-recorded voiceprint features in the voice interaction system, while non-target users have not. Therefore, to distinguish between target and non-target users, and thus respond quickly to target users and cautiously to non-target users, a smaller wake-up threshold (the first wake-up threshold) can be set for target users, and a larger wake-up threshold (the second wake-up threshold) can be set for non-target users. Unlike the aforementioned method, the first and second wake-up thresholds can also be set with initial values, which are continuously adjusted during the interaction process to stabilize them. For example, the first wake-up threshold can be set to 0.5, 0.6, etc., and the second wake-up threshold can be set to 0.8, 0.9, etc. By continuously adjusting the first wake-up threshold, the final first wake-up threshold is 0.7, and the final second wake-up threshold is 0.9. It is understood that the above method is only one possible case of wake-up threshold in actual application, and should not be used to limit the setting method and value used in actual application. The first wake-up threshold and the second wake-up threshold can be set according to the actual situation, and no specific limitation is made here.

[0024] In one implementation scenario, wake-up confidence can represent the degree of credibility of a user's voice containing a wake-up word, which refers to the word that switches the product from standby to working state. Understandably, the product can pre-set a nickname as the wake-up word; this nickname can be a reduplicated word (e.g., "Little X Little X") or an abbreviation (e.g., "Little X"), without limitation here.

[0025] In a specific implementation scenario, wake-up confidence can be determined by extracting acoustic parameters such as fundamental frequency, pitch, and harmonics from the user's speech, and then further detecting these extracted acoustic parameters. Alternatively, sample speech can be acquired first, and a network model can be trained using this sample speech to predict the wake-up confidence. The network model can include, but is not limited to, CNN (convolutional neural network) and RNN (recurrent neural network). The method for detecting wake-up confidence can be determined based on the actual situation and is not specifically limited here.

[0026] Furthermore, after detecting the wake-up confidence of the user's voice, the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold is analyzed in turn. Based on the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold, it is determined whether to enable voice interaction.

[0027] Step S12: In response to the wake-up confidence being no less than the second wake-up threshold, start voice interaction.

[0028] In one implementation scenario, as mentioned earlier, after obtaining the wake-up confidence score, it is necessary to analyze the relationship between the wake-up confidence score and the first wake-up threshold and the second wake-up threshold, respectively. Since the first wake-up threshold is less than the second wake-up threshold, when the wake-up confidence score is not less than the second wake-up threshold, it indicates that the wake-up confidence score is also not less than the first wake-up confidence score. For example, as mentioned earlier, in a car-riding scenario, passengers can generally be divided into frequent passengers and infrequent passengers. A first wake-up threshold is set for frequent passengers, and a second wake-up threshold is set for infrequent passengers. Since the first wake-up threshold is less than the second wake-up threshold, when the wake-up confidence score is not less than the second wake-up threshold, voice interaction can be directly initiated.

[0029] Step S13: In response to a wake-up confidence level that is not less than a first wake-up threshold and less than a second wake-up threshold, determine whether to enable voice interaction based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system.

[0030] In one implementation scenario, a voice interaction system can store the second voiceprint features of several users. These second voiceprint features can be acquired and stored continuously during the voice interaction process, or they can be acquired and stored in advance. The method of acquiring these second voiceprint features in the voice interaction system can be determined according to the actual situation, and no specific limitation is made here.

[0031] In one implementation scenario, when the wake-up confidence is not less than a first wake-up threshold and less than a second wake-up threshold, a similarity comparison can be further performed based on the first and second voiceprint features to determine whether to enable voice interaction. Specifically, the feature similarity between the first voiceprint feature and several second voiceprint features can be obtained first, that is, the feature similarity between the first voiceprint feature and each of the second voiceprint features can be calculated separately. For example, the feature similarity between the first and second voiceprint features can be calculated, or the feature distance between the first and second voiceprint features can be calculated, and then the feature similarity can be determined through the feature distance. The method of determining the feature similarity between the first and second voiceprint features can be selected according to the actual situation and is not specifically limited here. After obtaining the feature similarity between the first and second voiceprint features, voice interaction can be enabled if at least one feature similarity is greater than the voiceprint threshold; if none of the feature similarities are greater than the voiceprint threshold, voice interaction cannot be enabled. Alternatively, after determining the feature similarity between the first and second voiceprint features, the current feature similarity can be compared with a voiceprint threshold. If the current feature similarity is greater than the voiceprint threshold, voice interaction is initiated. That is, for each feature similarity obtained, it is first determined whether the feature similarity is greater than the voiceprint threshold. If the feature similarity is greater than the voiceprint threshold, voice interaction is initiated. After initiating voice interaction, it is unnecessary to obtain the feature similarity between the first voiceprint feature and the remaining second voiceprint features. If the feature similarity is not greater than the voiceprint threshold, feature similarity is continuously obtained and compared. The voiceprint threshold can be the minimum value required to initiate voice interaction, or any value between the minimum and maximum. The voiceprint threshold can be determined based on the actual situation and is not specifically limited here. In the above method, when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, by obtaining the feature similarity between the first voiceprint feature and several second voiceprint features and comparing the feature similarity with the voiceprint threshold to determine whether to initiate voice interaction, it helps reduce the false wake-up rate and further improve wake-up accuracy.

[0032] In one implementation scenario, voice interaction is not initiated in response to a wake-up confidence level that is less than the first wake-up threshold.

[0033] It should be noted that during the voice wake-up process, the step of determining whether to enable voice interaction based on the relationship between wake-up confidence and the first wake-up threshold and the second wake-up threshold is executed based on objective circumstances and does not restrict the execution order.

[0034] In one implementation scenario, after determining whether to initiate voice interaction based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system, it is possible to further determine whether the first wake-up threshold and the second wake-up threshold need to be adjusted. If adjustment is required, the first or second wake-up threshold is adjusted. Specifically, to determine whether the second wake-up threshold needs adjustment, a first ratio between a first value and a second value can be obtained. The first value represents the number of times voice interaction is initiated when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold; the second value represents the number of times the user attempts voice wake-up, including both successful and unsuccessful attempts. Then, based on the relationship between the first ratio and the first proportional threshold, it is determined whether to adjust the second wake-up threshold. The first ratio can be less than or equal to the first proportional threshold. The first proportional threshold can be set to 0.5, 0.6, etc., and can be determined according to the actual situation; no specific limitation is made here. Understandably, if the first ratio is less than the first ratio threshold, it indicates that voiceprint detection reduces the false wake-up rate while improving wake-up accuracy. However, if the first ratio is not less than the first ratio threshold, i.e., the first ratio is too large, it suggests that there may be some objective reasons causing a large number of situations where voiceprint detection is performed instead of direct voice interaction due to external environmental factors. For example, the microphone may be too far away, the user's voice may be too soft, or the user may be speaking too quickly. In this case, the second wake-up threshold can be fine-tuned to improve wake-up accuracy and wake-up rate.

[0035] Furthermore, in adjusting the second wake-up threshold, the first set of wake-up confidence scores can be clustered to obtain several first cluster sets. During this process, clustering methods can be used to cluster the wake-up confidence scores. These methods can include, but are not limited to, K-Means clustering, mean-shift clustering, density-based clustering, expectation-maximum (EM) clustering using a Gaussian mixture model (GMM), agglomerative hierarchical clustering, graph community detection, etc. Then, the first cluster set with the largest cluster center is selected as the first target set. Based on this first target set, the second wake-up threshold is adjusted. Specifically, the cluster centers of the first target set can be adjusted to the new second wake-up threshold, or the largest wake-up confidence score in the second target set can be adjusted to the second wake-up threshold. The method of adjusting the second wake-up threshold can be chosen according to the actual situation and is not specifically limited here. This method, by adjusting the second wake-up threshold, helps to improve wake-up response speed and wake-up accuracy.

[0036] In another implementation scenario, after determining whether to enable voice interaction based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system, it can be further determined whether the first wake-up threshold needs adjustment. This can be achieved by first obtaining a second ratio between a third value and a second value. The third value represents the number of times voice interaction is not enabled when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold; that is, the number of times voice interaction attempts fail. The second value represents the number of times the user attempts voice wake-up, including both successful and unsuccessful attempts. Then, based on the relationship between the second ratio and the second proportional threshold, it is determined whether to adjust the first wake-up threshold. The second ratio can be less than or equal to the second proportional threshold. The second proportional threshold can be set to 0.5, 0.6, etc., and can be determined according to actual circumstances; no specific limitation is made here. Understandably, if the second ratio is less than the second proportional threshold, it indicates that voiceprint detection reduces the false wake-up rate while improving wake-up accuracy. However, if the second ratio is not less than the second proportional threshold, i.e., the second ratio is too large, it means that the first wake-up threshold may be set unreasonably and needs to be adjusted to improve wake-up accuracy and wake-up rate.

[0037] Furthermore, in adjusting the first wake-up threshold, the third number of wake-up confidence scores can be clustered to obtain several second cluster sets. The clustering method can refer to the method in the above-disclosed embodiments, which will not be repeated here. Then, the second cluster set with the smallest cluster center is selected as the second target set, and the first wake-up threshold is adjusted based on the second target set. Specifically, the cluster center of the second target set can be adjusted to the new first wake-up threshold, or the smallest wake-up confidence score in the first target set can be adjusted to the first wake-up threshold, etc. The method of adjusting the first wake-up threshold can be selected according to the actual situation, and is not specifically limited here. The above method, by adjusting the first wake-up threshold, helps to improve the wake-up response speed and wake-up accuracy.

[0038] The above scheme detects the wake-up confidence of the user's voice and analyzes the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold, respectively, where the first wake-up threshold is less than the second wake-up threshold. In response to a wake-up confidence not being less than the second wake-up threshold, voice interaction is initiated. In response to a wake-up confidence not being less than the first wake-up threshold and less than the second wake-up threshold, based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system, it is determined whether to initiate voice interaction. On the one hand, setting dual wake-up thresholds, by using differentiated wake-up thresholds to determine whether to initiate voice interaction, helps reduce the probability of false wake-up. On the other hand, determining whether to initiate voice interaction based on the first and second voiceprint features when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold helps improve wake-up accuracy. Furthermore, during the voice wake-up process, by first judging the wake-up threshold and then, depending on the situation, detecting the voiceprint to determine whether to initiate voice interaction, the wake-up response speed and wake-up accuracy can be improved.

[0039] Please see Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the voice wake-up method of this application.

[0040] Specifically, this may include the following steps:

[0041] Step S21: Detect the wake-up confidence of the user's voice.

[0042] Specifically, the detection method described in the aforementioned public embodiments can be referred to, and will not be repeated here.

[0043] Step S22: Determine whether the wake-up confidence is not less than the first wake-up threshold; if not, proceed to step S23; otherwise, proceed to step S24.

[0044] In one implementation scenario, after detecting the wake-up confidence level of the user's voice, the wake-up confidence level can be compared with a first wake-up threshold to determine whether to enable voice interaction.

[0045] Step S23: Do not enable voice interaction.

[0046] In one implementation scenario, when the wake-up confidence level is less than the first wake-up threshold, the system may not respond to user voice commands, and of course, it will not initiate voice interaction.

[0047] Step S24: Determine whether the wake-up confidence is not less than the second wake-up threshold; if yes, proceed to step S25; otherwise, proceed to step S26.

[0048] In one implementation scenario, when the wake-up confidence is not less than the first wake-up threshold, it can be further determined whether the wake-up confidence is not less than the second wake-up threshold, thereby determining whether to enable voice interaction, which helps to reduce the false wake-up rate and improve wake-up accuracy.

[0049] Step S25: Enable voice interaction.

[0050] Step S26: Obtain the feature similarity between the first voiceprint feature and each of the second voiceprint features.

[0051] In one implementation scenario, the feature similarity between the first voiceprint feature and each of the second voiceprint features can be obtained sequentially, or only the feature similarity between the first voiceprint feature and the current second voiceprint feature can be obtained. The method of obtaining feature similarity can be determined according to the actual situation, and no specific limitation is made here.

[0052] Step S27: Determine whether the feature similarity is greater than the voiceprint threshold; if yes, proceed to step S28; otherwise, proceed to step S29.

[0053] In one implementation scenario, after obtaining the feature similarity between the first voiceprint feature and several second voiceprint features, it can be determined whether at least one feature similarity is greater than the voiceprint threshold. Alternatively, after obtaining the feature similarity, the feature similarity can be compared with the voiceprint threshold. The method for determining the relationship between feature similarity and voiceprint threshold can be chosen according to the actual situation, and no specific limitation is made here.

[0054] Step S28: Enable voice interaction.

[0055] In one implementation scenario, if after obtaining a feature similarity, the feature similarity is compared with the voiceprint threshold, and the feature similarity is greater than the voiceprint threshold, then there is no need to obtain the feature similarity between the first voiceprint feature and the other second voiceprint features, and voice interaction can be started directly.

[0056] Step S29: Do not enable voice interaction.

[0057] In one implementation scenario, if the feature similarity is not greater than the voiceprint threshold, then the voiceprint permissions do not match in the voiceprint detection, meaning that the user corresponding to the first voiceprint feature does not have the right to start voice interaction.

[0058] The above scheme detects the wake-up confidence of the user's voice and analyzes the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold, respectively, where the first wake-up threshold is less than the second wake-up threshold. In response to a wake-up confidence not being less than the second wake-up threshold, voice interaction is initiated. In response to a wake-up confidence not being less than the first wake-up threshold and less than the second wake-up threshold, based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system, it is determined whether to initiate voice interaction. On the one hand, setting dual wake-up thresholds, by using differentiated wake-up thresholds to determine whether to initiate voice interaction, helps reduce the probability of false wake-up. On the other hand, determining whether to initiate voice interaction based on the first and second voiceprint features when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold helps improve wake-up accuracy. Furthermore, during the voice wake-up process, by first judging the wake-up threshold and then, depending on the situation, detecting the voiceprint to determine whether to initiate voice interaction, the wake-up response speed and wake-up accuracy can be improved.

[0059] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0060] Please see Figure 3 , Figure 3 This is a schematic diagram of the framework of an embodiment of the voice wake-up device of this application. The voice wake-up device 30 includes a confidence detection module 31, a numerical analysis module 32, a first response module 33, and a second response module 34. The confidence detection module 31 is used to detect the wake-up confidence of the user's voice; the numerical analysis module 32 is used to sequentially analyze the relationship between the wake-up confidence and a first wake-up threshold and a second wake-up threshold, wherein the first wake-up threshold is less than the second wake-up threshold; the first response module 33 is used to initiate voice interaction in response to a wake-up confidence not being less than the second wake-up threshold; the second response module 34 is used to determine whether to initiate voice interaction in response to a wake-up confidence not being less than the first wake-up threshold and less than the second wake-up threshold, based on a first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system.

[0061] The above scheme, on the one hand, sets dual wake-up thresholds, and determines whether to enable voice interaction by using different wake-up thresholds, which helps to reduce the probability of false wake-up. On the other hand, when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, it determines whether to enable voice interaction based on the first voiceprint feature and the second voiceprint feature, which helps to improve wake-up accuracy. In addition, during the voice wake-up process, by first judging the wake-up threshold and then performing voiceprint detection as appropriate to determine whether to enable voice interaction, the wake-up response speed can be improved, and the wake-up accuracy can be improved at the same time.

[0062] In some disclosed embodiments, the voice wake-up device 30 includes a third response module, which is used to not initiate voice interaction in response to a wake-up confidence level being less than a first wake-up threshold.

[0063] In some disclosed embodiments, the second response module 34 includes an acquisition submodule, which is used to acquire the feature similarity between the first voiceprint feature and a plurality of second voiceprint features; the second response module 34 includes a first response submodule, which is used to enable voice interaction in response to the existence of at least one feature similarity greater than the voiceprint threshold; the second response module 34 further includes a second response submodule, which is used to not enable voice interaction in response to the existence of at least one feature similarity greater than the voiceprint threshold.

[0064] Therefore, when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, by obtaining the feature similarity between the first voiceprint feature and several second voiceprint features, and comparing the feature similarity with the voiceprint threshold, it is possible to determine whether to enable voice interaction, which helps to reduce the false wake-up rate and further improve wake-up accuracy.

[0065] In some disclosed embodiments, the voice wake-up device 30 includes a first acquisition module, which is used to acquire a first ratio between a first value and a second value; wherein, the first value represents the number of times voice interaction is initiated when the wake-up confidence is not less than a first wake-up threshold and less than a second wake-up threshold, and the second value represents the number of times the user attempts to wake up via voice; the voice wake-up device 30 also includes a first determination module, which is used to determine whether to adjust the second wake-up threshold based on the relationship between the first ratio and the first proportional threshold.

[0066] Therefore, by determining the relationship between the first ratio and the first proportional threshold, it is possible to decide whether to adjust the second wake-up threshold, thereby avoiding the situation where voiceprint detection is performed due to external environmental influences, thus reducing the false wake-up rate and improving wake-up accuracy.

[0067] In some disclosed embodiments, the first determining module includes a clustering submodule, which is used to cluster the first numerical wake-up confidence scores to obtain several first cluster sets; the first determining module includes a selection submodule, which is used to select the first cluster set with the largest cluster center as the first target set; the first determining module further includes an adjustment submodule, which is used to adjust the second wake-up threshold based on the first target set.

[0068] Therefore, adjusting the second wake-up threshold can help improve wake-up response speed and wake-up accuracy.

[0069] In some disclosed embodiments, the voice wake-up device 30 includes a second acquisition module, which is used to acquire a second ratio between a third value and a second value; wherein the third value represents the number of times the user determines that voice interaction will not be initiated when the wake-up confidence is not less than a first wake-up threshold and less than a second wake-up threshold, and the second value represents the number of times the user attempts to wake up via voice; the voice wake-up device 30 also includes a second determination module, which is used to determine whether to adjust the first wake-up threshold based on the relationship between the second ratio and the second proportional threshold.

[0070] Therefore, by determining the relationship between the second ratio and the second proportional threshold, it is possible to decide whether to adjust the first wake-up threshold, thereby avoiding the impact of an unreasonable first wake-up threshold setting on the wake-up response speed, and improving wake-up accuracy while reducing the false wake-up rate.

[0071] In some disclosed embodiments, the second determining module includes a clustering submodule, which is used to cluster the third numerical wake-up confidence scores to obtain several second cluster sets; the second determining module includes a selection submodule, which is used to select the second cluster set with the smallest cluster center as the second target set; the second determining module further includes an adjustment submodule, which is used to adjust the first wake-up threshold based on the second target set.

[0072] Therefore, adjusting the first wake-up threshold can help improve wake-up response speed and wake-up accuracy.

[0073] Please see Figure 4 , Figure 4This is a schematic diagram of a framework of an embodiment of the electronic device of this application. The electronic device 40 includes a memory 41 and a processor 42 coupled to each other. The memory 41 stores program instructions, and the processor 42 executes the program instructions to implement the steps in any of the above-described voice wake-up method embodiments. Specifically, the electronic device 40 may include, but is not limited to, desktop computers, laptops, servers, mobile phones, tablets, etc., and is not limited thereto. Furthermore, the electronic device may also include a microphone, which is coupled to the processor 42 and used to collect voice signals. There may be one or more microphones. Based on the microphone's directivity, the microphone may include, but is not limited to, omnidirectional microphones, bidirectional microphones, cardioid microphones, supercardioid microphones, shotgun microphones, etc.

[0074] Specifically, processor 42 controls itself and memory 41 to implement the steps in any of the above-described voice wake-up method embodiments. Processor 42 can also be referred to as a CPU (Central Processing Unit). Processor 42 may be an integrated circuit chip with signal processing capabilities. Processor 42 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 42 can be implemented using integrated circuit chips.

[0075] The above scheme, on the one hand, sets dual wake-up thresholds, and determines whether to enable voice interaction by using different wake-up thresholds, which helps to reduce the probability of false wake-up. On the other hand, when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, it determines whether to enable voice interaction based on the first voiceprint feature and the second voiceprint feature, which helps to improve wake-up accuracy. In addition, during the voice wake-up process, by first judging the wake-up threshold and then performing voiceprint detection as appropriate to determine whether to enable voice interaction, the wake-up response speed can be improved, and the wake-up accuracy can be improved at the same time.

[0076] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor. The program instructions 51 are used to implement the steps in any of the above-described voice wake-up method embodiments.

[0077] The above scheme, on the one hand, sets dual wake-up thresholds, and determines whether to enable voice interaction by using different wake-up thresholds, which helps to reduce the probability of false wake-up. On the other hand, when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, it determines whether to enable voice interaction based on the first voiceprint feature and the second voiceprint feature, which helps to improve wake-up accuracy. In addition, during the voice wake-up process, by first judging the wake-up threshold and then performing voiceprint detection as appropriate to determine whether to enable voice interaction, the wake-up response speed can be improved, and the wake-up accuracy can be improved at the same time.

[0078] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0079] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0080] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0081] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0082] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A voice wake-up method, characterized in that, include: The wake-up confidence of the user's voice is detected, and the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold is analyzed in turn; wherein the first wake-up threshold is less than the second wake-up threshold. In response to the wake-up confidence level being no less than the second wake-up threshold, voice interaction is initiated; In response to the wake-up confidence being not less than the first wake-up threshold and less than the second wake-up threshold, a method determines whether to enable voice interaction based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system; wherein, the method further includes at least one of the following: Obtain a first ratio between a first value and a second value; wherein the first value represents the number of times voice interaction is initiated when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, and the second value represents the number of times the user attempts to wake up via voice; based on the relationship between the first ratio and the first ratio threshold, determine whether to adjust the second wake-up threshold; Obtain a second ratio between a third value and a second value; wherein the third value represents the number of times voice interaction is determined not to be enabled when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, and the second value represents the number of times the user attempts to wake up via voice; based on the relationship between the second ratio and the second ratio threshold, determine whether to adjust the first wake-up threshold.

2. The method according to claim 1, characterized in that, The method further includes: If the wake-up confidence level is less than the first wake-up threshold, voice interaction is not initiated.

3. The method according to claim 1, characterized in that, The step of determining whether to enable voice interaction based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system includes: Obtain the feature similarity between the first voiceprint feature and the plurality of second voiceprint features; Voice interaction is initiated in response to the presence of at least one of the aforementioned features having a similarity greater than the voiceprint threshold. If the similarity of each of the aforementioned features is not greater than the voiceprint threshold, voice interaction is not enabled.

4. The method according to claim 1, characterized in that, When determining to adjust the second wake-up threshold based on the relationship between the first ratio and the first ratio threshold, the method further includes: The first numerical value of the wake-up confidence is clustered to obtain several first cluster sets; Select the first cluster set with the largest cluster center as the first target set; Based on the first target set, adjust the second wake-up threshold.

5. The method according to claim 1, characterized in that, When determining to adjust the first wake-up threshold based on the relationship between the second ratio and the second proportional threshold, the method further includes: The third number of wake-up confidence scores are clustered to obtain several second cluster sets; Select the second cluster set with the smallest cluster center as the second target set; Based on the second target set, adjust the first wake-up threshold.

6. A voice wake-up device, characterized in that, include: The confidence detection module is used to detect the wake-up confidence level of the user's voice. The numerical analysis module is used to sequentially analyze the relationship between the wake-up confidence and the first wake-up threshold and the second wake-up threshold, respectively; wherein the first wake-up threshold is less than the second wake-up threshold; The first response module is used to initiate voice interaction in response to the wake-up confidence level being not less than the second wake-up threshold. The second response module is configured to, in response to the wake-up confidence being not less than the first wake-up threshold and less than the second wake-up threshold, determine whether to initiate voice interaction based on the first voiceprint feature extracted from the user's voice and several second voiceprint features already stored in the voice interaction system; wherein, the device is further configured to perform at least one of the following: Obtain a first ratio between a first value and a second value; wherein the first value represents the number of times voice interaction is initiated when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, and the second value represents the number of times the user attempts to wake up via voice; based on the relationship between the first ratio and the first ratio threshold, determine whether to adjust the second wake-up threshold; Obtain a second ratio between a third value and a second value; wherein the third value represents the number of times voice interaction is determined not to be enabled when the wake-up confidence is not less than the first wake-up threshold and less than the second wake-up threshold, and the second value represents the number of times the user attempts to wake up via voice; based on the relationship between the second ratio and the second ratio threshold, determine whether to adjust the first wake-up threshold.

7. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores program instructions and the processor executes the program instructions to implement the voice wake-up method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the voice wake-up method according to any one of claims 1 to 5.