Always-on wakeless method, apparatus, device, and medium

By introducing a dual judgment mechanism of intent information and speech recognition confidence into the all-time wake-up-free technology, and setting multiple thresholds to filter user voice, the problem of high false trigger rate is solved, higher accuracy and stability are achieved, and the user interaction experience is improved.

CN119832901BActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411603035.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-11-18
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing always-on wake-up-free technologies have a high rate of false triggering, which affects user experience and may cause the device to perform incorrect operations.

Method used

By combining intent information and speech recognition confidence, a dual judgment mechanism is established, and multiple thresholds are set to determine whether to trigger full-time wake-up-free operation. This includes the relationship between the effective range of intent and the confidence level of speech recognition, ensuring that wake-up-free operation is triggered only when the intent is valid and the speech recognition result is reliable.

Benefits of technology

It significantly reduces the false trigger rate, improves the accuracy and stability of always-on wake-up-free operation, and provides a smoother user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832901B_ABST
    Figure CN119832901B_ABST
Patent Text Reader

Abstract

The application provides a full-time wake-up-free method, device, equipment and medium, wherein the method comprises: acquiring user voice; performing intent recognition on the user voice to obtain intent information, and performing speech recognition on the user voice to obtain a speech recognition confidence; determining whether to trigger full-time wake-up-free based on whether the intent information belongs to an effective intent range and a size relationship between the speech recognition confidence and a preset threshold. The full-time wake-up-free method, device, equipment and medium provided by the application determine whether the intent information of the user voice is within the effective range, and combine the speech recognition confidence, and only trigger the wake-up-free operation in the case that the intent is effective and the speech recognition result is reliable, thereby effectively improving the accuracy of full-time wake-up-free, significantly reducing the false trigger rate, improving the stability of voice interaction, and bringing a smoother full-time wake-up-free experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for all-time wake-up-free operation. Background Technology

[0002] The always-on wake-up-free feature allows smart devices to directly recognize and execute user commands without a wake word, significantly improving the convenience of voice interaction. This feature is widely used in smart homes, smart cockpits, and other fields, improving a natural and continuous human-computer interaction experience by reducing the need for repeated wake-up calls. However, controlling the false trigger rate remains a major challenge in the development of always-on wake-up-free technology.

[0003] Commonly used all-time wake-up-free technologies primarily control the false trigger rate through three schemes: multi-modal recognition, dynamic wake-up-free word lists, and semantic result verification. Multi-modal recognition, which combines non-voice devices such as cameras to assist in determining whether a user has issued a command, effectively controls the false trigger rate but increases hardware costs. Dynamic wake-up-free word lists switch wake-up-free words in different scenarios, resulting in poor applicability, limited flexibility, and an inability to effectively control the false trigger rate. Semantic result verification performs semantic matching on the recognized content to confirm its consistency with the current scenario, but it is highly dependent on semantic accuracy. Therefore, effectively reducing the false trigger rate of all-time wake-up-free functionality is a pressing issue that needs to be addressed. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and medium for all-time wake-up-free operation, in order to solve the problem of high false triggering rate in related technologies.

[0005] This invention provides a method for always-on wake-up-free operation, including...

[0006] Obtain user voice;

[0007] The user's voice is subjected to intent recognition to obtain intent information, and the user's voice is subjected to speech recognition to obtain speech recognition confidence.

[0008] Based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, it is determined whether to trigger full-time wake-up-free operation.

[0009] According to the present invention, a method for all-time wake-up-free operation is provided, wherein the preset threshold includes a first threshold;

[0010] The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, includes:

[0011] If the intent information falls within the valid range of the intent and the speech recognition confidence is greater than the first threshold, then it is determined to trigger full-time wake-up-free operation.

[0012] If the intent information does not fall within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined that full-time wake-up-free will not be triggered.

[0013] According to the present invention, the preset threshold further includes a second threshold, wherein the second threshold is greater than the first threshold;

[0014] The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, further includes:

[0015] If the intent information does not fall within the valid range of the intent and the speech recognition confidence is greater than the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the second threshold.

[0016] According to the present invention, a method for all-time wake-up-free operation is provided, wherein the preset threshold further includes a third threshold, and the third threshold is less than the first threshold;

[0017] The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, further includes:

[0018] If the intent information falls within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the third threshold.

[0019] According to the present invention, a method for all-time wake-up-free speech recognition is provided, wherein the preset threshold is determined based on the speech recognition confidence of positive samples and the speech recognition confidence of negative samples, wherein the positive samples are speech samples that are determined to trigger all-time wake-up-free speech recognition, and the negative samples are speech samples that are determined not to trigger all-time wake-up-free speech recognition.

[0020] The always-on wake-up-free method provided by the present invention further includes:

[0021] If it is determined that the full-time wake-up-free function is triggered, respond to the user's voice.

[0022] The present invention also provides a always-on wake-up-free device, comprising the following modules:

[0023] The voice acquisition unit is used to acquire the user's voice.

[0024] The model processing unit is used to perform intent recognition on the user's speech to obtain intent information, and to perform speech recognition on the user's speech to obtain speech recognition confidence.

[0025] The triggering judgment unit is used to determine whether to trigger full-time wake-up-free based on whether the intent information belongs to the valid range of intent and the relationship between the speech recognition confidence and the preset threshold.

[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the always-on wake-up-free method as described above.

[0027] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the full-time wake-up-free method as described above.

[0028] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the always-on wake-up-free method as described above.

[0029] The always-on wake-up-free method, apparatus, device, and medium provided by this invention determine whether the user's voice intent information is within the valid range and combine it with the voice recognition confidence level. The wake-up-free operation is triggered only when the intent is valid and the voice recognition result is reliable. This effectively improves the accuracy of always-on wake-up-free operation, significantly reduces the false trigger rate, thereby improving the stability of voice interaction and bringing a smoother always-on wake-up-free experience. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is one of the flowcharts of the always-on wake-up-free method provided by the present invention.

[0032] Figure 2 This is the second flowchart of the always-on wake-up-free method provided by the present invention.

[0033] Figure 3 This is a schematic diagram of the structure of the always-on wake-up-free device provided by the present invention.

[0034] Figure 4This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0036] With the widespread adoption of smart devices and voice interaction technologies, always-on wake-up-free functionality has become a key requirement for enhancing user experience in human-computer interaction. This functionality allows smart devices to directly recognize user voice commands without using a wake word, providing a more natural and fluid voice interaction experience, and is widely used in smart homes, in-vehicle systems, and other fields. However, the issue of false triggering in always-on wake-up-free mode has always been a bottleneck in technological development. A high false triggering rate not only affects user experience but may also lead to the device performing incorrect operations.

[0037] Existing all-time wake-up-free technologies primarily control false trigger rates through three schemes: multi-modal recognition, dynamic wake-up-free word lists, and semantic result verification. Multi-modal recognition technology, by combining cameras, sensors, and other devices with voice signals to determine user intent, significantly improves accuracy, but it has high hardware requirements, increasing equipment costs. The dynamic wake-up-free word list scheme switches specific command words according to different scenarios, limiting the system's recognition range. It is suitable for certain specific scenarios, but its flexibility is limited, and slight changes in user expression can easily trigger false triggers; furthermore, reducing the false trigger rate cannot be guaranteed. The semantic result verification scheme adds a semantic analysis step after speech recognition to ensure that the command conforms to the current operating context, further reducing false triggers. However, it relies on the generalization ability and accuracy of the semantic model; once recognition deviations occur, the false trigger problem remains serious.

[0038] To address the above problems, embodiments of the present invention provide a method for always-on wake-up-free operation. Figure 1 This is one of the flowcharts of the always-on wake-up-free method provided by the present invention, such as... Figure 1 As shown, the method includes:

[0039] Step 110: Obtain user voice recordings.

[0040] Specifically, the system continuously monitors the surrounding sound environment in a standby state using built-in microphones and other hardware to capture potential user voice input. During this process, noise reduction technology is employed to process the received voice signal, filtering out environmental noise and preserving the key features of the user's voice, thereby improving the clarity and accuracy of the user's voice input. This process ensures effective identification of the user's true intent and reduces misjudgments caused by environmental interference. For example, when a user says "Open navigation" in the car, the system extracts the user's voice signal despite background noise. After noise reduction processing, the "Open navigation" command is accurately obtained and analyzed and executed in subsequent steps, achieving operation without a wake-up word.

[0041] Step 120: Perform intent recognition on the user's voice to obtain intent information, and perform speech recognition on the user's voice to obtain speech recognition confidence.

[0042] Specifically, intent recognition and speech recognition are performed on the acquired user speech to ensure the accuracy and effectiveness of subsequent operations. First, intent recognition is a technology that uses natural language processing and deep learning to extract the user's true needs from their speech. The core of intent recognition lies in understanding the user's underlying operational intent, not just identifying specific words or phrases, but connecting these words with the context to determine the function the user wishes to achieve. For example, when a user says "I want to listen to some light music" or "Play some soothing songs," although the specific expressions are different, intent recognition technology can identify both commands as the intent to "play music." Intent recognition technology generally uses deep learning models pre-trained with a large amount of speech data and command samples to identify the same intent in different expressions. The output of intent recognition is intent information, which contains a classification of the user's true needs and specific operational instructions. Intent information is not just a simple summary of the user's speech, but a deep analysis of the user's request, capable of associating complex statements with corresponding functions, thereby helping the device respond quickly according to the user's intent. For example, in an in-vehicle system, intent information may include specific functions such as vehicle control, navigation, and playing music.

[0043] In addition to recognizing user intent from their speech, the system also performs speech recognition, converting audio signals into text for subsequent analysis and execution. For the converted text, a speech recognition confidence score is generated to measure the reliability of the recognized text. This score provides a probabilistic rating indicating the reliability of the recognized text. This score is typically based on various factors, including audio signal quality, the degree of matching in the language model, the decoding path during recognition, and background noise. Speech recognition confidence scores vary under different environments and conditions, helping to determine whether further action should be taken based on the recognition results. Speech recognition confidence scores are particularly important in practical applications. For example, when recognizing the command "Navigate to the company," a high confidence score indicates high accuracy, allowing for direct navigation execution. A low confidence score suggests the device is not entirely confident in the recognized text and may choose to re-ask the user or delay execution to reduce the risk of error.

[0044] By combining intent information and speech recognition confidence, user voice can be filtered more intelligently, ensuring the accuracy of operations. Even in complex environments, the validity of user voice can be reasonably judged using the dual criteria of intent information and speech recognition confidence, thereby effectively reducing false triggers and improving the reliability of the always-on wake-up-free function and user experience. This dual judgment mechanism not only understands the user's true needs but also enhances adaptability and accuracy in diverse usage scenarios.

[0045] Step 130: Based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, determine whether to trigger full-time wake-up-free operation.

[0046] An intent's effective scope is a set of intents designed to trigger a specific action, encompassing all intents that are allowed to be executed directly. This scope is typically configured based on the device's application scenario and user needs to ensure that a wake-free response is only given to specific command types, avoiding unnecessary false triggers. For example, in an in-vehicle system, the intent's effective scope might include navigation, vehicle control, and communication-related intent categories, such as "navigate to destination," "adjust air conditioning temperature," and "make a phone call." In a smart home scenario, the intent's effective scope might include commands closely related to home control, such as "turn on lights," "adjust temperature," and "play music." Setting the intent's effective scope helps filter out the types of commands that should be responded to in specific scenarios, thereby reducing the possibility of false triggers and improving the accuracy of user voice responses.

[0047] In addition to judging based on the effective range of intent, the decision to trigger full-time wake-up-free operation can be further determined by combining speech recognition confidence score. Speech recognition confidence score is a reliability rating of the text result recognized from the user's speech, used to indicate the accuracy of the recognition result. The speech recognition confidence score is compared with a preset threshold. If the confidence score is higher than the preset threshold, it indicates that the recognition result is relatively reliable and accurate, meeting the triggering conditions for full-time wake-up-free operation based on speech recognition confidence score; if it is lower than the preset threshold, it indicates that the reliability of the recognition result is insufficient, and the device may choose not to trigger full-time wake-up-free operation to avoid accidental operation.

[0048] By combining the determination of the effective range of intent and the confidence level of speech recognition, the accuracy and reliability of the always-on wake-up-free function can be significantly improved. On the one hand, the effective range of intent ensures that only valid commands are responded to, avoiding accidental triggering of invalid commands; on the other hand, the speech recognition confidence level filtering mechanism effectively reduces erroneous operations caused by inaccurate recognition. This determination mechanism greatly reduces the possibility of accidental triggering, ensuring accurate understanding and execution of user commands even in complex environments, and significantly improving the user's interactive experience.

[0049] The always-on wake-up-free method provided in this invention determines whether the user's voice intent information is within the valid range and combines it with the voice recognition confidence level. It only triggers the wake-up-free operation when the intent is valid and the voice recognition result is reliable. This effectively improves the accuracy of always-on wake-up-free operation, significantly reduces the false trigger rate, thereby improving the stability of voice interaction and bringing a smoother always-on wake-up-free experience.

[0050] Based on the above embodiments, the preset threshold includes a first threshold;

[0051] The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, includes:

[0052] If the intent information falls within the valid range of the intent and the speech recognition confidence is greater than the first threshold, then it is determined to trigger full-time wake-up-free operation.

[0053] If the intent information does not fall within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined that full-time wake-up-free will not be triggered.

[0054] Specifically, the preset threshold includes a first threshold, which is used to determine whether the confidence level of speech recognition is high enough to decide whether to trigger full-time wake-up-free operation. The significance of the first threshold is to set a reliability standard for speech recognition results, ensuring that the user's voice is only responded to under conditions of high confidence, thereby reducing erroneous operations caused by inaccurate or unclear speech recognition.

[0055] During the judgment process, if the intent information falls within the valid intent range and the speech recognition confidence level is greater than the first threshold, then full-time wake-up-free operation will be triggered. In this case, the user's voice command content conforms to the preset valid intent range, and the confidence level reaches or exceeds the first threshold, indicating that the speech recognition result has high credibility and the command content is valid, thus enabling a quick and accurate response to user needs. This dual verification of high confidence and valid intent ensures reduced false triggering of full-time wake-up-free operation, providing users with a smoother and more timely interactive experience. Conversely, when the intent information is not within the valid intent range and the speech recognition confidence level is lower than the first threshold, full-time wake-up-free operation will not be triggered. This means that the user's voice command neither conforms to the preset valid intent category nor has high speech recognition credibility; the command may not be the user's actual need or the intent may be unclear. In this case, the command will be ignored to avoid unnecessary misoperation and further improve the reliability and stability of the recognition process.

[0056] By using a first threshold for filtering, highly credible commands can be effectively selected while low-confidence content is ignored, thus improving the accuracy of speech recognition. Furthermore, the dual filtering mechanism based on the effective range of intent and the confidence level of speech recognition can ignore user speech containing invalid intents, making the always-on wake-up-free function more intelligent and accurate, and significantly reducing the false trigger rate.

[0057] Based on the above embodiments, the preset threshold further includes a second threshold, which is greater than the first threshold;

[0058] The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, further includes:

[0059] If the intent information does not fall within the valid range of the intent and the speech recognition confidence is greater than the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the second threshold.

[0060] Specifically, the preset threshold also includes a higher second threshold, used to further confirm the reliability of the speech recognition results and whether to trigger the all-time wake-up-free function under certain circumstances. Compared to the first threshold, the second threshold sets a higher confidence requirement, thereby ensuring that the all-time wake-up-free function can be conditionally triggered even when the intent is not within the valid range, in order to meet user needs.

[0061] When a user's voice intent is outside the valid range but the voice recognition confidence level is higher than the first threshold, the system does not decide whether to trigger full-time wake-up-free operation. Instead, it further compares the voice recognition confidence level with a second threshold to make a more accurate judgment. If the voice recognition confidence level exceeds the second threshold, it indicates that the recognition result has extremely high credibility. Even if the intent is not within the valid range, it can still be considered a valid command from the user. Therefore, in this case, full-time wake-up-free operation is triggered to respond to the user's needs. In this way, even if the command category is not within the preset valid range, if the confidence level is high enough, there is still a chance to trigger wake-up-free operation, ensuring that user needs are not overlooked when faced with highly credible commands.

[0062] Conversely, if the speech recognition confidence level falls between the first and second thresholds, it indicates that while the confidence level of the recognition result exceeds the basic judgment standard, it does not reach an extremely high confidence level. In this case, although the confidence level of the speech recognition is high, the full-time wake-up-free function will not be triggered because the user's intention information is not within the effective range, thus avoiding the execution of commands that may not meet the user's actual needs. This design ensures that when the intention is not within the effective range, the full-time wake-up-free function will only be allowed to be triggered when the recognition result reaches an extremely high confidence level, thereby minimizing the false trigger rate and ensuring the accuracy of the interactive response without affecting the user experience.

[0063] The purpose of this dual-threshold design is to introduce a more granular speech recognition confidence filtering mechanism. This allows the judgment process to not only determine the validity of commands based on the effective range of intent, but also, in special cases, respond to commands outside the effective range with extremely high speech recognition confidence. This mechanism not only improves the accuracy of always-on wake-up-free operation, but also significantly enhances the flexibility of interaction, further improving the user experience.

[0064] Based on the above embodiments, the preset threshold further includes a third threshold, which is less than the first threshold;

[0065] The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, further includes:

[0066] If the intent information falls within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the third threshold.

[0067] Specifically, the preset threshold also includes a third threshold, which is lower than the first threshold. The purpose of the third threshold is to provide a minimum standard for determining whether to trigger full-time wake-up-free operation when the speech recognition confidence is low. It introduces a more lenient judgment condition, ensuring that full-time wake-up-free operation can still be triggered even when the speech confidence does not reach the ideal standard, in order to meet the user's needs.

[0068] When the intent information falls within the valid range, but the speech recognition confidence level is below the first threshold, the system doesn't immediately decide not to trigger full-time wake-up control. Instead, it further compares the speech recognition confidence level with the third threshold. If the speech recognition confidence level is above the third threshold, full-time wake-up control will still be triggered even if the higher standard of the first threshold isn't met. This is to ensure that even with slightly lower confidence levels, commands that conform to valid intents are still executed, guaranteeing that the user's valid needs are not ignored due to slightly lower speech recognition confidence. This is particularly important in certain scenarios, such as when environmental noise causes a slight decrease in speech recognition confidence. In such cases, the user's commands can still be responded to, without affecting the continuity and convenience of operation. However, if the speech recognition confidence level is below the third threshold, it indicates insufficient credibility of the speech recognition result. Even if the user's speech intent information is within the valid intent range, due to the potential risk of misrecognition, full-time wake-up control will not be triggered, thus avoiding erroneous operations on commands with low credibility.

[0069] Through this two-layer screening mechanism, the third threshold effectively enhances the flexibility and adaptability of the interaction, enabling accurate judgment on whether to respond to user needs even in more complex environments. The setting of the third threshold not only reduces the strict reliance on high-confidence standards but also ensures timely responses to valid user commands in changing environments, resulting in a more intelligent and smoother interactive experience.

[0070] Based on the above embodiments, the preset threshold is determined based on the speech recognition confidence of positive samples and the speech recognition confidence of negative samples. The positive samples are speech samples that are determined to trigger full-time wake-up-free speech, and the negative samples are speech samples that are determined not to trigger full-time wake-up-free speech.

[0071] Specifically, positive samples refer to user voice samples that meet the conditions for triggering full-time wake-up-free operation; that is, voice samples that users expect to be recognized and responded to immediately, and their voice recognition confidence is usually high. Therefore, by analyzing the voice recognition confidence range of positive samples, the required voice recognition confidence standard when the triggering conditions are met can be clearly defined. This provides a basis for setting a threshold that can reliably trigger full-time wake-up-free operation, ensuring a fast and accurate response when a clear and valid user command is recognized.

[0072] Reverse samples are user voice samples that meet the conditions for not triggering full-time wake-up-free operation. These include invalid commands, noise, idle chatter, and other non-command voices that should be ignored, or ambiguous commands with low confidence. The confidence level of these samples is usually low. By analyzing the confidence level of reverse samples, a low confidence threshold can be derived to help prevent full-time wake-up-free operation from being triggered when receiving ambiguous or unreliable commands. Setting this baseline threshold effectively prevents false triggers, enabling the correct differentiation between commands requiring a response and invalid voices that should be ignored in complex environments.

[0073] This method for determining speech confidence thresholds based on both positive and negative samples provides a clear range of speech confidence, enabling effective determination of whether to trigger full-time wake-up-free operation in various scenarios. The dual analysis of positive and negative samples not only improves the accuracy of responses to valid commands but also reduces the false trigger rate, enhances adaptability in complex environments, and significantly improves the fluency and reliability of the user experience.

[0074] Based on the above embodiments, the always-on wake-up-free method further includes:

[0075] If it is determined that the full-time wake-up-free function is triggered, respond to the user's voice.

[0076] Specifically, once the system determines that a user's voice meets the conditions for triggering a wake-up-free experience through a dual assessment of intent information and voice recognition confidence, the user's command will be executed directly. For example, when a user utters a voice command like "Navigate to the nearest gas station," and it's determined that the user's voice meets the conditions for triggering a wake-up-free experience, the navigation function will automatically start without any additional user intervention. This design not only makes the interaction more natural but also avoids the cumbersome steps of multiple wake-ups or confirmations, significantly improving the user experience. This response mechanism is particularly important in complex environments or driving scenarios, allowing users to focus on the current task while the device efficiently executes commands in the background.

[0077] Based on the above embodiments, Figure 2 This is the second flowchart illustrating the always-on wake-up-free method provided in this invention. For example... Figure 2 As shown in the flowchart, this process illustrates a real-time wake-free decision-making process based on intent information and speech recognition confidence. It utilizes three different speech recognition confidence thresholds (first threshold, second threshold, and third threshold) to determine whether to trigger the wake-free function. The relationship between the three confidence thresholds is: second threshold > first threshold > third threshold. The entire process begins with user voice input. The voice signal is fed into the intent recognition engine and the speech recognition engine respectively, outputting intent information and speech recognition confidence for subsequent judgments.

[0078] After summarizing the user's voice intent information and voice recognition confidence level, the specific conditional judgment first checks whether the voice recognition confidence level is greater than a first threshold and whether the intent information is within a valid range. If the conditions are met, the always-on wake-up-free function is directly triggered to respond to the user's voice. If the conditions are not met, the next step is to check whether the voice recognition confidence level is greater than the first threshold and whether the intent information is not within a valid range. At this point, if the voice recognition confidence level exceeds a second threshold, the always-on wake-up-free function is still triggered; if the voice recognition confidence level does not reach the second threshold, the voice command is ignored.

[0079] When the speech recognition confidence level is below the first threshold but the intent information is within the valid range, the process will further determine whether the speech recognition confidence level is between the third threshold and the first threshold. If the speech recognition confidence level falls within this range, the always-on wake-up-free function is triggered; if the confidence level is below the third threshold, the instruction is ignored. In the last case, if the confidence level is below the first threshold and the intent information is not within the valid range, the instruction is directly ignored, and the always-on wake-up-free function is not triggered.

[0080] By using this multi-layered speech recognition confidence threshold and intent validity judgment, the process effectively balances the triggering conditions and false trigger control of the all-time wake-up-free function, ensuring that while reducing false triggers, it can flexibly respond to commands with high speech recognition confidence or valid intent, significantly improving the smoothness and accuracy of the user interaction experience.

[0081] The always-on wake-up-free device provided by the present invention is described below. The always-on wake-up-free device described below can be referred to in correspondence with the always-on wake-up-free method described above.

[0082] Figure 3 This is a schematic diagram of the structure of the always-on wake-up-free device provided by the present invention, as shown below. Figure 3 As shown, the device includes:

[0083] The voice acquisition unit 310 is used to acquire the user's voice.

[0084] The model processing unit 320 is used to perform intent recognition on the user's speech to obtain intent information, and to perform speech recognition on the user's speech to obtain speech recognition confidence.

[0085] The trigger judgment unit 330 is used to determine whether to trigger full-time wake-up-free based on whether the intent information belongs to the valid range of intent and the relationship between the speech recognition confidence and the preset threshold.

[0086] The always-on wake-up-free device provided in this embodiment of the invention determines whether the user's voice intent information is within the valid range and combines it with the voice recognition confidence level. It only triggers the wake-up-free operation when the intent is valid and the voice recognition result is reliable. This effectively improves the accuracy of always-on wake-up-free operation, significantly reduces the false trigger rate, thereby improving the stability of the interaction and bringing a smoother always-on wake-up-free experience.

[0087] Based on any of the above embodiments, the triggering determination unit is specifically used for:

[0088] The preset threshold includes a first threshold;

[0089] If the intent information falls within the valid range of the intent and the speech recognition confidence is greater than the first threshold, then it is determined to trigger full-time wake-up-free operation.

[0090] If the intent information does not fall within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined that full-time wake-up-free will not be triggered.

[0091] Based on any of the above embodiments, the triggering determination unit is specifically used for:

[0092] The preset threshold also includes a second threshold, which is greater than the first threshold;

[0093] If the intent information does not fall within the valid range of the intent and the speech recognition confidence is greater than the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the second threshold.

[0094] Based on any of the above embodiments, the triggering determination unit is specifically used for:

[0095] The preset threshold also includes a third threshold, which is less than the first threshold;

[0096] If the intent information falls within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the third threshold.

[0097] Based on any of the above embodiments, the triggering determination unit is specifically used for:

[0098] The preset threshold is determined based on the speech recognition confidence of positive samples and the speech recognition confidence of negative samples. The positive samples are speech samples that are determined to trigger full-time wake-up-free speech, and the negative samples are speech samples that are determined not to trigger full-time wake-up-free speech.

[0099] Based on any of the above embodiments, the triggering determination unit is specifically used for:

[0100] If it is determined that the full-time wake-up-free function is triggered, respond to the user's voice.

[0101] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a full-time wake-up-free method, which includes:

[0102] Obtain user voice;

[0103] The user's voice is subjected to intent recognition to obtain intent information, and the user's voice is subjected to speech recognition to obtain speech recognition confidence.

[0104] Based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, it is determined whether to trigger full-time wake-up-free operation.

[0105] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0106] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the full-time wake-up-free method provided by the above methods, the method comprising:

[0107] Obtain user voice;

[0108] The user's voice is subjected to intent recognition to obtain intent information, and the user's voice is subjected to speech recognition to obtain speech recognition confidence.

[0109] Based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, it is determined whether to trigger full-time wake-up-free operation.

[0110] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the full-time wake-up-free method provided by the methods described above, the method comprising:

[0111] Obtain user voice;

[0112] The user's voice is subjected to intent recognition to obtain intent information, and the user's voice is subjected to speech recognition to obtain speech recognition confidence.

[0113] Based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, it is determined whether to trigger full-time wake-up-free operation.

[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for always-on wake-up-free operation, characterized in that, include: Obtain user voice; The user's voice is subjected to intent recognition to obtain intent information, and the user's voice is subjected to speech recognition to obtain speech recognition confidence. Based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, it is determined whether to trigger full-time wake-up-free operation.

2. The always-on wake-up-free method according to claim 1, characterized in that, The preset threshold includes a first threshold; The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, includes: If the intent information falls within the valid range of the intent and the speech recognition confidence is greater than the first threshold, then it is determined to trigger full-time wake-up-free operation. If the intent information does not fall within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined that full-time wake-up-free will not be triggered.

3. The always-on wake-up-free method according to claim 2, characterized in that, The preset threshold also includes a second threshold, which is greater than the first threshold; The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, further includes: If the intent information does not fall within the valid range of the intent and the speech recognition confidence is greater than the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the second threshold.

4. The always-on wake-up-free method according to claim 2, characterized in that, The preset threshold also includes a third threshold, which is less than the first threshold; The step of determining whether to trigger full-time wake-up-free operation based on whether the intent information falls within the valid range of the intent, and the relationship between the speech recognition confidence level and the preset threshold, further includes: If the intent information falls within the valid range of the intent and the speech recognition confidence is less than or equal to the first threshold, it is determined whether to trigger full-time wake-up-free operation based on the relationship between the speech recognition confidence and the third threshold.

5. The full-time wake-up-free method according to any one of claims 1 to 4, characterized in that, The preset threshold is determined based on the speech recognition confidence of positive samples and the speech recognition confidence of negative samples. The positive samples are speech samples that are determined to trigger full-time wake-up-free speech, and the negative samples are speech samples that are determined not to trigger full-time wake-up-free speech.

6. The full-time wake-up-free method according to any one of claims 1 to 4, characterized in that, Also includes: If it is determined that the full-time wake-up-free function is triggered, respond to the user's voice.

7. A always-on, wake-up-free device, characterized in that, include: The voice acquisition unit is used to acquire the user's voice. The model processing unit is used to perform intent recognition on the user's speech to obtain intent information, and to perform speech recognition on the user's speech to obtain speech recognition confidence. The triggering judgment unit is used to determine whether to trigger full-time wake-up-free based on whether the intent information belongs to the valid range of intent and the relationship between the speech recognition confidence and the preset threshold.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the always-on wake-up-free method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the always-on wake-up-free method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the always-on wake-up-free method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice-combined intention recognition method and device

    CN111914563A

  • Wakeup-free word threshold adjustment method and device, vehicle and readable storage medium

    CN116052677A