Voice control method and device, storage medium and electronic equipment

By combining keyword detection and speech recognition technologies, the problem of high false recognition rate of wake-up words has been solved, achieving higher accuracy and efficiency in voice control.

CN119479634BActive Publication Date: 2026-03-24GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-10
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

The existing voice control methods have a high misrecognition rate of wake-word-free speech, resulting in insufficient accuracy of voice interaction.

Method used

By combining keyword detection technology and speech recognition technology, keywords are first obtained from the speech data through keyword detection, and then the keywords are confirmed using speech recognition technology to ensure the accurate execution of control commands.

Benefits of technology

It improves the accuracy of wake-word recognition, reduces the false wake-up rate, and enhances the accuracy and efficiency of voice control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479634B_ABST
    Figure CN119479634B_ABST
Patent Text Reader

Abstract

The application discloses a voice control method and device, a storage medium and an electronic device. The method comprises the following steps: obtaining to-be-detected voice data, performing keyword detection on the to-be-detected voice data to obtain a keyword in the to-be-detected voice data; obtaining voice data in which the keyword is located from the to-be-detected voice data to obtain to-be-recognized voice data; performing voice recognition on the to-be-recognized voice data to obtain a first voice recognition result; and if the keyword matches the first voice recognition result, determining a control instruction corresponding to the to-be-detected voice data and executing the control instruction. The application can reduce the misrecognition rate of a wake-up word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic technology, and in particular relates to a voice control method, device, computer-readable storage medium and electronic device. Background Technology

[0002] Voice wake-up-free technology is an intelligent voice technology that allows for voice interaction with electronic devices without requiring a wake-up word. When controlling electronic devices by voice, this technology provides greater convenience. However, existing voice control methods have a relatively high rate of misrecognition of wake-up words. Summary of the Invention

[0003] This application provides a voice control method, device, storage medium, and electronic device that can reduce the false recognition rate of wake-up words.

[0004] In a first aspect, embodiments of this application provide a voice control method, including:

[0005] Acquire the speech data to be detected, and perform keyword detection on the speech data to be detected to obtain the keywords in the speech data to be detected;

[0006] The speech data containing the keyword is obtained from the speech data to be detected, thus obtaining the speech data to be recognized;

[0007] The speech data to be recognized is subjected to speech recognition to obtain a first speech recognition result;

[0008] If the keyword matches the first speech recognition result, a control instruction corresponding to the speech data to be detected is determined and the control instruction is executed.

[0009] Secondly, embodiments of this application provide a voice control device, including:

[0010] The keyword detection module is used to acquire the speech data to be detected and to perform keyword detection on the speech data to be detected to obtain the keywords in the speech data to be detected.

[0011] The data acquisition module is used to acquire the speech data containing the keyword from the speech data to be detected, and obtain the speech data to be recognized.

[0012] The speech recognition module is used to perform speech recognition on the speech data to be recognized and obtain a first speech recognition result;

[0013] The instruction execution module is used to determine the control instruction corresponding to the speech data to be detected if the keyword matches the first speech recognition result, and to execute the control instruction.

[0014] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to perform the voice control method provided in embodiments of this application.

[0015] Fourthly, embodiments of this application also provide an electronic device, including a memory and a processor, wherein the processor executes the voice control method provided in embodiments of this application by calling a computer program stored in the memory.

[0016] In this embodiment, by acquiring speech data to be detected and performing keyword detection on the speech data to be detected, keywords in the speech data to be detected are obtained; speech data to be recognized is acquired from the speech data to be detected and performed on the speech data to be recognized to obtain a first speech recognition result; if the keyword matches the first speech recognition result, a control command corresponding to the speech data to be detected is determined and the control command is executed. Thus, by combining keyword detection technology and speech recognition technology to identify keywords in speech, i.e., wake-up-free words, compared with the scheme of identifying wake-up-free words in speech only by keyword detection technology, the accuracy of wake-up-free word recognition can be improved, thereby reducing the false wake-up rate of wake-up-free words. Attached Figure Description

[0017] The technical solution and its beneficial effects will become apparent from the following detailed description of specific embodiments of this application, in conjunction with the accompanying drawings.

[0018] Figure 1 This is a schematic diagram of the first type of voice control method provided in the embodiments of this application.

[0019] Figure 2 This is a schematic diagram of the second type of voice control method provided in the embodiments of this application.

[0020] Figure 3 This is a schematic diagram of a scenario for the voice control method provided in the embodiments of this application.

[0021] Figure 4 This is a schematic diagram of the structure of the voice control device provided in the embodiments of this application.

[0022] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0023] It should be noted that the terms "first," "second," and "third," etc., used in this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but some embodiments also include steps or modules not listed, or some embodiments also include other steps or modules inherent to these processes, methods, products, or devices.

[0024] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0025] This application provides a voice control method, a voice control device, a storage medium, and an electronic device. The entity executing the voice control method can be the voice control device provided in this application, or an electronic device integrating the voice control device, wherein the voice control device can be implemented in hardware or software. The electronic device can be a smartphone, tablet computer, PDA, laptop computer, or other device equipped with a processor and possessing voice control capabilities.

[0026] Please see Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the voice control method provided in this application. The process may include:

[0027] In step 101, the speech data to be detected is acquired, and keyword detection is performed on the speech data to be detected to obtain the keywords in the speech data to be detected.

[0028] The voice data to be detected is the sound in the environment in which the electronic device is located. This sound can be captured by the microphone built into the electronic device, or by an external device connected to the electronic device.

[0029] For example, you can use the phone's built-in microphone to collect the sound of the environment in which the phone is located, or you can use an external microphone connected to the phone to collect the sound of the environment in which the phone is located.

[0030] When an electronic device uses its built-in microphone to capture sounds from its environment, the microphone can be activated by first enabling the audio capture function. This audio capture function can either start automatically when the electronic device is powered on or be activated by the user.

[0031] For example, the electronic device may have the wake-up-free function enabled by default. Alternatively, the wake-up-free function may be enabled only when needed. The wake-up-free function and audio capture function may be enabled simultaneously, or the audio capture function may be enabled after the wake-up-free function is enabled; the order of their activation is not limited here. The wake-up-free function can be activated via any of the following methods: manual activation by the user, activation via user voice control, automatic activation triggered by preset conditions, or activation via user gesture.

[0032] When a user manually activates the wake-up-free function, for example, a touch button with the wake-up-free function can be displayed in the settings interface of an electronic device. The user can activate the wake-up-free function by operating the touch button. Of course, by configuring the existing touch button, it can also be configured to activate both the original function and the wake-up-free function. For example, the Bluetooth touch button can activate both the Bluetooth function and the wake-up-free function, that is, the Bluetooth function and the wake-up-free function can be activated simultaneously.

[0033] When a user activates the wake-up-free function via voice control, such as by using a voice assistant in related technologies to activate the wake-up-free function of an electronic device.

[0034] When the wake-up-free function of an electronic device is triggered by preset conditions, the preset conditions may be, for example, controlling the wake-up-free function to start automatically when the user launches a specified application, and the specified application may be one or more; the preset conditions may also be, controlling the wake-up-free function to start automatically when the electronic device is detected to be in a specified environment, and the specified environment may be, for example, a driving environment, a call environment, a file transfer environment, etc.

[0035] When a user activates the wake-up-free function with a gesture, it can be done by having the user draw a specified pattern on the screen of an electronic device using their knuckles, or by having the user display a specified gesture in front of a camera.

[0036] Since there are multiple ways to activate the wake-up-free function, they will not be listed here. Understandably, the above-mentioned activation methods can also be combined to activate the wake-up-free function, and after the wake-up-free function is activated, the electronic device enters the wake-up-free mode.

[0037] When the wake-free function is turned off, the wake-free mode is exited. After exiting the wake-free mode, the electronic device can still execute the voice control methods commonly used in related technologies. Therefore, by adding a wake-free function to the electronic device, this application enables more flexible voice control of the electronic device, allowing voice control of the electronic device through existing methods in related technologies as well as through the method provided in this application.

[0038] For example, after activating the wake-up-free function, the electronic device can collect voice data in real time through its built-in microphone to obtain the voice data to be detected. After obtaining the voice data to be detected, the electronic device can start the keyword detection engine to perform keyword detection on the voice data to obtain the keywords in the voice data to be detected.

[0039] It is understood that the keywords mentioned in the embodiments of this application are wake-up-free words. These keywords may include: "play next song", "play previous song", "navigate to", "make a phone call", etc.

[0040] Understandably, when no keyword is detected, the electronic device can continue to detect other voice data until a keyword is detected.

[0041] When performing keyword detection on the speech data to be detected using a keyword detection engine, the speech data can be divided into multiple sub-speech data sets. The keyword detection engine then performs keyword detection on each sub-speech data set. If a sub-speech data set contains any keyword except the last one, or if no keyword is present, the output of the keyword detection engine is 0. If the sub-speech data set contains the last keyword, the output of the keyword detection engine is a non-zero value. Each sub-speech data set is speech data of a first preset duration. The first preset duration can be determined based on the duration of the input speech data supported by the keyword detection engine; for example, the first preset duration could be 20ms, 30ms, etc.

[0042] In an optional embodiment, the speech data to be detected is in the form of a speech stream, and the input to the keyword detection engine is speech data of a first preset duration. Then, each time the electronic device collects speech data of the first preset duration, it performs keyword detection on the speech data through the keyword detection engine. When the speech data contains any keyword except the last keyword, or when no keyword is present, the output of the keyword detection engine is 0. When the speech data contains the last keyword, the output of the keyword detection engine is a non-zero value. The first preset duration can be determined according to the duration of the input speech data supported by the keyword detection engine; for example, the first preset duration can be 20ms, 30ms, etc.

[0043] In this context, 0 indicates that no keyword was detected, while non-zero values ​​indicate that a keyword was detected, and each non-zero value corresponds to a specific keyword. For example, if the keyword detection engine can detect a total of 8 keywords, the non-zero values ​​can range from 1 to 8, with each value uniquely corresponding to one of the 8 keywords.

[0044] In step 102, the speech data containing the keywords is obtained from the speech data to be detected, thus obtaining the speech data to be recognized.

[0045] When detecting the presence of keywords in speech data, the keyword detection engine calculates the probability of each keyword's presence and uses the highest probability as the confidence level output by the engine, i.e., the second confidence level. If the keyword engine outputs a non-zero value, and this second confidence level is not less than the second confidence level threshold, then the keyword corresponding to that non-zero value can be considered a keyword in the speech data to be detected.

[0046] To avoid situations where users need to repeat keywords multiple times to control electronic devices due to low keyword wake-up rates, the second confidence threshold is typically set relatively low, such as 0.35 or 0.38. However, this relatively low threshold can also lead to a higher false recognition rate for keywords, where a user utters a word similar to a keyword but it is mistakenly identified as such by the keyword detection engine. Therefore, to reduce the false recognition rate of wake-up-free words and thus the false wake-up rate of voice control, this embodiment, when detecting keywords from the voice data to be detected by the keyword detection engine, further retrieves the voice data containing the keyword to be recognized from the voice data to be detected, in order to further confirm the presence of the keyword in the voice data to be detected using speech recognition technology.

[0047] It is understandable that when users control electronic devices with their voice, the spoken words usually include control commands that include keywords. For example, if the user's voice output is "Call Mom," the keyword is "Call," and the control command is "Call Mom." If the user's voice output is "Navigate to xx restaurant," the keyword is "Navigate to," and the control command is "Navigate to xx restaurant." That is to say, when users control electronic devices with their voice, the spoken words include more than just keywords. If the entire spoken words were to be recognized, it would increase the power consumption of the electronic device and affect the speed of voice control. Therefore, in this embodiment, after detecting keywords from the voice data to be detected, the voice data containing the keywords is also obtained from the voice data to be detected, resulting in the voice data to be recognized. In other words, the portion of the voice data containing keywords in the voice data to be detected is determined as the voice data to be recognized.

[0048] For example, assuming the speech data to be detected is a 3-second speech data, and the speech data containing keywords is the speech data from the 0.5 second to the 2nd second of the speech data to be detected, then the speech data from the 0.5 second to the 2nd second of the speech data to be detected can be identified as the speech data to be recognized.

[0049] In step 103, speech recognition is performed on the speech data to be recognized to obtain the first speech recognition result.

[0050] In this embodiment, after obtaining the voice data to be recognized, the electronic device can also perform voice recognition on the voice data to be recognized through a voice recognition engine to obtain a first voice recognition result.

[0051] In an optional embodiment, when performing speech recognition on the speech data to be recognized using a speech recognition engine, the speech data to be recognized can be performed using a local speech recognition engine, thereby avoiding the time consumption of network transmission and improving the speed of keyword recognition.

[0052] Since the keyword detection engine consumes less power than the speech recognition engine, in this embodiment, only the low-power keyword detection engine is activated first. After the keyword detection engine detects a keyword, the high-power speech recognition engine is then activated. Furthermore, the high-power speech recognition engine only recognizes the speech data containing the keyword, typically for only a few seconds, and its impact on power consumption is negligible. This achieves a reduction in the keyword misidentification rate while ensuring low power consumption. The increased misidentification rate allows for a further reduction in the second confidence threshold, thereby further improving the keyword recognition rate. In other words, both the keyword recognition rate and misidentification rate performance are improved.

[0053] In step 103, if the keyword matches the first speech recognition result, the control command corresponding to the speech data to be detected is determined and executed.

[0054] In this embodiment, if the keyword matches the first speech recognition result, then the control command corresponding to the speech data to be detected can be determined and executed.

[0055] For example, an electronic device can use a speech recognition engine to perform speech recognition on the speech data to be detected, obtaining a second speech recognition result, which includes the text to be analyzed. Subsequently, the electronic device can perform semantic analysis on the text to be analyzed, obtaining a semantic analysis result; then, it can determine the control command corresponding to the semantic analysis result, use it as the control command corresponding to the speech to be detected, and execute the control command.

[0056] For example, if the control command corresponding to the voice data to be detected is "call mom", the electronic device can execute the command to call mom; if the control command corresponding to the voice data to be detected is "navigate to xx restaurant", the electronic device can execute the command to navigate to xx restaurant.

[0057] Understandably, if the keyword does not match the first speech recognition result, the electronic device can return to Execution 101 to perform keyword detection on the newly acquired speech data to be detected until the detected keyword matches the first speech recognition result.

[0058] In this embodiment, by acquiring the speech data to be detected and performing keyword detection on the speech data to be detected, the keywords in the speech data to be detected are obtained; speech data to be recognized is acquired from the speech data to be detected and speech recognition is performed on the speech data to be recognized to obtain a first speech recognition result; if the keyword matches the first speech recognition result, the control command corresponding to the speech data to be detected is determined and the control command is executed. Thus, by combining keyword detection technology and speech recognition technology to identify keywords in speech, i.e., wake-up words, compared with the scheme of identifying wake-up words in speech only by keyword detection technology, the accuracy of wake-up word recognition can be improved, thereby reducing the false wake-up rate of wake-up words.

[0059] In an optional embodiment, obtaining the speech data containing the keywords from the speech data to be detected to obtain the speech data to be recognized includes:

[0060] Determine the end time of the voice data containing the last keyword in the keyword list;

[0061] Based on the end time, obtain the speech data to be recognized from the speech data to be detected.

[0062] It should be noted that when the keyword detection engine performs keyword detection on the speech data to be detected, it usually inputs a portion of the speech data to be detected, such as speech data of a first preset duration, into the keyword detection engine in chronological order. The speech data of the preset duration usually only includes some keywords, such as only one or more keywords. The keyword detection engine only outputs a non-zero value and a second confidence level when it detects the speech data containing the last keyword. Thus, the electronic device can identify the keyword corresponding to the non-zero value as the keyword in the speech data to be detected when the second confidence level is not less than the second confidence level threshold.

[0063] Since the keyword detection engine outputs a non-zero value and a second confidence level when it detects the speech data containing the last keyword, the electronic device can determine the end time of the speech data and obtain the speech data to be recognized from the speech data to be detected based on the end time.

[0064] The end time of the speech data containing the last keyword can be the acquisition completion time of the last audio frame included in the speech data.

[0065] In one optional embodiment, the electronic device collects voice data for a first preset duration and then inputs the voice data into a keyword detection engine for keyword detection. During each collection of voice data for the first preset duration, the electronic device records the start and end times of the voice data collection. Therefore, after determining the voice data containing the last keyword, the electronic device can use the end time of that voice data collection as the end time of the voice data collection.

[0066] In an optional embodiment, before obtaining the speech data to be recognized from the speech data to be detected according to the end time, the method further includes:

[0067] Determine the target duration based on the number of words in the keywords; the target duration is positively correlated with the number of words.

[0068] Based on the end time, obtain the speech data to be recognized from the speech data to be detected, including:

[0069] Based on the end time, the speech data of the target duration is obtained from the speech data to be detected as the speech data to be recognized.

[0070] Understandably, the more characters a keyword has, the more time a user needs to speak it, resulting in a longer duration of the speech data containing that keyword. Therefore, in this embodiment, the duration of the speech data to be recognized, i.e., the target duration, can be determined based on the number of characters in the keyword. Then, based on this end time, speech data of the target duration is extracted from the speech data to be detected as the speech data to be recognized. This avoids performing speech recognition on the entire speech data to be detected to some extent, thereby reducing the power consumption of electronic devices and improving the speed of speech recognition, thus increasing the speed of keyword detection. The target duration is positively correlated with the number of characters; that is, the more characters, the longer the target duration, and vice versa.

[0071] For example, if the keyword is "play next song", the target duration can be 2.6 seconds; if the keyword is "navigate to", the target duration can be "1.2" seconds.

[0072] In an optional embodiment, voice data of each user speaking a keyword of a certain number of words can be collected in advance from multiple users to obtain multiple voice data; then the duration of each voice data is determined to obtain multiple durations; finally, the duration corresponding to the keyword of that number of words is determined based on the multiple durations; when the keyword of that number of words is obtained again in the future, the duration can be determined as the target duration corresponding to the keyword.

[0073] When determining the duration corresponding to a keyword with a certain number of characters based on multiple durations, the average of the multiple durations can be used as the duration corresponding to the keyword with that number of characters; the longest duration among the multiple durations can be used as the duration corresponding to the keyword with that number of characters; or the duration with the most occurrences among the multiple durations can be used as the duration corresponding to the keyword with that number of characters.

[0074] For example, suppose we have collected 1000 voice data points from multiple users, each user repeatedly uttering a keyword with 3 characters. Of these, 990 voice data points have a duration of 1.2 seconds. Therefore, 1.2 seconds can be determined as the duration corresponding to a keyword with 3 characters. When subsequent keywords also have 3 characters, the target duration for that keyword can be determined to be 1.2 seconds.

[0075] In an optional embodiment, before determining the target duration based on the number of characters in the keywords, the method further includes:

[0076] Speech rate detection is performed on the speech data to be detected to obtain the speech rate corresponding to the speech data to be detected;

[0077] Determine the target duration based on the number of words in the keywords, including:

[0078] Determine the target duration based on the number of words and speaking speed. The target duration is inversely related to the speaking speed.

[0079] It is understandable that, given the same number of words in a keyword, a user with a faster speaking speed will take less time to pronounce the keyword compared to a user with a slower speaking speed. Therefore, in this embodiment, the target duration can be determined by combining the number of words in the keyword and the speaking speed of the voice data to be detected. Specifically, the target duration is positively correlated with the number of words and inversely correlated with the speaking speed; that is, the more words and the slower the speaking speed, the longer the target duration, and vice versa.

[0080] In an optional embodiment, obtaining the speech data to be recognized from the speech data to be detected based on the end time includes:

[0081] The corresponding speech data before the end time in the speech data to be detected is identified as the speech data to be recognized.

[0082] For example, the speech data of the target duration before the end time in the speech data to be detected can be identified as the speech data to be recognized.

[0083] For example, the first duration of voice data before the end time in the voice data to be detected can be identified as the voice data to be recognized.

[0084] The first duration can be set in advance or determined based on the length of the keyword with the most characters among all the keywords that the keyword detection engine can detect.

[0085] In an optional embodiment, during the keyword detection process using the keyword detection engine, real-time voice data is simultaneously cached. Considering hardware costs, typically only voice data of a second preset duration is cached. This second preset duration is longer than the maximum target duration; for example, if the maximum target duration is 2.6 seconds, then the second preset duration can be 3 seconds. Therefore, when subsequent voice data to be recognized is needed, the cached voice data of the target duration before the end time can be used as the voice data to be recognized.

[0086] In an optional embodiment, obtaining the speech data to be recognized from the speech data to be detected based on the end time includes:

[0087] The corresponding speech data before the end time and the corresponding speech data after the end time in the speech data to be detected are identified as the speech data to be recognized.

[0088] For example, the second duration of voice data before the end time and the third duration of voice data after the end time can be identified as the voice data to be recognized.

[0089] The second duration is longer than the third duration.

[0090] In an optional embodiment, the second duration can be a target duration. For example, assuming the target duration is 1.2 seconds, the second duration can be 1.2 seconds, and the third duration can be 0.2 seconds.

[0091] In an optional embodiment, the sum of the second duration and the third duration can be a target duration. For example, assuming the target duration is 1.2 seconds, the second duration can be 1.1 seconds and the third duration can be 0.1 seconds.

[0092] In an optional embodiment, the first speech recognition result includes text; if a keyword matches the first speech recognition result, a control command corresponding to the speech data to be detected is determined, including:

[0093] If the keyword matches the text, then the control command corresponding to the speech data to be detected is determined.

[0094] For example, to further reduce the power consumption of speech recognition, when the speech recognition engine performs speech recognition on the speech data to be recognized, the first speech recognition result output by the speech recognition engine can include the recognized phonemes, without needing to output the text mapped to the phonemes. For example, the text can include "daohangdao", "dadianhuagei", etc.

[0095] For example, to further reduce the false recognition rate, when the speech recognition engine performs speech recognition on the speech data to be recognized, the text included in the first speech recognition result output by the speech recognition engine can be the text mapped to the recognized phonemes. For example, the text may include "navigate to" or "make a phone call".

[0096] It is understandable that, due to various factors, the voice data obtained when acquiring the keyword may include not only the keyword but also other content, such as the acquired voice data including "please call". Based on this, in this embodiment, when the text includes the keyword, it can be determined that the keyword matches the text. Then, the control command corresponding to the voice data to be detected can be further determined.

[0097] To facilitate matching, text and keywords can be expressed in the same way. For example, if the text is a phoneme, then the keyword is also a phoneme; if the text is the text mapped to a phoneme, then the keyword is also the corresponding text.

[0098] For example, assuming both the text and the keyword are in phoneme form, if the text includes "dadianhuagei" and the keyword is "dadianhuagei", then the text and keyword match. If both the text and the keyword are in corresponding text form, if the text includes "navigate to" and the keyword is "navigate to", then the text and keyword match.

[0099] In an optional embodiment, the first speech recognition result further includes a first confidence level; if a keyword matches the text, then a control command corresponding to the speech data to be detected is determined, including:

[0100] If the keyword matches the text and the first confidence level is greater than the first confidence level threshold, then the control command corresponding to the speech data to be detected is determined.

[0101] To further reduce the false recognition rate of keywords, this embodiment can also use the confidence level output by the speech recognition engine, i.e., the first confidence level, as an evaluation criterion to determine whether to perform voice control on the electronic device. For example, a confidence threshold can be preset as the first confidence threshold. If the keyword matches the text and the first confidence level is greater than the first confidence threshold, then voice control can be performed on the electronic device, that is, the control command corresponding to the speech data to be detected is determined and executed; if the keyword matches the text and the first confidence level is not greater than the first confidence threshold, then voice control is not performed on the electronic device.

[0102] The first confidence threshold can be set before or after leaving the factory. It can be set by professionals after testing before leaving the factory, by the user after leaving the factory, or by the electronic device based on certain rules.

[0103] In an alternative embodiment, the method further includes:

[0104] If the keyword matches the text, the first confidence level is greater than the first confidence level threshold, and the difference between the first confidence level and the first confidence level threshold is greater than the difference threshold, then the historical confidence level corresponding to each historical speech data to be identified in the historical speech data to be identified obtained in the first historical time period is obtained.

[0105] The target historical confidence level is determined from the historical confidence levels. The target historical confidence level is the historical confidence level that is greater than the first confidence level threshold and the difference between the target historical confidence level and the first confidence level threshold is greater than the difference threshold.

[0106] Determine the first number of the first confidence level and the target historical confidence level, and determine the second number of the first confidence level and the historical confidence level;

[0107] If the first quantity is not less than the first quantity threshold, and the ratio of the first quantity to the second quantity is not less than the ratio threshold, then the first confidence threshold is increased.

[0108] Considering that different users speak in different ways, but the first confidence threshold is the same for the same batch of electronic devices, the first confidence threshold is usually set low in order to ensure a high recognition rate for keywords of users with relatively non-standard speech. However, in order to ensure a low misrecognition rate for keywords of users with relatively standard speech, the first confidence threshold can be increased during the actual voice control process of the user.

[0109] For example, if an electronic device acquires a significant number of speech data points within a first historical time period whose confidence levels exceed a first confidence threshold, and the difference between these confidence levels and the first confidence threshold is greater than a difference threshold, then the first confidence threshold can be increased. The first historical time period can be preset. For example, the historical time period could be within one month, half a month, or 10 days, etc.

[0110] For example, assuming the first quantity threshold is 75 and the ratio threshold is 4 / 5, an electronic device acquires 99 historical voice data points to be recognized within a month. Among these, 79 historical voice data points have a historical confidence level greater than the first confidence threshold, and the difference between the historical confidence level and the first confidence threshold is greater than the difference threshold. If the text matches a keyword, the first confidence level is greater than the first confidence threshold, and the difference between the first confidence level and the first confidence threshold is greater than the difference threshold. Therefore, the total number of the first confidence level and the target historical confidence level can be determined, i.e., the first quantity is 79 + 1 = 80. The total number of the first confidence level and the historical confidence level, i.e., the second quantity, is 99 + 1 = 100. The ratio of the first quantity to the second quantity is 4 / 5. The electronic device can then increase the first confidence threshold. The first quantity threshold and the ratio threshold can be preset.

[0111] In an optional embodiment, increasing the first confidence threshold includes:

[0112] Based on the first confidence level and the target historical confidence level, increase the first confidence level threshold.

[0113] For example, an electronic device can determine the minimum confidence level from a first confidence level and the target historical confidence level, and set a first confidence level threshold that is less than the minimum confidence level, and the difference between the first confidence level and the minimum confidence level is less than a first difference threshold. The first difference threshold can be preset based on actual conditions.

[0114] In an alternative embodiment, the method further includes:

[0115] If the keyword matches the text and the first confidence level is not greater than the first confidence threshold, then the historical text corresponding to each historical speech data to be identified in the historical speech data to be identified obtained in the second historical time period is obtained.

[0116] If there is a target historical text that matches the text, and the number of target historical texts is not less than the second quantity threshold, then the first confidence threshold is reduced.

[0117] To ensure a low false recognition rate for keywords, the first confidence threshold is usually set relatively high. However, considering that different users speak differently, setting the first confidence threshold too high may result in a lower keyword recognition rate for users with relatively non-standard speech patterns. This could lead to users repeatedly uttering the same keyword within a certain period without it being recognized by the electronic device. To improve the keyword recognition rate for these users, the first confidence threshold can be lowered during their actual voice control process.

[0118] For example, if an electronic device repeatedly acquires historical text corresponding to the voice data to be recognized within a second historical time period, and all of these texts match the target text, then the first confidence threshold can be reduced. The second historical time period can be preset. For example, the second historical time period could be within 5 minutes, within 10 minutes, within 30 minutes, etc.

[0119] For example, assuming the second quantity threshold is 9, and an electronic device acquires a text matching the keyword corresponding to the voice data to be recognized at the 10th minute, and the confidence level (first confidence level) of the voice data to be recognized is not greater than the first confidence threshold, and among the historical voice data to be recognized acquired in the previous 9 minutes, 9 historical texts corresponding to historical voice data to be recognized match the keyword, and the number of historical texts corresponding to these 9 historical voice data to be recognized is not less than the second quantity threshold, then it can be determined that the user may have said the same keyword multiple times within these 10 minutes, but because the confidence threshold may have been set too high, it was not recognized by the electronic device. Therefore, the electronic device can lower the first confidence threshold so that it can recognize the keyword when the user says it again in the future. The second quantity threshold can be preset.

[0120] Specifically, a match is determined when both texts contain the same keyword. For example, if one text contains the keyword "make a phone call" and another text also contains the keyword "make a phone call," then the first text is considered a match for the second text.

[0121] In an alternative embodiment, reducing the confidence threshold includes:

[0122] Obtain the target historical confidence score corresponding to the target historical text;

[0123] Adjust the confidence threshold based on the target's historical confidence level and confidence level.

[0124] For example, an electronic device can determine the minimum confidence level from the existing confidence level and the target's historical confidence level, and set a confidence threshold lower than this minimum confidence level, with the difference between this threshold and the minimum confidence level being less than a second difference threshold. The second difference threshold can be preset based on actual circumstances.

[0125] In an optional embodiment, considering that a user may utter the same text multiple times within a certain period of time, and only a few texts have a confidence level lower than the first confidence level threshold, when the keyword matches the text and the first confidence level is not greater than the first confidence level threshold, the historical text and historical confidence level corresponding to each historical speech data to be recognized in the historical speech data to be recognized obtained in the second historical period can be obtained; if there is a target historical text that matches the text in the historical text, the historical confidence level corresponding to the target historical text is not greater than the first confidence level threshold, and the number of target historical texts is not less than the second number threshold, then the first confidence level threshold is reduced.

[0126] For example, suppose that an electronic device repeatedly acquires historical text corresponding to the voice data to be recognized within a second historical time period, and all of these texts match the target text, but the confidence level is less than the first confidence threshold, then the first confidence threshold can be reduced. The second historical time period can be preset. For example, the second historical time period could be within 5 minutes, within 10 minutes, within 30 minutes, etc.

[0127] For example, assuming the second quantity threshold is 9, and an electronic device acquires text corresponding to the speech data to be recognized at the 10th minute that matches the keyword, and the confidence level (first confidence level) of the speech data to be recognized is not greater than the first confidence threshold, and among the historical speech data to be recognized acquired in the previous 9 minutes, 9 historical texts corresponding to historical speech data to be recognized match the keyword, the historical confidence levels of these 9 historical texts corresponding to historical texts corresponding to historical speech data to be recognized are not greater than the first confidence threshold, and the number of historical texts corresponding to these 9 historical texts corresponding to historical speech data to be recognized is not less than the second quantity threshold, then it can be determined that the user said the same keyword multiple times within these 10 minutes, but because the confidence threshold was set too high, none of them were recognized by the electronic device. Therefore, the electronic device can reduce the first confidence threshold. The second quantity threshold can be preset.

[0128] It should be noted that when performing speech recognition on a certain speech data to be recognized, the resulting speech recognition result includes text and confidence score. The text can be used as the text corresponding to the speech data to be recognized, and the confidence score can be used as the confidence score corresponding to the speech data to be recognized, or the confidence score corresponding to the text.

[0129] For example, speech recognition is performed on historical speech data to be recognized, yielding historical speech recognition results. These results include historical text and historical confidence scores. The historical text is the historical text corresponding to the historical speech data to be recognized. The historical confidence score can be either the historical confidence score corresponding to the historical speech data to be recognized or the historical confidence score corresponding to the historical text.

[0130] In an optional embodiment, keyword detection is performed on the speech data to be detected to obtain keywords in the speech data, including:

[0131] The keyword detection engine performs keyword detection on the speech data to be detected, and obtains the keyword detection results, which include keywords and second confidence scores.

[0132] The speech data containing the keywords is obtained from the speech data to be detected, resulting in the speech data to be recognized, including:

[0133] If the second confidence level is not less than the second confidence level threshold and not greater than the third confidence level threshold, then the speech data containing the keyword is obtained from the speech data to be detected, and the speech data to be recognized is obtained.

[0134] To further reduce the power consumption of electronic devices, for the voice data to be detected with a second confidence level not less than the second confidence level threshold, if the second confidence level is relatively high, such as greater than the fourth confidence level threshold, a control command corresponding to the voice data to be detected can be directly determined and executed.

[0135] The fourth confidence threshold can be preset. For example, the fourth confidence threshold can be 0.9, 0.91, etc.

[0136] To reduce the false recognition rate, for the speech data to be detected with a second confidence level not less than the second confidence threshold, a step can be performed to obtain the speech data containing the keywords from the speech data to be detected when the second confidence level is relatively low, such as not greater than the third confidence threshold, in order to obtain the speech data to be recognized.

[0137] The third confidence threshold can be preset. The third confidence level is no greater than the fourth confidence threshold. For example, the third confidence threshold can be 0.7, 0.9, 0.91, etc.

[0138] In an optional embodiment, the third confidence thresholds for different keywords can be the same. For example, the third confidence threshold for different keywords can all be 0.9.

[0139] In an optional embodiment, considering that the confidence level output by the keyword detection engine is not the same for different keywords, such as the generally high confidence level output by the keyword detection engine for some keywords and the generally low confidence level output by the keyword detection engine for other keywords, this embodiment can set different confidence thresholds for different keywords. After obtaining the keyword and second confidence level corresponding to the speech data to be detected, the confidence threshold corresponding to the keyword of the speech data to be detected is determined as the third confidence threshold. Then, based on the second confidence level, the second confidence threshold and the third confidence threshold, it is determined whether to perform the step of obtaining the speech data containing the keyword from the speech data to be detected to obtain the speech data to be recognized.

[0140] For example, if the confidence level for a keyword or keywords is typically between 0.38 and 0.5, then the confidence threshold for that keyword or keywords can be set to 0.5. To further reduce the false positive rate, the confidence threshold for that keyword or keywords can also be set to a higher value, for example, 0.55.

[0141] For example, if the confidence level for a keyword or keywords is typically between 0.6 and 0.7, then the confidence threshold for that keyword or keywords can be set to 0.7. To further reduce the false positive rate, the confidence threshold for that keyword or keywords can also be set to a higher value, for example, 0.8.

[0142] In practical applications, if the determined third confidence threshold is less than the fourth confidence threshold, the step of obtaining the speech data containing the keywords from the speech data to be detected and obtaining the speech data to be recognized can be performed when the second confidence threshold is greater than the third confidence threshold but not greater than the fourth confidence threshold. Alternatively, the step of determining the control command corresponding to the speech data to be detected and executing the control command can be performed directly when the second confidence threshold is greater than the third confidence threshold but not greater than the fourth confidence threshold. The specific procedure depends on the actual requirements.

[0143] For example, if the determined third confidence threshold is less than the fourth confidence threshold, then when the second confidence threshold is greater than the third confidence threshold but not greater than the fourth confidence threshold, the remaining battery power and / or the remaining available resources of the electronic device's processor can be obtained; if the remaining battery power is not less than the battery power threshold and / or the remaining available resources are not less than the resource threshold, then the step of obtaining the speech data containing the keywords from the speech data to be detected to obtain the speech data to be recognized is executed; if the remaining battery power is less than the battery power threshold and / or the remaining available resources are less than the resource threshold, then the step of determining the control instruction corresponding to the speech data to be detected and executing the control instruction is executed.

[0144] In an optional embodiment, determining the control command corresponding to the voice data to be detected includes:

[0145] Speech recognition is performed on the speech data to be detected to obtain a second speech recognition result, which includes the text to be analyzed.

[0146] Perform semantic analysis on the text to be analyzed to obtain the semantic analysis results;

[0147] The control commands corresponding to the semantic analysis results are determined, and the control commands corresponding to the speech data to be detected are obtained.

[0148] Understandably, when keywords match text, the speech recognition engine can perform speech recognition on the speech data to be detected, obtaining a second speech recognition result, which includes the text to be analyzed. Then, the semantic analysis engine performs semantic analysis on the text to be recognized to obtain a semantic analysis result, and then determines the control instructions corresponding to the semantic analysis result.

[0149] Control commands may include, but are not limited to: playing the next song, playing a song, turning on an appliance, calling a person, or navigating to a location.

[0150] In an optional embodiment, performing speech recognition on the speech data to be detected to obtain a second speech recognition result includes:

[0151] If the keyword is a keyword of the target category, the local speech recognition engine will perform speech recognition on the speech data to be detected to obtain a second speech recognition result.

[0152] To avoid the leakage of user privacy, some keywords involving user privacy can be set as target category keywords in advance. For example, "make a phone call" and "send a text message" involve the acquisition of the user's address book. Since the user's address book is usually private data, the above keywords can be set as target category keywords. When the detected keyword is a target category keyword, only the local speech recognition engine is used to perform speech recognition on the speech data to be detected to obtain a second speech recognition result.

[0153] In an optional embodiment, performing speech recognition on the speech data to be detected to obtain a second speech recognition result includes:

[0154] If the keyword is not a keyword of the target category, then the speech data to be detected will be recognized by both the local speech recognition engine and the cloud speech recognition engine.

[0155] If the first candidate speech recognition result is obtained from the cloud speech recognition engine within the preset time period, the first candidate speech recognition result will be determined as the second speech recognition result.

[0156] If the first candidate speech recognition result is not obtained within the preset time period, the second candidate speech recognition result output by the local speech recognition engine will be determined as the second speech recognition result.

[0157] Due to limitations in storage space and processing resources of electronic devices, the accuracy of local cloud-based speech recognition engines is relatively lower than that of cloud-based speech recognition engines. Therefore, in this embodiment, when the detected keyword is not a keyword of the target category, both the cloud-based speech recognition engine and the local recognition engine can be used simultaneously to perform speech recognition on the speech data to be detected. Considering that interaction with the cloud-based speech recognition engine involves data communication, the cloud-based speech recognition engine usually does not promptly provide the first speech recognition result when network quality is poor. To ensure the response speed of voice control, in this embodiment, when the first candidate speech recognition result is obtained within the target time period, the first candidate speech recognition result can be determined as the second speech recognition result; when the first candidate speech recognition result is not received within the target time period, the second candidate speech recognition result output by the local speech recognition engine can be determined as the second speech recognition result.

[0158] The target time period can be set by the user or by the electronic device based on certain rules. To ensure the response speed of voice control, the target time period can be set to be relatively short. For example, the target time period can be 500 milliseconds, 1 second, etc.

[0159] In an optional embodiment, performing speech recognition on the speech data to be detected to obtain a second speech recognition result includes:

[0160] If the keyword is not a keyword of the target category, then determine whether the current network quality is below the quality threshold;

[0161] If the current network quality is lower than the quality threshold, the local speech recognition engine will perform speech recognition on the speech data to be detected to obtain a second speech recognition result.

[0162] If the current network quality is not lower than the quality threshold, the cloud-based speech recognition engine will perform speech recognition on the speech data to be detected to obtain a second speech recognition result.

[0163] Understandably, if the keyword is not a keyword of the target category, either a local speech recognition engine or a cloud-based speech recognition engine can be used for speech recognition. However, considering that network quality below the quality threshold will affect the speed at which the cloud-based recognition engine responds with speech recognition results, and thus affect the response speed of voice control, in this embodiment, if the keyword is not a keyword of the target category, when the current network quality is below the quality threshold, the local speech recognition engine is used to perform speech recognition on the speech data to be detected to obtain a second speech recognition result; when the current network quality is not below the quality threshold, the cloud-based speech recognition engine is used to perform speech recognition on the speech data to be detected to obtain a second speech recognition result.

[0164] Specifically, it can detect whether the current network quality is lower than a preset threshold based on network quality parameters.

[0165] Network quality parameters may include network signal strength, network transmission latency, or network packet loss rate. Based on this, the electronic device determines whether the current network quality is below a quality threshold using the following methods;

[0166] (1) Monitor in real time whether the network signal strength of the current network is lower than the signal strength threshold. When the network signal strength is detected to be lower than the signal strength threshold, determine that the current network quality is lower than the quality threshold.

[0167] (2) Monitor in real time whether the network transmission delay of the current network is greater than the network transmission delay threshold. When the network transmission delay is detected to be greater than the network transmission delay threshold, determine that the current network quality is lower than the quality threshold.

[0168] (3) Monitor in real time whether the current network packet loss rate is greater than the network packet loss rate threshold. When the network packet loss rate is detected to be greater than the network packet loss rate threshold, determine that the current network quality is lower than the quality threshold.

[0169] Among them, the signal strength threshold, network transmission delay threshold, and network packet loss rate threshold can be preset based on actual needs.

[0170] It should be noted that the methods for detecting whether the network quality is below the quality threshold are not limited to those described above.

[0171] In an optional embodiment, acquiring the speech data to be detected includes:

[0172] Determine the current use case;

[0173] If the current usage scenario is a preset usage scenario, the wake-up-free function will be activated and the voice data to be detected will be obtained. The preset usage scenarios include at least one of the following: driving scenario, audio playback scenario, call scenario, and shooting scenario.

[0174] The current usage scenario can be determined by the functions of the electronic device being used by the current application. For example, if you are using location services through navigation software, the current usage scenario is a driving scenario. Another example is if certain applications enable the electronic device's speaker for audio playback; these applications could be Kugou Music, QQ Music, NetEase Cloud Music, Kuwo Music, Douyin, Kuaishou, Tencent Video, Youku Video, iQiyi Video, etc. Yet another example is if certain applications enable the electronic device to connect to the network and activate the microphone for network communication; this indicates a call scenario, and the applications could be phone calls, WeChat, QQ, etc. Finally, if certain applications enable the electronic device's camera for image capture, this indicates a shooting scenario, and the applications could be the electronic device's built-in camera, WeChat, Alipay, etc.

[0175] Understandably, the above-mentioned preset usage scenarios are merely examples, intended to illustrate that when an electronic device enters such preset usage scenarios, the wake-up-free function can be automatically activated, thereby avoiding the need for users to manually operate the wake-up-free function, making it more convenient for users, and also improving the activation efficiency of the wake-up-free function.

[0176] It should also be noted that if the wake-up-free function has already been activated before the current usage scenario is detected, the wake-up-free function will not be activated again here. The wake-up-free function will only be activated if it has been deactivated before this.

[0177] In an optional embodiment, if the current usage scenario is an audio playback scenario, before performing keyword detection on the speech data to be detected to obtain the keywords in the speech data to be detected, the method further includes:

[0178] Get the audio data being played;

[0179] Echo cancellation is performed on the audio data to obtain the target speech data to be detected.

[0180] Keyword detection is performed on the speech data to be detected to obtain the keywords in the speech data, including:

[0181] Keyword detection is performed on the target speech data to be detected to obtain the keywords in the speech data;

[0182] Obtain the speech data containing the keywords from the speech data to be detected, including:

[0183] Obtain the speech data containing the keywords from the target speech data to be detected.

[0184] In this embodiment, echo cancellation can be performed on the voice data in the audio playback scenario to filter out the audio data played by the electronic device from the voice data, so that the voice data after echo cancellation no longer contains the audio data played by the current electronic device. This method can avoid the interference of the audio data played by the electronic device on the voice data during the voice recognition process, which is conducive to improving the accuracy of voice recognition.

[0185] In particular, when performing echo cancellation, the number of echo channels can be made consistent with the number of channels for the audio data played by the speaker, which helps to eliminate echoes in the speech data.

[0186] In some embodiments, if the current usage scenario is a driving scenario or a call scenario, before performing keyword detection on the voice data to be detected to obtain the keywords in the voice data to be detected, the method further includes: performing noise reduction processing on the voice data to obtain noise-reduced voice data; performing keyword detection on the voice data to be detected to obtain the keywords in the voice data to be detected, including: performing keyword detection on the noise-reduced voice data to obtain the keywords in the noise-reduced voice data; and obtaining the voice data containing the keywords from the voice data to be detected to obtain the voice data to be recognized, including: obtaining the voice data containing the keywords from the noise-reduced voice data to obtain the voice data to be recognized.

[0187] In this embodiment, by performing noise reduction processing on the voice data, external environmental noise in the voice data can be filtered out, which helps to improve the accuracy of the voice data when performing voice recognition.

[0188] External environmental noise can include, for example, car horns and engine sounds in a driving scenario, or other noises from other users and the environment in a call scenario.

[0189] In an optional embodiment, if the current usage scenario changes and the changed usage scenario is not the preset usage scenario, the wake-up-free function is disabled. A change in the current usage scenario may refer to the electronic device exiting the current usage scenario or other applications starting on the electronic device. When the current usage scenario changes, it can also be determined whether the changed usage scenario is the preset usage scenario. If so, the wake-up-free function remains enabled; otherwise, it is disabled.

[0190] In an alternative embodiment, the wake-up-free function can also be manually disabled by the user.

[0191] Understandably, when the wake-up-free function is turned off, the electronic device can disable the audio acquisition function, local speech recognition engine, cloud speech recognition engine, semantic analysis engine, etc. The specific implementation method can be selected by those skilled in the art according to actual needs, and is not limited here.

[0192] Please refer to the following: Figure 2 and Figure 3 , Figure 2 This is a schematic diagram of the second type of voice control method provided in the embodiments of this application. Figure 3 This is a schematic diagram of a scenario for the voice control method provided in this application embodiment. The process may include:

[0193] In step 201, the speech data to be detected is acquired, and the keyword detection engine is used to detect keywords in the speech data to obtain the keywords in the speech data to be detected.

[0194] In 202, the target duration is determined based on the number of words in the keywords, and the target duration is positively correlated with the number of words.

[0195] In 203, determine the end time of the speech data containing the last keyword in the keyword list.

[0196] In step 204, the speech data with the target duration before the end time in the speech data to be detected is identified as the speech data containing the keyword, thus obtaining the speech data to be recognized.

[0197] In step 205, the speech recognition engine performs speech recognition on the speech data to be recognized to obtain a first speech recognition result, which includes text and a first confidence level.

[0198] In step 206, if the keyword matches the text and the first confidence level is greater than the first confidence level threshold, then the speech recognition engine performs speech recognition on the speech data to be detected to obtain a second speech recognition result, which includes the text to be analyzed.

[0199] In version 207, the semantic analysis engine performs semantic analysis on the text to be analyzed, and obtains the semantic analysis results.

[0200] In step 208, the control instructions corresponding to the semantic analysis results are determined, the control instructions corresponding to the speech data to be detected are obtained, and the control instructions are executed.

[0201] The above process will be illustrated with specific examples below.

[0202] For example, assume that the user says to the electronic device: "Navigate to xx Restaurant", and the electronic device obtains the voice data to be detected. Then, the electronic device performs keyword detection on the voice data to be detected through a keyword detection engine and obtains the keyword "daohangdao".

[0203] Assume that the duration of the voice data to be detected is 3 seconds, the end time of the last keyword "dao" in the keyword "Navigate to" is the 1.5th second of the voice data to be detected, and the target duration is 1.2 seconds. Then, 1.2 seconds of voice data before the 1.5th second in the voice data to be detected can be obtained as the voice data to be recognized. After obtaining the voice data to be recognized, the electronic device performs voice recognition on the voice data to be recognized through a voice recognition engine and obtains a first voice recognition result. The first voice recognition result includes text and a first confidence level. Assume that the text is "daohangdao", the first confidence level is 0.9, and the first confidence level threshold is 0.85. It can be determined that the keyword matches the text and the first confidence level is greater than the first confidence level threshold. Then, the electronic device can perform voice recognition on the voice data to be detected through the voice recognition engine and obtain a second voice recognition result. The second voice recognition result includes the text to be analyzed.

[0204] Assume that the text to be analyzed is "Navigate to xx Restaurant". Then, the electronic device performs voice analysis on the text to be analyzed through a semantic analysis engine and obtains a semantic analysis result. It can be determined that the semantic analysis result indicates navigating to xx Restaurant. Then, the electronic device can determine that the control instruction is "Navigate to xx Restaurant". Then, the electronic device can navigate to xx Restaurant through the installed navigation application and display the navigation route to the destination.

[0205] ' Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the voice control device provided by an embodiment of the present application. The voice control device 300 includes: a keyword detection module 301, a data acquisition module 302, a voice recognition module 303, and an instruction execution module 304.

[0206] The keyword detection module 301 is configured to obtain the voice data to be detected and perform keyword detection on the voice data to be detected to obtain the keyword in the voice data to be detected.

[0207] The data acquisition module 302 is configured to obtain the voice data where the keyword is located from the voice data to be detected to obtain the voice data to be recognized.

[0208] The voice recognition module 303 is configured to perform voice recognition on the voice data to be recognized to obtain a first voice recognition result.

[0209] The instruction execution module 304 is used to determine the control instruction corresponding to the speech data to be detected if the keyword matches the first speech recognition result, and to execute the control instruction.

[0210] In an optional embodiment, the data acquisition module 302 may be used to: determine the end time of the speech data containing the last keyword in the keywords; and acquire the speech data to be recognized from the speech data to be detected based on the end time.

[0211] In an optional embodiment, the data acquisition module 302 may be used to: determine a target duration based on the number of characters in the keyword, wherein the target duration is positively correlated with the number of characters; and acquire the speech data of the target duration from the speech data to be detected as the speech data to be identified based on the end time.

[0212] In an optional embodiment, the data acquisition module 302 can be used to: perform speech rate detection on the speech data to be detected to obtain the speech rate corresponding to the speech data to be detected; determine a target duration based on the number of words and the speech rate, wherein the target duration is inversely correlated with the speech rate.

[0213] In an optional embodiment, the data acquisition module 302 may be used to: determine the corresponding voice data in the voice data to be detected before the end time as the voice data to be recognized.

[0214] In an optional embodiment, the data acquisition module 302 may be used to: determine the corresponding voice data before the end time and the corresponding voice data after the end time in the voice data to be detected as the voice data to be identified.

[0215] In an optional embodiment, the first speech recognition result includes text, and the instruction execution module 304 can be used to: if the keyword matches the text and the confidence level is greater than the confidence level threshold, then determine the control instruction corresponding to the speech data to be detected.

[0216] In an optional embodiment, the first speech recognition result further includes a first confidence level. The instruction execution module 304 can be used to: if the keyword matches the text and the first confidence level is greater than the first confidence level threshold, then determine the control instruction corresponding to the speech data to be detected.

[0217] In an optional embodiment, the voice control device 300 may further include a threshold adjustment module, which may be used to: if the keyword matches the text, the first confidence level is greater than a first confidence level threshold, and the difference between the first confidence level and the first confidence level threshold is greater than a difference threshold, then obtain the historical confidence level corresponding to each historical voice data to be recognized in the historical voice data to be recognized obtained in the first historical time period; determine a target historical confidence level from the historical confidence levels, wherein the target historical confidence level is a historical confidence level greater than the first confidence level threshold and the difference between the target historical confidence level and the first confidence level threshold is greater than the difference threshold; determine a first quantity of the first confidence level and the target historical confidence level, and determine a second quantity of the first confidence level and the historical confidence level; if the first quantity is not less than a first quantity threshold, and the ratio of the first quantity to the second quantity is not less than a ratio threshold, then increase the first confidence level threshold.

[0218] In an optional embodiment, the threshold adjustment module can be used to: if the keyword matches the text and the first confidence level is not greater than the first confidence level threshold, then obtain the historical text corresponding to each historical speech data to be identified in the historical speech data to be identified obtained in the second historical time period; if there is a target historical text matching the text in the historical text and the number of the target historical texts is not less than the second quantity threshold, then reduce the first confidence level threshold.

[0219] In an optional embodiment, the keyword detection module 301 can be used to: perform keyword detection on the speech data to be detected through a keyword detection engine to obtain a keyword detection result, wherein the keyword detection result includes keywords and a second confidence level;

[0220] The data acquisition module 302 can be used to: if the second confidence level is not less than the second confidence level threshold and not greater than the third confidence level threshold, then acquire the speech data containing the keyword from the speech data to be detected, and obtain the speech data to be recognized.

[0221] In an optional embodiment, the instruction execution module 304 may be used to: perform speech recognition on the speech data to be detected to obtain a second speech recognition result, the second speech recognition result including the text to be analyzed; perform semantic analysis on the text to be analyzed to obtain a semantic analysis result; determine a control instruction corresponding to the semantic analysis result to obtain a control instruction corresponding to the speech data to be detected.

[0222] In an optional embodiment, the instruction execution module 304 may be used to: if the keyword is a keyword of the target category, perform speech recognition on the speech data to be detected through a local speech recognition engine to obtain a second speech recognition result.

[0223] In an optional embodiment, the instruction execution module 304 can be used to: if the keyword is not a keyword of the target category, perform speech recognition on the speech data to be detected through a local speech recognition engine and a cloud speech recognition engine; if a first candidate speech recognition result output by the cloud speech recognition engine is obtained within a preset time period, determine the first candidate speech recognition result as the second speech recognition result; if the first candidate speech recognition result is not obtained within the preset time period, determine the second candidate speech recognition result output by the local speech recognition engine as the second speech recognition result.

[0224] In an optional embodiment, the instruction execution module 304 can be used to: if the keyword is not a keyword of the target category, determine whether the current network quality is lower than a quality threshold; if the current network quality is lower than the quality threshold, perform speech recognition on the speech data to be detected through a local speech recognition engine to obtain a second speech recognition result; if the current network quality is not lower than the quality threshold, perform speech recognition on the speech data to be detected through a cloud speech recognition engine to obtain a second speech recognition result.

[0225] In an optional embodiment, the keyword detection module 301 can be used to: determine the current usage scenario; if the current usage scenario is a preset usage scenario, then activate the wake-up-free function and acquire the voice data to be detected, wherein the preset usage scenario includes at least one of driving scenario, audio playback scenario, call scenario and shooting scenario.

[0226] In an optional embodiment, if the current usage scenario is an audio playback scenario, the keyword detection module 301 can be used to: acquire the playing audio data; perform echo cancellation on the speech data to be detected based on the audio data to obtain target speech data to be detected; and perform keyword detection on the target speech data to be detected to obtain the keywords in the target speech data to be detected.

[0227] The data acquisition module 302 can be used to: acquire the speech data containing the keyword from the target speech data to be detected.

[0228] This application provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed on a computer, it causes the computer to perform the voice control method provided in this embodiment.

[0229] This application also provides an electronic device, including a memory and a processor, wherein the processor executes the voice control method provided in this embodiment by calling a computer program stored in the memory.

[0230] For example, the aforementioned electronic device could be a mobile terminal such as a tablet or smartphone. See also... Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0231] The electronic device 400 may include components such as a processor 401 and a memory 402. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, electronic device 400 may also include a microphone.

[0232] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device through various interfaces and lines. By running or executing the application program stored in the memory 402 and calling the data stored in the memory 402, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0233] Memory 402 can be used to store applications and data. The applications stored in memory 402 contain executable code. Applications can be composed of various functional modules. Processor 401 executes various functional applications and data processing by running the applications stored in memory 402.

[0234] In this embodiment, the processor 401 in the electronic device loads the executable code corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402, thereby achieving:

[0235] Acquire the speech data to be detected, and perform keyword detection on the speech data to be detected to obtain the keywords in the speech data to be detected;

[0236] The speech data containing the keyword is obtained from the speech data to be detected, thus obtaining the speech data to be recognized;

[0237] The speech data to be recognized is subjected to speech recognition to obtain a first speech recognition result;

[0238] If the keyword matches the first speech recognition result, a control instruction corresponding to the speech data to be detected is determined and the control instruction is executed.

[0239] In an optional embodiment, when the processor 401 executes the step of obtaining the speech data containing the keyword from the speech data to be detected to obtain the speech data to be recognized, it may perform the following: determining the end time of the speech data containing the last keyword; and obtaining the speech data to be recognized from the speech data to be detected based on the end time.

[0240] In an optional embodiment, before the processor 401 executes the step of obtaining the speech data to be recognized from the speech data to be detected based on the end time, it may further execute: determining a target duration based on the number of characters in the keyword, wherein the target duration is positively correlated with the number of characters; when the processor 401 executes the step of obtaining the speech data to be recognized from the speech data to be detected based on the end time, it may execute: obtaining speech data of the target duration from the speech data to be detected as the speech data to be recognized based on the end time.

[0241] In an optional embodiment, before the processor 401 executes the step of determining the target duration based on the number of characters in the keyword, it may further execute: performing speech rate detection on the speech data to be detected to obtain the speech rate corresponding to the speech data to be detected; when the processor 401 executes the step of determining the target duration based on the number of characters in the keyword, it may execute: determining the target duration based on the number of characters and the speech rate, wherein the target duration is inversely correlated with the speech rate.

[0242] In an optional embodiment, when the processor 401 performs the step of obtaining the voice data to be recognized from the voice data to be detected according to the end time, it may perform the following: determining the corresponding voice data in the voice data to be detected before the end time as the voice data to be recognized.

[0243] In an optional embodiment, when the processor 401 performs the step of obtaining the voice data to be recognized from the voice data to be detected according to the end time, it may perform the following: determining the corresponding voice data before the end time and the corresponding voice data after the end time in the voice data to be detected as the voice data to be recognized.

[0244] In an optional embodiment, the first speech recognition result includes text. When the processor 401 executes the control instruction that determines the corresponding speech data to be detected if the keyword matches the first speech recognition result, it may execute: if the keyword matches the text, determine the control instruction corresponding to the speech data to be detected.

[0245] In an optional embodiment, the first speech recognition result further includes a first confidence level. When the processor 401 executes the control instruction that determines the corresponding speech data to be detected if the keyword matches the text, it may execute: if the keyword matches the text and the first confidence level is greater than the first confidence level threshold, then determine the control instruction corresponding to the speech data to be detected.

[0246] In an optional embodiment, the processor 401 may further perform the following: if the keyword matches the text, the first confidence level is greater than a first confidence level threshold, and the difference between the first confidence level and the first confidence level threshold is greater than a difference threshold, then obtain the historical confidence level corresponding to each historical speech data to be recognized in the historical speech data to be recognized obtained within the first historical time period; determine a target historical confidence level from the historical confidence levels, wherein the target historical confidence level is a historical confidence level greater than the first confidence level threshold and the difference between the target historical confidence level and the first confidence level threshold is greater than the difference threshold; determine a first quantity of the first confidence level and the target historical confidence level, and determine a second quantity of the first confidence level and the historical confidence level; if the first quantity is not less than a first quantity threshold, and the ratio of the first quantity to the second quantity is not less than a ratio threshold, then increase the first confidence level threshold.

[0247] In an optional embodiment, the processor 401 may further perform the following: if the keyword matches the text and the first confidence level is not greater than the first confidence level threshold, then obtain the historical text corresponding to each of the historical speech data to be identified in the historical speech data to be identified obtained in the second historical time period; if there is a target historical text that matches the text in the historical text and the number of the target historical text is not less than the second quantity threshold, then reduce the first confidence level threshold.

[0248] In an optional embodiment, when the processor 401 performs keyword detection on the speech data to be detected to obtain keywords in the speech data to be detected, it may perform the following: perform keyword detection on the speech data to be detected through a keyword detection engine to obtain keyword detection results, the keyword detection results including keywords and a second confidence level; when the processor 401 performs the step of obtaining the speech data containing the keywords from the speech data to be detected to obtain speech data to be recognized, it may perform the following: if the second confidence level is not less than a second confidence threshold and not greater than a third confidence threshold, then obtain the speech data containing the keywords from the speech data to be detected to obtain speech data to be recognized.

[0249] In an optional embodiment, when the processor 401 executes the control instruction corresponding to the speech data to be detected, it may perform the following: perform speech recognition on the speech data to be detected to obtain a second speech recognition result, the second speech recognition result including the text to be analyzed; perform semantic analysis on the text to be analyzed to obtain a semantic analysis result; determine the control instruction corresponding to the semantic analysis result to obtain the control instruction corresponding to the speech data to be detected.

[0250] In an optional embodiment, when the processor 401 performs speech recognition on the speech data to be detected to obtain a second speech recognition result, it may perform the following: if the keyword is a keyword of the target category, then perform speech recognition on the speech data to be detected through the local speech recognition engine to obtain a second speech recognition result.

[0251] In an optional embodiment, when the processor 401 performs speech recognition on the speech data to be detected to obtain a second speech recognition result, it may perform the following: if the keyword is not a keyword of the target category, then perform speech recognition on the speech data to be detected using a local speech recognition engine and a cloud speech recognition engine; if a first candidate speech recognition result output by the cloud speech recognition engine is obtained within a preset time period, then the first candidate speech recognition result is determined as the second speech recognition result; if the first candidate speech recognition result is not obtained within the preset time period, then the second candidate speech recognition result output by the local speech recognition engine is determined as the second speech recognition result.

[0252] In an optional embodiment, when the processor 401 performs speech recognition on the speech data to be detected to obtain a second speech recognition result, it may perform the following: if the keyword is not a keyword of the target category, determine whether the current network quality is lower than a quality threshold; if the current network quality is lower than the quality threshold, perform speech recognition on the speech data to be detected through a local speech recognition engine to obtain a second speech recognition result; if the current network quality is not lower than the quality threshold, perform speech recognition on the speech data to be detected through a cloud-based speech recognition engine to obtain a second speech recognition result.

[0253] In an optional embodiment, when the processor 401 executes the acquisition of voice data to be detected, it may perform the following: determine the current usage scenario; if the current usage scenario is a preset usage scenario, then activate the wake-up-free function and acquire the voice data to be detected, wherein the preset usage scenario includes at least one of driving scenario, audio playback scenario, call scenario and shooting scenario.

[0254] In an optional embodiment, if the current usage scenario is an audio playback scenario, before the processor 401 performs keyword detection on the speech data to be detected to obtain the keywords in the speech data to be detected, it may also perform: acquiring the playing audio data; performing echo cancellation on the speech data to be detected based on the audio data to obtain the target speech data to be detected; when the processor 401 performs keyword detection on the speech data to be detected to obtain the keywords in the speech data to be detected, it may perform: performing keyword detection on the target speech data to be detected to obtain the keywords in the target speech data to be detected; when the processor 401 performs the step of obtaining the speech data containing the keywords from the speech data to be detected, it may perform: obtaining the speech data containing the keywords from the target speech data to be detected.

[0255] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed description of the voice control method above, which will not be repeated here.

[0256] The voice control device provided in this application embodiment belongs to the same concept as the voice control method in the above embodiment. Any of the methods provided in the voice control method embodiment can be run on the voice control device. For details of its implementation process, please refer to the voice control method embodiment, which will not be repeated here.

[0257] It should be noted that, for the voice control method of this application embodiment, those skilled in the art will understand that all or part of the process of implementing the voice control method of this application embodiment can be accomplished by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium, such as a memory, and executed by at least one processor. During execution, it can include the process of the embodiment of the voice control method. The computer-readable storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), etc.

[0258] It is understood that in the specific implementation of this application, user information, such as application usage behavior data, logs and other related data, is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0259] For the voice control device of this application embodiment, its functional modules can be integrated into a processing chip, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0260] The above provides a detailed description of a voice control method, apparatus, storage medium, and electronic device provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A voice control method, characterized by, The method comprises: obtaining to-be-detected voice data, and performing keyword detection on the to-be-detected voice data to obtain a keyword in the to-be-detected voice data, comprising: performing keyword detection on the to-be-detected voice data by a keyword detection engine to obtain a keyword detection result, the keyword detection result comprising a keyword and a second confidence level; obtaining voice data in which the keyword is located from the to-be-detected voice data to obtain to-be-recognized voice data, comprising: if the second confidence level is not less than a second confidence level threshold and not greater than a third confidence level threshold, obtaining voice data in which the keyword is located from the to-be-detected voice data to obtain to-be-recognized voice data; performing voice recognition on the to-be-recognized voice data to obtain a first voice recognition result; if the keyword matches the first voice recognition result, determining a control instruction corresponding to the to-be-detected voice data, and executing the control instruction.

2. The voice control method of claim 1, wherein, The method of obtaining voice data in which the keyword is located from the to-be-detected voice data to obtain to-be-recognized voice data comprises: determining an end time of voice data in which a last keyword in the keyword is located; obtaining the to-be-recognized voice data from the to-be-detected voice data according to the end time.

3. The voice control method of claim 2, wherein, Before the step of obtaining the to-be-recognized voice data from the to-be-detected voice data according to the end time, the method further comprises: determining a target time length according to a number of words of the keyword, the target time length being positively correlated with the number of words. The step of obtaining the to-be-recognized voice data from the to-be-detected voice data according to the end time comprises: obtaining voice data of the target time length from the to-be-detected voice data as the to-be-recognized voice data according to the end time.

4. The voice control method of claim 3, wherein, Before the step of determining a target time length according to a number of words of the keyword, the method further comprises: performing speech rate detection on the to-be-detected voice data to obtain a speech rate corresponding to the to-be-detected voice data; The step of determining a target time length according to a number of words of the keyword comprises: determining the target time length according to the number of words and the speech rate, the target time length being inversely correlated with the speech rate.

5. The voice control method of claim 2, wherein, The step of obtaining the to-be-recognized voice data from the to-be-detected voice data according to the end time comprises: determining corresponding voice data of the to-be-detected voice data before the end time as the to-be-recognized voice data.

6. The voice control method of claim 2, wherein, The step of obtaining the to-be-recognized voice data from the to-be-detected voice data according to the end time comprises: determining corresponding voice data of the to-be-detected voice data before the end time and corresponding voice data of the to-be-detected voice data after the end time as the to-be-recognized voice data.

7. The voice control method of claim 1, wherein, The first voice recognition result comprises text, and the step of determining a control instruction corresponding to the to-be-detected voice data if the keyword matches the first voice recognition result comprises: determining a control instruction corresponding to the to-be-detected voice data if the keyword matches the text.

8. The voice control method of claim 7, wherein, The first voice recognition result further comprises a first confidence level, and the step of determining a control instruction corresponding to the to-be-detected voice data if the keyword matches the text comprises: If the keyword matches the text and the first confidence is greater than a first confidence threshold, a control instruction corresponding to the voice data to be detected is determined.

9. The voice control method of claim 8, wherein, The method further comprises: If the keyword matches the text, the first confidence is greater than the first confidence threshold, and a difference between the first confidence and the first confidence threshold is greater than a difference threshold, a historical confidence corresponding to each of historical voice data to be recognized obtained in a first historical period is obtained. A target historical confidence is determined from the historical confidence, the target historical confidence being a historical confidence greater than the first confidence threshold and having a difference from the first confidence threshold greater than the difference threshold. A first number of the first confidence and the target historical confidence is determined, and a second number of the first confidence and the historical confidence is determined. If the first number is not less than a first number threshold and a ratio of the first number to the second number is not less than a ratio threshold, the first confidence threshold is increased.

10. The voice control method of claim 8, wherein, The method further comprises: If the keyword matches the text and the first confidence is not greater than the first confidence threshold, a historical text corresponding to each of historical voice data to be recognized obtained in a second historical period is obtained. If there is a target historical text matching the text in the historical text and a number of the target historical text is not less than a second number threshold, the first confidence threshold is decreased.

11. The voice control method of claim 1, wherein, The determination of the control instruction corresponding to the voice data to be detected comprises: Voice recognition is performed on the voice data to be detected to obtain a second voice recognition result, the second voice recognition result comprising text to be analyzed; Semantic analysis is performed on the text to be analyzed to obtain a semantic analysis result; A control instruction corresponding to the semantic analysis result is determined to obtain the control instruction corresponding to the voice data to be detected.

12. The voice control method of claim 11, wherein, The voice recognition on the voice data to be detected to obtain the second voice recognition result comprises: If the keyword is a keyword of a target category, the voice recognition on the voice data to be detected is performed by a local voice recognition engine to obtain the second voice recognition result; wherein a keyword related to user privacy is set as a keyword of the target category.

13. The voice control method of claim 11, wherein, The voice recognition on the voice data to be detected to obtain the second voice recognition result comprises: If the keyword is not a keyword of a target category, the voice recognition on the voice data to be detected is performed by a local voice recognition engine and a cloud voice recognition engine; wherein a keyword related to user privacy is set as a keyword of the target category; If a first candidate voice recognition result output by the cloud voice recognition engine is obtained within a preset period, the first candidate voice recognition result is determined as the second voice recognition result; If the first candidate voice recognition result is not obtained within the preset period, a second candidate voice recognition result output by the local voice recognition engine is determined as the second voice recognition result.

14. The voice control method of claim 11, wherein, The voice recognition on the to-be-detected voice data is performed to obtain a second voice recognition result, and the second voice recognition result comprises: If the keyword is not a keyword of a target category, it is determined whether the current network quality is lower than a quality threshold; wherein the keyword related to user privacy is set as the keyword of the target category; If the current network quality is lower than the quality threshold, the to-be-detected voice data is recognized by a local voice recognition engine to obtain a second voice recognition result; If the current network quality is not lower than the quality threshold, the to-be-detected voice data is recognized by a cloud voice recognition engine to obtain a second voice recognition result.

15. The voice control method of any one of claims 1 to 14, characterized in that, The to-be-detected voice data is obtained, and the obtaining comprises: A current use scenario is determined; If the current use scenario is a preset use scenario, an awakening-free function is started, and to-be-detected voice data is obtained, and the preset use scenario comprises at least one of a driving scenario, an audio playing scenario, a call scenario, and a shooting scenario.

16. The voice control method of claim 15, wherein, If the current use scenario is the audio playing scenario, before the keyword detection on the to-be-detected voice data is performed to obtain a keyword in the to-be-detected voice data, the method further comprises: Audio data being played is obtained; The to-be-detected voice data is subjected to echo cancellation according to the audio data to obtain target to-be-detected voice data; The keyword detection on the to-be-detected voice data to obtain the keyword in the to-be-detected voice data comprises: The keyword detection on the target to-be-detected voice data is performed to obtain a keyword in the target to-be-detected voice data; The voice data in which the keyword is located is obtained from the target to-be-detected voice data. The method comprises:

17. A voice control device, comprising: A keyword detection module is configured to obtain to-be-detected voice data, and perform keyword detection on the to-be-detected voice data to obtain a keyword in the to-be-detected voice data, comprising: performing keyword detection on the to-be-detected voice data by a keyword detection engine to obtain a keyword detection result, wherein the keyword detection result comprises the keyword and a second confidence degree; A data obtaining module is configured to obtain voice data in which the keyword is located from the to-be-detected voice data to obtain to-be-recognized voice data, comprising: if the second confidence degree is not less than a second confidence degree threshold and not greater than a third confidence degree threshold, obtaining voice data in which the keyword is located from the to-be-detected voice data to obtain to-be-recognized voice data; A voice recognition module is configured to perform voice recognition on the to-be-recognized voice data to obtain a first voice recognition result; An instruction execution module is configured to, if the keyword matches the first voice recognition result, determine a control instruction corresponding to the to-be-detected voice data, and execute the control instruction. The storage medium stores a computer program, and when the computer program runs on a computer, the computer executes the voice control method in any one of claims 1 to 16.

18. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program runs on a computer, the computer executes the voice control method in any one of claims 1 to 16.

19. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory storing a computer program, and the processor is configured to execute the voice control method according to any one of claims 1-16 by invoking the computer program stored in the memory.

20. A computer program product, characterised in that, The computer program product stores a computer program, and the computer program is adapted to be loaded by a processor to execute the method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Speech recognition apparatus, method and electronic equipment

    CN104978963A

  • Voice recognition method and device, electronic equipment and storage medium

    CN111862943A

  • Voice control method and device, storage medium and electronic equipment

    CN115472156A