Method for checking wake-up voice, electronic device and computer readable storage medium

By performing speech recognition and pinyin matching on the wake-up voice, the accuracy problem of wake-up voice recognition was solved, ensuring the reliability of voice interaction functions and improving the user experience.

CN119274553BActive Publication Date: 2026-07-24NIO TECH ANHUI CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NIO TECH ANHUI CO LTD
Filing Date
2024-09-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify and verify whether a wake-up voice is the correct one, impacting the wake-up capability of voice interaction functions and leading to a decline in user experience.

Method used

The wake-up voice is recognized by speech recognition, converted into Chinese text and its pinyin is obtained. The pinyin of the wake-up word is matched with the pinyin of the preset Chinese wake-up word. Multi-level matching is performed considering the user's pronunciation habits and the similarity of pinyin to determine the correctness of the wake-up voice.

Benefits of technology

It achieves accurate recognition of wake-up voice, avoids the problem of inaccurate model testing and voice recognition method verification in existing technologies, and improves the reliability of voice interaction functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274553B_ABST
    Figure CN119274553B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of speech recognition processing, and particularly provides a checking method for a wake-up speech, an electronic device and a computer readable storage medium, and aims to solve the problem of accurately checking whether the wake-up speech is a correct wake-up speech. To achieve the purpose, the method provided by the application comprises the following steps: performing speech recognition on the wake-up speech to obtain text information of the wake-up speech, wherein the text information comprises Chinese text; obtaining a first pinyin of the text information according to the Chinese text; obtaining a preset Chinese wake-up word and converting the Chinese wake-up word into a second pinyin; matching the first pinyin with the second pinyin; if the matching is successful, it is determined that the wake-up speech is a correct wake-up speech; and if the matching fails, it is determined that the wake-up speech is an incorrect wake-up speech. Based on the above method, the wake-up speech does not need to be verified by using a speech wake-up model, and the semantics of the text information does not need to be considered, so that whether the wake-up speech is a correct wake-up speech can be determined conveniently and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition and processing technology, specifically to a method for checking wake-up speech, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With the continuous development of artificial intelligence technology, more and more electronic devices (such as mobile phones) are equipped with voice interaction functions. Users can control electronic devices by voice, reducing the need to press buttons and greatly improving the convenience of using electronic devices. For example, mobile phones have weather forecast software installed, and users can control the electronic devices to automatically play weather forecasts by voice.

[0003] For electronic devices equipped with voice interaction capabilities, users can pre-set a wake-up word for the voice interaction function. When a user wants to use the voice interaction function, they can speak the wake-up voice containing the wake-up word to the electronic device. The electronic device can capture this wake-up voice and use a voice wake-up model (hereinafter referred to as the first voice wake-up model) to identify whether the wake-up voice contains the wake-up word. If the wake-up word is present, the voice interaction function is activated. Only after the voice interaction function is activated can the user control the electronic device normally via voice. In this process, the recognition of the wake-up word is crucial. If the wake-up word cannot be accurately recognized, the voice interaction function cannot be activated, and the user will not be able to control the electronic device via voice, affecting the user experience. Therefore, it is necessary to accurately test or evaluate the ability to activate the voice interaction function (hereinafter referred to as voice wake-up capability).

[0004] Currently, the main methods for testing voice wake-up capabilities include model testing methods and speech recognition methods. These two methods will be explained below.

[0005] The model testing method involves acquiring a more complex and higher-performing voice wake-up model than the first voice wake-up model (hereinafter referred to as the second voice wake-up model), and using the second voice wake-up model to test the voice wake-up capability of the first voice wake-up model. The second voice wake-up model can be a larger model than the first, or it can incorporate scrambling pinyin processing during its construction. As described above, the model testing method relies on the second voice wake-up model, which is trained using the wake-up speech and its included wake words. In other words, the second voice wake-up model is trained using known wake words. If the wake word set by the user does not belong to these known wake words, the second voice wake-up model may not accurately recognize the wake word, thus failing to accurately test the voice wake-up capability of the first voice wake-up model. Furthermore, scrambling pinyin processing mainly involves pre-setting several pinyin scrambling rules and processing the scrambling according to these rules. However, the rules are finite and cannot cover all possible pinyin scrambling situations, which will affect the accuracy of the second voice wake-up model.

[0006] The speech recognition method involves performing speech recognition on the wake-up speech, converting the wake-up speech into text to obtain a wake-up word, and then using this wake-up word to verify the wake-up word recognized by the first speech wake-up model. However, in practical applications, wake-up speech is usually short and has weak semantic information. Therefore, when converting the wake-up speech into text through speech recognition, an accurate wake-up word may not be obtained, thus making it impossible to accurately test the speech wake-up capability of the first speech wake-up model.

[0007] Accordingly, a new technical solution is needed in this field to solve the above problems. Summary of the Invention

[0008] In order to overcome the above-mentioned deficiencies, this application is made to solve or at least partially solve the technical problem of accurately checking or verifying whether the wake-up voice is the correct wake-up voice.

[0009] In a first aspect, a method for checking wake-up voice is provided, the method comprising:

[0010] The wake-up voice is subjected to speech recognition to obtain the text information of the wake-up voice, and the text information includes Chinese text;

[0011] Based on the Chinese text, obtain the first pinyin of the text information;

[0012] Obtain a preset Chinese wake-up word and convert the Chinese wake-up word into a second pinyin;

[0013] Match the first pinyin with the second pinyin;

[0014] If the match is successful, the wake-up voice is determined to be the correct wake-up voice;

[0015] If the match fails, the wake-up voice is determined to be an incorrect wake-up voice.

[0016] In one technical solution of the above-mentioned wake-up voice checking method, obtaining the first pinyin of the text information based on the Chinese text includes:

[0017] The Chinese text is converted to Pinyin to obtain the Pinyin of the Chinese text;

[0018] If the text information also includes English text, then the first pinyin is obtained based on the pinyin of the Chinese text and the English text;

[0019] Otherwise, the first pinyin is obtained based on the pinyin of the Chinese text.

[0020] In one technical solution of the above-mentioned wake-up voice checking method, when the first pinyin fails to match the second pinyin, the method further includes determining whether the wake-up voice is a correct wake-up voice through the following steps:

[0021] Step S1: Determine whether the first pinyin contains English text;

[0022] If not included, the first pinyin is taken as the pinyin to be converted, and step S2 is executed;

[0023] If included, then obtain the pinyin syllable that is the same as the pronunciation of the English text, replace the English text in the first pinyin with the pinyin syllable, and obtain the third pinyin; and,

[0024] The third pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the third pinyin is used as the pinyin to be converted, and step S2 is executed.

[0025] Step S2: Obtain the first initial and / or the first final in the pinyin to be converted, replace the first initial with the second initial and / or replace the first final with the second final to obtain the converted pinyin, wherein the second initial is similar in pronunciation to the first initial, and the second final is similar in pronunciation to the first final; and,

[0026] The converted pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the wake-up voice is determined to be the incorrect wake-up voice.

[0027] In one technical solution of the above-mentioned wake-up voice detection method, step S2 includes:

[0028] Step S21: Obtain the first initial consonant in the pinyin to be converted, replace the first initial consonant with the second initial consonant to obtain the first converted pinyin; and,

[0029] The first converted pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, step S22 is executed.

[0030] Step S22: Obtain the first vowel in the pinyin to be converted, replace the first vowel with the second vowel to obtain the second converted pinyin; and,

[0031] The second converted pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the wake-up voice is determined to be the incorrect wake-up voice.

[0032] In one technical solution of the above-mentioned wake-up voice detection method, the method further includes obtaining the second initial consonant and the second final vowel by means of:

[0033] Obtain a first correspondence and a second correspondence, wherein the first correspondence is the correspondence between initials with similar pronunciations, and the second correspondence is the correspondence between finals with similar pronunciations;

[0034] Based on the first correspondence, obtain a second initial consonant that is similar in pronunciation to the first initial consonant;

[0035] Based on the second correspondence, a second vowel that is similar in pronunciation to the first vowel is obtained.

[0036] In one technical solution of the above-described wake-up voice checking method, before performing step S2, the method further includes:

[0037] Obtain the number of the first syllables of the pinyin to be converted;

[0038] Get the number of second syllables in the second pinyin;

[0039] If the number of the first syllables is less than the number of the second syllables, then the wake-up voice is determined to be an incorrect wake-up voice, and step S2 is no longer executed.

[0040] In one technical solution of the above-mentioned wake-up voice checking method, when the number of the first syllables is less than the number of the second syllables, the method further includes determining whether the wake-up voice is an erroneous wake-up voice through the following steps:

[0041] Obtain the string edit distance between the pinyin to be converted and the second pinyin;

[0042] If the string edit distance is greater than the set threshold, the wake-up voice is determined to be an incorrect wake-up voice, and step S2 is no longer executed.

[0043] In one technical solution of the above-described wake-up voice checking method, before performing step S2, the method further includes:

[0044] Obtain the string edit distance between the pinyin to be converted and the second pinyin;

[0045] If the string edit distance is greater than the set threshold, the wake-up voice is determined to be an incorrect wake-up voice, and step S2 is no longer executed.

[0046] In a second aspect, an electronic device is provided, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program, which, when executed by the at least one processor, implements the method described in any of the above-described wake-up voice checking methods.

[0047] In a third aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored therein, the program codes being adapted to be loaded and run by a processor to perform the method described in any of the above-described technical solutions for the wake-up voice checking method.

[0048] The above-described technical solutions of this application have at least one or more of the following beneficial effects:

[0049] In one technical solution of the wake-up voice checking method provided in this application, the wake-up voice can be subjected to speech recognition to obtain the text information of the wake-up voice, which includes Chinese text. A first pinyin is obtained from the Chinese text. A preset Chinese wake-up word is obtained and converted into a second pinyin. The first pinyin and the second pinyin are matched. If the match is successful, the wake-up voice is determined to be correct; if the match fails, the wake-up voice is determined to be incorrect. Based on the above implementation scheme, it is not necessary to use a voice wake-up model to verify the wake-up voice, thus avoiding the inaccurate verification problem that may occur when using model testing methods in existing technologies. Furthermore, after performing speech recognition on the wake-up voice, the above implementation scheme converts the text information of the wake-up voice into the first pinyin and then matches the first pinyin with the second pinyin of the Chinese wake-up word. In this process, the semantics of the text information do not need to be considered, thereby overcoming the problem of inaccurate verification caused by inaccurate semantic recognition when using speech recognition methods to verify the wake-up voice in existing technologies.

[0050] In one technical solution of the wake-up voice checking method provided in this application, multi-level matching can be performed on the first and second pinyin to accurately determine whether the wake-up voice is the correct wake-up voice. Specifically, the first and second pinyin are first matched; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the following steps are used to determine whether the wake-up voice is the correct wake-up voice:

[0051] Step S1: Determine whether the first pinyin contains English text; if not, use the first pinyin as the pinyin to be converted and proceed to step S2; if it does contain English text, obtain the pinyin syllables that are pronounced the same as the English text, replace the English text in the first pinyin with the pinyin syllables to obtain the third pinyin; and match the third pinyin with the second pinyin; if the match is successful, determine that the wake-up voice is the correct wake-up voice; if the match fails, use the third pinyin as the pinyin to be converted and proceed to step S2.

[0052] Step S2: Obtain the first initial and / or the first final in the pinyin to be converted, replace the first initial with the second initial and / or replace the first final with the second final to obtain the converted pinyin. The second initial is similar in pronunciation to the first initial, and the second final is similar in pronunciation to the first final. Also, match the converted pinyin with the second pinyin. If the match is successful, the wake-up voice is determined to be the correct wake-up voice. If the match fails, the wake-up voice is determined to be the incorrect wake-up voice.

[0053] Because different users have different pronunciation habits, the wake-up voice spoken by different users will vary to some extent for the same wake-up word. Different wake-up voices may lead to slight differences in the text information obtained by speech recognition, resulting in a failure to match the first and second pinyin. However, the above implementation scheme takes into account the situation where Chinese and English pronunciations are the same due to different users' pronunciation habits, and the similarity of vowel sounds. Based on this, the first pinyin is converted, and then the converted first pinyin (i.e., the third pinyin or the converted pinyin) is matched with the second pinyin, thereby accurately determining whether the wake-up voice is correct or incorrect. Attached Figure Description

[0054] The disclosure of this application will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Wherein:

[0055] Figure 1 This is a schematic flowchart of the main steps of a wake-up voice checking method according to an embodiment of this application;

[0056] Figure 2This is a flowchart illustrating the main steps of determining whether the wake-up voice is an incorrect wake-up voice when the first pinyin fails to match the second pinyin, according to an embodiment of this application.

[0057] Figure 3 This is a schematic flowchart illustrating the main steps of obtaining the pinyin to be converted and matching the pinyin to be converted with a second pinyin according to an embodiment of this application;

[0058] Figure 4 This is a schematic diagram of the overall flow of a wake-up voice checking method according to an embodiment of this application;

[0059] Figure 5 This is a detailed flowchart illustrating a wake-up voice checking method according to an embodiment of this application;

[0060] Figure 6 This is a schematic diagram of the main structure of an electronic device according to an embodiment of this application.

[0061] Figure label:

[0062] 11: Memory; 12: Processor. Detailed Implementation

[0063] Some embodiments of this application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this application and are not intended to limit the scope of protection of this application.

[0064] In the description of this application, "processor" can include hardware, software, or a combination of both. A processor can be a central processing unit, microprocessor, graphics processor, digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor can be implemented in software, in hardware, or a combination of both. Computer-readable storage media includes any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B.

[0065] The relevant user personal information that may be involved in the various embodiments of this application is processed in strict accordance with the requirements of laws and regulations, following the principles of legality, legitimacy, and necessity, based on the reasonable purpose of the business scenario, and includes personal information that users actively provide or that is generated as a result of using the product / service, as well as personal information obtained with user authorization.

[0066] The personal information processed in this application will vary depending on the specific product / service scenario and will be based on the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, driving information, vehicle information, or other related information. This application will treat the user's personal information and its processing with the utmost diligence.

[0067] This application attaches great importance to the security of users' personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect users' information and prevent unauthorized access, disclosure, use, modification, damage or loss of personal information.

[0068] The following describes an embodiment of the wake-up voice checking method provided in this application.

[0069] See appendix Figure 1 , Figure 1 This is a schematic flowchart illustrating the main steps of a wake-up voice checking method according to an embodiment of this application. Figure 1 As shown, the wake-up voice checking method in this application embodiment mainly includes the following steps S101 to S106.

[0070] Step S101: Perform speech recognition on the wake-up voice to obtain the text information of the wake-up voice.

[0071] Wake-up voice refers to voice information captured by an electronic device, specifically voice information containing a wake-up word recognized by the device. Specifically, electronic devices are equipped with voice-interactive applications. When a user needs to control the electronic device via voice, they can first wake up this voice-interactive application and then use it to control the device. Users can set a wake-up word for this application and speak the wake-up word to the device. The electronic device captures the user's voice and performs voice recognition to determine if it contains a wake-up word. If it does, the application is activated; this voice information is considered a wake-up voice. If it does not contain a wake-up word, the application is not activated; this voice information is considered a non-wake-up voice.

[0072] Electronic devices may include mobile phones, tablets, desktops, laptops, handheld computers, laptops, in-vehicle devices, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc., and this application does not limit them.

[0073] The wake-up word can be a Chinese wake-up word; therefore, the wake-up voice must include at least Chinese speech information, and the text information obtained through speech recognition must also include at least Chinese text, that is, at least Chinese characters. In this embodiment, conventional automatic speech recognition technology can be used to perform speech recognition on the wake-up voice to obtain text information. This embodiment does not specifically limit the speech recognition method.

[0074] Step S102: Obtain the first pinyin of the text information based on the Chinese text.

[0075] Specifically, based on Chinese phonetic alphabets, the Chinese text in the text information is converted into pinyin, and the first pinyin is obtained based on the pinyin of the Chinese text.

[0076] Step S103: Obtain the preset Chinese wake-up word and convert the Chinese wake-up word into the second pinyin.

[0077] The preset Chinese wake-up word is a wake-up word set by the user in advance, and the preset Chinese wake-up word has the same meaning as the wake-up word in the aforementioned step S101. In addition, as can be seen from the aforementioned step S101, the wake-up voice is the voice information containing the wake-up word recognized by the electronic device, and the wake-up voice contains the same wake-up word as the preset Chinese wake-up word mentioned above.

[0078] The preset Chinese wake word is actually Chinese text. Therefore, it can also be converted into pinyin based on Chinese phonetic alphabets. This pinyin is the second pinyin to be obtained.

[0079] Step S104: Match the first pinyin with the second pinyin.

[0080] If the match is successful, it means that the first pinyin is the same as the second pinyin, or the first pinyin includes the second pinyin, which means that the wake-up voice contains the preset Chinese wake-up word, that is, the wake-up voice contains the correct wake-up word, and the wake-up voice is correct. At this time, we can proceed to step S105.

[0081] If the matching fails, it indicates that the first pinyin is different from the second pinyin, and the first pinyin does not include the second pinyin. This indicates that the wake-up voice does not contain the preset Chinese wake-up word, that is, the wake-up voice does not contain the correct wake-up word, and the wake-up voice is incorrect. At this time, we can proceed to step S106.

[0082] The following explains the method for matching the first and second pinyin.

[0083] In addition to syllables, pinyin can also include the tones of syllables. In this embodiment, the tones of syllables can be ignored when matching the first and second pinyin. In some implementations, only the pinyin syllables can be obtained when obtaining the first and second pinyin, without obtaining the tones of the syllables.

[0084] Specifically, if the first and second pinyin syllables contain the same pinyin syllables and the pinyin syllables are arranged in the same order, then the first and second pinyin syllables can be determined to be a successful match; otherwise, the match fails.

[0085] Alternatively, if the first pinyin includes all the pinyin syllables of the second pinyin, and the order of these pinyin syllables in the first pinyin is the same as the order of these pinyin syllables in the second pinyin, then the first and second pinyin can be determined to be a successful match; otherwise, the match fails.

[0086] For example, the preset Chinese wake word is "Makka Bakka," and the second pinyin of this wake word is "mǎ kǎ bā kǎ." If the first pinyin obtained from the wake-up voice is "mā ka bā ka," since the first and second pinyins contain the same syllables and the syllables are arranged in the same order, it can be determined that the first and second pinyins match successfully. If the first pinyin obtained from the wake-up voice is "nǐ hǎo mā ka bā ka," since the first pinyin includes all the syllables of the second pinyin (i.e., "ma ka ba ka"), and "ma ka ba ka" is arranged in the same order in the first and second pinyins, it can also be determined that the first and second pinyins match successfully.

[0087] Step S105: Confirm that the wake-up voice is the correct wake-up voice.

[0088] Step S106: Determine that the wake-up voice is an incorrect wake-up voice.

[0089] Based on the methods described in steps S101 to S106 above, it is not necessary to use a voice wake-up model to verify the wake-up voice, thus avoiding the inaccurate verification problem that may occur when using model testing methods in existing technologies. Furthermore, after performing speech recognition on the wake-up voice, the above method converts the text information of the wake-up voice into first pinyin, and then matches the first pinyin with the second pinyin of the Chinese wake-up word. In this process, the semantics of the text information do not need to be considered, thereby overcoming the problem of inaccurate verification caused by inaccurate semantic recognition when using speech recognition methods to verify the wake-up voice in existing technologies.

[0090] The following provides further explanation of steps S102 and S106.

[0091] I. Explanation of step S102.

[0092] In some embodiments of step S102 above, the first pinyin of the text information can be obtained from the Chinese text through the following steps S1021 to S1024.

[0093] Step S1021: Convert the Chinese text in the text information into pinyin to obtain the pinyin of the Chinese text.

[0094] Step S1022: Determine whether the text information includes English text; if it does, proceed to step S1023; if it does not, proceed to step S1024.

[0095] Step S1023: Obtain the first pinyin based on the pinyin of the Chinese text and the English text. Specifically, based on the order of the Chinese and English texts in the above text information, combine the pinyin of the Chinese text and the English text to form the first pinyin. For example, if the wake-up voice text information is "Hi Makka Pakka", the first pinyin could be "Hi mǎ kǎ bā kǎ".

[0096] Step S1024: Obtain the first pinyin based on the pinyin of the Chinese text. Specifically, the pinyin of the Chinese text can be used as the first pinyin.

[0097] Based on the method described in steps S1021 to S1024 above, the first pinyin can be obtained using the wake-up voice when the wake-up voice includes English.

[0098] II. Explanation of step S106.

[0099] In some embodiments of step S106 above, when the first pinyin fails to match the second pinyin, it can be done by... Figure 2The following steps S201 to S209 determine whether the wake-up voice is the correct wake-up voice. That is, replace step S106 with the following steps.

[0100] Step S201: Determine whether the first pinyin contains English text; if not, proceed to step S202; if it does, proceed to step S203.

[0101] Step S202: Take the first pinyin as the pinyin to be converted, and then proceed to step S206.

[0102] Step S203: Obtain the pinyin syllables that are the same as the pronunciation of the English text, replace the English text in the first pinyin with the pinyin syllables, and obtain the third pinyin.

[0103] "Same pronunciation" can be understood as the English text being read aloud sounding the same as the pinyin syllables. For example, if the English text is "hi," the pinyin syllable that has the same pronunciation as "hi" could be "hai"; if the English text is "hey," the pinyin syllable that has the same pronunciation as "hey" could be "hei"; and if the English text is "how," the pinyin syllable that has the same pronunciation as "how" could be "hào."

[0104] Step S204: Match the third and second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice, and therefore, proceed to step S209; if the match fails, proceed to step S205. The method for matching the third and second pinyin is the same as the method for matching the first and second pinyin in step S104 above, and will not be described again here.

[0105] Because different users have different pronunciation habits, the wake-up voice spoken by different users will vary to some extent when using the same wake-up word. In some cases, the wake-up voice may contain the correct wake-up word and be correct, but due to the user's pronunciation habits, the spoken Chinese may be recognized as similar-sounding English, leading to a failure to match the first and second pinyin syllables and thus identifying the wake-up voice as incorrect. The methods described in steps S203 to S204 above can replace the English text in the text information with the same-sounding pinyin syllables, avoiding misjudgment of the wake-up voice due to the above situation.

[0106] Step S205: Take the third pinyin as the pinyin to be converted, and then proceed to step S206.

[0107] Step S206: Obtain the first initial and / or the first final in the pinyin to be converted, replace the first initial with the second initial and / or replace the first final with the second final to obtain the converted pinyin. The second initial is similar in pronunciation to the first initial, and the second final is similar in pronunciation to the first final.

[0108] By consulting the initial consonant table of Chinese phonetic alphabets, we can determine which letters in the pinyin to be converted are initial consonants and use these initial consonants as the first initial consonants; by consulting the final vowel table of Chinese phonetic alphabets, we can determine which letters in the pinyin to be converted are final vowels and use these final vowels as the first final vowels.

[0109] The similarity between the second initial consonant and the first initial consonant can be understood as the similarity in the sounds produced when the first and second initial consonants are pronounced; similarly, the similarity between the second final vowel and the first final vowel can be understood as the similarity in the sounds produced when the first and second final vowels are pronounced.

[0110] In this embodiment, those skilled in the art can conduct pronunciation experiments on all initials and finals in Chinese phonetic alphabets to determine which initials and finals are pronounced similarly. When it is necessary to obtain a second initial that is pronounced similarly to the first initial, the second initial can be retrieved based on the experimental results; similarly, when it is necessary to obtain a second final that is pronounced similarly to the first final, the second final can be retrieved based on the experimental results. It should be noted that this embodiment does not specifically limit the method for determining the similarity of initial and final pronunciations; any method that can determine which initials and finals are pronounced similarly is acceptable.

[0111] For example, in initials, b and p are pronounced similarly, f and p are pronounced similarly, k and h are pronounced similarly, and g and h are pronounced similarly; in finals, uo and o are pronounced similarly, ing and in are pronounced similarly, and ui and ei are pronounced similarly.

[0112] Step S207: Match the converted Pinyin with the second Pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice, and therefore, proceed to step S209; if the match fails, the wake-up voice is determined to be the incorrect wake-up voice, and therefore, proceed to step S208.

[0113] The method for matching the converted pinyin and the second pinyin is the same as the method for matching the first and second pinyin in step S104 above, and will not be repeated here.

[0114] Because different users have different pronunciation habits, the wake-up voices spoken by different users will vary to some extent when using the same wake-up word. In some cases, the wake-up voice may contain the correct wake-up word and is correct, but due to the user's pronunciation habits, they may confuse similar initials or finals, leading to a failure to match the first and second pinyin, and thus identifying the wake-up voice as incorrect. The method described in steps S206 to S207 above can avoid misjudging the wake-up voice due to the above situation.

[0115] Step S208: Determine that the wake-up voice is an incorrect wake-up voice.

[0116] Step S209: Confirm that the wake-up voice is the correct wake-up voice.

[0117] Based on the method described in steps S201 to S209 above, the influence of the user's pronunciation habits on the first and second pinyin matching results is taken into account, which can effectively avoid misjudgment of the wake-up voice.

[0118] The methods described in steps S201 to S209 above will be further explained below.

[0119] In some embodiments of step S206 above, the second initial consonant and the second final vowel can be obtained through the following steps S2061 to S2063.

[0120] Step S2061: Obtain the first correspondence and the second correspondence. The first correspondence is the correspondence between initials with similar pronunciations, and the second correspondence is the correspondence between finals with similar pronunciations. In this embodiment, those skilled in the art can conduct pronunciation experiments on all initials and finals in Chinese phonetic alphabets to determine which initials and finals are similar in pronunciation. The first correspondence is constructed based on the initials with similar pronunciations, and the second correspondence is constructed based on the finals with similar pronunciations. It should be noted that this embodiment does not specifically limit the method for determining the similarity of initial and final pronunciations; it is sufficient to determine which initials and finals are similar in pronunciation.

[0121] Step S2062: Based on the first correspondence, obtain the second initial consonant that is similar in pronunciation to the first initial consonant. Specifically, the first correspondence can be queried to see if it records which initial consonants are similar in pronunciation to the first initial consonant; if it records initial consonants that are similar in pronunciation, then that initial consonant is used as the second initial consonant; if it does not record initial consonants that are similar in pronunciation, then the second initial consonant is not obtained, and thus the first initial consonant is not replaced.

[0122] Step S2063: Based on the second correspondence, obtain the second vowel that is similar in pronunciation to the first vowel. Specifically, the second correspondence can be queried to see if the first vowel is recorded as having similar pronunciation to any other vowels; if a similar vowel is recorded, then that vowel is used as the second vowel; if no similar vowel is recorded, then the second vowel is not obtained, and consequently, the first vowel is not replaced.

[0123] Based on the method described in steps S2061 to S2063 above, the second initial consonant and the second final vowel can be conveniently obtained by utilizing the first and second correspondence.

[0124] In some embodiments of the method described in steps S201 to S209 above, steps S206 and S207 can be replaced with... Figure 3 The following steps S301 to S306 are shown.

[0125] Step S301: Obtain the first initial consonant in the pinyin to be converted, replace the first initial consonant with the second initial consonant, and obtain the first converted pinyin.

[0126] Step S302: Match the first converted Pinyin with the second Pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice, and therefore, proceed to step S305; if the match fails, proceed to step S303.

[0127] The method for matching the first converted pinyin and the second pinyin is the same as the method for matching the first and second pinyin in step S104 above, and will not be repeated here.

[0128] Step S303: Obtain the first vowel in the pinyin to be converted, replace the first vowel with the second vowel, and obtain the second converted pinyin.

[0129] Step S304: Match the second converted Pinyin with the second Pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice, and therefore, proceed to step S305; if the match fails, the wake-up voice is determined to be the incorrect wake-up voice, and therefore, proceed to step S306.

[0130] The method for matching the second converted pinyin and the second pinyin is the same as the method for matching the first and second pinyin in step S104 above, and will not be repeated here.

[0131] Step S305: Confirm that the wake-up voice is the correct wake-up voice.

[0132] Step S306: Determine that the wake-up voice is an incorrect wake-up voice.

[0133] Based on the method described in steps S301 to S306 above, initials and finals can be replaced sequentially. If the wake-up voice is determined to be correct based on the first converted pinyin obtained after the initial replacement, steps S303 and S304 do not need to be executed again, which improves the efficiency of wake-up voice judgment and saves computing resources.

[0134] In some embodiments of the method described in steps S201 to S209 above, before executing step S206, steps 11 to 13 can be executed first, and then the result of executing step 13 can be used to determine whether to continue executing step S206 and its subsequent steps.

[0135] Step 11: Obtain the number of the first syllables of the pinyin to be converted.

[0136] The number of the first syllable is the number of syllables in the pinyin to be converted. For example, if the pinyin to be converted is nǐ hào mā ka bā ka, then the number of the first syllable is 6.

[0137] Step 12: Obtain the number of second syllables in the second pinyin.

[0138] The number of second syllables refers to the number of syllables included in the second pinyin. For example, if the second pinyin is mǎ kǎ bākǎ, then the number of second syllables is 4.

[0139] Step 13: Determine if the number of the first syllable is less than the number of the second syllable;

[0140] If the number of first syllables is less than the number of second syllables, it indicates that the wake-up speech does not include the correct Chinese wake-up word, and the wake-up speech is incorrect. Therefore, step S206 and its subsequent steps can be discontinued.

[0141] If the number of first syllables is greater than or equal to the number of second syllables, it indicates that the wake-up voice may contain the correct Chinese wake-up word. At this point, step S206 and subsequent steps can be executed to finally determine whether the wake-up voice is correct or incorrect.

[0142] Based on the method described in steps 11 to 13 above, the number of syllables can be used to accurately determine whether the wake-up voice is an incorrect wake-up voice. If it is determined to be an incorrect wake-up voice, there is no need to execute step S206 and its subsequent steps, which further improves the efficiency of wake-up voice judgment and saves computing resources.

[0143] The methods described in steps 11 to 13 above will be further explained below.

[0144] In some embodiments of the method described in steps 11 to 13 above, when it is determined that the number of the first syllables is less than the number of the second syllables, it can also be determined whether the wake-up voice is an incorrect wake-up voice through the following steps 131 to 132.

[0145] Step 131: Obtain the string edit distance between the pinyin to be converted and the second pinyin. In this embodiment, a pinyin syllable is treated as a character, and the string edit distance between the pinyin to be converted and the second pinyin is obtained using a conventional string edit distance acquisition method.

[0146] For example, if the second pinyin is mǎ kǎ bā kǎ, and the pinyin to be converted is mǎ kuǎn bā kǎ, then when converting the pinyin to the second pinyin, the letters u and n need to be deleted from the syllable kuǎn. In other words, an editing operation is required. Therefore, the string edit distance between the pinyin to be converted and the second pinyin is 1.

[0147] Step 132: Determine if the string edit distance is greater than the set threshold;

[0148] If the string editing distance is greater than the set threshold, it indicates that more editing operations are needed to convert the pinyin to be converted into the second pinyin. In other words, the difference between the pinyin to be converted and the second pinyin is relatively large, and it can be determined that the wake-up voice is an incorrect wake-up voice. Step S206 and its subsequent steps will not be executed.

[0149] If the string edit distance is less than or equal to the set threshold, it indicates that there is no need for a lot of editing operations on the pinyin to be converted to convert it into the second pinyin. In other words, the difference between the pinyin to be converted and the second pinyin is not significant, and the wake-up voice may be correct. Therefore, step S206 and subsequent steps can be executed to finally determine whether the wake-up voice is correct or incorrect.

[0150] The threshold is set as a string edit distance value. When setting this threshold value, those skilled in the art can obtain multiple different correct wake-up voice pinyin (hereinafter described as correct pinyin), statistically analyze the string edit distances between these correct pinyin and the second pinyin, obtain the largest string edit distance, and set the threshold based on this largest string edit distance. For example, the threshold can be 3.

[0151] Based on the method described in steps 131 to 132 above, when it is determined that the number of the first syllables is less than the number of the second syllables, the string editing distance between the pinyin to be converted and the second pinyin can be used to further determine whether the wake-up voice is an incorrect wake-up voice, thereby improving the accuracy of judging incorrect wake-up voices.

[0152] In some embodiments of the method described in steps S201 to S209 above, before executing step S206, the following steps 21 to 22 can be executed first, and then the result of executing step 22 can be used to determine whether to continue executing step S206 and its subsequent steps.

[0153] Step 21: Obtain the string edit distance between the pinyin to be converted and the second pinyin.

[0154] The method for this step is the same as that for step 131 mentioned above, and will not be repeated here.

[0155] Step 22: Determine if the string edit distance is greater than the set threshold;

[0156] If the string edit distance is greater than the set threshold, the wake-up voice is determined to be incorrect, and step S206 and its subsequent steps are not executed. If the string edit distance is less than or equal to the set threshold, it indicates that the wake-up voice may be correct. Therefore, step S206 and its subsequent steps can be executed to ultimately determine whether the wake-up voice is correct or incorrect.

[0157] The method for this step is the same as that for step 132 mentioned above, and will not be repeated here.

[0158] Based on the method described in steps 21 to 22 above, the string edit distance can be used to accurately determine whether the wake-up voice is an incorrect wake-up voice. If it is determined to be an incorrect wake-up voice, there is no need to execute step S206 and its subsequent steps, which further improves the efficiency of wake-up voice judgment and saves computing resources.

[0159] The following is in conjunction with the appendix Figure 4 and attached Figure 5 This application describes the method for checking the wake-up voice. (Appendix) Figure 4 An exemplary illustration shows the overall flow of the wake-up voice check method, with appendix. Figure 5 The detailed process of the wake-up voice check method is illustrated by example.

[0160] First, please refer to the appendix. Figure 4The electronic device is equipped with an online voice wake-up system and a voice interaction application. The online voice wake-up system can collect the user's voice information and perform voice recognition to determine whether the voice information contains a wake word; if it contains a wake word, it will activate the voice interaction application, through which the user can control the electronic device by voice. Figure 4 The wake-up voice in the system refers to the voice information containing the wake-up word recognized by the online voice wake-up system. After the wake-up voice is acquired, the following steps S401 to S405 can be used to check whether the wake-up voice is the correct wake-up voice.

[0161] Step S401: Obtain the user-defined wake word, which is a Chinese wake word. Convert the Chinese characters in this wake word into pinyin (i.e., the second pinyin in the aforementioned method embodiment).

[0162] Step S402: Perform speech recognition on the wake-up voice to obtain the text information of the wake-up voice (i.e., Figure 4 (Semantic recognition results). Text information includes Chinese text.

[0163] Step S403: Convert the Chinese characters in the text information into pinyin, and obtain the first pinyin based on the pinyin.

[0164] Step S404: Convert the first pinyin.

[0165] Step S405: Perform pinyin matching on the converted first and second pinyin to obtain the check result of the wake-up voice.

[0166] It should be noted that when performing steps S404 to S405, the first and second pinyin can be matched first. If the match is successful, the wake-up voice is determined to be the correct wake-up voice. If the match fails, the method described in steps S201 to S209 of the aforementioned method embodiment is used to convert the first pinyin and perform pinyin matching on the converted first and second pinyin to obtain the check result of the wake-up voice. The detailed steps of the method described in steps S404 to S405 are detailed in the appendix. Figure 5 A detailed description will be provided in the following section.

[0167] See appendix Figure 5 First, the wake-up voice is processed by speech recognition and pinyin conversion to obtain the first pinyin, and the wake-up word is processed by pinyin conversion to obtain the second pinyin. Then, through the following steps S501 to S510, the wake-up voice is determined to be correct or incorrect based on the first and second pinyin.

[0168] Step S501: Match the first and second pinyin; if the match is successful, the wake-up voice is confirmed to be the correct wake-up voice; if the match fails, proceed to step S502.

[0169] Step S502: Replace the English text in the first pinyin with pinyin syllables that sound similar, to obtain the third pinyin.

[0170] Step S503: Match the third and second pinyin; if the match is successful, the wake-up voice is confirmed to be the correct wake-up voice; if the match fails, proceed to step S504.

[0171] Step S504: Obtain the number of the first syllables of the third pinyin, obtain the number of the second syllables of the second pinyin, and determine whether the number of the first syllables is less than the number of the second syllables; if it is less, determine that the wake-up voice is an incorrect wake-up voice; otherwise, proceed to step S505.

[0172] Step S505: Obtain the string edit distance between the third and second pinyin.

[0173] Step S506: Determine whether the string editing distance is greater than the set threshold; if it is, determine that the wake-up voice is an incorrect wake-up voice; otherwise, proceed to step S507.

[0174] Step S507: Replace the initial consonant in the third pinyin with another initial consonant that sounds similar to it to obtain the first converted pinyin.

[0175] Step S508: Match the first converted Pinyin with the first Pinyin; if the match is successful, the wake-up voice is confirmed to be the correct wake-up voice; if the match fails, proceed to step S509.

[0176] Step S509: Replace the finals in the third pinyin with other finals that have similar pronunciations to obtain the second converted pinyin.

[0177] Step S510: Match the second converted pinyin with the first pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the wake-up voice is determined to be the wrong wake-up voice.

[0178] The method described in steps S501 to S510 above can be understood as an appendix. Figure 4 The detailed steps described in steps S404 to S405 above.

[0179] The technical effects of the method described in steps S501 to S510 are explained below. In this embodiment, the wake-up word includes a default wake-up word and a custom wake-up word. An online voice wake-up system is a voice wake-up model. To enable it to recognize wake-up words, it is typically trained using wake-up voices and their included wake-up words. The default wake-up word can be understood as one or more of the wake-up words used during training. A custom wake-up word is defined by the user and is not part of the wake-up words used during training. Table 1 below shows the inspection results obtained by checking correct and incorrect wake-up voices using the above method.

[0180] The test results include the error between the hard wake-up rate of the default wake-up word and the actual hard wake-up rate, the error between the false wake-up rate of the default wake-up word and the actual false wake-up rate, the error between the hard wake-up rate of the custom wake-up word and the actual hard wake-up rate, and the error between the false wake-up rate of the custom wake-up word and the actual false wake-up rate.

[0181] The hard wake-up rate can be obtained as the percentage between the correctly wake-up voice that was misjudged and all correctly wake-up voices; the false wake-up rate can be obtained as the percentage between the incorrectly wake-up voice that was misjudged and all incorrectly wake-up voices.

[0182] Table 1

[0183]

[0184] As shown in Table 1, the error between the hard wake-up rate obtained by the above method and the actual hard wake-up rate is very small, and the error between the false wake-up rate and the actual false wake-up rate is also very small. Therefore, the above method has high accuracy. The inspection results obtained by the above method can be used to identify which types of wake-up words or wake-up speech the voice wake-up model has poor recognition ability. Based on the results, the voice wake-up model can be updated or upgraded to improve its recognition ability.

[0185] It should be noted that although the steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of this application, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders. These adjusted solutions are equivalent to the technical solutions described in this application and therefore will also fall within the protection scope of this application.

[0186] Those skilled in the art will understand that all or part of the processes in the method of the above-described embodiment can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0187] Another aspect of this application provides a computer-readable storage medium.

[0188] In one embodiment of a computer-readable storage medium according to this application, the computer-readable storage medium may be configured to store a program that performs the wake-up voice checking method of the above-described method embodiments. This program may be loaded and run by a processor to implement the wake-up voice checking method. For ease of explanation, only the parts related to the embodiments of this application are shown; for specific technical details not disclosed, please refer to the method section of the embodiments of this application. The computer-readable storage medium may be a storage device comprising various electronic devices. Optionally, in the embodiments of this application, the computer-readable storage medium is a non-transitory computer-readable storage medium.

[0189] Another aspect of this application provides an electronic device.

[0190] In one embodiment of an electronic device according to this application, the electronic device may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program that, when executed by the at least one processor, implements the methods described in any of the above embodiments. See Appendix Figure 6 , Figure 6 The image exemplarily illustrates a communication connection between memory 11 and processor 12 via a bus.

[0191] In some embodiments of this application, the electronic device may further include at least one sensor for sensing information. The sensor is communicatively connected to any type of processor mentioned in this application. The electronic device described in this application may be, but is not limited to, mobile phones, tablets, desktop computers, laptops, handheld computers, notebook computers, in-vehicle devices, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc., and this application does not limit the scope of the embodiments.

[0192] The technical solution of this application has been described above with reference to one embodiment shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. Without departing from the principles of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of this application.

Claims

1. A method for checking wake-up voice, characterized in that, The method includes: The wake-up voice is subjected to speech recognition to obtain the text information of the wake-up voice, and the text information includes Chinese text; Based on the Chinese text, obtain the first pinyin of the text information; Obtain a preset Chinese wake-up word and convert the Chinese wake-up word into a second pinyin; Match the first pinyin with the second pinyin; If the match is successful, the wake-up voice is determined to be the correct wake-up voice; If the match fails, the following steps are taken to determine whether the wake-up voice is the correct wake-up voice: Step S1: Determine whether the first pinyin contains English text; If not included, the first pinyin is taken as the pinyin to be converted, and step S2 is executed; If included, then obtain the pinyin syllable that is the same as the pronunciation of the English text, replace the English text in the first pinyin with the pinyin syllable, and obtain the third pinyin; and, The third pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the third pinyin is used as the pinyin to be converted, and step S2 is executed. Step S2: Obtain the first initial and / or the first final in the pinyin to be converted, replace the first initial with the second initial and / or replace the first final with the second final to obtain the converted pinyin, wherein the second initial is similar in pronunciation to the first initial and the second final is similar in pronunciation to the first final; and match the converted pinyin with the second pinyin; if the match is successful, determine that the wake-up voice is the correct wake-up voice; if the match fails, determine that the wake-up voice is the incorrect wake-up voice.

2. The method according to claim 1, characterized in that, Step S2 includes: Step S21: Obtain the first initial consonant in the pinyin to be converted, replace the first initial consonant with the second initial consonant to obtain the first converted pinyin; and, The first converted pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, step S22 is executed. Step S22: Obtain the first vowel in the pinyin to be converted, replace the first vowel with the second vowel to obtain the second converted pinyin; and, The second converted pinyin is matched with the second pinyin; if the match is successful, the wake-up voice is determined to be the correct wake-up voice; if the match fails, the wake-up voice is determined to be the incorrect wake-up voice.

3. The method according to claim 1, characterized in that, The method further includes obtaining the second initial consonant and the second final vowel by means of: Obtain a first correspondence and a second correspondence, wherein the first correspondence is the correspondence between initials with similar pronunciations, and the second correspondence is the correspondence between finals with similar pronunciations; Based on the first correspondence, obtain a second initial consonant that is similar in pronunciation to the first initial consonant; Based on the second correspondence, a second vowel that is similar in pronunciation to the first vowel is obtained.

4. The method according to any one of claims 1 to 3, characterized in that, Before performing step S2, the method further includes: Obtain the number of the first syllables of the pinyin to be converted; Get the number of second syllables in the second pinyin; If the number of the first syllables is less than the number of the second syllables, then the wake-up voice is determined to be an incorrect wake-up voice, and step S2 is no longer executed.

5. The method according to claim 4, characterized in that, When the number of the first syllables is less than the number of the second syllables, the method further includes determining whether the wake-up voice is an incorrect wake-up voice through the following steps: Obtain the string edit distance between the pinyin to be converted and the second pinyin; If the string edit distance is greater than the set threshold, the wake-up voice is determined to be an incorrect wake-up voice, and step S2 is no longer executed.

6. The method according to any one of claims 1 to 3, characterized in that, Before performing step S2, the method further includes: Obtain the string edit distance between the pinyin to be converted and the second pinyin; If the string edit distance is greater than the set threshold, the wake-up voice is determined to be an incorrect wake-up voice, and step S2 is no longer executed.

7. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores a computer program, which, when executed by the at least one processor, implements the wake-up voice checking method according to any one of claims 1 to 6.

8. A computer-readable storage medium storing a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by a processor to perform the method for checking the wake-up voice as described in any one of claims 1 to 6.