Voice wake-up method and apparatus, electronic device, and computer storage medium

By determining the target phonemes and candidate paths using the state likelihood data of the speech to be processed, the problem of decreased speech wake-up recognition rate and increased false wake-up rate is solved, achieving higher recognition accuracy and lower false wake-up probability.

CN115019774BActive Publication Date: 2025-12-19ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210761146.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-12-19
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

In existing technologies, as the number of wake words increases, the recognition rate of voice wake-up will gradually decrease, and the number of false wake-ups will gradually increase.

Method used

The target phonemes of each speech to be processed are determined by the state likelihood data of the speech to be processed, and candidate paths are obtained by the target phonemes. Then, based on the speech recognition results of the candidate paths, it is determined whether to wake up the target device.

Benefits of technology

It improves the recognition rate of voice wake-up and reduces the false wake-up rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019774B_ABST
    Figure CN115019774B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice wake-up method, device, electronic equipment and computer storage medium. For improving the recognition rate of voice wake-up and reducing the false wake-up situation. It includes: periodically obtaining the voice in the first time length as the to-be-processed voice in the first time length as the period; based on the target phonemes corresponding to the plurality of to-be-processed voices successively obtained, a plurality of arrangement paths are obtained; wherein the target phoneme is determined based on the state likelihood value array of the corresponding to-be-processed voice, and the state likelihood value array contains the state likelihood value corresponding to each basic phoneme in the to-be-processed voice in each specified state, and the target phoneme is the phoneme in the basic phoneme; and based on the state likelihood value of the target phoneme in the plurality of arrangement paths, a candidate path is determined from the plurality of arrangement paths; performing voice recognition on the candidate path, and determining whether to wake up the target device according to the voice recognition result of the candidate path.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, and particularly relates to a speech wake-up method and device, electronic equipment and computer storage medium. BACKGROUND

[0002] With the popularization of intelligent sound box products, speech wake-up technology is applied to various products to meet the needs of users to control products through voice. Unlike being bound to a specific device by a single wake-up word, users prefer to wake up a certain device through multiple wake-up words. However, as the number of wake-up words increases, the recognition rate of speech wake-up will gradually decrease, and the false wake-up situation will gradually increase. SUMMARY

[0003] In an exemplary embodiment of the present disclosure, a speech wake-up method, device, electronic equipment and computer storage medium are provided to improve the recognition rate of speech wake-up and reduce the false wake-up situation.

[0004] A first aspect of the present disclosure provides a speech wake-up method, comprising:

[0005] Periodically obtaining, as a cycle of a first time length, speech in the first time length as to-be-processed speech;

[0006] Obtaining a plurality of alignment paths based on target phonemes corresponding to a plurality of to-be-processed speeches successively obtained; wherein the target phonemes are determined based on a state likelihood value array of the corresponding to-be-processed speech, and the state likelihood value array contains state likelihood values corresponding to each basic phoneme in the to-be-processed speech in each specified state, and the target phonemes are phonemes in the basic phonemes; and

[0007] Determining a candidate path from the plurality of alignment paths based on state likelihood values of target phonemes in the plurality of alignment paths;

[0008] Performing speech recognition on the candidate path, and determining whether to wake up a target device according to a result of speech recognition of the candidate path.

[0009] In the present embodiment, the target phonemes of each to-be-processed speech are determined through the state likelihood value data of the to-be-processed speech, and the candidate path is obtained through the target phonemes, and then whether to wake up the target device is determined based on the result of speech recognition on the candidate path. Therefore, in the present embodiment, speech recognition is performed on the to-be-processed speech through state level information, phoneme information and path information of the to-be-processed speech, so that the recognition rate of speech wake-up is improved and the false wake-up rate is reduced.

[0010] In one embodiment, before the step of obtaining the plurality of alignment paths based on the target phonemes corresponding to the plurality of acquired to-be-processed speeches, the method further comprises:

[0011] For each to-be-processed speech in the plurality of to-be-processed speeches, the following operations are performed:

[0012] A maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech is determined, and a base phoneme corresponding to the maximum value is determined as the target phoneme corresponding to the to-be-processed speech.

[0013] In this embodiment, the base phoneme corresponding to the maximum value of the state likelihood values in the state likelihood value array of the to-be-processed speech is determined as the target phoneme of the to-be-processed speech. Thus, the target phoneme is obtained by the maximum value of the state likelihood values, so that the target phoneme is determined as the optimal phoneme, and the accuracy of the target phoneme recognition is ensured.

[0014] In one embodiment, before the step of determining the maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech, the method further comprises:

[0015] For any specified state of any base phoneme in the state likelihood value array of the to-be-processed speech, a smoothed state likelihood value of the base phoneme in the specified state is obtained according to the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in each state likelihood value array of the plurality of to-be-processed speeches; and

[0016] Based on the smoothed state likelihood values of each specified state of each base phoneme, a smoothed state likelihood value array corresponding to the to-be-processed speech is obtained.

[0017] In this embodiment, the state likelihood value array is smoothed to avoid the problem that the state likelihood values of the state likelihood value array in some specified states are too high or too low, thereby affecting the accuracy of the target phoneme.

[0018] In one embodiment, the step of obtaining the smoothed state likelihood value corresponding to the specified state of the base phoneme according to the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in each state likelihood value array of the plurality of to-be-processed speeches comprises:

[0019] add the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches, to obtain a total state likelihood value of the specified state of the base phoneme;

[0020] divide the total state likelihood value by the total duration of each to-be-processed speech to obtain a smoothed state likelihood value of the specified state of the base phoneme.

[0021] In the embodiment, the smoothed state likelihood value of the specified state of the base phoneme is determined by the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches, thereby improving the accuracy of the determined smoothed state likelihood value.

[0022] In one embodiment, the plurality of to-be-processed speeches corresponding to the target phoneme are continuously obtained to obtain a plurality of alignment paths, including:

[0023] For any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches, the target phonemes of the two adjacent to-be-processed speeches are connected in the order of acquisition time to obtain an alignment path corresponding to the two adjacent to-be-processed speeches.

[0024] The plurality of alignment paths are obtained based on the alignment paths corresponding to each pair of adjacent to-be-processed speeches.

[0025] In the embodiment, the plurality of alignment paths are determined by the target phonemes of any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches. Thus, the situation of missing a certain alignment path is avoided, so as to further improve the recognition rate of speech wake-up.

[0026] In one embodiment, the state likelihood value of the target phoneme in the plurality of alignment paths is used to determine a candidate path from the plurality of alignment paths, including:

[0027] The alignment path with the highest probability value in the plurality of alignment paths is determined as the candidate path.

[0028] In the embodiment, the candidate path is determined by the alignment path with the highest probability value in the plurality of alignment paths. Thus, the accuracy of the result of speech recognition based on the candidate path is ensured, and the recognition rate of speech wake-up is further improved.

[0029] In one embodiment, the result of speech recognition of the candidate path is obtained in the following manner:

[0030] For any one candidate path, a confidence degree of the candidate path is determined based on a probability value of the candidate path; and

[0031] If the confidence degree of the candidate path is greater than a specified threshold, each target phoneme corresponding to the candidate path is spliced according to the candidate path to obtain a result of speech recognition of the candidate path.

[0032] In this embodiment, the result of speech recognition is obtained by splicing the candidate paths with the confidence degrees greater than the specified threshold, so as to ensure the accuracy of the result of speech recognition of the word determination.

[0033] In one embodiment, the confidence degree of the candidate path is determined based on the probability value of the candidate path, including:

[0034] For any one candidate path, the probability value of the candidate path is determined as the confidence degree of the candidate path; or,

[0035] For any one candidate path, the confidence degree of the candidate path is obtained by using the probability value of the candidate path and each probability value of other multiple arrangement paths corresponding to the candidate path.

[0036] In this embodiment, the confidence degree of the candidate path is determined by the probability value of the candidate path, so that the determined confidence degree is more accurate.

[0037] In one embodiment, the confidence degree of the candidate path is determined by the following formula:

[0038]

[0039] wherein K1 is the confidence degree of the candidate path, P1 is the probability value of the candidate path, Pm is a probability value of the mth path in other arrangement paths corresponding to the candidate path, and N is the total number of other arrangement paths and the candidate path. m

[0040] In one embodiment, the probability value of the arrangement path is determined by the following method:

[0041] For any one arrangement path, a maximum value in each state likelihood value corresponding to each target phoneme in the arrangement path is determined, and the maximum values corresponding to each target phoneme in the arrangement path are added to obtain the probability value of the arrangement path; or,

[0042] For any one arrangement path, an exponential function is used to determine a base probability value of the maximum value in each state likelihood value corresponding to each target phoneme in the arrangement path, and each base probability value is added to obtain the probability value of the arrangement path.

[0043] ​In the embodiment, the maximum value in each state likelihood value corresponding to each target phoneme in the arrangement path is arranged to determine the probability value of the arrangement path. Thus, the accuracy of the probability value of the arrangement path is improved.

[0044] In one embodiment, the determining whether to wake up the target device according to the result of the speech recognition of the candidate path comprises:

[0045] The target speech recognition result is obtained based on the result of the speech recognition of the candidate path and the results of the speech recognitions of a plurality of continuous candidate paths adjacent to the candidate path.

[0046] The determining whether to wake up the target device according to the target speech recognition result.

[0047] In the embodiment, the target speech recognition result is obtained based on the result of the speech recognition of the candidate path and the results of the speech recognitions of a plurality of continuous candidate paths adjacent to the candidate path, and the determining whether to wake up the target device is based on the target speech recognition result. Thus, the determining whether to wake up the target device is based on the results of the speech recognitions of a plurality of candidate paths, the recognition rate of the speech wake-up is improved, and the probability of false wake-up is reduced.

[0048] In one embodiment, the result of the target speech recognition comprises a word-level recognition result and / or a word-level recognition result.

[0049] The determining whether to wake up the target device according to the target speech recognition result comprises:

[0050] If the result of the target speech recognition comprises a word-level speech recognition result, and the word-level speech recognition result matches any one of the preset wake-up words, it is determined to wake up the target device; or,

[0051] If the result of the target speech recognition comprises a word-level speech recognition result, and it is determined that there is a same word in the word-level speech recognition result as any one of the preset wake-up words, it is determined to wake up the target device; or,

[0052] If the result of the target speech recognition comprises a word-level speech recognition result and a word-level speech recognition result, and the word-level speech recognition result matches any one of the preset wake-up words, and there is a same word in the word-level speech recognition result as any one of the preset wake-up words, it is determined to wake up the target device; or,

[0053] If the result of the target speech recognition includes multiple word-level speech recognition results, and it is determined that the number of word-level speech recognition results in which the same word as the preset any wake-up word exists is greater than a specified threshold, and the wake-up words corresponding to the word-level speech recognition results greater than the specified threshold are the same, the target device is determined to be woken up; or,

[0054] If the result of the target speech recognition includes multiple word-level speech recognition results and character-level speech recognition results, and the character-level speech recognition results match any one of the wake-up words, and the number of word-level speech recognition results in which the same word as the preset any wake-up word exists among the word-level speech recognition results is greater than a specified threshold, and the wake-up words corresponding to the word-level speech recognition results greater than the specified threshold are the same, the target device is determined to be woken up.

[0055] In different cases, the embodiments determine whether to wake up the target device through the character-level recognition result and / or the word-level recognition result, thereby avoiding the situation of false wake-up or the situation of not waking up when it should be woken up, improving the recognition rate of voice wake-up, and reducing the probability of false wake-up.

[0056] The second aspect of the present disclosure provides a voice wake-up device, and the device comprises:

[0057] The acquisition module is configured to periodically acquire, as a to-be-processed voice, a voice in a first time length as a period.

[0058] The arrangement path determination module is configured to obtain multiple arrangement paths based on target phonemes corresponding to the multiple to-be-processed voices acquired continuously, wherein the target phonemes are determined based on a state likelihood value array of the corresponding to-be-processed voice, and the state likelihood value array contains state likelihood values of each basic phoneme included in the to-be-processed voice in each specified state, and the target phoneme is a phoneme in the basic phonemes.

[0059] The candidate path determination module is configured to determine a candidate path from the multiple arrangement paths based on the state likelihood values of the target phonemes in the multiple arrangement paths.

[0060] The speech recognition module is configured to perform speech recognition on the candidate path, and determine whether to wake up a target device according to a result of the speech recognition of the candidate path.

[0061] In one embodiment, the device further comprises:

[0062] The target phoneme determination module is configured to, before obtaining the multiple arrangement paths based on the target phonemes corresponding to the multiple to-be-processed voices acquired continuously, perform the following operations on each to-be-processed voice in the multiple to-be-processed voices:

[0063] determining a maximum value of the state likelihood values included in the state likelihood value array of the to-be-processed speech, and determining a base phoneme corresponding to the maximum value as a target phoneme corresponding to the to-be-processed speech.

[0064] In an embodiment, the apparatus further includes:

[0065] a smoothed state likelihood value determination module, configured to, before the determination of the maximum value of the state likelihood values included in the state likelihood value array of the to-be-processed speech, obtain, for any specified state of any base phoneme in the state likelihood value array of the to-be-processed speech, a smoothed state likelihood value of the base phoneme in the specified state according to a state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and state likelihood values corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches;

[0066] a state likelihood value array smoothing module, configured to obtain a smoothed state likelihood value array corresponding to the to-be-processed speech based on the smoothed state likelihood values of the specified states of the base phonemes.

[0067] In an embodiment, the smoothed state likelihood value determination module is specifically configured to:

[0068] add the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and the state likelihood values corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches to obtain a total state likelihood value of the specified state of the base phoneme;

[0069] divide the total state likelihood value by a total time length of the to-be-processed speeches to obtain the smoothed state likelihood value of the specified state of the base phoneme.

[0070] In an embodiment, the arrangement path determination module is specifically configured to:

[0071] for any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches, connect the target phonemes of the two adjacent to-be-processed speeches in a time sequence to obtain an arrangement path corresponding to the two adjacent to-be-processed speeches;

[0072] obtain the plurality of arrangement paths based on the arrangement paths corresponding to the adjacent two to-be-processed speeches.

[0073] In an embodiment, the candidate path determination module is specifically configured to:

[0074] determine the candidate path with the highest probability value in the plurality of arrangement paths as the candidate path.

[0075] In one embodiment, the speech recognition module obtains the result of speech recognition of the candidate path by:

[0076] For any one candidate path, determine the confidence of the candidate path based on the probability value of the candidate path; and,

[0077] If the confidence of the candidate path is greater than a specified threshold, splice each target phoneme corresponding to the candidate path according to the candidate path to obtain the result of speech recognition of the candidate path.

[0078] In one embodiment, the speech recognition module performs determination of the confidence of the candidate path based on the probability value of the candidate path, and is specifically used for:

[0079] For any one candidate path, determine the confidence of the candidate path as the probability value of the candidate path; or,

[0080] For any one candidate path, obtain the confidence of the candidate path by using the probability value of the candidate path and each probability value of a plurality of arrangement paths corresponding to the candidate path.

[0081] In one embodiment, the speech recognition module determines the confidence of the candidate path by the following formula:

[0082]

[0083] wherein K1 is the confidence of the optimal path, P1 is the probability value of the optimal path, P m is the probability value of the mth path in the arrangement paths corresponding to the first specified time length, and N is the total number of arrangement paths corresponding to the first specified time length.

[0084] In one embodiment, the apparatus further comprises:

[0085] The probability value determination module is configured to determine the probability value of the arrangement path by:

[0086] For any one arrangement path, determine the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path, and add the maximum values corresponding to each target phoneme in the arrangement path to obtain the probability value of the arrangement path; or,

[0087] For any one arrangement path, use an exponential function to determine the base probability value of the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path, and add each base probability value to obtain the probability value of the arrangement path.

[0088] In one embodiment, the voice recognition module performs the determining whether to wake up the target device according to the result of the voice recognition of the candidate path, specifically for:

[0089] obtaining a target voice recognition result based on the result of the voice recognition of the candidate path and the results of the voice recognitions of a plurality of continuous candidate paths adjacent to the candidate path;

[0090] determining whether to wake up the target device according to the target voice recognition result.

[0091] In one embodiment, the result of the target voice recognition includes a word-level recognition result and / or a word-level recognition result;

[0092] The voice recognition module performs the determining whether to wake up the target device according to the result of the target voice recognition, specifically for:

[0093] if the result of the target voice recognition includes a word-level voice recognition result, and the word-level voice recognition result matches any one of the preset wake-up words, it is determined that the target device is woken up; or,

[0094] if the result of the target voice recognition includes a word-level voice recognition result, and it is determined that there is a word in the word-level voice recognition result that is the same as any one of the preset wake-up words, it is determined that the target device is woken up; or,

[0095] if the result of the target voice recognition includes a word-level voice recognition result and a word-level voice recognition result, and the word-level voice recognition result matches any one of the preset wake-up words, and there is a word in the word-level voice recognition result that is the same as any one of the preset wake-up words, it is determined that the target device is woken up; or,

[0096] if the result of the target voice recognition includes a plurality of word-level voice recognition results, and it is determined that the number of word-level voice recognition results in the word-level voice recognition results that are the same as any one of the preset wake-up words is greater than a specified threshold, and the wake-up words corresponding to the word-level voice recognition results greater than the specified threshold are the same, it is determined that the target device is woken up; or,

[0097] If the result of the target speech recognition includes multiple word-level speech recognition results and character-level speech recognition results, and the character-level speech recognition results match any of the preset wake-up words, and the number of word-level speech recognition results in the word-level speech recognition results that have the same word as the preset wake-up word is greater than a specified threshold, and the wake-up words corresponding to the word-level speech recognition results greater than the specified threshold are the same, the target device is determined to be woken up.

[0098] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising:

[0099] at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executed by the at least one processor; the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.

[0100] According to a fourth aspect of the embodiments of the present disclosure, a computer storage medium is provided, the computer storage medium stores a computer program for executing the method of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0101] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.

[0102] Figure 1 One of the flowcharts of the voice wake-up method according to an embodiment of the present disclosure;

[0103] Figure 2 The flowchart of determining the result of the voice recognition of the candidate path according to an embodiment of the present disclosure;

[0104] Figure 3 The flowchart of smoothing the state likelihood value array according to an embodiment of the present disclosure;

[0105] Figure 4 The second flowchart of the voice wake-up method according to an embodiment of the present disclosure;

[0106] Figure 5 The voice wake-up device according to an embodiment of the present disclosure;

[0107] Figure 6 The structural schematic diagram of the electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0108] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0109] In this disclosure, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0110] The application scenarios described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided in this disclosure. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided in this disclosure are also applicable to similar technical problems. In the description of this disclosure, unless otherwise stated, "multiple" means two or more.

[0111] In existing technologies, as the number of wake words increases, the recognition rate of voice wake-up will gradually decrease, and the number of false wake-ups will gradually increase.

[0112] Therefore, this disclosure provides a voice wake-up method that determines the target phonemes of each voice to be processed by using the state likelihood value data of the voice to be processed, obtains candidate paths by using the target phonemes, and then determines whether to wake up the target device based on the result of speech recognition of the candidate paths. Thus, this disclosure achieves speech recognition of the voice to be processed by using state-level information, phoneme information, and path information. Therefore, this embodiment improves the recognition rate of voice wake-up and reduces the false wake-up rate by performing speech recognition on multi-level information of the voice to be processed. The solution of this disclosure will now be described in detail with reference to the accompanying drawings.

[0113] like Figure 1 The diagram shown is a flowchart of the voice wake-up method of this disclosure, which may include the following steps:

[0114] Step 101: Periodically acquire the speech within the first duration as the speech to be processed, using the first duration as the period.

[0115] The first duration can be set according to the actual situation, and this embodiment does not limit the specific value of the first duration.

[0116] Step 102: obtaining a plurality of alignment paths based on a plurality of target phonemes corresponding to a plurality of to-be-processed speeches acquired continuously; wherein the target phoneme is determined based on a state likelihood value array of the corresponding to-be-processed speech, and the state likelihood value array contains state likelihood values corresponding to each basic phoneme in the to-be-processed speech in each specified state, and the target phoneme is a phoneme in the basic phonemes;

[0117] Wherein, the state likelihood value corresponding to any one basic phoneme in any one specified state represents the probability of the basic phoneme being recognized in the specified state. And the basic phoneme is a preset Chinese pinyin letter.

[0118] It should be noted that the state likelihood value array of any one to-be-processed speech is obtained by inputting the to-be-processed speech into a pre-trained state likelihood value array training model. In this embodiment, the number of specified states corresponding to each basic phoneme is three, which are the start state, the continuous state and the end state. However, the specified state of the basic phoneme is not limited in this embodiment, which can be set according to the specific situation.

[0119] For example, the basic phonemes included in the state likelihood value array are {x y a b c d e f g h I j k}, and the form of the state likelihood value array corresponding to the to-be-processed speech determined based on the to-be-processed speech can be Wherein, the first row and the first column correspond to the state likelihood values of the three specified states of the basic phoneme x, the first row and the second column correspond to the state likelihood values of the three specified states of the basic phoneme y, the first row and the third column correspond to the state likelihood values of the three specified states of the basic phoneme a, and so on.

[0120] Wherein, the state likelihood value array described in the foregoing is only used for illustration and does not limit the state likelihood value array. In this embodiment, the basic phonemes included in each state likelihood value array are the same, and the specific basic phonemes can be set according to the actual situation, and the state likelihood value array in this embodiment does not limit the basic phonemes contained in the state likelihood value array. And the state likelihood values of the specified states of each basic phoneme are determined based on the to-be-processed language.

[0121] In one embodiment, the target phoneme is determined in the following manner:

[0122] For each to-be-processed speech in the plurality of to-be-processed speeches, the following operation is performed: determining the maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech, and determining the basic phoneme corresponding to the maximum value as the target phoneme corresponding to the to-be-processed speech.

[0123] For example, the state likelihood value array corresponding to to-be-processed speech 1 is And the basic phonemes corresponding to the state likelihood value array are {x y a b c d e f g h I j k}, it is determined that the maximum value of the state likelihood value contained in the state likelihood value array is 5 in the sixth column of the third row, and the basic phoneme corresponding to the state likelihood value is the basic phoneme d, so the basic phoneme d is determined as the target phoneme corresponding to the to-be-processed speech.

[0124] After introducing the determination method of the target phoneme, the specific method of determining the plurality of arrangement paths is introduced in detail. In an embodiment, for any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches, the target phonemes of the two adjacent to-be-processed speeches are connected in the order of acquisition time to obtain an arrangement path corresponding to the two adjacent to-be-processed speeches; and the plurality of arrangement paths are obtained based on the arrangement paths corresponding to each adjacent two to-be-processed speeches.

[0125] Wherein, the to-be-processed speeches are arranged in the order of time.

[0126] For example, the plurality of to-be-processed speeches are arranged in the order of time as to-be-processed speech 1, to-be-processed speech 2, to-be-processed speech 3 and to-be-processed speech 4. If the target phoneme corresponding to the to-be-processed speech 1 is e, the target phoneme corresponding to the to-be-processed speech 2 is x, the target phoneme corresponding to the to-be-processed speech 3 is iao, and the target phoneme corresponding to the to-be-processed speech 4 is l, the target phoneme e corresponding to the to-be-processed speech 1 is connected with the target phoneme x corresponding to the to-be-processed speech 3 to obtain the arrangement path e-x. The target phoneme x corresponding to the to-be-processed speech 2 is connected with the target phoneme iao corresponding to the to-be-processed speech 3 to obtain the arrangement path x-iao. The target phoneme iao corresponding to the to-be-processed speech 3 is connected with the target phoneme l corresponding to the to-be-processed speech 4 to obtain the arrangement path iao-l. Therefore, the plurality of arrangement paths corresponding to the to-be-processed speech 1 to the to-be-processed speech 4 are e-x, x-iao and iao-l respectively.

[0127] Step 103: determining a candidate path from the plurality of arrangement paths based on the state likelihood values of the target phonemes in the plurality of arrangement paths;

[0128] In an embodiment, step 103 can be implemented as: determining the arrangement path with the highest probability value in the plurality of arrangement paths as the candidate path.

[0129] Wherein, the probability value of the arrangement path is determined by the following two methods:

[0130] The first way is: for any one of the arrangement paths, determining the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path; and adding the maximum values corresponding to each target phoneme in the arrangement path to obtain the probability value of the arrangement path.

[0131]

[0132] Wherein, P is the probability value of the arrangement path, L i is the maximum value of each state likelihood value corresponding to the target phoneme i in the arrangement path, and n is the number of target phonemes in the arrangement path.

[0133] The second way is: for any one of the arrangement paths, using an exponential function to respectively determine the base probability value of the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path, and adding the base probability values to obtain the probability value of the arrangement path. Wherein, the probability value of the arrangement path can be determined by formula (2):

[0134]

[0135] Wherein, P is the probability value of the arrangement path, L i is the maximum value of each state likelihood value corresponding to the target phoneme i in the arrangement path, and n is the number of target phonemes in the arrangement path.

[0136] Step 104: performing speech recognition on the candidate path, and determining whether to wake up the target device according to the result of speech recognition of the candidate path.

[0137] Wherein, the target device in the embodiment is a device for obtaining the to-be-processed speech.

[0138] Next, the specific way of determining the result of speech recognition of the candidate path will be described in detail, such as Figure 2 The flowchart for determining the result of speech recognition of the candidate path is shown in FIG. 2, which includes the following steps:

[0139] Step 201: for any one of the candidate paths, determining the confidence degree of the candidate path based on the probability value of the candidate path.

[0140] In one embodiment, the confidence degree of the optimal path is determined by the following two ways:

[0141] The first way is: for any one of the candidate paths, determining the probability value of the candidate path as the confidence degree of the candidate path.

[0142] The second way is: for any one of the candidate paths, using the probability value of the candidate path and the probability values of other arrangement paths corresponding to the candidate path to obtain the confidence degree of the candidate path.

[0143] wherein the confidence of the candidate path can be determined by formula (3):

[0144]

[0145] wherein K1 is the confidence of the candidate path, P1 is the probability value of the candidate path, Pm is the probability value of the mth path in the other multiple arrangement paths corresponding to the candidate path, and N is the total number of paths of the candidate path and the other multiple arrangement paths corresponding to the candidate path. m wherein K1 is the confidence of the candidate path, P1 is the probability value of the candidate path, Pm is the probability value of the mth path in the other multiple arrangement paths corresponding to the candidate path, and N is the total number of paths of the candidate path and the other multiple arrangement paths corresponding to the candidate path.

[0146] It should be noted that the other multiple arrangement paths corresponding to the candidate path are each arrangement path other than the candidate path in the multiple arrangement paths obtained based on the target phonemes corresponding to the multiple continuously acquired to-be-processed speeches.

[0147] Step 202: If the confidence of the candidate path is greater than a specified threshold, then the target phonemes corresponding to the candidate path are spliced according to the candidate path to obtain the result of speech recognition of the candidate path.

[0148] For example, the candidate path 1 is greater than the specified threshold, and the candidate path 1 is x-iao, then the splicing result of the candidate path 1 is xiao, and the corresponding Chinese characters of xiao are determined by using the preset corresponding relationship between the splicing result and the Chinese characters. If it is determined that the Chinese characters corresponding to xiao include “small”, “dawn”, “Xiao”, “school”, “laugh”, and “Xiao”, then “small”, “dawn”, “Xiao”, “school”, “laugh”, and “Xiao” are determined as the result of speech recognition of the candidate path 1.

[0149] In an embodiment, the step of determining whether to wake up the target device according to the result of speech recognition of the candidate path in step 104 can be implemented as: obtaining a target speech recognition result based on the result of speech recognition of the candidate path and the results of speech recognition of multiple continuous candidate paths adjacent to the candidate path; and determining whether to wake up the target device according to the target speech recognition result.

[0150] wherein the target speech recognition result includes a character-level recognition result and / or a word-level recognition result;

[0151] In an embodiment, the step of determining whether to wake up the target device according to the target speech recognition result includes the following cases:

[0152] Case 1: If the target speech recognition result includes a character-level speech recognition result, and the character-level speech recognition result matches any one of the preset wake-up words, then it is determined to wake up the target device.

[0153] For example, the preset wake-up words include "Xiaole" and "Xiaoming". If the character-level speech recognition result includes [Xiao- Le], [Xiao-Le], [Xiao-Le], and [Xiao-Le], it is determined that [Xiao-Le] in the character-level speech recognition result matches "Xiaole" in the wake-up word, and the target device is determined to be woken up.

[0154] Case 2: If the result of the target speech recognition includes a word-level speech recognition result, and it is determined that there is a word in the word-level speech recognition result that is the same as any of the preset wake-up words, the target device is determined to be woken up.

[0155] For example, the preset wake-up words include "Xiaole" and "Xiaoming". If the word-level speech recognition result includes Xiaole, Xiaoer, Xiaole, and Xiaole, it is determined that "Xiaole" in the word-level speech recognition result matches the preset wake-up word "Xiaole", and the target device is determined to be woken up.

[0156] Case 3: If the result of the target speech recognition includes a character-level speech recognition result and a word-level speech recognition result, and the character-level speech recognition result matches any of the preset wake-up words, and there is a word in the word-level speech recognition result that is the same as any of the preset wake-up words, the target device is determined to be woken up.

[0157] For example, the preset wake-up words include "Xiaole" and "Xiaoming". If the word-level speech recognition result includes Xiaole, Xiaoer, Xiaole, and Xiaole, it is determined that "Xiaole" in the word-level speech recognition result matches the preset wake-up word "Xiaole", and if the character-level speech recognition result includes [Xiao-Le], [Xiao-Le], [Xiao-Le], and [Xiao-Le], it is determined that [Xiao-Le] in the character-level speech recognition result matches "Xiaole" in the wake-up word, and the target device is determined to be woken up.

[0158] Case 4: If the result of the target speech recognition includes multiple word-level speech recognition results, and it is determined that the number of word-level speech recognition results in which there is a word that is the same as any of the preset wake-up words is greater than a specified threshold, and the wake-up words corresponding to the word-level speech recognition results that are greater than the specified threshold are the same, the target device is determined to be woken up.

[0159] For example, if word-level speech recognition result 1 includes Xiaole, Xiaoer, Xiaole, and Xiaole, word-level speech recognition result 2 includes Xiaoai, Xiaoai, and Xiaoai, and word-level speech recognition result 3 includes Xiaole, Xiaoer, Xiaole, and Xiaole, and the specified threshold is 1, and the preset wake-up words include "Xiaole" and "Xiaoming", it is determined that the number of word-level speech recognition results in which there is a word that is the same as any of the preset wake-up words is 2, which is greater than the specified threshold, and the target device is woken up.

[0160] Case 5: If the result of the target speech recognition includes multiple word-level speech recognition results and character-level speech recognition results, and the character-level speech recognition results match any of the preset wake-up words, and the number of word-level speech recognition results in which the same word as the preset wake-up word exists in each of the word-level speech recognition results is greater than a specified threshold, and the wake-up words corresponding to each of the word-level speech recognition results greater than the specified threshold are the same, it is determined to wake up the target device.

[0161] It should be noted that the specified threshold in the present embodiment can be set according to actual conditions, and the present embodiment does not limit this. For example, it can be set according to the type of the device and the environment in which it is located. If it is in a noisy environment, the specified number can be increased to improve the accuracy of recognition. If it is in a quiet environment, the specified number can be decreased.

[0162] To ensure the accuracy of the target phonemes corresponding to the to-be-processed speech, before step 102 is performed, the state likelihood value arrays corresponding to each to-be-processed language need to be smoothed, as shown in FIG. 3, which is a flowchart for smoothing the state likelihood value arrays. The process can include the following steps: Figure 3

[0163] Step 301: For any specified state of any base phoneme in the state likelihood value array of the to-be-processed speech, a smoothed state likelihood value of the base phoneme in the specified state is obtained according to the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in each state likelihood value array of the plurality of to-be-processed speeches.

[0164] In one embodiment, step 301 can be specifically implemented as follows: adding the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in each state likelihood value array of the plurality of to-be-processed speeches to obtain a total state likelihood value of the specified state of the base phoneme; dividing the total state likelihood value by the total duration of each to-be-processed speech to obtain the smoothed state likelihood value of the specified state of the base phoneme.

[0165] Step 302: Based on the smoothed state likelihood values of each specified state of each base phoneme, a smoothed state likelihood value array corresponding to the to-be-processed speech is obtained.

[0166] In one embodiment, the state likelihood values of each specified state of each base phoneme in the state likelihood value array are replaced by the smoothed state likelihood values of each specified state of each base phoneme to obtain the smoothed state likelihood value array corresponding to the to-be-processed speech.​

[0167] For further understanding of the technical solutions of the present disclosure, the following will combine the Figure 4 be described in detail, which can include the following steps:

[0168] Step 401: periodically acquire the voice in the first time length as the voice to be processed in the first time length as the period;

[0169] Step 402: for any two adjacent voice to be processed in the plurality of voice to be processed, connect the target phonemes of the two adjacent voice to be processed in the order of acquisition time to obtain the arrangement path corresponding to the two adjacent voice to be processed;

[0170] Step 403: based on the arrangement path corresponding to each adjacent two voice to be processed, obtain the plurality of arrangement paths;

[0171] Step 404: determine the arrangement path with the highest probability value in the plurality of arrangement paths as the candidate path;

[0172] Step 405: for any one candidate path, based on the probability value of the candidate path, determine the confidence of the candidate path;

[0173] Step 406: determine whether the confidence of the candidate path is greater than a specified threshold, if yes, execute step 407, if not, execute step 410;

[0174] Step 407: splice each target phoneme corresponding to the candidate path according to the candidate path to obtain the result of voice recognition of the candidate path;

[0175] Step 408: based on the result of voice recognition of the candidate path and the results of voice recognition of the plurality of continuous candidate paths adjacent to the candidate path, obtain the target voice recognition result;

[0176] Step 409: determine whether to wake up the target device according to the target voice recognition result;

[0177] Step 410: delete the candidate path.

[0178] Based on the same disclosure concept, the voice wake-up method as described above can also be implemented by a voice wake-up device. The effect of the voice wake-up device is similar to that of the foregoing method, which will not be described here.

[0179] Figure 5 The structural schematic diagram of the voice wake-up device according to one embodiment of the present disclosure.

[0180] As Figure 5As shown, the voice wake-up device 500 of the present disclosure can include an acquisition module 510, an arrangement path determination module 520, a candidate path determination module 530, and a speech recognition module 540.

[0181] The acquisition module 510 is configured to periodically acquire, with a first time length as a period, speech in the first time length as to-be-processed speech.

[0182] The arrangement path determination module 520 is configured to obtain a plurality of arrangement paths based on target phonemes corresponding to a plurality of to-be-processed speeches acquired continuously; wherein the target phoneme is determined based on a state likelihood value array of the corresponding to-be-processed speech, and the state likelihood value array contains state likelihood values of each basic phoneme included in the to-be-processed speech in each specified state, and the target phoneme is a phoneme in the basic phoneme.

[0183] The candidate path determination module 530 is configured to determine a candidate path from the plurality of arrangement paths based on the state likelihood value of the target phoneme in the arrangement path.

[0184] The speech recognition module 540 is configured to perform speech recognition on the candidate path, and determine whether to wake up a target device according to a result of the speech recognition of the candidate path.

[0185] In one embodiment, the device further comprises:

[0186] The target phoneme determination module 550 is configured to, before the plurality of arrangement paths are obtained based on the target phonemes corresponding to the plurality of to-be-processed speeches, perform the following operations for each to-be-processed speech in the plurality of to-be-processed speeches:

[0187] Determine the maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech, and determine the basic phoneme corresponding to the maximum value as the target phoneme corresponding to the to-be-processed speech.

[0188] In one embodiment, the device further comprises:

[0189] The smoothed state likelihood value determination module 560 is configured to, before the maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech is determined, for any one specified state of any one basic phoneme in the state likelihood value array of the to-be-processed speech, obtain a smoothed state likelihood value of the basic phoneme in the specified state according to the state likelihood value corresponding to the specified state of the basic phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the basic phoneme in each state likelihood value array of the plurality of to-be-processed speeches.

[0190] The state likelihood value array smoothing module 570 is configured to obtain a smoothed state likelihood value array corresponding to the to-be-processed speech based on the smoothed state likelihood value of each specified state of each base phoneme.

[0191] In an embodiment, the smoothed state likelihood value determination module 560 is specifically configured to:

[0192] add the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches, to obtain a total state likelihood value of the specified state of the base phoneme;

[0193] divide the total state likelihood value by the total duration of each to-be-processed speech, to obtain the smoothed state likelihood value of the specified state of the base phoneme.

[0194] In an embodiment, the arrangement path determination module 520 is specifically configured to:

[0195] for any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches, connect the target phonemes of the two adjacent to-be-processed speeches in the order of acquisition time, to obtain an arrangement path corresponding to the two adjacent to-be-processed speeches;

[0196] obtain the plurality of arrangement paths based on the arrangement paths corresponding to each pair of adjacent to-be-processed speeches.

[0197] In an embodiment, the candidate path determination module 530 is specifically configured to:

[0198] determine the arrangement path with the highest probability value in the plurality of arrangement paths as the candidate path.

[0199] In an embodiment, the speech recognition module 540 obtains the speech recognition result of the candidate path by:

[0200] for any one candidate path, determines a confidence degree of the candidate path based on the probability value of the candidate path; and,

[0201] if the confidence degree of the candidate path is greater than a specified threshold, concatenates each target phoneme corresponding to the candidate path according to the candidate path, to obtain the speech recognition result of the candidate path.

[0202] In an embodiment, the speech recognition module 540 determines the confidence degree of the candidate path based on the probability value of the candidate path, and is specifically configured to:

[0203] For any one candidate path, the probability value of the candidate path is determined as the confidence of the candidate path; or,

[0204] For any one candidate path, the probability value of the candidate path and the probability values of other multiple arrangement paths corresponding to the candidate path are used to obtain the confidence of the candidate path.

[0205] In one embodiment, the speech recognition module 540 determines the confidence of the candidate path by the following formula:

[0206]

[0207] wherein K1 is the confidence of the optimal path, P1 is the probability value of the optimal path, P m is the probability value of the mth path in the arrangement path corresponding to the first specified time length, and N is the total number of arrangement paths corresponding to the first specified time length.

[0208] In one embodiment, the apparatus further comprises:

[0209] The probability value determination module 580 is configured to determine the probability value of the arrangement path by the following manner:

[0210] For any one arrangement path, the maximum value in the state likelihood values corresponding to each target phoneme in the arrangement path is determined; and the maximum values corresponding to each target phoneme in the arrangement path are added to obtain the probability value of the arrangement path; or,

[0211] For any one arrangement path, the exponential function is used to respectively determine the base probability value of the maximum value in the state likelihood values corresponding to each target phoneme in the arrangement path, and the base probability values are added to obtain the probability value of the arrangement path.

[0212] In one embodiment, the speech recognition module 540 performs the determination of whether to wake up the target device according to the result of the speech recognition of the candidate path, and is specifically configured to:

[0213] Based on the result of the speech recognition of the candidate path and the results of the speech recognition of multiple continuous candidate paths adjacent to the candidate path, a target speech recognition result is obtained;

[0214] According to the target speech recognition result, it is determined whether to wake up the target device.

[0215] In one embodiment, the result of the target speech recognition includes a word-level recognition result and / or a word-level recognition result.

[0216] The voice recognition module 540 performs the determining whether to wake up the target device according to the target voice recognition result, and specifically is configured to:

[0217] If the target voice recognition result includes a word-level voice recognition result, and the word-level voice recognition result matches any one of the preset wake-up words, it is determined that the target device is woken up; or,

[0218] If the target voice recognition result includes a word-level voice recognition result, and it is determined that there is a word in the word-level voice recognition result that is the same as any one of the preset wake-up words, it is determined that the target device is woken up; or,

[0219] If the target voice recognition result includes a word-level voice recognition result and a word-level voice recognition result, and the word-level voice recognition result matches any one of the preset wake-up words, and there is a word in the word-level voice recognition result that is the same as any one of the preset wake-up words, it is determined that the target device is woken up; or,

[0220] If the target voice recognition result includes multiple word-level voice recognition results, and it is determined that the number of word-level voice recognition results in which there is a word that is the same as any one of the preset wake-up words in each word-level voice recognition result is greater than a specified threshold, and the wake-up words corresponding to the word-level voice recognition results greater than the specified threshold are the same, it is determined that the target device is woken up; or,

[0221] If the target voice recognition result includes multiple word-level voice recognition results and a word-level voice recognition result, and the word-level voice recognition result matches any one of the preset wake-up words, and the number of word-level voice recognition results in which there is a word that is the same as any one of the preset wake-up words in each word-level voice recognition result is greater than a specified threshold, and the wake-up words corresponding to the word-level voice recognition results greater than the specified threshold are the same, it is determined that the target device is woken up.

[0222] After introducing a voice wake-up method and device of an example embodiment of the present disclosure, next, an electronic device according to another example embodiment of the present disclosure is introduced.

[0223] Those skilled in the art can understand that each aspect of the present disclosure can be implemented as a system, a method or a program product. Therefore, each aspect of the present disclosure can be embodied as a whole hardware embodiment, a whole software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.

[0224] In some possible implementations, the electronic device according to the present disclosure can include at least one processor, and at least one computer storage medium. Among them, the computer storage medium stores program codes, when the program codes are executed by the processor, the processor executes the steps in the voice wake-up method according to various exemplary embodiments of the present disclosure described above in the specification. For example, the processor can execute the steps 101-104 as shown in Figure 1 FIG. 1.

[0225] The electronic device 600 according to this embodiment of the present disclosure will be described below with reference to Figure 6 FIG. 6. Figure 6 The electronic device 600 shown is merely an example and should not impose any limitation on the function and scope of use of the embodiments of the present disclosure.

[0226] As shown in Figure 6 FIG. 6, the electronic device 600 is in the form of a general electronic device. The components of the electronic device 600 can include, but are not limited to, the at least one processor 601 described above, the at least one computer storage medium 602 described above, and the bus 603 connecting different system components, including the computer storage medium 602 and the processor 601.

[0227] The bus 603 represents one or more of several types of bus structures, including a computer storage medium bus or computer storage medium controller, a peripheral bus, a processor bus, or a local bus using any of a variety of bus structures.

[0228] The computer storage medium 602 can include readable media in the form of volatile computer storage medium, such as random access computer storage medium (RAM) 621 and / or cache storage medium 622, and can further include read-only computer storage medium (ROM) 623.

[0229] The computer storage medium 602 can further include program / utility 625 having a set of (at least one) program modules 624, such as an operating system, one or more application programs, other program modules, and program data, each of which can include or be included in the implementation of a network environment, or some combination thereof.

[0230] The electronic device 600 can also communicate with one or more external devices 604 such as a keyboard or a pointing device, by way of Input / Output (I / O) interfaces 605. Further, the electronic device 600 can communicate with one or more devices that enable user interaction with the electronic device 600, and / or one or more devices that enable communication of the electronic device 600 with one or more other electronic devices. This communication can be via the I / O interfaces 605. Still yet, the electronic device 600 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or the public network, such as the Internet, by way of the network adapter 606. As illustrated, the network adapter 606 communicates with the other components of the electronic device 600 by way of the bus 603. It should be appreciated that the electronic device 600 can be a part of a larger system, including but not limited to a distributed computing environment, a grid computing environment, a client-server computing environment, or any other computing environment. It should also be appreciated that the electronic device 600 can have different architectures and / or components than those illustrated in FIG. 6. For example, the electronic device 600 can be a stand-alone system, or it can be a distributed system that includes multiple electronic devices that are connected to each other via a network.

[0231] In some possible implementation, each aspect of the voice wake-up method provided by the present disclosure can also be implemented as a program product in the form of a computer-readable medium, which includes program code for causing a computer device to execute the steps of the voice wake-up method according to various exemplary embodiments of the present disclosure described above in the specification when the program product is run on the computer device.

[0232] The program product can employ any combination of one or more computer-readable media. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical computer storage device, a magnetic computer storage device, or any suitable combination of the above.

[0233] The program product of the voice wake-up method of the embodiments of the present disclosure can employ a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on an electronic device. However, the program product of the present disclosure is not limited thereto, and in this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0234] A readable signal medium can be any medium that can be read by a machine (e.g., a computer) and can contain various kinds of information. Examples of readable signal mediums include, but are not limited to, a floppy disk, a hard disk, a magnetic tape, an optical data storage medium, a magneto-optical data storage medium, a solid state RAM, a flash memory, a CD-ROM, a CD-R, a CD-RW, a DVD, a Blu-ray™ disk, a memory stick, a memory card, a ROM, and the like.

[0235] The program code embodied on the readable medium can be transmitted by any medium, including, but not limited to, wireless, wired, optical, RF, and the like, or any suitable combination of the foregoing.

[0236] Program code, used by or in connection with the described embodiments, can be written in any of a number of suitable programming languages and can be stored in any suitable manner, or implemented as read-only code. Further, such program code, when stored in a computer-readable medium, can direct a computer or other configurable electrical data processing apparatus, when processing the executable instructions stored in the computer-readable medium, to fabricate the aspects of the various embodiments.

[0237] It should be noted that, although the above detailed description refers to several modules of the apparatus, such a division is merely exemplary and not mandatory. Indeed, according to an embodiment of the present disclosure, the features and functionalities of two or more modules described above can be embodied in one module. Conversely, the features and functionalities of one module described above can be further divided into modules.

[0238] Moreover, while operations of the methods of the present disclosure are described in a particular order in the drawings, this is not required or implied, and the operations can be performed in any order, or not all of the illustrated operations can be performed, or additional operations can be performed, to achieve the desired results. Additionally or alternatively, certain steps can be combined into a single step, and / or a single step can be divided into multiple steps.

[0239] Those skilled in the art will appreciate that embodiments of the disclosure can be devised for a variety of applications. It is therefore intended that the disclosure be considered as in all respects only illustrative and not restrictive. Those skilled in the art will further appreciate that the disclosure can be used for a variety of applications. Accordingly, the disclosure is intended to embrace all alternatives, modifications and variations of the present disclosure that have been disclosed, suggested and / or can be apparent in light of the disclosure to those skilled in the art, and the present disclosure intends to embrace all alternatives, modifications and variations that fall within the scope of the claims and their equivalents. Those skilled in the art will further appreciate that the disclosure can be used for a variety of applications. Accordingly, the disclosure is intended to embrace all alternatives, modifications and variations of the present disclosure that have been disclosed, suggested and / or can be apparent in light of the disclosure to those skilled in the art, and the present disclosure intends to embrace all alternatives, modifications and variations that fall within the scope of the claims and their equivalents.

[0240] The present disclosure is described in reference to the drawings using a flowchart and / or a block diagram of the method, apparatus (system) and computer program product according to the present disclosure. It will be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0241] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0242] These computer program instructions can also be loaded onto a computer or other programmable data processing device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process such that the instructions which execute on the computer or other programmable device provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0243] Obviously, numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore intended that the disclosure be considered in all respects as only illustrative and not restrictive.

Claims

1. A voice wake-up method, characterized by, The method comprises: Periodically obtaining speech in a first time length as to-be-processed speech in a cycle of the first time length; For each to-be-processed speech in the plurality of to-be-processed speeches, determining the maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech, and determining the base phoneme corresponding to the maximum value as the target phoneme corresponding to the to-be-processed speech; Based on the target phonemes corresponding to the plurality of to-be-processed speeches obtained in succession, a plurality of alignment paths are obtained; wherein the target phoneme is determined based on the state likelihood value array of the corresponding to-be-processed speech, and the state likelihood value array contains the state likelihood values corresponding to each base phoneme in the to-be-processed speech in each specified state, the target phoneme is a phoneme in the base phonemes, the state likelihood value corresponding to any one base phoneme in any one specified state represents the probability of the base phoneme being recognized in the specified state, and the base phoneme is each Chinese pinyin letter; the state likelihood value array of any one to-be-processed speech is obtained by inputting the to-be-processed speech into a pre-trained state likelihood value array training model, and the specified states corresponding to each base phoneme are respectively the start state, the continuous state and the end state; and Based on the state likelihood values of the target phonemes in the plurality of alignment paths, a candidate path is determined from the plurality of alignment paths; Performing speech recognition on the candidate path, and determining whether to wake up a target device according to the speech recognition result of the candidate path.

2. The method of claim 1, wherein, Before determining the maximum value of the state likelihood values contained in the state likelihood value array of the to-be-processed speech, the method further comprises: For any one specified state of any one base phoneme in the state likelihood value array of the to-be-processed speech, a smoothed state likelihood value of the base phoneme in the specified state is obtained according to the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in each state likelihood value array of the plurality of to-be-processed speeches; and Based on the smoothed state likelihood values of each specified state of each base phoneme, a smoothed state likelihood value array corresponding to the to-be-processed speech is obtained.

3. The method of claim 2, wherein, The method further comprises: Adding the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and each state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the plurality of to-be-processed speeches to obtain the total state likelihood value of the specified state of the base phoneme; Dividing the total state likelihood value by the total time length of each to-be-processed speech to obtain the smoothed state likelihood value of the specified state of the base phoneme.

4. The method of claim 1, wherein, The plurality of to-be-processed phonemes corresponding to the plurality of to-be-processed speeches are obtained based on continuous acquisition, and a plurality of arrangement paths are obtained, including: For any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches, target phonemes of the two adjacent to-be-processed speeches are connected in the order of acquisition time to obtain an arrangement path corresponding to the two adjacent to-be-processed speeches; Based on the arrangement paths corresponding to each of the two adjacent to-be-processed speeches, the plurality of arrangement paths are obtained.

5. The method of claim 1, wherein, The candidate path is determined from the plurality of arrangement paths based on the state likelihood values of the target phonemes in the plurality of arrangement paths, including: The arrangement path with the highest probability value in the plurality of arrangement paths is determined as the candidate path.

6. The method of claim 1, wherein, The speech recognition result of the candidate path is obtained in the following manner: For any one candidate path, the confidence of the candidate path is determined based on the probability value of the candidate path; and If the confidence of the candidate path is greater than a specified threshold, each target phoneme corresponding to the candidate path is spliced according to the candidate path to obtain the speech recognition result of the candidate path.

7. The method of claim 6, wherein, The confidence of the candidate path is determined based on the probability value of the candidate path, including: For any one candidate path, the probability value of the candidate path is determined as the confidence of the candidate path; or For any one candidate path, the confidence of the candidate path is obtained by using the probability value of the candidate path and the probability values of other plurality of arrangement paths corresponding to the candidate path.

8. The method according to claim 6 or 7, characterized in that, The confidence of the candidate path is determined by the following formula: wherein K1 is the confidence of the candidate path, P1 is the probability value of the candidate path, P m is the probability value of the mth path in other permutation paths corresponding to the candidate path, and N is the total number of other permutation paths and the candidate path.

9. The method according to any one of claims 5 to 7, characterized in that, The probability value of the arrangement path is determined in the following manner: For any one arrangement path, the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path is determined; and the maximum values corresponding to each target phoneme in the arrangement path are added to obtain the probability value of the arrangement path; or For any one arrangement path, the base probability value of the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path is determined by using an exponential function, and each base probability value is added to obtain the probability value of the arrangement path.

10. The method of claim 1, wherein, The determination of whether to wake up the target device according to the speech recognition result of the candidate path includes: A target speech recognition result is obtained based on the speech recognition result of the candidate path and the speech recognition results of a plurality of continuous candidate paths adjacent to the candidate path; Whether to wake up the target device is determined according to the target speech recognition result.

11. The method of claim 10, wherein, The target speech recognition result includes a word-level recognition result and / or a word-level recognition result; The determination of whether to wake up the target device according to the target speech recognition result includes: If the target speech recognition result includes a word-level speech recognition result, and the word-level speech recognition result matches any one of the preset wake-up words, it is determined that the target device is woken up; or If the target speech recognition result includes a word-level speech recognition result, and it is determined that there is a word in the word-level speech recognition result that is the same as any one of the preset wake-up words, it is determined that the target device is woken up; or If the result of the target speech recognition includes a word-level speech recognition result and a character-level speech recognition result, and the character-level speech recognition result matches any preset wake-up word, and there is a word in the word-level speech recognition result that is the same as the preset wake-up word, the target device is determined to be woken up; or, If the result of the target speech recognition includes multiple word-level speech recognition results, and the number of word-level speech recognition results in which there is a word that is the same as any preset wake-up word is greater than a specified threshold, and the wake-up words corresponding to the word-level speech recognition results that are greater than the specified threshold are the same, the target device is determined to be woken up; or, If the result of the target speech recognition includes multiple word-level speech recognition results and a character-level speech recognition result, and the character-level speech recognition result matches any wake-up word, and the number of word-level speech recognition results in which there is a word that is the same as any preset wake-up word is greater than a specified threshold, and the wake-up words corresponding to the word-level speech recognition results that are greater than the specified threshold are the same, the target device is determined to be woken up.

12. A voice wake-up apparatus, characterized by comprising: The apparatus comprises: An acquisition module configured to periodically acquire, with a first time length as a period, speech in the first time length as to-be-processed speech; For each to-be-processed speech in the multiple to-be-processed speeches, the following operations are performed: determining a maximum value of state likelihood values contained in a state likelihood value array of the to-be-processed speech, and determining a base phoneme corresponding to the maximum value as a target phoneme corresponding to the to-be-processed speech; An arrangement path determination module configured to obtain multiple arrangement paths based on target phonemes corresponding to multiple to-be-processed speeches that are acquired continuously; wherein the target phonemes are determined based on state likelihood value arrays of the corresponding to-be-processed speeches, and the state likelihood value arrays contain state likelihood values of each base phoneme contained in the to-be-processed speech in each specified state, the target phonemes are phonemes in the base phonemes, the state likelihood value of any base phoneme in any specified state represents the probability that the base phoneme is recognized in the specified state, and the base phonemes are preset Chinese pinyin letters; the state likelihood value array of any to-be-processed speech is obtained by inputting the to-be-processed speech into a pre-trained state likelihood value array training model, and the specified states corresponding to the base phonemes are respectively a start state, a continuous state, and an end state; A candidate path determination module configured to determine a candidate path from the multiple arrangement paths based on state likelihood values of target phonemes in the multiple arrangement paths; A speech recognition module configured to perform speech recognition on the candidate path, and determine whether to wake up a target device according to a result of the speech recognition of the candidate path.

13. The apparatus of claim 12, wherein, The apparatus further comprises: The smoothing state likelihood value determination module is configured to, before determining the maximum value of the state likelihood value included in the state likelihood value array of the to-be-processed speech, obtain a smoothing state likelihood value of a specified state of any base phoneme in the state likelihood value array of the to-be-processed speech according to the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and the state likelihood values corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches. The state likelihood value array smoothing module is configured to obtain the smoothed state likelihood value array corresponding to the to-be-processed speech based on the smoothing state likelihood values of the specified states of the base phonemes.

14. The apparatus of claim 13, wherein, The smoothing state likelihood value determination module is specifically configured to: add the state likelihood value corresponding to the specified state of the base phoneme in the state likelihood value array of the to-be-processed speech and the state likelihood values corresponding to the specified state of the base phoneme in the state likelihood value arrays of the plurality of to-be-processed speeches to obtain a total state likelihood value of the specified state of the base phoneme; and divide the total state likelihood value by the total duration of each to-be-processed speech to obtain the smoothing state likelihood value of the specified state of the base phoneme.

15. The apparatus of claim 12, wherein, The arrangement path determination module is specifically configured to: connect the target phonemes of any two adjacent to-be-processed speeches in the plurality of to-be-processed speeches in the order of acquisition time to obtain an arrangement path corresponding to the two adjacent to-be-processed speeches; and obtain the plurality of arrangement paths based on the arrangement paths corresponding to each pair of adjacent to-be-processed speeches.

16. The apparatus of claim 12, wherein, The candidate path determination module is specifically configured to: determine the arrangement path with the highest probability value in the plurality of arrangement paths as the candidate path.

17. The apparatus of claim 12, wherein, The speech recognition module obtains the speech recognition result of the candidate path by: determining the confidence of the candidate path based on the probability value of the candidate path; and if the confidence of the candidate path is greater than a specified threshold, concatenating the target phonemes corresponding to the candidate path in the order of the candidate path to obtain the speech recognition result of the candidate path.

18. The apparatus of claim 17, wherein, The speech recognition module determines the confidence of the candidate path based on the probability value of the candidate path, and is specifically configured to: determine the probability value of the candidate path as the confidence of the candidate path for any candidate path; or obtain the confidence of the candidate path by using the probability value of the candidate path and the probability values of other arrangement paths corresponding to the candidate path for any candidate path.

19. The apparatus of claim 17 or 18, wherein, The speech recognition module determines the confidence of the candidate path by the following formula: Wherein, K1 is the confidence of the optimal path, P1 is the probability value of the optimal path, P m is the probability value of the mth path in the arrangement path corresponding to the first specified time length, and N is the total number of arrangement paths corresponding to the first specified time length.

20. The apparatus of any of claims 16-18, wherein The device further includes: The probability value determination module is configured to determine the probability value of the arrangement path by: For any one of the arrangement paths, the maximum value of each state likelihood value corresponding to each target phoneme in the arrangement path is determined; and the maximum values corresponding to each target phoneme in the arrangement path are added to obtain a probability value of the arrangement path; or, For any one of the arrangement paths, the base probability values of the maximum values of each state likelihood value corresponding to each target phoneme in the arrangement path are determined by using an exponential function, respectively, and the base probability values are added to obtain a probability value of the arrangement path.

21. The apparatus of claim 12, wherein, The speech recognition module performs the result of the speech recognition according to the candidate path to determine whether to wake up the target device, specifically for: Based on the result of the speech recognition of the candidate path and the results of the speech recognitions of a plurality of continuous candidate paths adjacent to the candidate path, a target speech recognition result is obtained; According to the target speech recognition result, it is determined whether to wake up the target device.

22. The apparatus of claim 21, wherein, The result of the target speech recognition includes a word-level recognition result and / or a word-level recognition result; The speech recognition module performs the result of the target speech recognition to determine whether to wake up the target device, specifically for: If the result of the target speech recognition includes a word-level speech recognition result, and the word-level speech recognition result matches any one of the preset wake-up words, it is determined to wake up the target device; Or, If the result of the target speech recognition includes a word-level speech recognition result, and it is determined that there is a word in the word-level speech recognition result that is the same as any one of the preset wake-up words, it is determined to wake up the target device; Or, If the result of the target speech recognition includes a word-level speech recognition result and a word-level speech recognition result, and the word-level speech recognition result matches any one of the preset wake-up words, and there is a word in the word-level speech recognition result that is the same as any one of the preset wake-up words, it is determined to wake up the target device; Or, If the result of the target speech recognition includes a plurality of word-level speech recognition results, and it is determined that the number of word-level speech recognition results in which there is a word that is the same as any one of the preset wake-up words in each word-level speech recognition result is greater than a specified threshold, and the wake-up words corresponding to each word-level speech recognition result greater than the specified threshold are the same, it is determined to wake up the target device; Or, If the result of the target speech recognition includes a plurality of word-level speech recognition results and a word-level speech recognition result, and the word-level speech recognition result matches any one of the wake-up words, and the number of word-level speech recognition results in which there is a word that is the same as any one of the preset wake-up words in each word-level speech recognition result is greater than a specified threshold, and the wake-up words corresponding to each word-level speech recognition result greater than the specified threshold are the same, it is determined to wake up the target device.

23. An electronic device, comprising: The device comprises at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executed by the at least one processor; the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-11.

24. A computer storage medium, comprising, The computer storage medium stores a computer program for executing the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Voice recognition method and device, electronic equipment and storage medium

    CN111862943A

  • Speech recognition method, electronic device, program product and storage medium

    CN114255754A