Method and device for recognizing keywords in speech

By generating a fuzzy pronunciation space and matching search with the keyword set, the problem of low full-rate recognition and search in the prior art is solved, and a more efficient voice matching effect is achieved.

CN114255739BActive Publication Date: 2025-05-06CHINA MOBILE GROUP DESIGN INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010996191.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-21
Publication Date
2025-05-06
Estimated Expiration
2040-09-21

AI Technical Summary

Technical Problem

In the prior art, the query rate of keyword recognition in pronunciation is low, especially when facing the accent, speech fuzzy and similar expressions of different voice actors, the matching effect is not good.

Method used

By inputting the speech to be recognized into the speech recognition model, a fuzzy pronunciation space is generated, and a pre-established keyword set is searched to output the matching keywords. This method can deal with similar expressions of pronunciation, word swallowing and inaccurate pronunciation, and improve the search rate.

Benefits of technology

This method successfully improves the full-scale check rate of speech matching, can more accurately identify keywords in speech, and is suitable for handling various mutations in speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255739B_ABST
    Figure CN114255739B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method and device for recognizing keywords in speech, wherein the method comprises: inputting the speech to be recognized into a speech recognition model, outputting a fuzzy pronunciation space corresponding to the speech to be recognized; searching a keyword set according to the fuzzy pronunciation space, and obtaining recognition results of keywords corresponding to the speech to be recognized; wherein the fuzzy pronunciation space is used to represent a variety of speech recognition results corresponding to the speech to be recognized. The method and device for recognizing keywords in speech provided by the embodiment of the present invention recognizes the speech to be recognized through a speech recognition model, obtains a variety of possible speech recognition results, forms a fuzzy pronunciation space, matches and searches the fuzzy pronunciation space with a pre-established keyword set, and outputs the matched keywords. The method of searching the fuzzy pronunciation space can successfully handle similar expressions in speech, swallowing words in speech, and inaccurate pronunciation in speech, and can improve the recall rate of speech matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for recognizing keywords in speech. Background Art

[0002] The existing methods for recognizing keywords in speech mainly include three categories, namely, speech signal recognition, speech-to-text recognition, and speech-to-pinyin recognition. In speech signal recognition, the speech signal is first segmented and then compared with the keyword speech signal; in speech-to-text recognition, the speech is converted into text using a neural network model, and then the keyword is retrieved from the database to return the result; in speech-to-pinyin recognition, the speech is converted into pinyin, and the keyword corresponding to the pinyin is retrieved from the dictionary.

[0003] However, the existing methods have the following shortcomings: First, they have strict requirements on pronunciation accuracy. For some voice data, the speaker’s individual habits may result in different accents, which may lead to matching failures. Second, the fuzzy matching ability is poor. For similar sentence expressions, the continuous word-by-word matching method may lead to matching failures, affecting the recall rate.

[0004] In summary, the existing methods have the disadvantage of low recall rate of keyword recognition in speech. Summary of the invention

[0005] The embodiment of the present invention provides a method and device for recognizing keywords in speech, so as to solve the defect of low recall rate of keyword recognition in speech in the prior art and improve the recall rate.

[0006] An embodiment of the present invention provides a method for recognizing keywords in speech, comprising:

[0007] Inputting the speech to be recognized into the speech recognition model, and outputting the fuzzy pronunciation space corresponding to the speech to be recognized;

[0008] Searching a keyword set according to the fuzzy pronunciation space to obtain a recognition result of the keyword corresponding to the speech to be recognized;

[0009] Among them, the speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent multiple speech recognition results corresponding to the speech to be recognized.

[0010] According to a method for recognizing keywords in speech according to an embodiment of the present invention, the specific steps of inputting the speech to be recognized into a speech recognition model and outputting the fuzzy pronunciation space corresponding to the speech to be recognized include:

[0011] Segmenting the speech to be recognized into a plurality of syllables;

[0012] The candidate pronunciation groups of each syllable included in the speech to be recognized are obtained to form the fuzzy pronunciation space.

[0013] According to an embodiment of the present invention, the method for recognizing keywords in speech, the specific steps of searching the keyword set according to the fuzzy pronunciation space to obtain the recognition result of the keyword corresponding to the speech to be recognized include:

[0014] Searching the keyword set according to each candidate syllable in the fuzzy pronunciation space to obtain a plurality of candidate keywords;

[0015] The fuzzy pronunciation space is matched with each of the candidate keywords, and a recognition result of the keyword corresponding to the speech to be recognized is obtained according to the matching result.

[0016] According to a method for recognizing keywords in speech according to an embodiment of the present invention, the specific steps of searching the keyword set according to each candidate syllable in the fuzzy pronunciation space to obtain a plurality of candidate keywords include:

[0017] Search the keyword set according to each candidate syllable to obtain a plurality of keywords in the keyword set that include the candidate syllable;

[0018] According to the number of the candidate syllables contained in each keyword in the keyword set, a plurality of the keywords are determined as the plurality of candidate keywords.

[0019] According to an embodiment of the present invention, the method for recognizing keywords in speech, the specific steps of matching the fuzzy pronunciation space with each candidate keyword and obtaining the recognition result of the keyword corresponding to the speech to be recognized according to the matching result include:

[0020] Matching the fuzzy pronunciation space with each candidate keyword to obtain a matching degree corresponding to each candidate keyword;

[0021] According to the matching degree corresponding to each of the candidate keywords, the recognition result of the keyword corresponding to the speech to be recognized is obtained.

[0022] According to a method for recognizing keywords in speech according to an embodiment of the present invention, the specific steps of matching the fuzzy pronunciation space with each candidate keyword and obtaining the matching degree corresponding to each candidate keyword include:

[0023] Matching each syllable in the candidate keyword with each of the candidate syllables to obtain the number of matching syllables;

[0024] According to the number of the matched syllables and the number of syllables included in the candidate keyword, the matching degree corresponding to the candidate keyword is obtained.

[0025] According to an embodiment of the present invention, in the method for recognizing keywords in speech, the specific steps of obtaining the recognition result of the keyword corresponding to the speech to be recognized according to the matching degree corresponding to each candidate keyword include:

[0026] If it is determined that the maximum value of the matching degrees corresponding to the candidate keywords is greater than a preset matching threshold, the candidate keyword corresponding to the maximum value is used as the recognition result of the keyword corresponding to the speech to be recognized.

[0027] The embodiment of the present invention further provides a device for recognizing keywords in speech, comprising:

[0028] A speech recognition module, used for inputting the speech to be recognized into the speech recognition model, and outputting the fuzzy pronunciation space corresponding to the speech to be recognized;

[0029] A space search module, used to search a keyword set according to the fuzzy pronunciation space to obtain a recognition result of the keyword corresponding to the speech to be recognized;

[0030] Among them, the speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent multiple speech recognition results corresponding to the speech to be recognized.

[0031] An embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-described methods for recognizing keywords in speech are implemented.

[0032] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of any of the above methods for recognizing keywords in speech are implemented.

[0033] The method and device for recognizing keywords in speech provided by the embodiment of the present invention recognize the speech to be recognized through a speech recognition model, obtain multiple possible speech recognition results, form a fuzzy pronunciation space, match and search the fuzzy pronunciation space and a pre-established keyword set, and output the matched keywords. The method of using the fuzzy pronunciation space search can successfully handle similar expressions in speech, word swallowing in speech, and inaccurate pronunciation in speech, and can improve the recall rate of speech matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 It is a flowchart of a method for recognizing keywords in speech provided by an embodiment of the present invention;

[0036] Figure 2 is a schematic diagram of the structure of a device for recognizing keywords in speech provided by an embodiment of the present invention;

[0037] Figure 3 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] In the description of the embodiments of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the embodiments of the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.

[0040] In the description of the embodiments of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0041] In order to overcome the above-mentioned problems of the prior art, an embodiment of the present invention provides a method and device for identifying keywords in speech. The inventive concept is to identify the speech to be identified, obtain multiple possible speech recognition results, form a fuzzy pronunciation space, match and search the fuzzy pronunciation space with a pre-established keyword set, and output the matched keywords. It can adaptively adapt to fuzzy and non-standard pronunciation situations and automatically match similar expressions of sentences, thereby improving the recall rate of speech matching.

[0042] Figure 1 1 is a flow chart of a method for recognizing keywords in speech provided by an embodiment of the present invention. Figure 1 The method for recognizing keywords in speech according to an embodiment of the present invention is described. Figure 1 As shown, the method includes: step S101, inputting the speech to be recognized into the speech recognition model, and outputting the fuzzy pronunciation space corresponding to the speech to be recognized.

[0043] Among them, the speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent multiple speech recognition results corresponding to the speech to be recognized.

[0044] Specifically, the speech to be recognized is speech data to be recognized (or detected).

[0045] The speech recognition model is used to perform speech recognition on the speech to be recognized and identify the pronunciation corresponding to the speech to be recognized.

[0046] The speech recognition model can be a model established based on machine learning such as various artificial neural networks (ANN for short).

[0047] In the embodiment of the present invention, after the speech to be recognized is input into the speech recognition model, the speech recognition model does not output a unique recognition result, but outputs a plurality of possible speech recognition results to form a fuzzy pronunciation space.

[0048] Possible speech recognition results can be expressed in pinyin.

[0049] It is understandable that the speech to be recognized may include multiple syllables. Accordingly, the speech recognition model recognizes each syllable in the speech to be recognized, obtains the confidence of various pronunciations corresponding to the syllable, and outputs multiple possible pronunciations of the syllable, thereby obtaining multiple speech recognition results for the speech to be recognized.

[0050] Outputting multiple possible speech recognition results is to adapt to ambiguity and non-standard pronunciation. By considering multiple possible pronunciations of speech, the adaptive matching of ambiguity and non-standard pronunciation is strengthened. For example, "China Mobile Communications" is said as "zong China Mobile Communications".

[0051] It is understandable that before step S101, the sample speech signal of the speech sample can be used as a training sample, and the pronunciation corresponding to the sample speech signal of the speech sample can be used as a label of the training sample to train the speech recognition model to obtain a trained speech recognition model. The trained speech recognition model can be used to perform speech recognition on the speech to be recognized in step S101.

[0052] According to the fuzzy pronunciation space, the keyword set is searched to obtain the recognition result of the keyword corresponding to the speech to be recognized.

[0053] Specifically, a keyword set is a keyword list including multiple keywords.

[0054] A keyword is a target that needs to be detected in speech, which can be a phrase or a sentence.

[0055] In the keyword list, a keyword can be represented by its pinyin.

[0056] By searching the keyword set according to the fuzzy pronunciation space, the keyword with the highest matching degree with the fuzzy pronunciation space can be obtained as the matched keyword.

[0057] According to the matched keywords, the recognition results of the keywords corresponding to the speech to be recognized are obtained.

[0058] The embodiment of the present invention recognizes the speech to be recognized through a speech recognition model, obtains multiple possible speech recognition results, forms a fuzzy pronunciation space, matches and searches the fuzzy pronunciation space with a pre-established keyword set, and outputs the matched keywords. The method of using the fuzzy pronunciation space search can successfully handle similar expressions in speech, word swallowing in speech, and inaccurate pronunciation in speech, and can improve the recall rate of speech matching.

[0059] Based on the contents of the above embodiments, the specific steps of inputting the speech to be recognized into the speech recognition model and outputting the fuzzy pronunciation space corresponding to the speech to be recognized include: dividing the speech to be recognized into a plurality of syllables.

[0060] Specifically, a syllable is the smallest unit of phonetic structure composed of one or more phonemes. Generally speaking, in Chinese, a syllable is the pronunciation of a Chinese character.

[0061] The speech recognition model may include two sub-models: a syllable segmentation sub-model and a syllable recognition sub-model.

[0062] The syllable segmentation sub-model is used to segment the speech to be recognized into several syllables based on the speech knowledge and / or the characteristics of the speech to be recognized (such as the half-wave difference spectrum).

[0063] The syllable segmentation sub-model may be a model established based on machine learning such as various artificial neural networks (ANN for short).

[0064] The speech to be recognized is input into the syllable segmentation sub-model, and a number of syllables are output as the syllable segmentation (or segmentation) result of the speech to be recognized.

[0065] The candidate pronunciation groups of each syllable included in the speech to be recognized are obtained to form a fuzzy pronunciation space.

[0066] Specifically, each syllable obtained by the syllable segmentation sub-model is sequentially input into the syllable recognition sub-model, and multiple possible pronunciations of the syllable are output to form an alternative pronunciation group of the syllable.

[0067] The alternative pronunciation group may include multiple alternative pronunciations. Possible pronunciations are alternative pronunciations.

[0068] The syllable recognition submodel is used to obtain multiple possible pronunciations of the input syllable

[0069] The syllable recognition sub-model may be a model established based on machine learning such as various artificial neural networks (ANN for short).

[0070] According to the alternative pronunciation groups of each syllable, a fuzzy pronunciation space can be formed.

[0071] The number of alternative pronunciations included in the alternative pronunciation group of each syllable can be a preset first number. The preset first number is the length of the alternative pronunciation group.

[0072] The preset first number is greater than one and can be set according to actual conditions. The embodiment of the present invention does not limit the specific value of the preset first number.

[0073] Accordingly, the fuzzy pronunciation space can be a matrix, in which the elements are pinyin (i.e., alternative pronunciations).

[0074] The fuzzy pronunciation space has a phonetic dimension and a time dimension. The time dimension describes how the pronunciation of a speech changes over time, corresponding to the horizontal direction of the matrix. The phonetic dimension is the alternative pronunciation group at a certain point in time, corresponding to the vertical direction of the matrix. Selecting a pronunciation in each column of the fuzzy pronunciation space will give a speech recognition result from the time dimension.

[0075] It should be noted that the syllable recognition sub-model can obtain the confidence of the predicted syllable as a certain pronunciation, and determine the pronunciation of the first number with the highest confidence as the alternative pronunciation. Therefore, an attribute can also be added by fuzzying the cell background color of the pronunciation space or the element in the cell to indicate the confidence of the alternative pronunciation.

[0076] The embodiment of the present invention can successfully handle the inaccurate pronunciation phenomenon in the speech by dividing the speech to be recognized into several syllables, recognizing each syllable, obtaining the alternative pronunciation group of each syllable, and forming a fuzzy pronunciation space, thereby improving the recall rate of speech matching.

[0077] Based on the contents of the above embodiments, the specific steps of searching the keyword set according to the fuzzy pronunciation space and obtaining the recognition result of the keyword corresponding to the speech to be recognized include: searching the keyword set according to each alternative syllable in the fuzzy pronunciation space to obtain multiple candidate keywords.

[0078] Specifically, the candidate syllables in the fuzzy pronunciation space are traversed, and for each candidate syllable, a keyword set is searched to obtain a keyword matching the candidate syllable.

[0079] According to the keywords matching each candidate syllable, the preset second-number keywords can be obtained as candidate keywords.

[0080] The preset second number is greater than one and can be set according to actual conditions. The embodiment of the present invention does not limit the specific value of the preset second number.

[0081] The fuzzy pronunciation space is matched with each candidate keyword, and the recognition result of the keyword corresponding to the speech to be recognized is obtained according to the matching result.

[0082] Specifically, for each candidate keyword, a deep matching search is performed to match the candidate keyword with the fuzzy pronunciation space.

[0083] According to the matching results of each candidate keyword and the fuzzy pronunciation space, the candidate keyword with the highest matching degree with the fuzzy pronunciation space can be obtained as the matched candidate keyword.

[0084] According to the matched candidate keywords, the recognition results of the keywords corresponding to the speech to be recognized are obtained.

[0085] The embodiment of the present invention searches a keyword set according to each candidate syllable to obtain multiple candidate keywords, matches each candidate keyword with a fuzzy pronunciation space, and obtains a recognition result of the keyword corresponding to the speech to be recognized based on the matching result. This can shorten the time of the matching process, obtain the recognition result of the keyword corresponding to the speech to be recognized more quickly, and achieve a real-time detection effect.

[0086] Based on the contents of the above embodiments, the keyword set is searched according to each candidate syllable in the fuzzy pronunciation space, and the specific steps of obtaining multiple candidate keywords include: searching the keyword set according to each candidate syllable, and obtaining several keywords in the keyword set that contain the candidate syllable.

[0087] Specifically, for each candidate syllable, a keyword set may be searched according to the candidate syllable to determine the keyword containing the candidate syllable, and obtain a keyword list corresponding to the candidate syllable.

[0088] The keywords in the keyword set can be converted into pinyin, and an index can be created based on the pinyin of the keywords in the keyword set. Each keyword is marked with a label.

[0089] For each pinyin obtained by the conversion, an inverted index of the pinyin is established according to the keywords including the pinyin. The key value of the inverted index is used to represent the keywords including the pinyin.

[0090] For example, the keyword "Are you Boss Wang?" is labeled 1, and its pinyin is nin shi wang lao ban ma; the keyword "Laobai dou di zhu" is labeled 31, and its pinyin is lao bai xing dou di zhu; the keyword "Reward five hundred happy beans" is labeled 40, and its pinyin is jiang li wu bai huan le dou; the pronunciation of "dou" appears in both keywords 31 and 40, and the key value of "dou" in the inverted index is [31,40]; similarly, the key value of "shi" is [1], the key value of "le" is

[40] , the key value of "zhu" is

[31] , and the key value of "bai" is [31,40].

[0091] When searching for a keyword set based on an alternative syllable, the inverted index can be searched to obtain which (or which) keyword contains the alternative syllable based on the key value corresponding to the alternative syllable, so that the candidate keywords can be obtained faster, thereby shortening the search time.

[0092] According to the number of candidate syllables contained in each keyword in the keyword set, a plurality of keywords are determined as a plurality of candidate keywords.

[0093] Specifically, after determining the keywords containing each candidate syllable, the number of candidate syllables contained in each keyword can be counted, that is, the keywords can be voted (or scored) to obtain the voting results of the fuzzy pronunciation space for all keywords, and the keywords with the second largest number of votes can be selected as candidate keywords.

[0094] The embodiment of the present invention obtains the keyword containing each alternative syllable in the keyword set, and determines multiple keywords as multiple candidate keywords according to the number of alternative syllables contained in each keyword in the keyword set. The candidate keywords can be obtained more quickly, thereby shortening the time of the matching process, and the recognition results of the keywords corresponding to the speech to be recognized can be obtained more quickly, thereby achieving the effect of real-time detection.

[0095] Based on the contents of the above embodiments, the fuzzy pronunciation space is matched with each candidate keyword, and the specific steps of obtaining the recognition result of the keyword corresponding to the speech to be recognized according to the matching result include: matching the fuzzy pronunciation space with each candidate keyword, and obtaining the matching degree corresponding to each candidate keyword.

[0096] Specifically, a deep matching search is performed for each candidate keyword, the candidate keyword is matched with the fuzzy pronunciation space, and the matching degree between the candidate keyword and the fuzzy pronunciation space is determined as the matching degree corresponding to the candidate keyword.

[0097] According to the matching degree of each candidate keyword, the recognition result of the keyword corresponding to the speech to be recognized is obtained.

[0098] Specifically, according to the matching degree corresponding to each candidate keyword, it can be determined which candidate keyword has the highest matching degree with the fuzzy pronunciation space, and the candidate keyword is used as the matched candidate keyword.

[0099] According to the matched candidate keywords, the recognition results of the keywords corresponding to the speech to be recognized are obtained.

[0100] The embodiment of the present invention matches each candidate keyword with the fuzzy pronunciation space to obtain the matching degree corresponding to each candidate keyword, and obtains the recognition result of the keyword corresponding to the speech to be recognized based on the matching degree corresponding to each candidate keyword. It can successfully handle similar expressions in speech, word swallowing in speech, and inaccurate pronunciation in speech, and can improve the recall rate of speech matching.

[0101] Based on the contents of the above embodiments, the fuzzy pronunciation space is matched with each candidate keyword, and the specific steps of obtaining the matching degree corresponding to each candidate keyword include: matching each syllable in the candidate keyword with each alternative syllable to obtain the number of matching syllables.

[0102] Specifically, for each candidate keyword, a deep matching search is performed to match the candidate keyword with the fuzzy pronunciation space. Each syllable in the candidate keyword can be matched with each alternative syllable in the fuzzy pronunciation space to determine whether the syllable is the same as a certain alternative syllable.

[0103] If they are the same, it means they match; if they are different, it means they do not match.

[0104] According to the matching results of each syllable in the candidate keyword, the number of syllables matched by the candidate keyword can be determined, and the confidence of each matched syllable can also be obtained.

[0105] According to the confidence of each matched syllable, the total confidence of each matched syllable can be obtained.

[0106] The number of matching syllables may be obtained based on whether the syllables match or not. The number of matching syllables and the sum of the confidences of the matching syllables may also be obtained using a search matrix.

[0107] The search matrix has a total of m rows and n columns, where m is the number of syllables in the candidate keyword and n is the number of syllables in the speech to be recognized or the number of candidate pronunciation groups in the fuzzy pronunciation space.

[0108] For example, the pronunciation sequence of the candidate keyword "Add you on WeChat" is "jia yi xia nin wei xin", which has 6 pronunciations in total, so m=6; the speech to be recognized is "Is it convenient to scan ning WeChat?", which has 8 syllables in total, so n=8.

[0109] The i-th row and j-th column of the search matrix represent the sum of the confidences of the pronunciations that are successfully matched after deep matching of the 1st to 1st pronunciations (i.e., syllables) in the candidate keyword with the 1st to 1st candidate pronunciation groups in the fuzzy pronunciation space. The deep matching process can be transformed into the process of filling in the search matrix.

[0110] When the search matrix is ​​completely filled, it means that the deep search is finished. The element located in the lower right corner of the search matrix is ​​counted as the number of confidence change points, which is the number of syllables that successfully match the candidate keyword with the fuzzy pronunciation space.

[0111] The specific steps to fill in the search matrix include:

[0112] The method for filling in the first column of the search matrix is: if the i-th pronunciation of the candidate keyword appears in the first column of the fuzzy pronunciation space, then the value of the i-th row of the first column of the search matrix is ​​the confidence of the pronunciation in the fuzzy pronunciation space, otherwise it is 0.

[0113] Method for filling the first row of the search matrix: If the pronunciation of the first character of the candidate keyword is in the j-th column of the fuzzy pronunciation space, the value in the j-th column of the first row of the search matrix is the confidence level in the fuzzy space of the pronunciation; otherwise, it is 0.

[0114] Method for filling the remaining elements of the search matrix: When the i-th pronunciation of the candidate keyword is in the j-th alternative pronunciation group in the fuzzy pronunciation space, the value in the j-th column of the i-th row of the search matrix is the value in the (i - 1)-th row and (j - 1)-th column plus the confidence level of this pronunciation in the fuzzy space. When the i-th pronunciation of the candidate keyword is not in the j-th alternative pronunciation group in the fuzzy space, the value in the j-th column of the i-th row of the search matrix is the maximum value of the value in the (i - 1)-th row and j-th column and the value in the i-th row and (j - 1)-th column.

[0115] For example, the pronunciation sequence of the candidate keyword "Add your WeChat" is "jia yi xia nin wei xin". The quick search of the fuzzy pronunciation space according to this candidate keyword is shown in Table 1.

[0116] Table 1 Schematic diagram of quick search in the fuzzy pronunciation space

[0117]

[0118] As shown in Table 1, the first to fourth rows in Table 1 are the fuzzy pronunciation space, and the fifth to tenth rows are the search matrix; the pinyin on the left side of Table 1 is the pronunciation sequence of the candidate keyword. By backtracking the confidence level change points from the lower right corner of the search matrix, the number of matched pronunciations can be obtained as 4.

[0119] The specific backtracking method includes: The pointer first points to the lower right corner of the search matrix, and then observes the elements on its left and above respectively. If there is an element equal to the current element, the pointer moves in the direction corresponding to this element; when the element pointed to by the pointer is not equal to the elements on the left and above, mark this position as the confidence level change point, and the pointer moves to the upper left. And so on until the pointer moves to the upper left corner of the search matrix. For example: The value of the element in the lower right corner of the search matrix is 1.43, and the element on its left is also 1.43, then the pointer moves to the left.

[0120] It can be seen that the candidate keyword is "Add your WeChat", and the input voice is "Is it convenient to scan your ning WeChat?". Among them, "add" in the keyword and "scan" in the voice belong to similar expressions of the same sentence. "Yi xia" in the keyword and "xia" in the voice belong to the omission of a character. "Nin" in the keyword and "ning" in the voice belong to inaccurate pronunciation. This example can prove that this search algorithm can complete the voice matching under the above three interference situations.

[0121] The embodiment of the present invention can automatically match similar expressions of sentences, can detect sentences in the input speech that are similar to the expressions of the sentences to be matched, can correctly handle the situations of swallowing words and multiple words, and can improve the recall rate. For example, "add WeChat" can be recognized as "add WeChat" or "add me WeChat". In addition, using the fuzzy pronunciation space fast search method, the optimal matching result can be searched from two dimensions of pronunciation proximity and matching degree at the same time, and the parameters can be adjusted to achieve different matching effects and meet the needs of different scenarios.

[0122] According to the number of matched syllables and the number of syllables included in the candidate keyword, the matching degree corresponding to the candidate keyword is obtained.

[0123] Specifically, the ratio of the number of matched syllables to the number of syllables included in the candidate keyword can be used as the matching degree corresponding to the candidate keyword. At this time, the matching degree can also be called matching completeness.

[0124] According to Table 1, the matching degree of the candidate keywords is 4 / 6=60%.

[0125] The matching degree corresponding to the candidate keyword may also be obtained according to the number of matched syllables and the total confidence level, and the number of syllables included in the candidate keyword.

[0126] For example, the number of matched syllables can be multiplied by the total confidence of the matched syllables, and then divided by the number of syllables included in the candidate keyword to obtain the matching degree corresponding to the candidate keyword. According to Table 1, the matching degree corresponding to the candidate keyword can be obtained as 1.43×4 / 6=0.95.

[0127] The embodiment of the present invention obtains the number of syllables that match the candidate keyword with the fuzzy pronunciation space, and obtains the matching degree corresponding to the candidate keyword according to the number of matched syllables and the number of syllables included in the candidate keyword. It can successfully handle similar expressions in speech, word swallowing in speech, and inaccurate pronunciation in speech, and can improve the recall rate of speech matching.

[0128] Based on the contents of the above embodiments, according to the matching degree corresponding to each candidate keyword, the specific steps of obtaining the recognition result of the keyword corresponding to the speech to be recognized include: if it is determined that the maximum value of the matching degree corresponding to each candidate keyword is greater than a preset matching threshold, the candidate keyword corresponding to the maximum value is used as the recognition result of the keyword corresponding to the speech to be recognized.

[0129] Specifically, the candidate keyword corresponding to the maximum value of the matching degree corresponding to each candidate keyword is most likely to be the keyword in the speech to be recognized. However, if the maximum value of the matching degree is small, it means that the candidate keyword corresponding to the maximum value is not likely to be the keyword in the speech to be recognized, and may not be the keyword in the speech to be recognized.

[0130] Therefore, it can be determined whether the maximum value of the matching degree corresponding to each candidate keyword is greater than a preset threshold.

[0131] If it is greater than, the candidate keyword corresponding to the maximum value can be used as the recognition result of the keyword corresponding to the speech to be recognized, and the candidate keyword corresponding to the maximum value and the matching degree corresponding to the candidate keyword can be output; if it is less than, a null value can be output as the recognition result of the keyword corresponding to the speech to be recognized, indicating that the speech to be recognized does not contain the content in the keyword set.

[0132] The preset threshold value can be set according to actual conditions. The embodiment of the present invention does not impose any specific restrictions on the specific value of the threshold value.

[0133] For example, when the ratio of the number of matched syllables to the number of syllables included in the candidate keyword is used as the matching degree corresponding to the candidate keyword, the threshold value may be set to 50% or 60%.

[0134] The embodiment of the present invention can obtain a more accurate recognition result by using the candidate keyword corresponding to the maximum value as the recognition result of the keyword corresponding to the speech to be recognized when the maximum value among the matching degrees corresponding to the candidate keywords is greater than a preset matching threshold.

[0135] The following is a description of an apparatus for recognizing keywords in speech provided by an embodiment of the present invention. The apparatus for recognizing keywords in speech described below and the method for recognizing keywords in speech described above can be referred to each other.

[0136] Figure 2 is a schematic diagram of the structure of an apparatus for recognizing keywords in speech according to an embodiment of the present invention. Figure 2 As shown, the device includes a speech recognition module 201 and a space search module 202, wherein:

[0137] The speech recognition module 201 is used to input the speech to be recognized into the speech recognition model and output the fuzzy pronunciation space corresponding to the speech to be recognized;

[0138] A space search module 202, used to search the keyword set according to the fuzzy pronunciation space to obtain the recognition result of the keyword corresponding to the speech to be recognized;

[0139] Among them, the speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent multiple speech recognition results corresponding to the speech to be recognized.

[0140] Specifically, the speech recognition module 201 and the space search module 202 are electrically connected.

[0141] The speech recognition module 201 inputs the speech to be recognized into the speech recognition model. The speech recognition model does not output a unique recognition result, but outputs multiple possible speech recognition results to form a fuzzy pronunciation space.

[0142] The speech recognition module 201 may include a syllable segmentation submodule and a syllable recognition submodule.

[0143] The syllable segmentation submodule is used to segment the speech to be recognized into several syllables.

[0144] The syllable recognition submodule is used to obtain a candidate pronunciation group for each syllable included in the speech to be recognized to form a fuzzy pronunciation space.

[0145] The space search module 202 searches the keyword set according to the fuzzy pronunciation space, and can obtain the keyword with the highest matching degree with the fuzzy pronunciation space as the matched keyword; and obtains the recognition result of the keyword corresponding to the speech to be recognized according to the matched keyword.

[0146] The spatial search module 202 may include a voting submodule and a depth matching submodule.

[0147] The voting submodule is used to search the keyword set according to each candidate syllable in the fuzzy pronunciation space to obtain multiple candidate keywords.

[0148] The deep matching submodule is used to match the fuzzy pronunciation space with each candidate keyword, and obtain the recognition result of the keyword corresponding to the speech to be recognized based on the matching result.

[0149] The voting submodule may include a searching unit and a voting unit.

[0150] A search unit, used for searching a keyword set according to each candidate syllable, and obtaining a number of keywords in the keyword set that contain the candidate syllable;

[0151] The voting unit is used to determine multiple keywords as multiple candidate keywords according to the number of candidate syllables contained in each keyword in the keyword set.

[0152] The depth matching submodule may include a matching unit and an output unit.

[0153] A matching unit, used to match the fuzzy pronunciation space with each candidate keyword to obtain a matching degree corresponding to each candidate keyword;

[0154] The output unit is used to obtain the recognition result of the keyword corresponding to the speech to be recognized according to the matching degree corresponding to each candidate keyword.

[0155] The matching unit includes a matching search subunit and a matching degree acquisition subunit.

[0156] The matching search subunit is used to match each syllable in the candidate keyword with each alternative syllable to obtain the number of matching syllables.

[0157] The matching degree obtaining subunit is used to obtain the matching degree corresponding to the candidate keyword according to the number of matched syllables and the number of syllables included in the candidate keyword.

[0158] The output unit is specifically used to use the candidate keyword corresponding to the maximum value as the recognition result of the keyword corresponding to the speech to be recognized if it is determined that the maximum value among the matching degrees corresponding to the candidate keywords is greater than a preset matching threshold.

[0159] The device for recognizing keywords in speech provided in an embodiment of the present invention is used to execute the method for recognizing keywords in speech provided in the above-mentioned embodiments of the present invention. The specific methods and processes for implementing corresponding functions of each module included in the device for recognizing keywords in speech are detailed in the embodiments of the method for recognizing keywords in speech above, and will not be repeated here.

[0160] The device for recognizing keywords in speech is used in the method for recognizing keywords in speech in the above embodiments. Therefore, the description and definition in the method for recognizing keywords in speech in the above embodiments can be used for understanding each execution module in the embodiments of the present invention.

[0161] The embodiment of the present invention recognizes the speech to be recognized through a speech recognition model, obtains multiple possible speech recognition results, forms a fuzzy pronunciation space, matches and searches the fuzzy pronunciation space with a pre-established keyword set, and outputs the matched keywords. The method of using the fuzzy pronunciation space search can successfully handle similar expressions in speech, word swallowing in speech, and inaccurate pronunciation in speech, and can improve the recall rate of speech matching.

[0162] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3As shown, the electronic device may include: a processor 301, a memory 302 and a bus 303; wherein the processor 301 and the memory 302 communicate with each other through the bus 303; the processor 301 is used to call computer program instructions stored in the memory 302 and can be run on the processor 301, so as to execute the method for recognizing keywords in speech provided by the above-mentioned method embodiments, the method comprising: inputting the speech to be recognized into the speech recognition model, and outputting the fuzzy pronunciation space corresponding to the speech to be recognized; according to the fuzzy pronunciation space, searching the keyword set to obtain the recognition result of the keyword corresponding to the speech to be recognized; wherein the speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent a variety of speech recognition results corresponding to the speech to be recognized.

[0163] In addition, the logic instructions in the above-mentioned memory 302 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0164] On the other hand, an embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method for recognizing keywords in speech provided by the above-mentioned method embodiments, the method including: inputting the speech to be recognized into a speech recognition model, and outputting a fuzzy pronunciation space corresponding to the speech to be recognized; according to the fuzzy pronunciation space, searching a keyword set to obtain recognition results of keywords corresponding to the speech to be recognized; wherein the speech recognition model is obtained after training based on a sample speech signal of a speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one to the sample speech signal; the fuzzy pronunciation space is used to represent a variety of speech recognition results corresponding to the speech to be recognized.

[0165] On the other hand, an embodiment of the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the method for recognizing keywords in speech provided in the above embodiments, the method comprising: inputting the speech to be recognized into a speech recognition model, and outputting a fuzzy pronunciation space corresponding to the speech to be recognized; searching a keyword set according to the fuzzy pronunciation space to obtain recognition results of the keywords corresponding to the speech to be recognized; wherein the speech recognition model is obtained after training based on a sample speech signal of a speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one to the sample speech signal; the fuzzy pronunciation space is used to represent a variety of speech recognition results corresponding to the speech to be recognized.

[0166] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0167] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for recognizing keywords in speech, characterized in that: include: Inputting the speech to be recognized into the speech recognition model, and outputting the fuzzy pronunciation space corresponding to the speech to be recognized; Searching a keyword set according to the fuzzy pronunciation space to obtain a recognition result of the keyword corresponding to the speech to be recognized; The speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent a variety of speech recognition results corresponding to the speech to be recognized; The specific steps of searching the keyword set according to the fuzzy pronunciation space to obtain the recognition result of the keyword corresponding to the speech to be recognized include: Searching the keyword set according to each candidate syllable in the fuzzy pronunciation space to obtain a plurality of candidate keywords; The fuzzy pronunciation space is matched with each of the candidate keywords, and a recognition result of the keyword corresponding to the speech to be recognized is obtained according to the matching result.

2. The method for recognizing keywords in speech according to claim 1, characterized in that: The specific steps of inputting the speech to be recognized into the speech recognition model and outputting the fuzzy pronunciation space corresponding to the speech to be recognized include: Segmenting the speech to be recognized into a plurality of syllables; The candidate pronunciation groups of each syllable included in the speech to be recognized are obtained to form the fuzzy pronunciation space.

3. The method for recognizing keywords in speech according to claim 1, characterized in that: The specific steps of searching the keyword set according to each candidate syllable in the fuzzy pronunciation space to obtain a plurality of candidate keywords include: Search the keyword set according to each candidate syllable to obtain a plurality of keywords in the keyword set that include the candidate syllable; According to the number of the candidate syllables contained in each keyword in the keyword set, a plurality of the keywords are determined as the plurality of candidate keywords.

4. The method for recognizing keywords in speech according to claim 1, characterized in that: The specific steps of matching the fuzzy pronunciation space with each of the candidate keywords and obtaining the recognition result of the keyword corresponding to the speech to be recognized according to the matching result include: Matching the fuzzy pronunciation space with each candidate keyword to obtain a matching degree corresponding to each candidate keyword; According to the matching degree corresponding to each of the candidate keywords, the recognition result of the keyword corresponding to the speech to be recognized is obtained.

5. The method for recognizing keywords in speech according to claim 4, characterized in that: The specific steps of matching the fuzzy pronunciation space with each candidate keyword and obtaining the matching degree corresponding to each candidate keyword include: Matching each syllable in the candidate keyword with each of the candidate syllables to obtain the number of matching syllables; According to the number of the matched syllables and the number of syllables included in the candidate keyword, the matching degree corresponding to the candidate keyword is obtained.

6. The method for recognizing keywords in speech according to claim 4, characterized in that: The specific steps of obtaining the recognition result of the keyword corresponding to the speech to be recognized according to the matching degree corresponding to each candidate keyword include: If it is determined that the maximum value of the matching degrees corresponding to the candidate keywords is greater than a preset matching threshold, the candidate keyword corresponding to the maximum value is used as the recognition result of the keyword corresponding to the speech to be recognized.

7. A device for recognizing keywords in speech, characterized in that: include: A speech recognition module, used for inputting the speech to be recognized into the speech recognition model, and outputting the fuzzy pronunciation space corresponding to the speech to be recognized; A space search module, used to search a keyword set according to the fuzzy pronunciation space to obtain a recognition result of the keyword corresponding to the speech to be recognized; The spatial search module includes a voting submodule and a depth matching submodule; The voting submodule is used to search the keyword set according to each candidate syllable in the fuzzy pronunciation space to obtain multiple candidate keywords; The deep matching submodule is used to match the fuzzy pronunciation space with each candidate keyword, and obtain the recognition result of the keyword corresponding to the speech to be recognized according to the matching result; Among them, the speech recognition model is obtained after training based on the sample speech signal of the speech sample and the corresponding pronunciation; the pronunciation is predetermined based on the speech sample and corresponds one-to-one with the sample speech signal; the fuzzy pronunciation space is used to represent multiple speech recognition results corresponding to the speech to be recognized.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for recognizing keywords in speech as described in any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for recognizing keywords in speech as claimed in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Voice recognition method and device, server and storage medium

    CN110310631A

  • Speech recognition program medium, speech recognition apparatus, and speech recognition method

    US20180240460A1