Speech recognition real-time enhancement method and system based on words in finite word library

By constructing a finite lexicon library and using the pinyin distance matching strategy, the problem of low speech recognition accuracy is solved, and higher recognition accuracy and error correction capabilities are achieved.

CN120089132AActive Publication Date: 2025-06-03GUANGZHOU UNIVERSITY OF CHINESE MEDICINE
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510098097.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-03
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The accuracy of speech recognition in the prior art is not high, especially when pronunciation is not standard and contains proper nouns.

Method used

By constructing a finite lexicon library, storing multiple words in the library and their corresponding pinyin strings, and using the pinyin distance matching strategy, the accuracy of speech recognition is enhanced in real time.

Benefits of technology

It improves the accuracy of speech recognition, reduces recognition errors caused by unstandard pronunciation and homophones, and can correct wrongly recognized words when there is a deviation in the preliminary recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089132A_ABST
    Figure CN120089132A_ABST
Patent Text Reader

Abstract

The invention discloses a speech recognition real-time enhancement method and system based on words in an endless word library, and the method comprises the following steps: constructing the endless word library, the endless word library comprises a plurality of words in the library and a plurality of first pinyin strings, and each first pinyin string is a pinyin string corresponding to one word in the library; according to user input voice, performing preliminary voice recognition to obtain output words; converting the output word into a pinyin string to obtain a second pinyin string; matching the second pinyin string in the lexicon with the exhaustion to obtain a first pinyin string with the nearest pinyin distance; and according to the obtained first pinyin string with the nearest pinyin distance, outputting the corresponding words in the library. Compared with a traditional speech recognition method, the speech recognition method has the advantages that recognition errors caused by nonstandard pronunciation, homophones and the like are reduced, and the speech recognition accuracy is improved; and meanwhile, the complexity of retrieval matching is controlled, and the real-time performance of identification is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition enhancement, and more specifically, to a real-time speech recognition enhancement method and system based on words within a finite vocabulary. Background Art

[0002] As an important research field in the field of intelligent recognition, speech recognition technology has a development history of more than 60 years. Speech recognition is the process of recognizing sound into text, and Chinese speech recognition is to convert speech into Chinese text according to the pronunciation of the speaker. However, due to individual pronunciation differences, non-standard Chinese pinyin pronunciations will greatly reduce the accuracy of speech recognition.

[0003] With the development of technology, both the accuracy and speed of speech recognition have made great progress. In quiet environments, with standard accents and common vocabulary scenarios, the accuracy rate of speech recognition has reached a level similar to that of humans. However, if the speaker's pronunciation is not standard and the conversation contains proper nouns, such as specific names of people, places, organizations, etc., the recognition accuracy rate will be greatly reduced. Summary of the Invention

[0004] One of the objectives of the present invention is to provide a real-time speech recognition enhancement method based on words within a finite vocabulary to solve the technical problem of low speech recognition accuracy in the prior art; another objective of the present invention is to provide a real-time speech recognition enhancement system based on words within a finite vocabulary.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] The first aspect of the present invention provides a real-time speech recognition enhancement method based on words within a finite vocabulary, including the following steps:

[0007] Construct a finite vocabulary, which includes storing a plurality of words within the vocabulary and a plurality of first pinyin strings, and each of the first pinyin strings is a pinyin string corresponding to a word within the vocabulary;

[0008] Based on the user input speech, perform preliminary speech recognition to obtain an output word;

[0009] Convert the output word into a pinyin string to obtain a second pinyin string;

[0010] Match the second pinyin string in the finite vocabulary to obtain the first pinyin string with the closest pinyin distance;

[0011] Based on the obtained first pinyin string with the closest pinyin distance, output the corresponding word within the vocabulary.

[0012] Further, the construction of the finite vocabulary includes:

[0013] Store a finite number of in-library words as a first word list, and sort them in ascending order according to the internal machine code of Chinese characters. After sorting, form a second word list; make a deep copy of the second word list to form a third word list.

[0014] Store all the first pinyin strings as a first pinyin string list, and generate a multi-level subset system based on the first pinyin string list to construct a multi-level pinyin space.

[0015] Further, generating a multi-level subset system based on the first pinyin string list to construct a multi-level pinyin space includes:

[0016] 3A: Record the position numbers of all the first pinyin strings in the first pinyin string list to obtain a first position list;

[0017] 3B: All the first pinyin strings form an initial set;

[0018] 3C: Count and record the number of first pinyin strings in the current set. If the number of first pinyin strings in the current set is less than the first quantity threshold, there is no need to split; if the number of first pinyin strings in the current set is not less than the first quantity threshold, execute steps 3D to 3I;

[0019] 3D: In the current set, take the two first pinyin strings with the largest pinyin distance as the initial two subset center points; judge whether the split number threshold is 2. If so, skip steps 3E to 3G and execute step 3H;

[0020] 3E: Take any first pinyin string in the current set that is not listed as a subset center point as the third pinyin string, calculate the pinyin distance between the third pinyin string and each existing subset center, and then count the sum of all pinyin distances and the coefficient of variation to obtain the pinyin distance sum and the pinyin distance coefficient of variation;

[0021] 3F: When the pinyin distance coefficient of variation is less than the first coefficient of variation threshold, take the third pinyin string corresponding to the largest pinyin distance sum as the new subset center point;

[0022] 3G: Judge whether the existing subset center points have reached the split number threshold. If not, return to step 3E;

[0023] 3H: Calculate the pinyin distance between all the first pinyin strings in the current set and each subset center point, and according to the nearest principle, classify each first pinyin string into one of the existing subsets;

[0024] 3I: Take each subset as the current set and execute step 3C.

[0025] Further, the pinyin distance includes:

[0026] Define a pinyin converted from a character as a pinyin unit, and a pinyin unit includes an initial, a final and a tone;

[0027] Calculate the absolute value p of the difference in the number of pinyin units between two pinyin strings, and record PinyinDistance as 7*p;

[0028] Calculate the gap between the pinyin units at each same position of the two pinyin strings, and accumulate it into PinyinDistance to obtain the pinyin distance between the two pinyin strings.

[0029] Further, the construction of the finite vocabulary also includes actively adding the binary group of the recognized word and the real word, including:

[0030] Obtain the user input voice and recognize the output word;

[0031] Judge whether the output word is one in the second word list. If so, end the process; if not, provide several words in the library, and the user selects the word in the library that the user input voice really wants to express as the third word, and inserts the output word into the second word list in ascending order of the machine code of the Chinese characters, record the insertion position, and insert the third word into the same position in the third word list.

[0032] Further, after initially performing speech recognition on the user input voice to obtain the output word, it also includes recording the current time.

[0033] Further, before matching the second pinyin string in the finite vocabulary to obtain the first pinyin string with the closest pinyin distance, it also includes:

[0034] Perform a binary search on the output word in the second word list. If found, use the word in the same position in the third word list as the output to end the real-time enhancement of speech recognition. If not found, perform the subsequent steps.

[0035] Further, matching the second pinyin string in the finite vocabulary to obtain the first pinyin string with the closest pinyin distance includes:

[0036] 7A: If the number of the first pinyin strings in the current set is less than the first quantity threshold, perform steps 7B to 7C; if the number of the first pinyin strings in the current set is not less than the first quantity threshold, perform steps 7D to 7E;

[0037] 7B: Calculate the pinyin distance between the second pinyin string and each first pinyin string in the current set;

[0038] 7C: Using the pinyin distance from smallest to largest as the primary sorting condition. When the pinyin distances are the same, for the output word and the in-library word corresponding to the first pinyin string, according to the principle of matching each character one by one and finally calculating the number of different Chinese characters, obtain the Chinese character string distance. Using the Chinese character string distance from smallest to largest as the secondary sorting condition, obtain the first matchNum pinyin strings in the current set that are the closest to the second pinyin string, and output the matchNum first pinyin strings.

[0039] 7D: Calculate the pinyin distances between the second pinyin string and the centers of multiple subsets in the current set. According to the smallest K distance values, obtain K candidate subsets.

[0040] 7E: Take each of the K candidate subsets in turn as the current set, and return to step 7A; each time, obtain matchNum first pinyin strings. Using the pinyin distance from smallest to largest as the primary sorting condition and the Chinese character string distance from smallest to largest as the secondary sorting condition, select the first matchNum first pinyin strings that are the closest to the second pinyin string from all the obtained first pinyin strings, and output the matchNum first pinyin strings.

[0041] Further, according to the first pinyin strings with the closest pinyin distances obtained, output the corresponding in-library words, including:

[0042] 9A: In the first word list, obtain the in-library words corresponding to the matchNum first pinyin strings output in step 7E.

[0043] 9B: Compare the current time with the stored time. If the output words are the same for both times and the archive time difference is not less than the preset time threshold, then execute step 9C; if the output words are different for both times, or the archive time difference is less than the preset time threshold, then execute step 9D.

[0044] 9C: Provide the user with matchNum in-library word options in the form of text or voice, and finally determine the unique in-library word according to the user's answer and output it; before output, create a new subprocess. In this subprocess, if it is monitored to be idle, insert the output word into the second word list in the order of the machine internal code of the Chinese characters from smallest to largest.

[0045] 9D: Use the one ranked first among the matchNum in-library words as the output.

[0046] The second aspect of the present invention provides a real-time enhancement system for speech recognition based on in-library words of a finite word library, including:

[0047] A word library module, which constructs a finite word library. The finite word library includes storing a finite number of in-library words and multiple first pinyin strings, and each first pinyin string is the pinyin string corresponding to an in-library word.

[0048] A preliminary recognition module, which preliminarily recognizes the input speech of the user to obtain an output word through preliminary speech recognition;

[0049] A conversion module, which converts the output word into a pinyin string to obtain a second pinyin string;

[0050] A matching module, which matches the second pinyin string in the finite vocabulary to obtain a first pinyin string with the closest pinyin distance;

[0051] An output module, which outputs a corresponding in-library word according to the obtained first pinyin string with the closest pinyin distance.

[0052] Compared with the prior art, the beneficial effect of the technical solution of the present invention is:

[0053] By constructing a finite vocabulary and converting the output word obtained through preliminary speech recognition into a pinyin string and then matching it in the vocabulary, the present invention can more accurately find an in-library word that conforms to the actual intention of the user. Compared with traditional speech recognition methods, the present invention reduces recognition errors caused by inaccurate pronunciation, homophones, etc., and improves the accuracy of speech recognition. Using the matching strategy of the closest pinyin distance, when there are deviations in the preliminary recognition results, the closest correct pinyin string can be found, thereby correcting the misrecognized words. This matching method takes into account problems such as inaccurate pronunciation and accent differences that may occur during the speech recognition process, further improving the recognition accuracy. Description of the Drawings

[0054] Figure 1 It is a schematic flowchart of a real-time enhancement method for speech recognition of in-library words based on a finite vocabulary provided by an embodiment of the present invention;

[0055] Figure 2 It is a schematic diagram of the modules of a real-time enhancement system for speech recognition of in-library words based on a finite vocabulary provided by an embodiment of the present invention. Detailed Embodiments

[0056] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0057] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;

[0058] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0059] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0060] Embodiment 1

[0061] This embodiment provides a real-time enhancement method for speech recognition based on words in a finite word library, as Figure 1 shown, which includes the following steps:

[0062] Construct a finite word library, where the finite word library includes a storage of multiple in-library words and multiple first pinyin strings, and each of the first pinyin strings is a pinyin string corresponding to an in-library word;

[0063] Based on the user input speech, perform preliminary speech recognition to obtain an output word;

[0064] Convert the output word into a pinyin string to obtain a second pinyin string;

[0065] Match the second pinyin string in the finite word library to obtain the first pinyin string with the closest pinyin distance;

[0066] According to the obtained first pinyin string with the closest pinyin distance, output the corresponding in-library word.

[0067] In a further embodiment, the construction of the finite word library includes:

[0068] Store a finite number of in-library words as a first word list, and sort them in ascending order according to the machine code of Chinese characters. After sorting, form a second word list; make a deep copy of the second word list to form a third word list.

[0069] Store all the first pinyin strings as a first pinyin string list, and generate a multi-level subset system based on the first pinyin string list to construct a multi-level pinyin space.

[0070] In a specific embodiment, the rule for converting each Chinese character into pinyin is: the pinyin of each character is divided into three parts: initial, final, and tone. If there are polyphonic characters, the system prompts manual correction and storage. In this embodiment, the number of in-library words is 1273, and after storage, they are recorded in the first word list.

[0071] In a further embodiment, generating a multi-level subset system based on the first pinyin string list to construct a multi-level pinyin space includes:

[0072] 3A: Record the position numbers of all the first pinyin strings in the first pinyin string list to obtain a first position list;

[0073] 3B: All the first pinyin strings form an initial set;

[0074] 3C: Count and record the number of the first pinyin strings in the current set. If the number of the first pinyin strings in the current set is less than the first quantity threshold, there is no need to split; if the number of the first pinyin strings in the current set is not less than the first quantity threshold, then execute steps 3D to 3G. In this embodiment, the first quantity threshold is default set to 1000, and the default number of pinyin strings or in-library words returned by manual setting is 4. Since the number of in-library words in this embodiment is 1273, which is greater than the first quantity threshold, therefore, in the first step, execute steps 3D to 3I first;

[0075] 3D: In the current set, take the two first pinyin strings with the largest pinyin distance as the initial two subset center points. Specifically:

[0076] Specifically:

[0077] The first center point is Isodon tenuiflorus (xi4hua1xian4wen2xiang1cha2cai4), and the second center point is Turpinia arguta var. acutata (rong2mao2rui4jian1shan1xiang1yuan2). The distance between the two is calculated according to the method set for "the pinyin distance between two pinyin strings" in this patent. After calculation, it is obtained that: the length difference weight = 0, the initial consonant difference = 42, the final consonant difference = 42, the tone difference = 15, and the total distance = 99;

[0078] Supplementary note: xi4hua1xian4wen2xiang1cha2cai4 represents the pinyin of seven characters. x is the initial consonant, the final consonant is i, and the pinyin string with the 4th tone. And so on for others.

[0079] Judge whether the split number threshold is 2, and the result is no. Continue to execute in sequence. In this embodiment, the split number threshold can be adjusted according to the actual situation, and it is set to 4 in this embodiment;

[0080] 3E: Take any first pinyin string in the current set that is not listed as a subset center point as the third pinyin string, calculate the pinyin distance between the third pinyin string and each existing subset center, and then count the sum of all pinyin distances and the coefficient of variation to obtain the pinyin distance sum and the pinyin distance coefficient of variation;

[0081] 3F: When the pinyin distance coefficient of variation is less than the first coefficient of variation threshold, take the third pinyin string corresponding to the largest pinyin distance sum as the new subset center point;

[0082] 3G: Judge whether the existing subset center points have reached the split number threshold. If not, return to step 3E; as mentioned before, the split number threshold is set to 4 in this embodiment;

[0083] All the center points of this split include:

[0084] Initial central points: Rabdosia tenuiflora, Turpinia arguta var. acutata

[0085] Newly added central points: Liparis viridiflora, Liriope cymbispatha, a total of 4 central points.

[0086] 3H: Calculate the pinyin distance between all the first pinyin strings in the current set and the pinyin of each subset central point. According to the nearest principle, classify each first pinyin string into one of the existing subsets.

[0087] For example, the entry "chrysanthemum" (ju2hua1):

[0088] Distance from Rabdosia tenuiflora: 50

[0089] Distance from Turpinia arguta var. acutata: 62

[0090] Distance from Liparis viridiflora: 45

[0091] Distance from Liriope cymbispatha: 51

[0092] The nearest central point is Liparis viridiflora, so "chrysanthemum" is assigned to subset 3 (the subset with "Liparis viridiflora" as the central point).

[0093] 3I: Take each subset as the current set and execute step 3C.

[0094] The number of words within each subset finally:

[0095] Subset 1 (central point Rabdosia tenuiflora): 439 words.

[0096] Subset 2 (central point Turpinia arguta var. acutata): 63 words.

[0097] Subset 3 (Liparis viridiflora (Changjing Yangersuan)): 190 words.

[0098] Subset 4 (Liriope muscari (Duan Ting Shan Maidong)): 581 words.

[0099] Since the number of words in each subset is less than the first quantity threshold of 1000, this split is completed and no further recursive splitting is required.

[0100] In a further embodiment, the pinyin distance includes:

[0101] Defining the pinyin converted from a character as a pinyin unit, and a pinyin unit includes an initial, a final and a tone;

[0102] Calculating the absolute value p of the difference in the number of pinyin units of two pinyin strings, and recording PinyinDistance as 7*p;

[0103] Calculating the difference between the pinyin units at each same position of two pinyin strings and accumulating it into PinyinDistance to obtain the pinyin distance between the two pinyin strings.

[0104] In a specific embodiment, calculating the difference between the pinyin units at each same position of two pinyin strings has the following rules:

[0105] 1) If the tones are different, the difference is 3.

[0106] 2) If the initials are different but belong to a pair of ambiguous sounds, such as z and zh, the difference is 2. If the initials are different and do not belong to a pair of ambiguous sounds, the difference is 6. The definition rule of the pair of ambiguous sounds is in 10.4.6.

[0107] 3) If the finals are different but belong to a pair of ambiguous sounds, such as an and ang, the difference is 2. If the finals are different and do not belong to a pair of ambiguous sounds, the difference is 6. The definition rule of the pair of ambiguous sounds is in 6).

[0108] 4) For a pinyin without an initial, the initial is recorded as empty. For example: the pinyin of "er". The empty initial and the non-empty initial are also different, and the difference is 6.

[0109] 5) Accumulating the differences between a pinyin unit and another pinyin unit in the three parts of tone, initial and final as the difference between them.

[0110] 6) z and zh, c and ch, s and sh, an and ang, en and eng, in and ing, ian and iang, uan and uang. The user can customize the pair of ambiguous sounds.

[0111] In the above rules, the specific values of each difference can be set as needed. Only one example is given in this embodiment.

[0112] Meanwhile, the embodiment of the present invention also gives the following specific process for calculating the Pinyin distance:

[0113] In the first step, calculate the length difference weight. The number of syllables of the two Pinyin strings are respectively:

[0114] Rabdosia tenuiflora: xi4, hua1, xian4, wen2, xiang1, cha2, cai4, a total of 7 syllables.

[0115] Turpinia arguta var. acutisepala: rong2, mao2, rui4, jian1, shan1, xiang1, yuan2, a total of 7 syllables.

[0116] Then the length difference p is:

[0117] |7 - 7| = 0

[0118] Then record PinyinDistance as 7 * p = 7 * 0 = 0;

[0119] In the second step, calculate the initial consonant difference. The calculation process is:

[0120] Comparison of initial consonants of corresponding syllables and calculation of differences:

[0121]

[0122]

[0123] Total initial consonant difference: 6 + 6 + 6 + 6 + 6 + 6 + 6 = 42 Result: 42;

[0124] In the third step, calculate the final consonant difference. The calculation process is: Comparison of finals of corresponding syllables and calculation of differences:

[0125]

[0126] Total final consonant difference: 6 + 6 + 6 + 6 + 6 + 6 + 6 = 42 Result: 42

[0127] In the fourth step, calculate the tone difference:

[0128] Comparison of tones of corresponding syllables and calculation of differences:

[0129]

[0130]

[0131] Total tone difference: 3 + 3 + 0 + 3 + 0 + 3 + 3 = 15

[0132] Then the total pinyin distance = 0 (length difference weight) + 42 (initial consonant difference) + 42 (final consonant difference) + 15 (tone difference) = 99.

[0133] In a further embodiment, the construction of the finite vocabulary further includes actively adding the binary group of the recognition word and the real word, and dynamically adjusting the corresponding relationship between the recognition word and the real word under interaction, including:

[0134] Obtain the user input voice and recognize the output word;

[0135] Judge whether the output word is one of the words in the second word list. If so, end the process; if not, provide several words in the library, and the user selects the word in the library that the user input voice really wants to express as the third word, and inserts the output word into the second word list in ascending order of the internal code of Chinese characters, record the insertion position, and insert the third word into the same position in the third word list. Specifically:

[0136] The user orally states the target word "ginseng" to the system through voice. The system receives the voice signal and generates the output word "human body" through the voice recognition module. Check whether the output word "human body" is one of the words in the second word list. The result is "no". Then, the system successively provides ginseng, scrophularia, lotus, and salvia miltiorrhiza for the user to select. The user selects "ginseng". Then the system sorts in ascending order of the internal code of Chinese characters and determines the insertion position 1270 of "human body" in the second word list through the binary search algorithm. Immediately, insert "human body" into the 1270th position of the second word list, and then insert "ginseng" into the 1270th position of the third word list.

[0137] By setting the third word list, it is ensured that the words in the third word list are all correct words. When the user orally states ginseng again next time and the system recognizes it as human body, the system can immediately find human body in the second word list and find the correct word ginseng in the third word list.

[0138] Embodiment 2

[0139] On the basis of Embodiment 1, this embodiment further discloses the following content:

[0140] In this embodiment, after using the third-party voice recognition module to recognize the user input voice and obtaining the output word, it further includes recording the current time.

[0141] In a further embodiment, before matching the second pinyin string in the finite vocabulary to obtain the first pinyin string with the closest pinyin distance, it further includes:

[0142] Perform a binary search on the output word in the second word list. If found, use the found word in the library as the output and end the real-time enhancement of voice recognition. If not found, execute the subsequent steps.

[0143] Specifically:

[0144] The user orally states the target word "Coptis chinensis" through voice, and the system calls a third-party voice recognition module to recognize and obtain the output word "Coptis chinensis". Archive the current output word "Coptis chinensis" and record the current archiving time. Perform a binary search for the output word in the second word list. The result is "not found". Execute the subsequent steps.

[0145] Convert the output word Coptis chinensis from a Chinese character string to a pinyin string huang2lian2.

[0146] In a further embodiment, match the second pinyin string in the finite vocabulary to obtain the first pinyin string with the closest pinyin distance, including:

[0147] 7A: If the number of first pinyin strings in the current set is less than the first quantity threshold, execute steps 7B to 7C; if the number of first pinyin strings in the current set is not less than the first quantity threshold, execute steps 7D to 7E; specifically, the number of pinyin strings in the current set is 1273, which is not less than the first quantity threshold, steps 7D to 7E;

[0148] 7B: Calculate the pinyin distance between the second pinyin string and each first pinyin string in the current set;

[0149] 7C: Take the pinyin distance from smallest to largest as the primary sorting condition. When the pinyin distances are the same, according to the principle of matching the output word and the in-library words corresponding to the first pinyin string character by character and finally calculating the number of different Chinese characters, obtain the Chinese character string distance, and take the Chinese character string distance from smallest to largest as the secondary sorting condition to obtain the first matchNum first pinyin strings in the current set that are closest to the second pinyin string, and output the matchNum first pinyin strings;

[0150] 7D: Calculate the pinyin distance between the second pinyin string and the centers of multiple subsets in the current set, and obtain K candidate subsets according to the smallest K distance values; specifically:

[0151] Calculate the pinyin distance between Coptis chinensis (huang2lian2) and the centers of 4 subsets in the current set:

[0152] Distance from Rabdosia tenuiflora (xi4hua1xian4wen2xiang1cha2cai4): 65

[0153] Distance from Turpinia arguta var. velutina (rong2mao2rui4jian1shan1xiang1yuan2): 59

[0154] Distance from Liparis viridiflora (chang2jing1yang2er3suan4): 44

[0155] Distance from Liriope muscari (Decne.) Baily var. abbreviata Y. T. Ma: 40

[0156] Based on the smallest first K (specifically, K is set to 1 in this embodiment) distance values, determine 1 candidate subset - the subset with Liriope muscari (Decne.) Baily var. abbreviata Y. T. Ma as the center point.

[0157] 7E: Take each of the K candidate subsets one by one as the current set, and return to step 7A; each time, obtain matchNum first pinyin strings. Using the pinyin distance from smallest to largest as the primary sorting condition and the Chinese character string distance from smallest to largest as the secondary sorting condition, select the first matchNum first pinyin strings that are the closest to the second pinyin string from all the obtained first pinyin strings, and output the matchNum first pinyin strings. Specifically:

[0158] Take the subset with Liriope muscari (Decne.) Baily var. abbreviata Y. T. Ma as the center point as the current set, and execute step 7A. The steps are as follows:

[0159] 7A: The number of pinyin strings in the current set is 581, which is less than the first quantity threshold. Execute steps 7B to 7C to obtain the output.

[0160] 7B: Calculate the pinyin distance between Coptis chinensis Franch. and each pinyin string in the current set;

[0161] 7C: Use the pinyin distance from smallest to largest as the primary sorting condition. When the pinyin distances are the same, according to the principle of calculating the number of different Chinese characters by matching each character one by one between the output word and the in-library word corresponding to the pinyin string, obtain the Chinese character string distance. Use the Chinese character string distance from smallest to largest as the secondary sorting condition. Obtain the first 4 pinyin strings huang2lan2, cong1lian2, huang2jin3, hong2liao that are the closest to huang2lian2 in the current set. Output these 4 pinyin strings and return.

[0162] In a further embodiment, according to the first pinyin strings with the closest pinyin distance obtained, output the corresponding in-library words, including:

[0163] 9A: In the first word list, obtain the in-library words corresponding to the matchNum first pinyin strings output in step 7E; specifically:

[0164] In the first word list, read out the in-library words corresponding to the 4 pinyin strings returned in step 7E;

[0165] 9B: Compare the current time with the stored time. If the output words are the same for both times and the difference in the archive times is not less than the preset time threshold, then execute step 9C; if the output words are different for both times or the difference in the archive times is less than the preset time threshold, then execute step 9D. Specifically:

[0166] Compare the output words and archive times of this time and the last time. The output words are different for both times and the difference in the archive times is greater than 60 seconds. Take the word ranked first among the 4 in-library words as the output. One system real-time service ends;

[0167] 9C: Provide the user with matchNum in-library word options in the form of text or voice, and finally determine the unique in-library word according to the user's answer and output it; before output, create a new subprocess. In this subprocess, if it is detected that it is idle, insert the output word into the second word list in the order of the machine codes of Chinese characters from small to large;

[0168] 9D: Take the word ranked first among the matchNum in-library words as the output.

[0169] Embodiment 3

[0170] Based on the speech recognition real-time enhancement method of in-library words of finite word libraries in Embodiment 1 and Embodiment 2, this embodiment gives another specific embodiment of recognition enhancement:

[0171] The user orally states the target word "ginseng" through voice, and the system calls a third-party speech recognition module to recognize and obtain the output word "human body". Archive the output word "human body" of this time and record the archive time of this time. Perform a binary search for the output word "human body" in the second word list. The result is "found", and the position number 1270 is obtained. Take the word "ginseng" at the 1270th position of the third word list as the system output. End this method. One system real-time service ends.

[0172] Embodiment 4

[0173] This embodiment provides a speech recognition real-time enhancement system based on in-library words of a finite word library, as Figure 2 shown. The speech recognition real-time enhancement system applies the speech recognition real-time enhancement method of in-library words of finite word libraries described in Embodiments 1 and 2, and includes:

[0174] A word library module that constructs a finite word library, and the finite word library includes a plurality of in-library words and a plurality of first pinyin strings, and each first pinyin string is a pinyin string corresponding to an in-library word;

[0175] A preliminary recognition module that performs preliminary speech recognition on the user input voice to obtain an output word;

[0176] A conversion module that converts the output word into a pinyin string to obtain a second pinyin string;

[0177] A matching module that matches the second pinyin string in the finite vocabulary to obtain a first pinyin string with the closest pinyin distance;

[0178] An output module that outputs the corresponding in-library word according to the obtained first pinyin string with the closest pinyin distance.

[0179] Identical or similar reference numerals correspond to identical or similar components;

[0180] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;

[0181] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A real-time enhancement method for speech recognition based on words in a finite vocabulary, characterized in that: The following steps are involved: Constructing a finite word library, wherein the finite word library includes storing a plurality of words in the library and a plurality of first pinyin strings, each of the first pinyin strings being a pinyin string corresponding to a word in the library; According to the user's input voice, preliminary voice recognition is performed to obtain output words; Convert the output word into a pinyin string to obtain a second pinyin string; Matching the second pinyin string in the finite word library to obtain the first pinyin string with the nearest pinyin distance; According to the obtained first pinyin string of the nearest neighbor, the corresponding word in the library is output.

2. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 1, characterized in that: The constructing of a finite vocabulary library comprises: The finite words in the library are stored as the first word list, and are sorted from small to large according to the machine internal code of the Chinese characters to form the second word list after sorting; the second word list is deeply copied to form the third word list. All first Pinyin strings are stored as a first Pinyin string list, and a multi-level subset system is generated according to the first Pinyin string list to construct a multi-level Pinyin space.

3. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 2, characterized in that: Generate a multi-level subset system according to the first pinyin string list to construct a multi-level pinyin space, including: 3A: Record the position numbers of all the first pinyin strings in the first pinyin string list to obtain a first position list; 3B: All the first pinyin strings form an initial set; 3C: Count and record the number of first pinyin strings in the current set. If the number of first pinyin strings in the current set is less than the first quantity threshold, there is no need to split; if the number of first pinyin strings in the current set is not less than the first quantity threshold, execute steps 3D to 3I; 3D: In the current set, take the two first pinyin strings with the largest pinyin distance as the initial two subset center points; determine whether the split number threshold is 2, if so, skip step 3E to step 3G, and execute step 3H; 3E: Take any first pinyin string in the current set that is not listed as the center point of the subset as the third pinyin string, calculate the pinyin distance between the third pinyin string and each existing subset center, and then count the sum of all pinyin distances and the coefficient of variation to obtain the sum of the pinyin distances and the coefficient of variation of the pinyin distances; 3F: When the coefficient of variation of the phonetic distance is less than the first coefficient of variation threshold, the third phonetic string with the largest phonetic distance is used as the new subset center point; 3G: Determine whether the center point of the existing subset reaches the split number threshold, if not, return to step 3E; 3H: Calculate the pinyin distance between all the first pinyin strings in the current set and the center point of each subset, and classify each first pinyin string into one of the existing subsets according to the nearest principle; 3I: Take each subset as the current set and execute step 3C.

4. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 3, characterized in that: The pinyin distance includes: The pinyin converted from a character is called a pinyin unit, which includes initial consonants, finals and tones. Calculate the absolute value p of the difference in the number of pinyin units of two pinyin strings, and record PinyinDistance as 7*p; Calculate the difference between each pinyin unit at the same position of the two pinyin strings, and add it to PinyinDistance to obtain the pinyin distance between the two pinyin strings.

5. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 4, characterized in that: The constructing of the finite word library further includes actively adding recognition word and real word bigrams, including: Get the user's input voice and recognize it to get the output word; Determine whether the output word is one of the words in the second word list. If so, end the process; if not, provide several words in the library, and the user selects the word in the library that the user input voice really wants to express as the third word, and insert the output word into the second word list in ascending order according to the machine code of the Chinese characters, record the insertion position, and insert the third word into the same position in the third word list.

6. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 5, characterized in that: According to the user's input voice, after the preliminary voice recognition obtains the output word, it also includes recording the current time.

7. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 6, characterized in that: Before matching the second pinyin string in the finite word library to obtain the first pinyin string with the nearest pinyin distance, the method further includes: The output word is subjected to a binary search in the second word list. If found, the word at the same position in the third word list is used as the output, and the real-time enhancement of speech recognition is terminated. If not found, the subsequent steps are executed.

8. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 7, characterized in that: Matching the second pinyin string in the finite word library to obtain the first pinyin string with the nearest pinyin distance includes: 7A: If the number of first pinyin strings in the current set is less than the first quantity threshold, execute steps 7B to 7C; if the number of first pinyin strings in the current set is not less than the first quantity threshold, execute steps 7D to 7E; 7B: Calculate the pinyin distance between the second pinyin string and each first pinyin string in the current set; 7C: Take the pinyin distance from small to large as the primary sorting condition. When the pinyin distances are the same, match the output word and the word in the library corresponding to the first pinyin string word by word and finally calculate the number of different Chinese characters to obtain the Chinese character string distance. Take the Chinese character string distance from small to large as the secondary sorting condition, obtain the first matchNum first pinyin strings that are the nearest neighbors to the second pinyin string in the current set, and output the matchNum first pinyin strings; 7D: Calculate the pinyin distance between the second pinyin string and the center points of multiple subsets of the current set, and obtain K candidate subsets according to the smallest first K distance values; 7E: Take each of the K candidate subsets one by one as the current set and return to step 7A; obtain matchNum first pinyin strings each time, use the pinyin distance from small to large as the primary sorting condition, and use the Chinese character string distance from small to large as the secondary sorting condition, and select the first matchNum first pinyin strings that are closest to the second pinyin string from all the obtained first pinyin strings, and output the matchNum first pinyin strings.

9. The method for real-time enhancement of speech recognition based on words in a finite vocabulary according to claim 8, characterized in that: According to the obtained first pinyin string of the nearest neighbor, the corresponding words in the library are output, including: 9A: In the first word list, obtain the words in the library corresponding to the matchNum first pinyin strings output in step 7E; 9B: Compare the current time with the stored time. If the two output words are the same and the difference between the archived time is not less than the preset time threshold, execute step 9C; if the two output words are not the same, or the difference between the archived time is less than the preset time threshold, execute step 9D; 9C: Provide the user with matchNum library word options in text or voice, and finally determine the unique library word according to the user's answer and output it; before outputting, create a new subprocess, in which, if idle is detected, the output word is inserted into the second word list in the order of the machine internal code of the Chinese characters from small to large; 9D: Take the first ranked word among the matchNum words in the database as the output.

10. A real-time speech recognition enhancement system based on words in a finite vocabulary, characterized in that: include: A word library module, wherein the word library module constructs a finite word library, wherein the finite word library includes a plurality of words in the library and a plurality of first pinyin strings, each of which is a pinyin string corresponding to a word in the library; A preliminary recognition module, wherein the preliminary recognition module obtains output words based on preliminary speech recognition of user input speech; A conversion module, wherein the conversion module converts the output word into a pinyin string to obtain a second pinyin string; A matching module, wherein the matching module matches the second pinyin string in the finite word library to obtain a first pinyin string with a pinyin distance closest to the first pinyin string; An output module is provided, wherein the output module outputs the corresponding in-library word according to the first pinyin string of the nearest neighbor of the obtained pinyin distance.

Citation Information

Patent Citations

  • Method and device for recognizing natural speech

    CN102867512A

  • A method and a device for establishing an error correction thesaurus

    CN109271037A

  • ASR language model identification annotation and optimization method and device

    CN110942767A

  • HMM-based pinyin completion training method, HMM-based pinyin completion model, HMM-based pinyin completion method and HMM-based pinyin completion input method

    CN111144096A

  • Pinyin-based dual-stage decoupling Chinese speech recognition model

    CN114743544A