An intelligent expansion similar word model system
By intelligently expanding the similar word model system and automatically selecting and generating similar word models, the problem of time-consuming and labor-intensive manual selection of similar words in the existing technology is solved, and the efficiency and accuracy of the speech recognition system are improved.
Patent Information
- Application Number
- CN202211255240.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2022-10-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-10-13
AI Technical Summary
Existing speech recognition systems require a lot of manual intervention to select similar words when learning keywords, which is difficult to automate and is time-consuming and difficult, especially for non-native developers.
By intelligently expanding the similar word model system, using the text analysis unit, candidate word generation unit, recognition rate processing unit and false call rate processing unit, similar word models are automatically selected and generated to reduce manual intervention, including the automated processing of keyword acoustic models, candidate word acoustic models and similar word acoustic models.
It realizes the automatic selection of similar words, reduces manpower consumption, improves the operational convenience for non-native developers, and improves the efficiency and accuracy of the speech recognition system.
Smart Images

Figure CN115985300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent expansion similar word model system and method thereof, and in particular to an intelligent expansion similar word model system and method thereof which can automatically add similar word models with similar keywords when keyword expansion is required in the field of speech recognition. Background Art
[0002] With the development and popularization of artificial intelligence, the demand for voice-controlled applications such as speech recognition control is also increasing. In speech recognition control, keywords are needed as a necessary means to wake up the system. The speech recognition system must learn to identify various keywords, and various similar words need to be added to train the system to avoid being mistakenly called out by words and phrases similar to keywords.
[0003] This learning process requires first collecting a vast amount of data to facilitate subsequent semantic analysis and classification. Next, speech recognition through audio conversion requires matching with semantic analysis to find the most relevant matches. This process requires a significant amount of computation and traditionally requires manual creation of similar words, making it extremely difficult for non-native developers. Furthermore, the selection of similar words also requires manual decision-making, which is quite time-consuming. Summary of the Invention
[0004] In view of the above shortcomings of the prior art, the purpose of the present invention is to provide an intelligent expansion similar word model system and method, which aims to automate the process of selecting similar words and reduce manpower consumption. It only requires preparing audio data of "keywords" and audio data of "misrecognition test" to operate, and non-native developers can also easily get started.
[0005] To achieve the above-mentioned purpose and other related purposes, the present invention provides an intelligent expansion similar word model system, which operates in a database system host, including: a text analysis unit, which is used to generate a plurality of keyword acoustic models based on a keyword text, and merge the plurality of keyword acoustic models with an interference sound keyword test set into a keyword forward test module; a candidate word generation unit, which is telegraphed to the text analysis unit, and generates a plurality of candidate word temporary acoustic models based on the plurality of keyword acoustic models; a recognition rate processing unit, which is telegraphed to the candidate word generation unit, and is used to execute a speech recognition simulation program to obtain a keyword forward test module. a keyword recognition rate of a candidate word forward test module and a candidate word recognition rate of a candidate word forward test module; when a recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than a recognition threshold value, the candidate word acoustic model that will reduce the recognition rate in a candidate word acoustic model in the candidate word forward test module is deleted to generate a first candidate word acoustic model; a false call rate processing unit is telecommunication-linked to the recognition rate processing unit to execute a false call rate simulation program to obtain a candidate word false call rate of a first candidate word false call test module and a keyword false call rate of a keyword false call test module; when the candidate word false call rate is less than the keyword false call rate, the false call rate of the candidate word is generated. When the false call rate is increased, the plurality of candidate word acoustic models that will reduce the false call rate are selected to generate a second candidate word acoustic model; and an adjustment unit, which is electrically connected to the text analysis unit, the candidate word generation unit, the recognition rate processing unit, and the false call rate processing unit, is used to merge the plurality of keyword acoustic models and the plurality of candidate word temporary acoustic models into the candidate word acoustic model, and merge the candidate word acoustic model and the interference sound keyword test set into the candidate word forward test module; and combine the first candidate word acoustic model with a false wakeup test set into the first candidate word false call test module; and then A plurality of keyword acoustic models are combined with the false wake-up test set to form the keyword false call test module; finally, the plurality of keyword acoustic models are combined with the second candidate word acoustic model to form a similar word acoustic model; wherein, the interference sound keyword test set further includes an interference sound keyword audio data; wherein, the interference sound keyword test set in the keyword forward test module and the keyword acoustic models are used to execute the speech recognition simulation program to obtain the keyword recognition rate; wherein, the interference sound keyword test set in the candidate word forward test module and the candidate word acoustic models are used to execute the speech recognition simulation program to obtain the candidate word recognition rate.
[0006] Furthermore, the intelligent expansion similar word model system further includes a keyword acoustic model processing unit for combining a keyword test set corresponding to a plurality of keyword acoustic models with an interference sound to generate the interference sound keyword test set.
[0007] Furthermore, the intelligent expansion similar word model system further includes a keyword extraction unit for obtaining the keyword test set.
[0008] Furthermore, the intelligent expansion similar word model system further includes a storage unit for storing the keyword acoustic model, the candidate word temporary acoustic model, the candidate word acoustic model, the first candidate word acoustic model, the second candidate word acoustic model and the similar word acoustic model.
[0009] Furthermore, the text analysis unit generates the keyword acoustic model through a monophone modeling method or a triphone modeling method.
[0010] Furthermore, the candidate word generating unit uses a syllable correspondence table to replace each syllable of the keyword text one by one to generate a plurality of candidate word temporary acoustic models.
[0011] Furthermore, the recognition rate difference in the recognition rate processing unit is a natural number.
[0012] Furthermore, the number of candidate word acoustic models in the candidate word forward test module is greater than or equal to the number of candidate word acoustic models in the first number of candidate word acoustic models; the number of candidate word acoustic models in the first number of candidate word acoustic models is greater than or equal to the number of candidate word acoustic models in the second number of candidate word acoustic models; and the number of candidate word acoustic models in the second number of candidate word acoustic models is greater than or equal to the number of candidate word acoustic models in the similar word acoustic models.
[0013] Furthermore, the identification threshold is a system default value or a user-set value.
[0014] In order to achieve the above-mentioned purpose and other related purposes, the present invention also provides a method for intelligently expanding similar word models, including: through a text analysis unit, generating a plurality of keyword acoustic models based on a keyword text, and merging the plurality of keyword acoustic models with an interference sound keyword test set into a keyword forward test module; through a candidate word generation unit, based on the plurality of keyword acoustic models, using a syllable correspondence table to replace each syllable of the keyword text one by one to generate a plurality of candidate word temporary acoustic models; and then through an adjustment unit, merging the plurality of keyword acoustic models with the plurality of candidate word temporary acoustic models. The acoustic model is merged into a plurality of candidate word acoustic models, and the plurality of candidate word acoustic models and the interference sound keyword test set are merged into a candidate word forward test module; through a recognition rate processing unit, a speech recognition simulation program is executed to obtain a keyword recognition rate of the keyword forward test module and a candidate word recognition rate of the candidate word forward test module. When a recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than a recognition threshold value, the candidate word acoustic model that will reduce the recognition rate in the candidate word acoustic model in the candidate word forward test module is deleted to generate A first candidate word acoustic model; combining the first candidate word acoustic model and a false awakening test set into a first candidate word false awakening test module through the adjustment unit; combining the plurality of keyword acoustic models and the false awakening test set into a keyword false awakening test module through the adjustment unit; executing a false awakening rate simulation program through a false awakening rate processing unit to obtain a candidate word false awakening rate of the first candidate word false awakening test module and a keyword false awakening rate of the keyword false awakening test module, and when the candidate word false awakening rate is less than the keyword false awakening rate, selecting the plurality of candidate words that will reduce the false awakening rate Acoustic model, generating a second candidate word acoustic model; combining the plurality of keyword acoustic models and the second candidate word acoustic model into a similar word acoustic model through the adjustment unit; wherein the interference sound keyword test set further includes an interference sound keyword audio data; wherein, the interference sound keyword test set in the keyword forward test module and the keyword acoustic models execute the speech recognition simulation program to obtain the keyword recognition rate; wherein, the interference sound keyword test set in the candidate word forward test module and the candidate word acoustic models execute the speech recognition simulation program to obtain the candidate word recognition rate.
[0015] As described above, the present invention has the following beneficial effects: by automatically generating a large number of similar word candidates, then using keyword test sentences to automatically delete similar words from the candidate list that affect recognition rate, and using test audio files with false positive rates to gradually select the most effective similar words from the candidate list to establish the keyword model required for speech recognition, this effectively improves the traditional method of manually creating similar words, which is very difficult for non-native developers. In addition, the selection of similar words also requires manual decision-making, which is quite time-consuming. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0017] Figure 1 A block diagram of a system for intelligently expanding similar word models according to an embodiment of the present invention;
[0018] Figure 2 A flowchart of the steps of a method for intelligently expanding similar word models according to an embodiment of the present invention;
[0019] Figure 3 A flow chart of an embodiment of a method for intelligently expanding similar word models according to an embodiment of the present invention;
[0020] Figure 4 A flow chart of another embodiment of the method for intelligently expanding similar word models according to an embodiment of the present invention;
[0021] Figure 5 Schematic diagram of an embodiment of the method for intelligently expanding similar word models according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following describes the implementation of the present invention through specific embodiments. People skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification.
[0023] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they have no substantial technical significance. Any modification of the structure, change in the proportional relationship or adjustment of the size should still fall within the scope of the technical content disclosed by the present invention without affecting the efficacy and purpose that can be achieved by the present invention. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of description and are not used to limit the scope of the implementation of the present invention. Changes or adjustments in their relative relationships should also be regarded as the scope of the implementation of the present invention without substantially changing the technical content.
[0024] Figure 1 This is a block diagram of an intelligent expansion similar word model system according to the present invention. Figure 1In the invention, an intelligent expansion similarity model system operates in a database system host 100, comprising: a text analysis unit 110 for generating a plurality of keyword acoustic models based on a keyword text, and combining the plurality of keyword acoustic models with an interference sound keyword test set into a keyword forward test module; a candidate word generation unit 120, which is electronically linked to the text analysis unit 110, and replaces each syllable of the keyword text one by one with a syllable correspondence table according to the plurality of keyword acoustic models to generate a plurality of candidate word temporary acoustic models; a recognition rate processing unit 130, which is electronically linked to the candidate word generation unit 120, and performs a speech recognition A simulation program is provided to obtain a keyword recognition rate of a keyword forward test module and a candidate word recognition rate of a candidate word forward test module. When a recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than a recognition threshold value, the candidate word acoustic model that will reduce the recognition rate in the plurality of candidate word acoustic models in the candidate word forward test module is deleted to generate a first candidate word acoustic model; a false call rate processing unit 140 is telecommunication-linked to the recognition rate processing unit 130 to execute a false call rate simulation program to obtain a candidate word false call rate of a first candidate word false call test module and a keyword false call rate of a keyword false call test module. When the candidate word false call rate is greater than a recognition threshold value, the candidate word acoustic model that will reduce the recognition rate is deleted to generate a first candidate word acoustic model. When the call rate is less than the keyword false call rate, the plurality of candidate word acoustic models that will reduce the false call rate are selected to generate a second candidate word acoustic model; and an adjustment unit 150, which is connected to the text analysis unit 110, the candidate word generation unit 120, the recognition rate processing unit 130, and the false call rate processing unit 140, is used to merge the plurality of keyword acoustic models and the plurality of candidate word temporary acoustic models into the candidate word acoustic model, and merge the candidate word acoustic model and the interference sound keyword test set into the candidate word forward test module; and combine the first candidate word acoustic model with a false wakeup test set into the first candidate word acoustic model. word false call test module; then combining the plurality of keyword acoustic models with the false wake-up test set into the keyword false call test module; finally combining the plurality of keyword acoustic models with the second candidate word acoustic model into a similar word acoustic model; wherein, the interference sound keyword test set further includes an interference sound keyword audio data; wherein, the interference sound keyword test set in the keyword forward test module and the keyword acoustic models execute the speech recognition simulation program to obtain the keyword recognition rate; wherein, the interference sound keyword test set in the candidate word forward test module and the candidate word acoustic models execute the speech recognition simulation program to obtain the candidate word recognition rate.
[0025] In this embodiment, the keyword forward testing module includes an interference sound keyword test set having interference sound keyword audio data and a plurality of keyword acoustic models.
[0026] In this embodiment, the candidate word forward testing module includes an interference sound keyword test set and a plurality of candidate word acoustic models.
[0027] In this embodiment, each candidate word acoustic model includes a plurality of keyword acoustic models and a plurality of candidate word temporary acoustic models.
[0028] In this embodiment, the candidate word forward testing module can be regarded as a keyword forward testing module plus a plurality of candidate word temporary acoustic models.
[0029] In this embodiment, the plurality of keyword acoustic models and the plurality of candidate word temporary acoustic models are merged into the candidate word acoustic model, so that the recognition target can cover the contents of the two sets of acoustic models.
[0030] In this embodiment, the intelligent expansion similar word model system further includes a keyword acoustic model processing unit for combining a keyword test set corresponding to a plurality of keyword acoustic models with an interference sound to generate the interference sound keyword test set.
[0031] In this embodiment, the intelligent expansion similar word model system further includes a keyword extraction unit for obtaining the keyword test set.
[0032] In this embodiment, the intelligent expansion similar word model system further includes a storage unit for storing the keyword acoustic model, the candidate word temporary acoustic model, the candidate word acoustic model, the first candidate word acoustic model, the second candidate word acoustic model and the similar word acoustic model.
[0033] In this embodiment, the text analysis unit generates the keyword acoustic model through a monophone modeling method or a triphone modeling method.
[0034] In this embodiment, the candidate word generating unit uses a syllable correspondence table to replace each syllable of the keyword text one by one to generate a plurality of candidate word temporary acoustic models.
[0035] In this embodiment, the recognition rate difference in the recognition rate processing unit is a natural number.
[0036] In this embodiment, the number of candidate word acoustic models in the candidate word forward test module is greater than or equal to the number of candidate word acoustic models in the first candidate word acoustic model number; the number of candidate word acoustic models in the first candidate word acoustic model number is greater than or equal to the number of candidate word acoustic models in the second candidate word acoustic model number; and the number of candidate word acoustic models in the second candidate word acoustic model number is greater than or equal to the number of candidate word acoustic models in the similar word acoustic model.
[0037] In this embodiment, the identification threshold is a system default value or a user-set value.
[0038] Figure 2 This is a flowchart of a method for intelligently expanding similar word models according to the present invention. The steps are as follows:
[0039] Step S210: Generate a plurality of keyword acoustic models according to the keyword text through the text analysis unit, and combine the plurality of keyword acoustic models and the interference sound keyword test set into a keyword forward testing module.
[0040] The interference sound keyword audio data contained in this test module is sent to the keyword acoustic model contained in this test module for recognition simulation program to obtain the result.
[0041] For example, if there are 1,000 sentences of keyword audio data, and 900 sentences are successfully recognized after being fed into the model, the keyword recognition rate is 90%.
[0042] Step S220: Generate a plurality of temporary acoustic models of candidate words according to a plurality of keyword acoustic models through the candidate word generation unit.
[0043] Step S230: The adjustment unit is then used to merge the plurality of keyword acoustic models and the plurality of candidate word temporary acoustic models into a plurality of candidate word acoustic models, and the plurality of candidate word acoustic models and the interference sound keyword test set are merged into a candidate word forward test module.
[0044] As a preferred method, the interference sound keyword audio data contained in the test module is sent to the candidate word acoustic model contained in the test module for recognition simulation program to obtain the result.
[0045] For example, if there are 1000 audio sentences with keywords, 700 of them are successfully identified as keywords after being fed into the model, and 250 are incorrectly identified as candidate words. Since the candidate words are partially misidentified, the keyword recognition rate is 70%.
[0046] Step S240: Execute a speech recognition simulation program through the recognition rate processing unit to obtain the keyword recognition rate of the keyword forward test module and the candidate word recognition rate of the candidate word forward test module. When the recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than the recognition threshold value, delete the candidate word acoustic model in the candidate word forward test module that will reduce the recognition rate, and generate a first candidate word acoustic model.
[0047] Step S250: The first candidate word acoustic model and the false awakening test set are combined into a first candidate word false awakening test module through the adjustment unit.
[0048] Step S260: Combining a plurality of keyword acoustic models and a false alarm test set into a keyword false alarm test module through an adjustment unit.
[0049] Step S270: Execute a false call rate simulation program through the false call rate processing unit to obtain the candidate word false call rate of the first candidate word false call test module and the keyword false call rate of the keyword false call test module. When the candidate word false call rate is lower than the keyword false call rate, select multiple candidate word acoustic models that will reduce the false call rate to generate a second candidate word acoustic model.
[0050] Step S280: Combine the plurality of keyword acoustic models and the second candidate word acoustic model into a similar word acoustic model through the adjustment unit.
[0051] As a preferred embodiment, the first candidate word false alarm test module includes the first candidate word acoustic model and the false alarm test set. The false alarm test set audio data is fed into the first candidate word acoustic model included in the test module to perform recognition simulation program to obtain the result.
[0052] For example, the false awakening test set has a total of 48 hours of audio data. After being fed into the model, the keyword awakening appears twice, and the false awakening rate of the candidate word is 1 time / 24 hours.
[0053] Furthermore, the aforementioned test results may also include candidate word wakeups (e.g., 8 candidate word wakeups in 48 hours).
[0054] As a preferred method, the keyword false call test module includes a keyword acoustic model and a false wakeup test set. The false wakeup test set audio data is fed into the keyword acoustic model included in the test module for recognition simulation program to obtain the result.
[0055] For example, the false wakeup test set contains 48 hours of audio data. After being fed into the model, 10 false wakeups occur, and the keyword false call rate is 5 times / 24 hours.
[0056] As a preferred method, the first candidate word acoustic model in the first candidate word miscall test module contains the keyword acoustic model and the screened candidate word acoustic model; the keyword acoustic model in the keyword miscall test module does not contain other candidate word acoustic models.
[0057] In this embodiment, the acoustic model specifically refers to the text portion, which further includes speech parameters generated through the text.
[0058] In this embodiment, the test set specifically refers to voice signals.
[0059] In this embodiment, the acoustic model and the test set are combined into a test module, wherein the test module is used to execute a speech recognition simulation program and a false call rate simulation program.
[0060] Among them, the keyword acoustic model and the keyword test set are combined into a keyword forward testing module.
[0061] Among them, the keyword acoustic model and the false wakeup test set are combined into the keyword false wakeup test module.
[0062] Among them, the keyword acoustic model, the candidate word acoustic model and the keyword test set are combined into a candidate word forward test module.
[0063] Among them, the keyword acoustic model, the candidate word acoustic model and the false awakening test set are combined into a candidate word false awakening test module.
[0064] In this embodiment, a large number of keyword acoustic models similar to the keyword text are generated by the text analysis unit.
[0065] In this embodiment, the text analysis unit generates the keyword acoustic model through a monophone modeling method or a triphone modeling method.
[0066] Among them, the monophone modeling method uses a monophone to establish a keyword acoustic model.
[0067] The triphone modeling method establishes a keyword acoustic model using left and right related phonemes, and further includes configuring the multilingual mixed dictionary according to a single language dictionary corresponding to each different language.
[0068] In this embodiment, the candidate word generating unit uses a syllable correspondence table to replace each syllable of the keyword text one by one.
[0069] Among them, the intelligent expansion similar word model system further includes a keyword syllable replacement correspondence table.
[0070] In this embodiment, through the recognition rate processing unit, using the test sentence of the keyword acoustic model, the acoustic model of the negatively affecting candidate word that "will affect the recognition rate" is deleted from the above-mentioned candidate word forward test module, and what remains is the first candidate word acoustic model of "similar word that does not affect the recognition rate".
[0071] In this embodiment, a false call rate processing unit utilizes a false call rate test audio file to select the candidate word acoustic model that best reduces the false call rate from the aforementioned "similar words that do not affect the recognition rate" candidate word acoustic model. When the candidate word false call rate is lower than the keyword false call rate, the plurality of candidate word acoustic models that reduce the false call rate are selected to generate a second candidate word acoustic model.
[0072] As a preferred method, the difference between the candidate word false call rate and the keyword false call rate is in the acoustic model, such as the first candidate word acoustic model in the first candidate word false call test module contains the keyword acoustic model and the screened candidate word acoustic model; the keyword acoustic model in the keyword false call test module does not contain other candidate word acoustic models.
[0073] Figure 3 This is a flow chart of an embodiment of a method for intelligently expanding similar word models according to the present invention. Figure 3 In the method of intelligently expanding the similar word model, step S2401 is further included: executing a speech recognition simulation program through a recognition rate processing unit to obtain the keyword recognition rate of the keyword forward test module and the candidate word recognition rate of the candidate word forward test module. When the recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than the recognition threshold value, step S241 is executed; otherwise, step S242 is executed.
[0074] Step S241: After deleting the candidate word acoustic model in the candidate word forward test module that will reduce the recognition rate, the updated candidate word forward test module repeats step S2401.
[0075] Step S242: Generate a first candidate word acoustic model.
[0076] In this embodiment, a keyword forward testing module is first used, including interference tone keyword audio data and a keyword acoustic model, to obtain a keyword recognition rate of 90% as an example.
[0077] Then, the candidate word forward testing module, including the interference sound keyword audio data and the candidate word acoustic model, obtains the candidate word recognition rate, taking 70% as an example.
[0078] Among them, the candidate word acoustic model includes the keyword acoustic model and the candidate word temporary acoustic model.
[0079] The identification threshold is taken as 20% as an example.
[0080] If the difference between the recognition rate of the previous keyword and the recognition rate of the candidate word is less than or equal to 20%, the acoustic model of the candidate word is the acoustic model of the first candidate word.
[0081] If the difference between the previous keyword recognition rate and the candidate word recognition rate is greater than 20%, it means that some candidate word temporary acoustic models will interfere with the recognition results. The affected candidate word temporary acoustic models must be removed from the candidate word acoustic models in the candidate word forward test module and used as a new candidate word forward test module.
[0082] Figure 4 FIG4 is a flow chart of another embodiment of a method for intelligently expanding similar word models according to the present invention. In FIG4 , the method further includes executing step S2701: executing a false call rate simulation program through a false call rate processing unit to obtain a candidate false call rate from a first candidate false call test module and a keyword false call rate from a keyword false call test module. If the candidate false call rate is less than the keyword false call rate and the number of selected candidate acoustic models is less than a false call threshold, executing step S271; otherwise, executing step S280.
[0083] Step S271: Select multiple candidate word acoustic models that will reduce the false call rate to generate a second candidate word acoustic model. Simultaneously, move the multiple candidate word acoustic models from the first candidate word acoustic model in the first candidate word false call test module to the keyword acoustic model in the keyword false call test module. Repeat step S2701 using the updated first candidate word false call test module and keyword false call test module.
[0084] In this embodiment, the false awakening test set audio data has a total of 48 hours. After being fed into the model, the keyword awakening occurs twice, and the candidate word false awakening rate is 1 time / 24 hours.
[0085] Furthermore, the aforementioned test results may also include candidate word wakeups (e.g., 8 candidate word wakeups in 48 hours).
[0086] As a preferred method, the 8 candidate words awakened at this time can be regarded as "multiple candidate word acoustic models that will reduce the false call rate". These 8 candidate word acoustic models are noted as the second candidate word acoustic models, and are moved from the first candidate word acoustic model to the keyword acoustic model.
[0087] If the false call rate is still greater than the threshold at this time, the merged acoustic model and the remaining candidate word models are used as recognition targets, and the same false call test set is used as input. The candidate word false call rate and keyword false call rate are calculated again using the simulation recognition process.
[0088] Similarly, "a plurality of candidate word acoustic models that will reduce the false call rate" are selected and added to the second candidate word acoustic model, and this process is repeated until the false call rate is lower than the threshold value.
[0089] The number of acoustic models in the second candidate word acoustic model will gradually increase.
[0090] Among them, the false call threshold value of the false call rate is calculated as the number of false triggers per unit time. For example, when the false call threshold value is 30, if the first candidate word false call test module is falsely triggered more than 30 times within 24 hours, the multiple candidate word acoustic models that will reduce the false call rate are selected.
[0091] The false call threshold is a system default value or a user-set value.
[0092] Figure 5 500 is a schematic diagram of an embodiment of a method for intelligently expanding similar word models according to the present invention. Figure 5 In the example, the keyword "query weather" is used as an example. The text analysis unit is used to perform text analysis based on the keyword text to generate a keyword positive test module. Then, the candidate word generation unit is used to generate 2000 candidate word voice signals including "query weather", "query weather", ..., "query weather in Ba" according to the keyword positive test module, and generate a candidate word acoustic model.
[0093] Then, through the recognition rate processing unit, a speech recognition simulation program is executed to obtain the keyword recognition rate of the keyword forward test module and the candidate word recognition rate of the candidate word forward test module. When the recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than the recognition threshold value, the candidate word acoustic model that will reduce the recognition rate in the candidate word forward test module is deleted to generate a first candidate word acoustic model.
[0094] In this embodiment, through the recognition rate processing unit, using the test sentence of the keyword acoustic model, the acoustic model of the negatively affecting candidate word that "will affect the recognition rate" is deleted from the above-mentioned candidate word forward test module, and what remains is the first candidate word acoustic model of "similar word that does not affect the recognition rate".
[0095] In this embodiment, taking the recognition threshold value equal to 0 as an example, the recognition is first performed on "weather query". Assuming there are 100 test data sentences, 95 sentences are recognized, and the keyword recognition rate is 95%.
[0096] "Query weather" and 2000 candidate words are identified together. Among the 100 sentences, 70 are weather queries, 10 are candidate word A, 10 are candidate word B, and 5 are candidate word C. The candidate word recognition rate is 70%. When the recognition rate difference with the keyword recognition rate is greater than 0, the candidate words A, candidate word B, and candidate word C that are identified incorrectly more often are deleted. The candidate word acoustic model has 1997 candidate words remaining.
[0097] Repeat the above steps until the recognition rate difference is less than or equal to the recognition threshold, and then generate the remaining candidate word voice signals, such as "Query Weather Bus", ..., "Query Weather Bus", etc., as the first candidate word model.
[0098] The recognition rate difference is a system default value or a user-set value.
[0099] A false call rate simulation program is executed through the false call rate processing unit to obtain the candidate word false call rate of the candidate word false call test module and the keyword false call rate of the keyword false call test module. When the candidate word false call rate is lower than the keyword false call rate, multiple candidate word acoustic models that will reduce the false call rate are selected.
[0100] Using the test audio file of the false call rate, the candidate word acoustic model that can best reduce the false call rate is selected from the above-mentioned "similar words that do not affect the recognition rate" candidate word acoustic model. When the false call rate of the candidate word is less than the false call rate of the keyword, the multiple candidate word acoustic models that will reduce the false call rate are selected to generate a second candidate word acoustic model.
[0101] Among them, the false call threshold value of the false call rate is calculated as the number of false triggers per unit time. For example, when the false call threshold value is 30, if the first candidate word false call test module is falsely triggered more than 30 times within 24 hours, the multiple candidate word acoustic models that will reduce the false call rate are selected.
[0102] The false call threshold is a system default value or a user-set value.
[0103] In this embodiment, the remaining candidate word voice signals including "check weather", "food and weather", ..., "check popularity", etc. are established as the second candidate word acoustic model.
[0104] Finally, the plurality of keyword acoustic models and the second candidate word acoustic model are combined into a similar word acoustic model.
[0105] In summary, the speech recognition program in the present invention can be roughly decomposed into two inputs, including an acoustic model and a test set. Such a combination of an acoustic model and a test set is referred to as a module in this specification; and an output including a recognition result. The "test set" is compared with the "acoustic model" to obtain the "recognition result". By automatically generating a large number of similar word candidates, and then using keyword test sentences, similar words that affect the recognition rate in the similar word candidate list are automatically deleted. And through the test audio file of the false trigger rate, the most effective similar words are gradually selected from the similar word candidate list to establish the keyword model required for speech recognition, which effectively improves the traditional manual creation of similar words, which is very difficult for non-native developers. In addition, the selection of similar words also needs to be manually decided, which is a very time-consuming problem.
[0106] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. An intelligent expansion similar word model system, which operates in a database system host, characterized in that: include: a text analysis unit for generating a plurality of keyword acoustic models according to a keyword text, and combining the plurality of keyword acoustic models with a noise keyword test set into a keyword forward testing module; a candidate word generating unit, electrically connected to the text analyzing unit, for generating a plurality of candidate word temporary acoustic models based on the plurality of keyword acoustic models; a recognition rate processing unit, telecommunication-connected to the candidate word generation unit, configured to execute a speech recognition simulation program, obtain a keyword recognition rate of a keyword forward test module and a candidate word recognition rate of a candidate word forward test module, and when a recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than a recognition threshold, delete the candidate word acoustic model that reduces the recognition rate from a plurality of candidate word acoustic models in the candidate word forward test module, thereby generating a first candidate word acoustic model; a false call rate processing unit, telecommunication-connected to the recognition rate processing unit, configured to execute a false call rate simulation program to obtain a candidate word false call rate of a first candidate word false call testing module and a keyword false call rate of a keyword false call testing module; and when the candidate word false call rate is lower than the keyword false call rate, selecting the plurality of candidate word acoustic models that will reduce the false call rate to generate a second candidate word acoustic model; as well as an adjustment unit, telecommunication-linked to the text analysis unit, the candidate word generation unit, the recognition rate processing unit, and the false call rate processing unit, configured to merge the plurality of keyword acoustic models and the plurality of candidate word temporary acoustic models into the candidate word acoustic model, merge the candidate word acoustic model and the interference sound keyword test set into the candidate word forward test module, and combine the first candidate word acoustic model with a false call test set into the first candidate word false call test module; Then, the plurality of keyword acoustic models are combined with the false awakening test set to form the keyword false awakening test module; finally, the plurality of keyword acoustic models are combined with the second candidate word acoustic model to form a similar word acoustic model; The interference tone keyword test set further includes interference tone keyword audio data; wherein, executing the speech recognition simulation program with the interference sound keyword test set in the keyword forward test module and the keyword acoustic models to obtain the keyword recognition rate; The speech recognition simulation program is executed on the interference sound keyword test set in the candidate word forward test module and the candidate word acoustic models to obtain the candidate word recognition rate.
2. The intelligent expansion similar word model system according to claim 1, characterized in that: The intelligent expansion similar word model system further includes a keyword acoustic model processing unit for combining a keyword test set corresponding to a plurality of keyword acoustic models with an interference sound to generate the interference sound keyword test set.
3. The intelligent expansion similar word model system according to claim 2, characterized in that: The intelligent expansion similar word model system further includes a keyword extraction unit for obtaining the keyword test set.
4. The intelligent expansion similar word model system according to claim 1, characterized in that: The intelligent expansion similar word model system further includes a storage unit for storing the keyword acoustic model, the candidate word temporary acoustic model, the candidate word acoustic model, the first candidate word acoustic model, the second candidate word acoustic model and the similar word acoustic model.
5. The intelligent expansion similar word model system according to claim 1, characterized in that: The text analysis unit generates the keyword acoustic model through a monophone modeling method or a triphone modeling method.
6. The intelligent expansion similar word model system according to claim 1, characterized in that: The candidate word generating unit uses a syllable correspondence table to replace each syllable of the keyword text one by one to generate a plurality of candidate word temporary acoustic models.
7. The intelligent expansion similar word model system according to claim 1, characterized in that: The recognition rate difference in the recognition rate processing unit is a natural number.
8. The intelligent expansion similar word model system according to claim 1, characterized in that: The number of the candidate word acoustic models in the candidate word forward test module is greater than or equal to the number of the candidate word acoustic models in the first candidate word acoustic models; the number of the candidate word acoustic models in the first candidate word acoustic models is greater than or equal to the number of the candidate word acoustic models in the second candidate word acoustic models; And the number of candidate word acoustic models in the second candidate word acoustic models is greater than or equal to the number of candidate word acoustic models in the similar word acoustic models.
9. The intelligent expansion similar word model system according to claim 1, characterized in that: The identification threshold value is a system default value or a user-set value.
10. A method for intelligently expanding similar word models, characterized in that: include: A text analysis unit is used to generate a plurality of keyword acoustic models according to a keyword text, and the plurality of keyword acoustic models and a noise keyword test set are combined into a keyword forward testing module; generating a plurality of temporary acoustic models of candidate words according to the plurality of keyword acoustic models through a candidate word generating unit; Then, an adjustment unit is used to combine the plurality of keyword acoustic models and the plurality of candidate word temporary acoustic models into a plurality of candidate word acoustic models, and the candidate word acoustic models and the interference sound keyword test set are combined into a candidate word forward test module; A speech recognition simulation program is executed through a recognition rate processing unit to obtain a keyword recognition rate of the keyword forward test module and a candidate word recognition rate of the candidate word forward test module. When a recognition rate difference between the keyword recognition rate and the candidate word recognition rate is greater than a recognition threshold, the candidate word acoustic model that reduces the recognition rate is deleted from the candidate word acoustic models in the candidate word forward test module to generate a first candidate word acoustic model; Combining the first candidate word acoustic model and a false awakening test set into a first candidate word false awakening test module through the adjustment unit; Combining the plurality of keyword acoustic models with the false alarm test set into a keyword false alarm test module through the adjustment unit; executing a false call rate simulation program through a false call rate processing unit to obtain a candidate word false call rate of the first candidate word false call testing module and a keyword false call rate of the keyword false call testing module; and selecting the plurality of candidate word acoustic models that have a lower false call rate when the candidate word false call rate is lower than the keyword false call rate to generate a second candidate word acoustic model; Combining the plurality of keyword acoustic models and the second candidate word acoustic model into a similar word acoustic model through the adjustment unit; The interference tone keyword test set further includes interference tone keyword audio data; wherein, executing the speech recognition simulation program with the interference sound keyword test set in the keyword forward test module and the keyword acoustic models to obtain the keyword recognition rate; The speech recognition simulation program is executed on the interference sound keyword test set in the candidate word forward test module and the candidate word acoustic models to obtain the candidate word recognition rate.
Citation Information
Patent Citations
Acoustic model combination method and device, and voice identification method and system
CN104167206A
Method for training wake-up model and device thereof
CN111667818A