Speech Recognition and Training for Data Input
By segmenting the data input into n-gram blocks and generating language model variants, the language model is trained to generate alternative text, and the problem of inaccurateness of conventional speech recognition technologies when recognizing alphanumeric input is solved, achieving higher recognition accuracy.
Patent Information
- Application Number
- CN202180022530.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-20
- Filing Date
- 2021-02-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-02-23
AI Technical Summary
Conventional speech recognition techniques are inaccurate when recognizing alphanumeric input, especially when the discourse contains a combination of words and numbers, making it difficult to distinguish letters and numbers of similar syllables.
By segmenting the data input into sequential n-gram blocks and receiving metadata about the characteristics of the data input, generating language models and variations thereof, the language model is trained to generate alternative texts for the data input, thereby improving recognition accuracy.
The accuracy of speech recognition technology when recognizing alphanumeric input is improved, and the speech can be converted into text more accurately, solving the inaccuracy problem of conventional speech recognition technology in such inputs.
Smart Images

Figure CN115298736B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to speech-to-text technology, and more particularly, to training a system to recognize alphanumeric speech data input in a speech-to-text system. Background Art
[0002] With the increase in technological capabilities, speech-to-text capabilities are increasingly utilized. For example, when a user calls a help desk, service, etc., the user typically has to give an account number, social security number, birthday, password, etc. The user can use speech to input the required data (e.g., input data by voice over the phone). Speech recognition technology can be used to determine the content input via speech and transform or convert that input into text that can be recognized and processed by a computer system.
[0003] Speech recognition technology can be used in a variety of different usage scenarios. These usage scenarios may require speech recognition to process and recognize multiple types of utterances. For example, conventional speech recognition usage scenarios include using general utterances (from a user) to recognize intents and entities. Conventional speech recognition base models can be used to recognize common words and sentences and convert them into text. For example, a general utterance such as "My name is John Doe" can be recognized and converted into text by a conventional speech recognition model.
[0004] However, conventional speech recognition technology may have difficulties with more complex utterances. For example, it may be challenging for speech recognition technology to recognize data inputs such as IDs, dates, or other alphanumeric data inputs, and conventional speech recognition technology may not be very accurate in recognizing alphanumeric inputs. Alphanumeric data input can be data input that includes both letters / words and numbers. For example, it may be difficult to use conventional speech recognition technology to recognize the alphanumeric utterance "my date of birth is January 8, 1974" due to the combination of words and numbers in the utterance. Conventional speech recognition may not distinguish between "8" and "h"; between "f" and "s"; between "d" and "t"; between "m" and "n"; between "4" and the word "for"; between "to" and "too" and the number "2", etc. Thus, continuing with the previous example, conventional speech recognition may convert the above speech input into text such as "my date of birth is January H 1970 For", which is inaccurate.
[0005] Therefore, it is necessary to solve the above problems in the art. Summary of the Invention
[0006] From a first aspect, the present invention provides a computer-implemented method, comprising: splitting a data input into sequential n-gram blocks based on a predetermined rule, wherein the data input is received by speech recognition; receiving metadata regarding characteristics of the data input; generating a language model based on the metadata; generating a first set of language model variants of the data input; training the language model based at least on the first set of language model variants; using the trained language model to generate one or more alternatives to the data input; and sending an output comprising the one or more alternatives to the data input.
[0007] From another aspect, the present invention provides a system having one or more computer processors configured to: split input data into sequential n-gram blocks based on a predetermined rule, wherein the data input is a speech-to-text transcription; receive metadata regarding characteristics of the data input; generate a language model based on the metadata; generate a first set of language model variants of the data input; train the language model based at least on the first set of language model variants; use the trained language model to generate one or more alternatives to the data input; and send an output comprising the one or more alternatives to the data input.
[0008] From another aspect, the present invention provides a computer program product for identifying alphanumeric speech data input in a speech-to-text system, the computer program product comprising a computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method for performing the steps of the present invention.
[0009] From another aspect, the present invention provides a computer program stored on a computer-readable medium and loadable into the internal memory of a digital computer, the computer program comprising software code portions for performing the steps of the present invention when the program is run on the computer.
[0010] From another aspect, the present invention provides a computer program product comprising a computer-readable storage medium having program instructions implemented thereon, the program instructions executable by a server to cause the server to perform a method comprising: splitting input data into sequential n-gram blocks based on a predetermined rule, wherein the data input is a speech-to-text transcription; receiving metadata regarding characteristics of the data input; generating a language model based on the metadata; generating a first set of language model variants of the data input; training the language model based at least on the first set of language model variants; using the trained language model to generate one or more alternatives to the data input; and sending an output comprising the one or more alternatives to the data input.
[0011] The present disclosure provides computer-implemented methods, systems, and computer program products for identifying and training alphanumeric speech data inputs. The method may include segmenting a data input into sequential n-gram blocks based on a predetermined rule, wherein the data input is received via speech recognition. The method may also include receiving metadata regarding characteristics of the data input. The method may further include generating a language model based on the metadata. The method may also include generating a set of language model variants for a first set of data inputs. The method may further include training the language model based on at least the first set of language model variants. The method may also include using the trained language model to generate one or more alternatives for the data input. The method may also include sending an output that includes the one or more alternatives for the data input.
[0012] The system and computer program product may include similar steps.
[0013] The foregoing summary is not intended to describe every illustrated embodiment or every implementation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings included in this application are incorporated into the specification and form a part of the specification. The drawings illustrate embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. The drawings illustrate only certain embodiments and do not limit the present disclosure.
[0015] Figure 1 A flowchart depicting a set of operations for training speech recognition technology and using the technology to identify alphanumeric speech data inputs, according to some embodiments.
[0016] Figure 2 A flowchart depicting a set of operations for generating a set of language model variants, according to some embodiments.
[0017] Figure 3 A schematic diagram depicting an example speech recognition environment, according to some embodiments.
[0018] Figure 4 A schematic diagram depicting an example speech-to-text environment utilizing 4-gram blocks, according to some embodiments.
[0019] Figure 5 A schematic diagram depicting an example speech-to-text environment utilizing 5-gram blocks, according to some embodiments.
[0020] Figure 6 A block diagram depicting a sample computer system, according to some embodiments.
[0021] The present invention can be modified into different modifications and alternative forms, and its details are shown in the drawings by way of example and described in detail. However, it is to be understood that the present invention is not limited to the specific embodiments described. On the contrary, the present invention is intended to cover all modifications, equivalents and alternatives that fall within the scope of the present invention. DETAILED DESCRIPTION
[0022] Aspects of the invention relate to speech-to-text technology, and more particularly, to training systems to recognize alphanumeric voice data input in speech-to-text systems. Although the present disclosure is not necessarily limited to such applications, various aspects of the present disclosure may be understood through a discussion of different examples using this context.
[0023] The present disclosure provides computer-implemented methods, systems and computer program products for identifying and training alphanumeric voice data input. In some cases, after a speech recognition system receives data input (e.g., via a user's voice), the data may be segmented into blocks (i.e., n-gram blocks) and may be processed via the segmented blocks. Various rules and features (including input-specific rules, general voice rules and variants, etc.) may be used to generate a language model (and, in some instances, an acoustic model) based on rules and features. The model or models may be trained using various rules and features and various data inputs to improve the model. Data is segmented into blocks, and a language model is generated and trained using various rules and features, which can improve the accuracy of speech recognition technology for alphanumeric data input.
[0024] See now Figure 1 , describes a flow chart showing a method 100 for training speech recognition technology and using the technology to recognize alphanumeric speech data input according to some embodiments. In some embodiments, the method 100 is performed by a speech recognition system (e.g., speech recognition system 330 ( Figure 3 )、Speech recognition system 430( Figure 4 ) and / or speech recognition 530 ( Figure 5 )) on or connected to a server of the speech recognition system (e.g., computer system / server 602 ( Figure 6 In some embodiments, the method 100 is implemented as a computer system (e.g., computer system 600 ( Figure 6 )) is executed on or connected to a computer system (e.g., computer system 600 ( Figure 6 )) of a computer script or computer program (e.g., computer executable code). In some embodiments, the speech recognition system is on a computer system or connected to a computer system.
[0025] Method 100 includes operation 110 for receiving a data input. The data input can be a voice-to-text data input. In some embodiments, receiving the data input includes receiving a voice data input and converting the voice into text. In some embodiments, in operation 110, the voice has been converted into text before the system receives the voice. The conversion from voice to text can be accomplished using conventional methods.
[0026] For example, a client (e.g., a company) can use conventional speech recognition technology to convert user voice into text. However, this conventional speech recognition technology may inaccurately convert voice into text. Thus, the client can send the text data input (since the client has converted it into text) to the system that executes method 100 to correct the text data input and create an improved model for converting voice into text that the client can use in future conversions. For example, a user can input information by voice into a client system (e.g., an insurance company, a healthcare company, or any other company that utilizes speech recognition technology), the client system can convert the voice into text, and then can send the text of the data input to the system.
[0027] In some instances, there is no third-party client, and the system that executes method 100 is a system that receives voice directly from the user. In these instances, the system can initially use conventional methods to convert the voice into text in operation 110. For example, a user can input information by voice to the system (e.g., via a telephone), and the system can convert the voice / sound input into text.
[0028] Method 100 includes operation 115 for splitting the data input into sequential n-gram blocks based on a predetermined rule. After the data input is converted into text (discussed in operation 110), the data input can be split into n-gram blocks, where n is relative to the length of the data input (the data input that is received and established as text in operation 110). An n-gram block can be a sequence of n letters, numbers, words, or any combination of the three. In some embodiments, the n-gram block can be a 4-gram block or a 5-gram block, so splitting the data into sequential n-gram blocks can include splitting the data into sequential 4-gram blocks or 5-gram blocks. In some embodiments, splitting the data input into 4-gram blocks (or 5-gram blocks) includes decomposing the data input (in its text form) into sequential 4-character (or 5-character) blocks. In some embodiments, splitting the data input into 4-gram (or 5-gram) blocks includes decomposing the data input into sequential 4-word (or 5-word) blocks. Examples of splitting the data input into n-gram blocks are discussed below.
[0029] In some embodiments, the predetermined rule includes: for data inputs of 9 positions (e.g., characters, words, etc.) or fewer, the n-gram blocks are 4-gram blocks, and for data inputs of 10 positions or more, the n-gram blocks are 5-gram blocks. In some embodiments, the positions can be characters (e.g., letters or numbers). For example, a data input including "12345ABCD" has 9 positions and is thus split into 4-gram blocks. Splitting the example data input into 4-gram blocks can include splitting the input into blocks "1234", "2345", "345A", "45AB", "5ABC", and "ABCD". Thus, starting from the first character, the data input is split into consecutive 4-character blocks until all characters are split.
[0030] In another example, when the input includes words and numbers, the data input includes "my birthday isJanuary 1,1985 (My birthday is January 1, 1985)". In this example, each position can include my-birthday-is-January-1,-1985 (My-birthday-is-January-1,-1985), where the spaces indicate the separation between positions of the data input. In this example, there are 6 positions, and thus the data input is split into 4-gram blocks. Splitting this example data input into 4-gram blocks can include splitting the input into blocks "my birthday is January (My birthday is January)", "birthday is January 1 (Birthday is January 1)", "is January 1,1985 (Is January 1, 1985)".
[0031] In another example, a data input including "ABCDE98765" has 10 positions (i.e., characters) and is thus split into 5-gram blocks. These 5-gram blocks can include the split blocks "ABCDE", "BCDE9", "CDE98", "DE987", "E9876", and "98765". Extending to 5-gram blocks for longer inputs can improve the context coverage of the data input, especially for longer position (or character) data inputs. 5-gram blocks can give a more complete coverage of the data input because 5 positions of the data input are together in each block. 4-gram blocks can have 4 positions of the data input included in each block. However, for smaller data inputs, by splitting the input into 5-gram blocks (as opposed to 4-gram), the context coverage of the data input may be minimally improved or not improved at all. Thus, resources (e.g., bandwidth, etc.) can be saved by splitting smaller inputs (e.g., 9 positions or fewer) into 4-gram blocks.
[0032] Further discussed herein and described in Figure 4 is an exemplary data input segmented into 4-gram blocks. Further discussed herein and described in Figure 5 is the same example data input segmented into 5-gram blocks.
[0033] Method 100 includes operation 120 for receiving metadata regarding characteristics of a data input. In some embodiments, the metadata regarding characteristics of the data input may include specific requirements of the data input. For example, if the data input is a social security number, the specific requirement may be that the data input has 9 positions (e.g., 9 characters) and each position is a digit. In another example, if the data input is a birthday, the specific requirement may be that the data input includes a month, a day, and a year.
[0034] In another example, the data input may be a Medicare Beneficiary Identifier (MBI). The MBI may have specific characteristics or a specific rule set for the MBI number. These specific characteristics may be the metadata regarding characteristics of the data input. In the example of the MBI, there may be a specific requirement that the data input includes 11 positions or 11 characters. These 11 characters may have to have a specific format of C / A / AN / N / A / AN / N / A / A / N / N. There may be a specific requirement that the character in any C position must be a digit from 1 to 9. Another specific requirement may be that the character in any N position must be a digit from 0 to 9. Another specific requirement may be that the character in an A position must be an alphabetic character from A-Z but excluding S, L, O, I, B, and Z. Finally, there may be another requirement that any character in an AN position must be an A or an N. All of these example specific requirements for the MBI may be the metadata regarding characteristics of the data input received by the system.
[0035] In some embodiments, the metadata regarding characteristics of the data input is determined by a client. The client (e.g., an insurance company, a healthcare company, or any other company using speech recognition technology) may have specific rules (such as a healthcare ID, an insurance ID, etc.) for its data input and may send these rules or characteristics to the system that executes method 100. In some embodiments, the system that executes method 100 is also a client and thus may pre-determine these characteristics and input them into the system.
[0036] Method 100 includes operation 125 of generating a language model based on metadata. In some embodiments, the language model can be a model for converting speech input into text. As discussed herein, the data input received (in operation 110) may have been converted into text using conventional methods; however, these conventional methods may not be very accurate in their conversion, particularly for alphanumeric data input. Thus, a language model can be generated (and as discussed in operation 145, trained) to more accurately convert speech into text. The language model (especially after being trained) can be used in future speech recognition scenarios to convert speech into text. Additionally, the language model can be used to correct the conversion of the data input received in operation 110.
[0037] In some embodiments, the language model is a machine learning model. In some cases, the language model can be a regular expression (regex) algorithm. The regex algorithm can match each character in an expression to its meaning. For example, using the above example of MBI, the regex algorithm can match each character to its corresponding characteristic(s) (e.g., related to its position of C, A, N, or AN). In some instances, the language model can be a rule matching algorithm. The rule matching algorithm can effectively apply or match many rules or patterns to many characters. For example, again using the above MBI example, the rule matching algorithm can match the rule that the character must be a number to positions C and N. The rule algorithm can also match the rule that the character must be a letter to positions A and AN. In some instances, the rule matching algorithm can continue to match position rules until all rules are matched. In some embodiments, the language model is a speech-to-text grammar model.
[0038] In some embodiments, generating the language model includes creating a language model that shows the relationship between metadata (regarding characteristics of the data input) and the data input. For example, if the data input is a 4-digit PIN number, the characteristics of the data input can include that each character (or position) of the PIN number is a digit / number. The language model can show the relationship between the PIN number and the metadata (the character is a digit / number). In another example, if the data input is MBI, the language model can show the relationship between the first character and the metadata (indicating that the first character must be a number between 1 and 9). The language model can also show the relationship between the second character and the metadata (indicating that the second character must be a letter between A and Z) (and so on for each other character in MBI). In some embodiments, the language model can show the relationship between the metadata and the segmented n-grams of the data input. Generating the language model can include mapping the relationship between each component of the data input and the metadata.
[0039] Method 100 includes operation 130 for generating a first set of language model variants for data input. The first set of language model variants can be possible variations of characters and / or positions that can meet the criteria of metadata regarding characteristics. For example, the first set of language model variants can be possible variations of characters / positions that still follow the language model. For example, if the data input is a PIN number (e.g., a 4-digit code) "1123", each position of the PIN number can be a digit between 0 and 9. Since each of the four digits has 10 possible digits (0 - 9), there can be 10,000 (10 4 ) variations in the first set of language model variants. The set of 10,000 variations can include "1111", "1112", "1113", "1234", "2211", etc., until all possible variations of the 4-digit code are listed in the set. In some embodiments, the first set of model variants is generated at least based on metadata and predetermined rules.
[0040] The generation of the first set of language model variants for data input is further discussed herein and depicted in Figure 2 the [depiction].
[0041] In some embodiments, method 100 further includes operation 135 for generating a second set of language model variants. In some instances, the generation of the second set of language model variants uses common user metadata. The common user metadata can be common speech patterns, phrases, words, terms, etc. The second set of language model variants can include possible variations of data input based on the common speech pattern. Continuing the previous example, the data input can be the 4-digit PIN number "1123". As described above, the first set of language model variants can include a set of 10,000 variations. To determine the second set of language model variants, the common speech pattern can be used. The common speech pattern for 1123 can include stating "11" as "eleven"; stating "11" as "one, one"; stating "23" as "twenty-three"; stating "23" as "two, three"; stating "12" as "twelve", etc. Thus, in this example, the second set of language model variants includes at least "one, one, two, three", "eleven, two, three", "eleven, twenty-three", "one, one, twenty-three", "one, twelve, three", etc. In some embodiments, machine learning can be used to collect and learn the common user metadata.
[0042] In some embodiments, method 100 further includes operation 140 for creating pre-attached text and additional text for a subset of a first set of language model variants and a second set of language model variants. The pre-attached text can be text that is input before the data input. For example, using the above PIN number example, a user may state (via voice) "my PIN number is 1123". The text "my PIN number is" can be the pre-attached text. The additional text can be text that is input after the data input. For example, in the same PIN number example, the user may state "1123 is my PIN number". The text "is my PIN number" can be the additional text. Historical text data input (e.g., historical text data input obtained using conventional methods) can be used to create the pre-attached text and the additional text. In some instances, voice data regarding common voice patterns is used to create the pre-attached text and the additional text.
[0043] Creating the pre-attached text and the additional text can include creating sample pre-attached and additional text that can be used in the data input based on a subset of a first set of language model variants and a second set of language model variants. For example, if the data input is a 4-digit PIN number, the first set of language model variants can include ten thousand possible variations of the PIN number. For each of the ten thousand variations in the first set of language model variants, it may not be necessary to generate sample pre-attached text (such as "my PIN number is…") and sample additional text (such as "…is my PIN number"). If pre-attached text and additional text are generated for each of the ten thousand variants, the pre-attached text and the additional text may be very repetitive. Thus, for example, pre-attached text and additional text can be generated for a subset of one thousand randomly selected variants (from the ten thousand variants). One variant can have pre-attached text such as "this is my PIN number…", another variant can have "ready? Ok, it is…", and another variant can have simple pre-attached text of "It is". To increase the efficiency of creating the pre-attached and additional text, only a subset of the language model variants can be used (since fewer variants may need to be reviewed).
[0044] In some embodiments, creating pre-attached text and attached text includes analyzing common user metadata. As discussed herein, common user metadata can be common speech patterns, phrases, words, terms, etc. used by the user. Analyzing the common user metadata can indicate which phrases (e.g., pre-attached and / or attached numeric or alphanumeric data inputs) can be used with the data input.
[0045] In some embodiments, creating pre-attached text and attached text includes determining, based on the analysis, common pre-attached phrases and attached phrases to use when inputting speech. In some embodiments, the system can access historical input data. The historical input data can include past data about inputs submitted by the user (e.g., via speech). The historical input data can be used, along with the common user metadata, to determine common phrases for pre-attached and / or attached data inputs.
[0046] In some embodiments, creating pre-attached text and attached text includes generating a plurality of template sentences for a subset of a first set of language model variants and a second set of language model variants. In some embodiments, a random subset of the first set of language model variants and a random subset of the second set of language model variants can be selected. For each language model variant (from the subset), a template sentence (with pre-attached text and / or attached text) can be generated. For example, for a language model variant that is a subset of the first set of language model variants (e.g., PIN number 1111), the template sentence can be “eleven-eleven is my PIN”.
[0047] Method 100 includes operation 145 for training a language model based on at least a first set of language model variants. An originally generated language model (generated in operation 125) can be generated based on the data input itself, along with general rules, requirements, characteristics, etc. of the data input (as indicated by the metadata received in operation 120). The originally generated language model may not account for all possible specific variants of the data input. For example, the data input (when spoken) can be "one hundred seventy-five", and a characteristic of the data input can include that the data input is a 5-digit PIN number. An initial speech-to-text conversion (e.g., done using conventional methods) can convert the speech to "175". Since five digits are required (based on the metadata), the originally generated language model can correct the data input to "00175". However, in this example, the user may have meant to convey that the data input is "10075". In this example, all possible variants of the 5-digit PIN number can be the first set of language model variants. Training the language model based on these variants can include having the model learn each possible variant of the 5-digit PIN. In this way, the trained language model can identify that a possible alternative conversion of the speech input is "10075" (identified in operation 150, using the model trained in operation 145). This alternative may not have been identified using the original learning model, and thus in this example, the trained model may be more accurate. The first set of language variations (including possible variants for the data input) can help train the model to accurately identify and convert speech data inputs because the language model can learn all possible variations or at least a large number of possible variations for the data input.
[0048] In some embodiments, when method 100 includes operations 135 and 140, the language model can be further trained based on a second set of language model variants and pre-attached text and appended text. The second set of language variants can include variations of the data input based on common speech patterns, and in addition to the necessary data input, the pre-attached / appended text can also include possible text / speech. Training the language model using the first set of language model variants, the second set of language model variants, and the pre-attached / appended text can allow the language model to learn the possible variants and speech habits that can occur when the data is input, and thus can more accurately convert speech inputs to text.
[0049] In some embodiments, a first set of language variants, a second set of language variants, pre- and post- appended text, and data input may also be used to generate and train an acoustic model. The acoustic model may show the relationship between linguistic units such as phonemes and an audio input. In some embodiments, the acoustic model may be generated based on sound recordings and text transcripts of the data input. The model may then be trained using the first set of language variants, the second set of language variants, the pre- and post- appended text, and the data input. In some cases, a language model (generated in operation 125) may also be used to train the acoustic model.
[0050] Method 100 includes operation 150 for generating one or more alternatives of the data input using the trained language model. In some embodiments, operation 150 also uses the trained acoustic model. In some embodiments, one or more alternatives of the data input include one or more possible (other) conversions of a particular speech input. For example, a user may say "my dateof birth is January 8,1974" (e.g., over the phone). The initial data input may have converted the speech to text "my date of birth is January H 1970 for". The initial data input may have mistaken "8" for "h" and "4" for "for". Using the trained learning model, alternatives such as "January A,1974", "January8,1970", and "January 8,1974" may be generated. In some instances, as the language model is continuously learned and trained, the model may become increasingly accurate in properly identifying / converting the data input. In some instances, only the alternative "January 8,1974" may be generated (e.g., when the model is well- trained).
[0051] Method 100 includes operation 155 for sending an output that includes one or more alternatives of the data input. In some embodiments, the output may be sent to a client. The client may be an intermediary between the user and the speech recognition system (e.g., as Figure 3 shown). In some embodiments, the client is part of the system that executes method 100 and thus may send the output to different components of the system, e.g., which may directly receive speech input from various users.
[0052] See Figure 2, which shows a flowchart of a method 200 for generating a set of language model variants according to some embodiments. In some embodiments, method 200 may correspond to operation 130( Figure 1 ).
[0053] Method 200 includes operation 210 to determine a plurality of possible variants of the data input. In some embodiments, the number of possible variants is based on the total number of positions or characters required for the data input (e.g., as indicated by metadata regarding the characteristics of the data input). For example, a data input indicating a PIN number may have four required positions (since a PIN can be a 4 - digit PIN number). A PIN number can consist only of digits, so each position can have the possibility of a digit between 0 - 9 (i.e., ten possible digits). Thus, in this example, the total number of possible variants can be 10 4 or ten thousand variants.
[0054] Using the 11 - position MBI example from above, each MBI position can have different rules. For example, the first position (digits 1 - 9) can have nine possible digits, the second position can have twenty possible letters, the third position can have two possible letters, the fourth position can have ten possible digits, the fifth position can have twenty possible letters, the sixth position can have two possible letters, the seventh position can have ten possible digits, the eighth position can have twenty possible letters, the ninth position can have twenty possible letters, the tenth position can have ten possible digits, and the eleventh position can have ten possible digits (based on the specific MBI rules discussed above). This can result in a total of approximately 80 trillion possible variants.
[0055] Method 200 includes operation 220 to determine whether the number of variants is a small number of variants. In some embodiments, a small number of variants is based on a threshold pre - determined by the client. The "small" number of variants considered can depend on the usage scenario. For example, in some usage scenarios, a small number of variants can be a number of variants less than one million variants. In other usage scenarios, a small number of variants can be a number of variants less than 5 million variants.
[0056] In some embodiments, determining whether the number of possible variants is a small number of variants includes comparing the number of possible variants with a threshold number of variants. If the number of possible variants is greater than the threshold number of variants, then the number of variants may be a large (i.e., not small) number of variants.
[0057] If it is determined that the number of possible variants is a small number of variants, method 200 can proceed to operation 230 to generate all possible language model variants. For example, continuing with the previous PIN number example, since ten thousand possible PIN number variants are less than one million variants, ten thousand possible PIN number variants can be a small number of variants. Thus, all ten thousand possible PIN number variants can be generated. These can include variants such as "1111", "1122", "2222", "2233", "3333", "3344", "4444", "4455", "5555", "5566", "6666", "7777", "8888", "9999", etc.
[0058] If it is determined that the number of possible variants is a large number of variants or not a small number of variants, method 200 can proceed to operation 235 to generate a reduced number of variants for the data input based on n-gram blocks. For example, as discussed herein, the data input for MBI can have 80 trillion possible variants, which can be determined to be a large (i.e., not small) number of variations. Thus, not every possible variant out of the 80 trillion variants is generated. Instead, variants can be generated based on n-gram blocks. For example, the data input may have been segmented into 4-gram blocks. The MBI data input can be alphanumeric (i.e., the data can include both letters and numbers), so there can be 36 possible characters (26 letters (A-Z) and 10 digits (0-9)). The reduced number of variants for the alphanumeric data input segmented into 4-gram blocks can be 36 4 , or approximately 1.7 million variants. For data input segmented into 5-gram blocks, the reduced number of variants for the alphanumeric data input can be 36 5 , or approximately 60 million variants.
[0059] In some embodiments, when generating all possible language model variants (e.g., in operation 230), it is determined (after generating alternative words in operation 240) which language model variants that include numbers have alternative words. For each language model variant that includes a number, additional variants can be generated by replacing the number with an alternative word. For example, for each variant that includes 0, additional variants can be generated by replacing "0" with "oh" or the letter "O".
[0060] In some embodiments, when generating a reduced number of variants (e.g., in operation 235), it is determined (after generating alternative words in operation 240) which language model variants (from the reduced number of variants) include numbers with alternative words. For each language model variant that includes a number, additional variants can be generated by replacing the number with an alternative word.
[0061] Method 200 includes operation 240 to generate alternative words or text that can represent numbers (based on common number patterns). For example, the word "oh" can be used to represent the number 0. The common number pattern can be determined based on historical data. The common number pattern can be a common alternative pronunciation of a number and / or number sequence. For example, the number 0 can be pronounced as "oh" or "zero". The number 100 can be pronounced as "ahundred", "one hundred", etc. Generating alternative words can include generating additional language model variants with common alternative pronunciations.
[0062] Reference Figure 3 , depicts a schematic diagram of an example speech recognition environment 300 according to some embodiments. The speech recognition environment can include a user 305, a telephone 310, a client 320, and a speech recognition system 330. In some embodiments, the speech recognition system 330 includes a computer system 500 ( Figure 5 ). In some embodiments, the speech recognition system 330 can execute Figure 1 Method 100 (and Figure 2 Method 200). The speech recognition system 330 includes a speech-to-text module 332, a segmentation module 334, and a machine learning module 336.
[0063] In some embodiments, the user 305 speaks into the telephone 310. The speech can be sent from the telephone 310 to the client 320. The client 320 can send the speech input to the speech-to-text module 332. In some embodiments, the speech-to-text module 332 can receive the speech input (e.g., Figure 1 Operation 110) and can convert the speech input into text. Then, the text input can be sent to the segmentation module 334, and the segmentation module can perform, for example, Figure 1 Operations 115 and 120. The segmented data input (from operation 115 ( Figure 1 )) can be sent to the machine learning module 336, and the machine learning module 336 can perform, for example, operations 125 - 155 of method 100 ( Figure 1 ). In some embodiments, the machine learning module 336 can send the output back to the client 320. In some embodiments, the machine learning module also sends an updated (e.g., trained) learning model to the speech-to-text module 332.
[0064] In some embodiments, the client can be the owner of the speech recognition system 330. Thus, in some embodiments, the client 320 may not be included in the speech recognition environment 300. In these cases, the speech input (from the telephone 310) can go directly to the speech-to-text module 332.
[0065] Reference Figure 4 illustrates a schematic diagram of an example speech-to-text environment 400 utilizing 4-gram blocks according to some embodiments. The speech-to-text environment 400 may include a user 405 who states a speech input 407 (in this instance, it is MBI) to a telephone 410. In some embodiments, the user 405 and the telephone 410 may correspond to Figure 3 the user 305 and the telephone 310 in
[0066] . The speech input 407 may be sent to a speech-to-text module 432, and the speech-to-text module 432 may convert the speech input 407 into a data input (i.e., a text input) 433. The data input 433 is sent to a segmentation module 434, and the segmentation module may segment the data input 433 into 4-gram blocks. The 4-gram blocks of the data input 433 include 1AA0, AA0A, A0AA, 0AA0, AA0A, A0AA, 0AA0, and AA00. Segmenting the data input 433 into 4-gram blocks may include starting at a first position of the data input 433 and selecting the next three (excluding the first position) positions. This may be the first 4-gram block. Further segmenting the data input 433 may include starting at a second position of the data input 433 and selecting the next three positions (excluding the second position). This may be repeated until all positions of the data input 433 have been segmented. Figure 3 )
[0067] Reference Figure 5 illustrates a schematic diagram of an example speech-to-text environment 500 utilizing 5-gram blocks according to some embodiments. The speech-to-text environment 500 may include a user 505 who states a speech input 507 (in this instance, it is MBI) to a telephone 510. In some embodiments, the user 505 and the telephone 510 may correspond to Figure 3User 305 and telephone 310 therein. The voice input 507 can be sent to the voice-to-text module 532, and the voice-to-text module 532 can convert the voice input 507 into a data input (i.e., text input) 533. The data input 533 is sent to the segmentation module 534, and the segmentation module can segment the data input 533 into 5-gram chunks. The 5-gram chunks of the data input 533 include 1AA0A, AA0AA, A0AA0, 0AA0A, AA0AA, A0AA0, and 0AA00. Segmenting the data input 533 into 5-gram chunks can include starting at the first position of the data input 533 and selecting the next four (excluding the first position) positions. This can be the first 5-gram chunk. Segmenting the data input 533 can also include starting at the second position of the data input 533 and selecting the next four positions (excluding the second position). This can be repeated until all positions of the data input 533 have been segmented.
[0068] The voice-to-text module 532 and the segmentation module 534 can be part of the speech recognition system 530. In some embodiments, the voice-to-text module 532, the segmentation module 534, and the speech recognition system 530 respectively correspond to the voice-to-text module 332, the segmentation module 334, and the speech recognition system 330( Figure 3 ).
[0069] See Figure 6 , according to some embodiments, the computer system 600 is a computer system / server 602 shown in the form of a general-purpose computing device. In some embodiments, the computer system / server 602 is located on a linked device. In some embodiments, the computer system 602 is connected to a linked device. The components of the computer system / server 602 can include, but are not limited to, one or more processors or processing units 610, a system memory 660, and a bus 615 that couples the different system components including the system memory 660 to the processor 610.
[0070] The bus 615 represents one or more of any number of types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, this architecture includes Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0071] The computer system / server 602 generally includes a variety of computer system readable media. This media can be any available media accessible by the computer system / server 602, and it includes volatile and non-volatile media, removable and non-removable media.
[0072] The system memory 660 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 662 and / or cache memory 664. The computer system / server 602 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 665 may be provided for reading from and writing to a non-removable non-volatile magnetic medium (not shown and typically referred to as a "hard disk drive"). Although not shown, a disk drive may be provided for reading from or writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media. In such cases, each may be connected to the bus 615 by one or more data media interfaces. As will be further depicted and described below, the memory 660 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.
[0073] A program / utility 668 having a set (at least one) of program modules 669, as well as an operating system, one or more application programs, other program modules, and program data may be stored in the memory 660 by way of example and not limitation. Each or some combination of the operating system, one or more application programs, other program modules, and program data may include an implementation of a network environment. The program modules 669 generally execute the functions and / or methods of the embodiments of the present invention as described herein.
[0074] The computer system / server 602 can also communicate with the following: one or more external devices 640 (such as a keyboard, a pointing device, a display 630, etc.); one or more devices that enable a user to interact with the computer system, server 602; and / or any device that enables the computer system / server 602 to communicate with one or more other computing devices (for example, a network card, a modem, etc.). This communication can occur via the input / output (I / O) interface 620. In addition, the computer system / server 602 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (such as the Internet) via the network adapter 650. As depicted, the network adapter 650 communicates with other components of the computer system / server 602 via the bus 615. It should be understood that although not shown, other hardware and / or software components can be used in conjunction with the computer system / server 602. Examples include but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0075] The present invention can be a system, method, and / or computer program product with any possible degree of integration of technical details. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0076] The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (for example, an optical pulse passing through an optical fiber cable) or an electronic signal emitted through a wire.
[0077] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0078] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc. and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit, so as to carry out aspects of the present invention.
[0079] The present invention will be described below with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0080] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions comprises an article of manufacture including instructions embodying aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram. The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to some embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or combinations of special purpose hardware and computer instructions.
[0082] The description of the various embodiments of the present disclosure has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or technical improvement of technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method, comprising: segmenting a data input into sequential n-gram blocks based on a predetermined rule, wherein the data input is received via speech recognition; receiving metadata regarding characteristics of the data input; generating a language model based on the metadata; generating a first set of language model variants of the data input; generating a second set of language model variants in response to generating the first set of language model variants; generating a plurality of template sentences for a subset of the first set of language model variants and the second set of language model variants; training the language model based on at least the first set of language model variants; using the trained language model to generate one or more alternatives for the data input; and sending an output comprising the one or more alternatives for the data input.
2. The method according to claim 1, wherein the first set of language model variants is generated based at least on the metadata and the predetermined rule.
3. The method according to claim 2, wherein generating a plurality of template sentences is included in a process of creating prefixed text and appended text for a subset of the first set of language model variants and the second set of language model variants.
4. The method according to claim 3, wherein the language model is further trained based on the second set of language model variants, the prefixed text, and the appended text.
5. The method according to claim 3, wherein creating the prefixed text and appended text comprises: analyzing common user metadata; and determining common prefixed phrases and appended phrases to use when inputting speech based on the analysis.
6. The method according to any one of the preceding claims, wherein generating the first set of language model variants comprises: determining the number of possible variants of the data input.
7. The method according to claim 6, further comprising: determining that the number of possible variants of the data input is below a threshold number of variants; and in response to the determination, generating all possible language model variants of the data input.
8. The method according to claim 6, further comprising: determining that the number of possible variants of the data input is above a threshold number of variants; and in response to the determination, generating a reduced number of variants of the data input based on the n-gram blocks.
9. The method according to claim 6, wherein generating the first set of language model variants of the data input comprises: generating alternative words that can represent numbers based on a common number pattern.
10. The method according to any one of claims 1-5, wherein the n-gram block is one of a 4-gram block and a 5-gram block.
11. A computer system having one or more computer processors, the system configured to: segment a data input into sequential n-gram blocks based on a predetermined rule, wherein the data input is a speech-to-text transcription; receive metadata regarding characteristics of the data input; generate a language model based on the metadata; generate a first set of language model variants of the data input; In response to generating the first set of language model variants, generate a second set of language model variants; Generate a plurality of template sentences for a subset of the first set of language model variants and the second set of language model variants; Train the language model based on at least the first set of language model variants; Use the trained language model to generate one or more alternatives for the data input; and Send an output including the one or more alternatives for the data input.
12. The system according to claim 11, wherein, Generate the first set of language model variants based at least on the metadata and the predetermined rules.
13. The system according to claim 12, wherein, Generating a plurality of template sentences is included in the process of creating pre-attached text and additional text for a subset of the first set of language model variants and the second set of language model variants.
14. The system according to claim 13, wherein, Also train the language model based on the second set of language model variants, the pre-attached text, and the additional text.
15. The system according to claim 13 or 14, wherein, Creating the pre-attached text and additional text includes: Analyzing common user metadata; and Based on the analysis, determining common pre-attached phrases and additional phrases to use when inputting speech.
16. A computer program product for identifying alphanumeric speech data input in a speech-to-text system, the computer program product comprises: A computer-readable storage medium that can be read by a processing circuit and stores instructions for execution by the processing circuit to perform the method according to any one of claims 1 to 10.
17. A computer-readable storage medium having computer instructions stored thereon, the computer instructions being loadable into the internal memory of a digital computer, the computer instructions including a software code portion that, when the computer instructions are run on a computer, is used to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Phrase-based dialogue modeling with particular application to creating recognition grammars for voice-controlled user interfaces
US20020128821A1
Generic spelling mnemonics
US20060111907A1
Speech Recognition Using Context-Aware Recognition Models
US20120323557A1
Pronunciation learning from user correction
US20130090921A1
Systems and methods for providing metadata-dependent language models
US20140324434A1