Context-Aware Correction of User Input Recognition by Machine Learning
The context-aware correction module on-device improves input recognition accuracy and privacy by replacing mistranslated terms with non-generic terms in real-time, addressing the challenges of recognizing non-general terms and maintaining user privacy in existing systems.
Patent Information
- Application Number
- JP2024561885
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-21
- Filing Date
- 2023-04-18
- Publication Date
- 2025-05-09
AI Technical Summary
Existing input recognition systems face challenges in accurately recognizing user input, particularly when terms are not recognized as general terms, and may raise privacy concerns due to the need for user feedback or reliance on rewriting user-generated words.
A context-aware correction module is implemented on-device, which identifies candidate terms for substitution by accessing pairs of mistranslated and non-generic terms. This module replaces candidate terms with non-generic terms in real-time, improving input recognition accuracy while maintaining user privacy.
The solution enhances the accuracy and efficiency of input recognition by correcting mistranslations in real-time, reduces computational resources required, and ensures user privacy by processing data locally without sharing user information with servers.
Smart Images

Figure 2025514770000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS REFERENCE / INCORPORATION BY REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 363,319, filed April 21, 2022, which is incorporated by reference in its entirety herein. [Background technology]
[0002] Many modern computing devices, including mobile phones, personal computers, and tablets, include on-device input recognition capabilities: the device can capture and recognize text or voice input and provide the recognized text or voice to one or more downstream applications.
[0003] Mobile phone applications have a limited amount of available computational resources, so allowing users access to efficient input recognition tools while still utilizing fewer computational resources is a significant technological advancement. Summary of the Invention
[0004] Accurate on-device recognition of user input can be challenging when the input includes terms that are not recognized as common terms. For example, when the input is in audio form, one or more phrases may be pronounced in a manner that is highly speaker-dependent. In some examples, phonetic variations of phrases, or misspellings of terms, grammatical preferences, etc., may be personal to the user (e.g., names, places, proper nouns, user-entered). In other examples, the use of certain terms may raise privacy concerns, especially when user-specific models are generated. For example, medical terms such as "hernia" or "hemangioma" may be difficult to recognize in speech using aggregate speech recognition models, and such medical terms may also be considered to be highly personalized content.
[0005] Some existing techniques for improving input recognition require user feedback. For example, a user may be prompted to provide these terms as search terms, and personalization may occur based on such user feedback. However, such personalization does not occur in real-time when speech is being recognized. Other techniques may rely on user-generated rewrites of words or phrases to recognize correctly transcribed terms. However, such user-generated rewrites may not be appropriate for infrequent terms and / or terms specific to a particular domain. Also, for example, users may speak terms with different accents, dialects, etc., and different phonetic representations of a particular term may be difficult to distinguish. Thus, there is a need for an on-device input recognition tool that includes a context-aware correction module for automatic input recognition when input recognition occurs on the device.
[0006] In one aspect, a computer-implemented method is provided. The method includes receiving, by an on-device system operating on a computing device, an input from a user during interaction with the computing device. The method further includes receiving a transcription of the input from an input recognition model. The method additionally includes identifying, by the on-device system, a candidate term for replacement in the transcription of the input, the candidate term having a high probability of being mistranscribed. The method also includes accessing, by the on-device system, a plurality of pairs of mistranscribed terms and uncommon terms based on the candidate terms, the uncommon terms having a high probability of being mistranscribed, the mistranscribed terms being incorrect versions of the uncommon terms, and the mistranscribed terms being generated by a machine learning model. The method further includes replacing, by the on-device system, the candidate term with the uncommon term in the transcription of the input based on the plurality of pairs of mistranscribed terms and uncommon terms.
[0007] In another aspect, a computing device is provided. The computing device includes one or more processors and a data storage. The data storage has stored therein computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform a function. The feature includes receiving, by an on-device system operating on the computing device, an input from a user during an interaction with the computing device. The feature further includes receiving, from an input recognition model, a transcription of the input. The feature additionally includes identifying, by the on-device system, candidate terms for replacement in the transcription of the input, the candidate terms being likely to be mistranscribed. The feature also includes accessing, by the on-device system, a plurality of pairs of mistranscribed terms and uncommon terms based on the candidate terms, the uncommon terms being likely to be mistranscribed, the mistranscribed terms being incorrect versions of the uncommon terms, and the mistranscribed terms being generated by a machine learning model. The functionality further includes replacing, by the on-device system, the candidate term with the uncommon term in the transcription of the input based on a plurality of pairs of the incorrectly transcribed term and the uncommon term.
[0008] In another aspect, a product is provided. The product includes one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of the computing device, cause the computing device to perform a function. The function includes receiving, by an on-device system operating on the computing device, an input from a user during an interaction with the computing device. The function further includes receiving a transcription of the input from an input recognition model. The function additionally includes identifying, by the on-device system, a candidate term for replacement in the transcription of the input, the candidate term having a high probability of being mistranscribed. The function also includes accessing, by the on-device system, a plurality of pairs of mistranscribed terms and uncommon terms based on the candidate terms, the uncommon terms having a high probability of being mistranscribed, the mistranscribed terms being incorrect versions of the uncommon terms, the mistranscribed terms being generated by a machine learning model. The function further includes replacing, by the on-device system, the candidate term with the uncommon term in the transcription of the input based on the plurality of pairs of mistranscribed terms and uncommon terms.
[0009] In another aspect, a system is provided that includes means for receiving, by an on-device system operating on the computing device, an input from a user during interaction with the computing device, means for receiving a transcription of the input from an input recognition model, means for identifying, by the on-device system, candidate terms for replacement in the transcription of the input, where the candidate terms are likely to be mistranscribed, means for accessing, by the on-device system, a plurality of pairs of mistranscribed terms and uncommon terms based on the candidate terms, where the uncommon terms are likely to be mistranscribed, the mistranscribed terms are incorrect versions of the uncommon terms, where the mistranscribed terms have been generated by a machine learning model, and means for replacing, by the on-device system, the candidate terms with the uncommon terms in the transcription of the input based on the plurality of pairs of mistranscribed terms and uncommon terms.
[0010] The above summary is illustrative only and is not intended to be in any way limiting. In addition to the exemplary aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description, and accompanying drawings. [Brief description of the drawings]
[0011] [Figure 1] 1 illustrates an overall system of a context-aware correction model for input recognition, according to an example embodiment. [Diagram 2] 1 illustrates an exemplary input recognition system, according to an exemplary embodiment. [Diagram 3] 4 illustrates an exemplary context-aware correction for input recognition, according to an exemplary embodiment. [Figure 4] 5 illustrates an example training phase of a neural network for a context-aware correction model for input recognition, according to an example embodiment. [Diagram 5]1 illustrates an example neural network for a context-aware correction model for input recognition, according to an example embodiment. [Figure 6] FIG. 1 illustrates the training and inference stages of a machine learning model, according to an example embodiment. [Figure 7] 1 illustrates a distributed computing architecture in accordance with an illustrative embodiment. [Figure 8] FIG. 1 is a block diagram of a computing device in accordance with an exemplary embodiment. [Figure 9] 1 is a flowchart of a method according to an example embodiment. [Figure 10] 1 is a flowchart of a correction method according to an example embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] overview Examples described herein relate to context-aware correction by machine learning for input recognition (e.g., automatic speech recognition (ASR)). Specifically, a trained machine learning model (e.g., a convolutional neural network) can recognize potentially mistranscribed terms in speech transcribed by an ASR system and replace the potentially mistranscribed terms with another term in the transcribed speech. For example, pairs of uncommon terms and related mistranscribed terms may be pre-generated to enable transcription improvements in substantially real-time. In some embodiments, the uncommon terms may be terms observed during a previous interaction of a user with the device to capture related terms without overburdening the system. In some embodiments, the uncommon terms may be terms that are likely to be mistranscribed during transcription. For example, the uncommon terms may be terms that are likely to be mistranscribed. In some embodiments, the uncommon terms may be generated by a suitably trained machine learning algorithm. For example, the non-common term may also be a corrected version of an incorrectly transcribed term (e.g., a term previously corrected by a user and / or a term selected by a user). The incorrectly transcribed term may be incorrect or may be a more likely mistranscribed version of the non-common term. In some embodiments, the incorrectly transcribed term may be generated using a machine learning model. For example, one or more machine learning algorithms may be used to identify and correct the incorrect version of the non-common term. In some embodiments, the incorrectly transcribed term is generated prior to being generated by an input recognition system operating on the device. In general, the incorrectly transcribed term is a possible error that may occur during transcription.
[0013] The generated mistranscribed terms remain on the device, thereby maintaining the privacy of the user data, and for example, data generation and model training can occur entirely on the user's device, which provides important computational and privacy advantages for large service providers with millions of users potentially submitting billions of terms per hour.
[0014] The generated mistranscribed terms extend the capabilities of the existing input recognition system. The modified input recognition system uses terms that are well known to the user, so the total number of mistranscribed terms that the device must consider is limited to a small number. This reduces battery and computational power consumption, especially on resource-limited mobile devices. This technique provides improved latency performance and significantly improved privacy properties.
[0015] The generated incorrectly transcribed terms can be applied to an upstream input recognition system. This ensures that input recognition can be performed on the server or on the device depending on the best available resources. For example, available network resources may enable the device to communicate with the server, while in situations where the network is unavailable and / or the network bandwidth is limited, an on-device modified input recognition system can be used. Also, for example, input recognition can be performed on a remote server while a correction module can be applied by the on-device system.
[0016] Applying the correction module on the device has several advantages. For example, the generated incorrectly transcribed terms never leave the device (e.g., are not shared with servers, other devices, etc.) but are used for on-device processing. This has several advantages. For example, by restricting the content to the device, any user information that may be used for input recognition is maintained with appropriate privacy and / or security controls. For example, the content may be encrypted and stored in a dedicated memory location on the device, and so on. Such features also enhance the capabilities of the input recognition system, since input recognition can be appropriately personalized to a particular user, user preferences, etc., based on user control of the extent of personalization. Also, for example, the user can maintain control over what information about the user is collected, how that information is used, and what information is provided to the user. On-device processing is advantageous because it allows for faster processing, lower power consumption, lower latency, and reduced data transmission over the network. Additionally, on-device processing can continue to perform automatic input recognition in situations where, for example, the server may be unavailable, the network may be unavailable, a protected network may be unavailable, and / or network bandwidth may be limited.
[0017] In addition to the above, the system, program, or functionality described herein may provide the user with controls that allow the user to choose both whether and when the system, program, or functionality described herein may enable collection of user information (e.g., information regarding the user's social network, social contacts, or activities, the user's preferences, or the user's current location, etc.) and whether to transmit content or communications from the server to the user. Additionally, certain data may be processed in one or more ways, such as personal data being removed, protected, encrypted, etc., before being stored or used. For example, the user's identity may be processed such that the user's personally identifiable information cannot be determined, or if location information is obtained (such as at the city, zip code, or state level), the user's geographic location may be generalized such that the user's specific location cannot be determined. Thus, the user may control what information is collected about the user, how that information is used, and what information is provided to the user. In addition to user control, in embodiments where user information is used for input recognition, such user information is restricted to the user's device and is not shared with the server and / or with other devices. Also, for example, the user information may be deleted after use. For example, if the user consents to such use of the data, the data may be used to determine which terms were transcribed incorrectly, and then the data may be securely stored on the device and not shared with other devices, servers, etc.
[0018] In some examples, the trained machine learning model may function on a variety of computing devices, including, but not limited to, mobile computing devices (e.g., smartphones, tablet computers, mobile phones, laptop computers), stationary computing devices (e.g., desktop computers), and server computing devices.
[0019] A machine learning model, such as a convolutional neural network, may be trained using training data (e.g., audio data) to perform one or more aspects as described herein. In some examples, the neural network may be configured as an encoder / decoder neural network.
[0020] The trained machine learning model can process input data to predict output data including one or more words and / or phrases associated with the input data. In one example, (a copy of) the trained neural network can reside within a mobile computing device. In some embodiments, the mobile computing device can include a microphone that can capture the input voice data. In response, the trained neural network can generate predicted output words associated with the input voice data.
[0021] In some embodiments, the first trained machine learning model can perform input recognition to generate a transcription of the input, and the second trained machine learning model can perform auto-correction of the transcription. The auto-correction of the transcription can be performed by the on-device system, while the input recognition can be performed by a remote server. In such embodiments, the on-device system can receive the input and send the input to the remote server. The remote server can apply the first trained machine learning model to generate a transcription of the input. The on-device system can then receive the transcription from the remote server and perform auto-correction of the transcription.
[0022] As such, the techniques described herein can improve input recognition by applying a context-aware correction module, thereby improving the actual and / or perceived quality of the input recognition. Improving the actual and / or perceived quality of the input recognition can benefit downstream applications that rely on the output of the input recognition system (e.g., voice-enabled applications). These techniques are flexible, allowing them to accommodate a wide variety of user inputs, such as human voices, including various languages, dialects, and accents.
[0023] Localized voice correction An on-device correction module is described. The correction module is context-aware. For example, the correction module may be trained using synthetically generated web-scale data via a noisy channel simulator. The model may then be fine-tuned for different applications. Such techniques allow clients to fine-tune the same model, which may enable transfer learning while limiting information leakage between clients.
[0024] FIG. 1 illustrates an overall system 100 of a context-aware correction model for input recognition, according to an exemplary embodiment. A device 110 can receive input from a user during interaction with the device 110 by an on-device system 105 running on the device 110. The device 110 can include a computing device, such as a laptop, desktop computer, smart television, electronic reading device, streaming content device, game console, tablet device, or other related computing device configured to execute software instructions and application programs. The voice input issued by the user can be a voice command to the application program. For example, the application program can be a media playback application to play music, and the voice command can be an instruction to play a particular song. As another example, the voice command can be a search query to search the music library of the media playback application.
[0025] In another example, the application program may be a text editor and the voice command may be dictation to enter text. Also, for example, the voice command may be a command to the computing device to perform one or more operations. For example, the computing device may be a mobile device and the voice command may be an instruction to open an application program (e.g., maps, contacts, browser, etc.), find directions to a location, initiate a voice or video call with a user from a list of contacts, send an instruction to a digital assistant device (e.g., a home assistant device), etc. As another example, the computing device may be associated with a controller of an autonomous vehicle and the voice command may be an instruction related to the operation of the vehicle.
[0026] In some embodiments, the on-device system 105 can be configured to operate on an operating system of the device 110. The on-device system 105 can include an interface (e.g., an interface through an application programming interface (API)) for communicating with one or more application programs on the device 110. For example, the on-device system 105 can communicate with a first application program through a first application API 115a and with a second application program through a second application API 115b.
[0027] As used herein, the term "application program" may be any computer program configured to interact with a user of device 110. Exemplary application programs may include search applications, email applications, text message applications, instant messaging applications, web browsing applications, mapping applications, media playback applications, weather applications, phone applications, video communication applications, camera applications, applications related to service providers (e.g., finance, insurance, etc.), applications related to digital assistants (e.g., home assistants), or any other application program configured to receive user input, such as voice audio input, digital text input, alphanumeric input, character input, and / or digital image input.
[0028] The term "interaction" may broadly refer to any activity, active and / or passive, that a user performs with device 110 or an application program on device 110. For example, an interaction may include viewing content, listening to content, inputting (e.g., via keyboard, mouse, tapping, etc.), editing, and / or modifying content, sensory interaction (e.g., haptic, visual, auditory, tactile interaction, etc.), scrolling interaction, voice interaction, user selection, etc. In some embodiments, an interaction may not be a direct interaction of a user with content. For example, a user may listen to a song of a particular genre or watch a movie of a particular genre. A computing device may determine a user interaction with a particular song from a particular genre as an interaction with a song of the same genre in a library. Similarly, a computing device may determine a user interaction with a particular movie from a particular genre as an interaction with a movie of the same genre in a library. As another example, a user interaction with an email may be determined to be an interaction with an entire chain of emails and / or an interaction with multiple email exchanges with a particular sender of the email.
[0029] In some embodiments, a user interaction may be an interaction with a digital assistant (e.g., an intelligent digital assistant). For example, a user may send a voice command such as, for example, "Turn on the light in the study," "Play music by Jenna K.," "Play episode 2 of the series I watched yesterday," or "Set the thermometer to home." In some embodiments, a user interaction may be an interaction with a search assistant. For example, a user may enter text into a search field in a web browser. As another example, a user may use voice commands to enter search terms such as, for example, "Find the nearest Italian restaurant that was reviewed in the local daily newspaper last weekend."
[0030] In some embodiments, the user interaction may be an interaction with a map application. For example, a user may enter a street address as a text input into an address input field of a mapping application. Or, for example, a user may enter a destination in a navigation application using voice commands. For example, a user may say, "Take me home," or "Find a toll-free route," or "Is there public transportation to the Globe Theatre?"
[0031] The on-device system 105 can interface with one or more hardware components 130 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), memory, input and / or output devices of the device 110). For example, the on-device system 105 can interface with one or more input method editors, such as a voice editor for editing voice input from microphone(s) 135. In some embodiments, voice input by a user can be captured by the microphone(s) 135. The microphone(s) 135 can be part of the device 110 or can be a separate audio input device (e.g., a wired or wireless microphone) communicatively linked to the device 110. For example, a user can activate the microphone(s) 135 to enable voice dictation or to send voice commands to perform actions. Also, for example, the on-device system 105 can interface with an input method editor for text input. Text input can be by a digital or physical keyboard. For example, the touch screen of device 110 can display a digital keyboard. In some embodiments, different digital keyboards may be displayed corresponding to different languages. Also, different digital keyboards may be displayed corresponding to different layouts, designs, etc., for example.
[0032] The on-device system 105 can interface with an input recognition system 125. In some embodiments, the on-device system 105 can receive a transcription of a user input from the input recognition system 125. The input recognition system 125 can be configured to recognize the user input. For example, the input recognition system 125 can be a speech recognition system configured to recognize speech. Also, for example, the input recognition system 125 can be a text recognition system configured to recognize text. For example, the input recognition system 125 can include text recognition logic, programmed instructions, or algorithms to manage the identification, extraction, and analysis of characteristics of the text input. For example, the input recognition system 125 can execute comparison logic to compare spatial characteristics of the text input to various model parameters in a spatial model and / or a language model. For example, the spatial model can be used for text prediction by relating the spatial coordinates of characters or spatial relationships between characters entered by typing, swiping, or gestures.
[0033] In some embodiments, the input recognition system 125 can reside within the device 110. In some embodiments, the input recognition system 125 can be an interface (e.g., an application programming interface) at the device 110 for an input recognition system that resides on a remote server. For example, a first trained machine learning model can perform input recognition to generate a transcription of the input, and a second trained machine learning model can perform automatic correction of the transcription. The automatic correction of the transcription is performed by an on-device system (e.g., the context-aware correction module 120), while the input recognition can be performed by a remote server that interfaces with the device 110 by an API (e.g., the input recognition system 125). In such an embodiment, the on-device system can receive a user input and send the user input to the remote server. The remote server can apply the first trained machine learning model to generate a transcription of the input. The on-device system can then receive the transcription from the remote server and perform automatic correction of the transcription.
[0034] Some embodiments include receiving a transcription of the input from an input recognition model. For example, the input recognition system 125, such as a speech recognition system, can transcribe audio input received, for example, by the microphone(s) 135 and provide the transcription. As used herein, the term "speech recognition" can generally refer to any process of recognizing audio input and converting the audio input into text format. For example, the microphone(s) 135 can receive audio input in the form of a human voice, and the input recognition system 125 can transcribe the human voice into text. In some embodiments, the input recognition system 125 can include one or more machine learning models trained to recognize speech.
[0035] Recognizing input (e.g., speech, text, etc.) can be a challenging task. For example, there can be challenges due to the inherent complexities of various languages and / or dialects of those languages. In some embodiments, there can be variations in the way speech is uttered and / or text is entered by different individuals. For example, input recognition can be particularly difficult when a particular term is not recognized as a common term in a dictionary associated with the language. Although input recognition systems can be trained to recognize different languages, accents, dialects, etc., such input recognition systems may still be unable to recognize words or phrases spoken and / or entered by a user. Also, for example, such input recognition systems may be located on a server remote from the device, and various data processing limitations may limit proper personalization of such systems, and / or network limitations may limit access to such remote servers. Thus, on-device processing may be preferred for both privacy and / or security controls, as well as to enable faster processing, less power consumption, less latency, less data transmission over the network, etc.
[0036] In some embodiments, the on-device system 105 may identify candidate terms for replacement in the transcription of the input. For example, the on-device system 105 may perform one or more operations including scanning the transcription, filtering common terms, performing context analysis, etc. The on-device system 105 may then determine that the transcription is accurate (within an appropriate threshold of acceptable accuracy) and no corrections are made. In some embodiments, the on-device system 105 may identify one or more candidate terms that are likely to have been transcribed incorrectly. For example, the on-device system 105 may identify one or more uncommon terms that are likely to be transcribed incorrectly, one or more terms that have been previously transcribed incorrectly, one or more terms that have been previously corrected by a user in a past interaction with the computing device, etc.
[0037] In some embodiments, the accuracy of transcription of a user input (e.g., speech-to-text conversion, text input via a keyboard, etc.) can be improved based on the context of the input (e.g., email application, mapping application, home assistant, search query, etc.). For example, phonetically similar words may be transcribed into two different transcription versions based on the context of the speech. To that end, the on-device system 105 can include a context-aware correction model 120 that can correct the transcription of the input by the input recognition system 125. For example, the context-aware correction model 120 can interface with one or more application programs on the device 110 (e.g., interface with a first application program via a first application API 115a and interface with a second application program via a second application API 115b) to understand the underlying context of the user input and correct the transcribed terms generated by the input recognition system 125.
[0038] In some embodiments, the on-device system 105 accesses a plurality of pairs of mistranscribed terms and uncommon terms based on the candidate terms, where the uncommon terms are likely to be mistranscribed and the uncommon terms are likely to be mistranscribed, and the mistranscribed terms of the uncommon terms in the plurality of pairs are generated by a machine learning model. In some embodiments, the uncommon terms in the plurality of pairs may have been observed by the on-device system in one or more past interactions between the user and the computing device. In some embodiments, the uncommon terms in the plurality of pairs may have been synthetically generated (e.g., based on aggregate statistics of user inputs, trained machine learning models, etc.). Also, for example, the uncommon terms may be corrected versions of the mistranscribed terms (e.g., terms previously corrected by the user and / or terms selected by the user during an autocorrection process). The mistranscribed terms are commonly occurring mistranscribed versions of the uncommon terms. In some embodiments, the mistranscribed terms may be generated using a machine learning model. For example, one or more machine learning algorithms may be used to identify and correct incorrect choices. In some embodiments, the mistranscribed terms are generated prior to being generated by an input recognition system running on the device. In general, mistranscribed terms are possible errors that may occur during transcription. As described herein, pairs of uncommon terms and related mistranscribed terms may be pre-generated to allow for on-the-fly improvement of the transcription.
[0039] In some embodiments, based on multiple pairs of incorrectly transcribed terms and uncommon terms, the on-device system 105 can replace a candidate term with an uncommon term in the transcription of the input. For example, the on-device system 105 can compare the candidate term with one or more incorrectly transcribed terms and then select an uncommon term from the pairs of incorrectly transcribed terms and uncommon terms (e.g., based on a ranking of the pairs). The on-device system 105 can then replace the candidate term with the selected uncommon term.
[0040] 2 illustrates an exemplary input recognition system 200, according to an exemplary embodiment. In some embodiments, speech input 210 uttered by a user 205 during interaction with a computing device 220 may be received by an on-device system 215 operating on the computing device 220. The speech input 210 may be received in the form of an audio signal 225. Feature extraction 230 may be configured to extract one or more features of the audio signal 225. An acoustic model 235 may be configured to associate relationships between the audio signal 225 and the phonemes or other linguistic characteristics that form the speech audio. For example, the acoustic model 235 may be configured to identify and associate particular received utterances that exhibit acoustic characteristics that match the acoustics associated with a spoken word or phrase.
[0041] The language model 240 can be configured to specify or identify particular word combinations or sequences. In some implementations, the language model 240 can be configured to generate word sequence probability factors that can be used to indicate the likely occurrence or presence of particular word sequences or word combinations. The identified word sequences may correspond primarily to sequences characteristic of a speech corpus rather than a written corpus.
[0042] The speech recognition model 245 may be configured to receive inputs from the acoustic model 235 and the language model 240 to generate a transcript of the speech input 210. For example, the speech recognition model 245 may be configured to include speech recognition logic, programmed instructions, and / or algorithms executed by one or more processors to transcribe the speech input 210. For example, the speech recognition model 245 may execute program code that manages the identification, extraction, and analysis of characteristics of the received audio signal 225. Additionally, the speech recognition model 245 may execute comparison logic to compare characteristics of the received audio signal 225 to various model parameters stored in the acoustic model 235 and the language model 240. The comparison may result in a text transcription output that substantially corresponds to the speech input 210 provided by the user 205 of the computing device 220.
[0043] In some embodiments, the on-device system 215 may have access to multiple pairs of incorrectly transcribed terms and uncommon terms. The uncommon terms in the multiple pairs may have been observed by the on-device system 215 in one or more past interactions between the user 205 and the computing device 220. In general, an "uncommon term" as used herein may refer to any term that may have a high probability of being incorrectly transcribed in a speech-to-text transcription process. In some embodiments, the uncommon terms may have different phonetic versions and / or may be different transcribed versions of the phonetic terms.
[0044] In some embodiments, the user 205 may have viewed one or more medically related documents and the on-device system 215 may have extracted uncommon terms that appeared in the medically related one or more documents. As another example, the user 205 may have listened to one or more songs associated with a music genre and the on-device system 215 may have extracted uncommon terms that appeared in the one or more songs, or transcripts thereof. Also, for example, the user 205 may have interacted with a text editor to enter and / or modify text and the on-device system 215 may have extracted uncommon terms that appeared in the text.
[0045] The mistranscribed terms of the uncommon term in the plurality of pairs may have been generated by a machine learning model. In general, the "mistranscribed terms" may refer to different versions of the uncommon term, where the different versions correspond to likely phonetic variations of the uncommon term when spoken in speech, different text versions, misspelled versions, and / or likely mistranscribed versions of the uncommon term. The mistranscribed terms may be synthetically generated by a trained machine learning model. In some embodiments, the mistranscribed terms may be generated based on training data using speech spoken by various individuals.
[0046] In some embodiments, the on-device system 215 can identify candidate terms for replacement in the transcription of the input. For example, the on-device system 215 can identify candidate terms 250 that are likely to have been transcribed incorrectly by the speech recognition model 245. The context-aware correction model 255 can thus be configured to determine whether the candidate terms 250 have been transcribed incorrectly, and if so, to replace the candidate terms 250 with corrected terms 260 in the transcription of the input. In some embodiments, the context-aware correction model 255 can be configured to communicate with the computing device 220 to identify one or more uncommon terms observed by the on-device system 215 in one or more past interactions between the user 205 and the computing device 220.
[0047] 2 are shown on computing device 220 for illustrative purposes only, and various additional and / or alternative embodiments are contemplated. For example, context-aware correction model 255 resides within computing device 220, while one or more of the other components may reside within computing device 220, on a remote server, or both. For example, speech recognition model 245 may reside within computing device 220. In some embodiments, speech recognition model 245 may reside on a remote server (e.g., a cloud server). Also, for example, certain portions of speech recognition model 245 may reside within computing device 220 and other portions may reside on a remote server.
[0048] For example, the on-device system 215 can receive the audio signal 225 and send the audio signal 225 to a remote server for processing. Feature extraction 230 can be performed by the remote server. Also, for example, at the remote server, a speech recognition model 245 can utilize an acoustic model 235 and a language model 240 to generate a transcription based on the audio signal 225 and send the transcription to the computing device 220. The on-device system 215 can then identify candidate terms 250, apply a context-aware correction model 255, and identify corrected terms 260 to correct the transcription.
[0049] Although FIG. 2 illustrates an exemplary embodiment for speech recognition, a similar approach can be applied to text recognition. For example, the speech recognition model 245 can be a text recognition model that can utilize a keyboard model to recognize text input. In some embodiments, the keyboard model can include a spatial model, a language model, a keyboard input mode editor, and the like. For example, the keyboard model can be configured to receive touch input and / or physical input corresponding to letters, numbers, symbols, emoticons, and / or characters. The incorrectly transcribed terms can include spelling variations, typical misspellings, common grammatical errors, capitalization variations, and the like.
[0050] FIG. 3 illustrates an exemplary context-aware correction 300 for input recognition, according to an exemplary embodiment. For example, a user 305, while interacting with a device 310, can utter speech 340 such as "Find items related to nurnia" that includes the term "Nurnia." Although the device 310 is shown as a mobile device, the device 310 can be any device (e.g., smart TV, digital content delivery device) configured to interact with the user 305. In some embodiments, the device 310 can receive speech uttered by the user 305, and a speech recognition system can transcribe the uttered speech into a transcription of the speech input, such as "Find items related to nurnia." In some embodiments, the context-aware correction model 325 can identify "nurnia" as a candidate term for replacement.
[0051] In some embodiments, the context-aware correction model 325 may have access to multiple pairs of mistranscribed terms and uncommon terms. For example, the context-aware correction model 325 may have access to an on-device repository 360 of pairs of mistranscribed terms and uncommon terms. As an illustrative example, the pairs of mistranscribed terms and uncommon terms may be "(haarnia, hernia)", "(narnia, hernia)", and "(hair near, hernia)", where the mistranscribed terms "haarnia", "narnia", and "hair near" are paired with the uncommon term "hernia". As shown in block 355, the context-aware correction model 325 may compare the transcribed term "nurnia" 350 to the mistranscribed terms and determine that "nurnia" 350 matches the mistranscribed term "narnia". As used herein, the term "match" may generally refer to a match within a similarity threshold. In some embodiments, the term "match" may refer to an exact match.
[0052] In some embodiments, the context-aware correction model 325 may replace a candidate term with a non-common term in the transcription of the input based on a plurality of pairs of the mistranscribed term and the non-common term. For example, if it is determined that "nurnia" 350 matches the mistranscribed term "narnia", the context-aware correction model 325 may determine that the non-common term "hernia" paired with "narnia" is the correct transcription of the term "nurnia" 350 as "(narnia, hernia)". For example, the selection of the pair may be based on the confidence level of the pair. Thus, the non-common term "hernia" 365 is used to replace the candidate term "nurnia" 350 in the transcribed text. For example, the context-aware correction model 325 corrects the transcription "find items related to nurnia" generated by the speech recognition system to "find items related to hernia". Although this example shows the replacement of a single term, "nurnia," the same techniques can be applied to correct multiple candidate terms in a transcription.
[0053] In some embodiments, replacing the candidate term includes comparing, by the on-device system, the candidate term to one or more mistranscribed terms from the plurality of pairs. In general, a first comparison may be made between the candidate term and the plurality of mistranscribed terms based on a similarity determination or matching, and a second comparison may be made between the pairs including the mistranscribed terms based on a confidence level for each of the pairs. In some embodiments, a final selection of the pair may be based on a combination of the first and second comparisons.
[0054] In some embodiments, the context-aware correction model 325 can be configured to parse (e.g., observe) text that the user 305 is viewing and / or editing. This can include text inputs of emails, short messages, query terms, and / or text that the user 305 is viewing on the device 310, such as, for example, web documents, inbound messages, text within applications, etc. This content is used for on-device processing without leaving the device (e.g., without being shared with servers, other devices, etc.). This has several advantages. For example, by restricting the content to the device, any user information that may be used for input recognition is maintained with appropriate privacy controls. Such a feature also enhances the capabilities of the input recognition system, since input recognition can be appropriately personalized to a particular user, user preferences, etc., based on user control of the extent of personalization. Also, for example, the user can maintain control over what information about the user is collected, how that information is used, and what information is provided to the user. On-device processing is advantageous because it allows for faster processing, lower power consumption, lower latency, and reduced data transmission over the network, and can continue to perform automatic input recognition in situations where, for example, a server may be unavailable, a network may be unavailable, a protected network may be unavailable, and / or network bandwidth may be limited.
[0055] In some embodiments, identifying candidate terms for replacement in the transcription of the input by the on-device system may include filtering common terms and retaining uncommon terms. For example, the context-aware correction model 325 may be configured to filter common terms and retain uncommon terms. The candidate terms are likely to have been transcribed incorrectly.
[0056] For example, the user 305 may have previously interacted with the device 310. For example, the user 305 may have viewed content related to medical information. Accordingly, the context-aware correction model 325 may have filtered out the non-common term “hernia” that appears in the content previously viewed by the user 305. In some embodiments, the context-aware correction model 325 may have generated one or more mistranscribed terms (e.g., by using a trained machine learning model). The mistranscribed terms may be determined to be text versions of phonetically similar utterances of the non-common term “hernia”. For example, the context-aware correction model 325 may have generated “haarnia”, “narnia”, and “hair near” as mistranscribed terms corresponding to the non-common term “hernia”. Thus, the mistranscribed term and non-common term pairs “(haarnia, hernia)”, “(narnia, hernia)”, and “(hair near, hernia)” may be stored in the on-device repository 360.
[0057] In some embodiments, the user 305's past interactions with the device 310 may have included viewing media content. For example, the user 305 may have viewed a movie titled "Narnia" or may have listened to the soundtrack of the movie "Narnia." As another example, the user 305 may have read a review of the movie "Narnia." Accordingly, the context-aware correction model 325 may have filtered out the non-generic term "narnia" that appears in the content previously viewed by the user 305. In some embodiments, the context-aware correction model 325 may have generated one or more mistranscribed terms (e.g., by using a trained machine learning model). The mistranscribed terms may be determined to be text versions of phonetically similar utterances of the non-generic term "narnia." For example, the context-aware correction model 325 may have generated "haarnia," "hernia," and "hair near" as mistranscribed terms that correspond to the non-generic term "narnia." Thus, the pairs of incorrectly transcribed terms and uncommon terms “(haarnia, narnia)”, “(hernia, narnia)”, and “(hair near, narnia)” may be stored in the on-device repository 360.
[0058] In some embodiments, each pair of mistranscribed terms and uncommon terms stored in the on-device repository 360 may be associated with a respective confidence level that represents the similarity of the particular mistranscribed term to the particular uncommon term paired with the particular mistranscribed term. In some embodiments, the similarity may be a measure of phonetic similarity between the mistranscribed term and the uncommon term.
[0059] The confidence level may be determined based on one or more factors. For example, the confidence level may be based on the source of the mistranscribed term. In some embodiments, a voice sample of the user may be used to select or deselect a term as a mistranscribed term. Also, for example, one or more previous voice interactions of the user may be used to select or deselect a term as a mistranscribed term. In such an embodiment, a confidence level associated with a pair including a selected mistranscribed term and an associated non-common term may be determined to be high.
[0060] In some embodiments, the frequency of utterances of the mistranscribed term by a user can be used to determine the confidence level. For example, if the mistranscribed term is uttered in the same manner with a high frequency, the confidence level associated with a pair including the mistranscribed term and the related non-common term can be determined to be high. For example, a user may pronounce the non-common term "Narnia" as "naaarnia," and the computing device may determine that the frequency of utterances of "Narnia" as "naaarnia" is higher than a frequency threshold. Thus, the confidence level associated with the pair (Naaarnia, Narnia) can be determined to be high. The frequency of utterances of a particular mistranscribed term can be based on several factors, including term frequency, term frequency-inverse literature frequency (TF-IDF), etc.
[0061] In some embodiments, the frequency of adherence of the uncommon term by the user may determine the confidence level. For example, the computing device may determine the frequency of occurrence of a particular uncommon term in content viewed by the user. The viewed content may be based on a single document, a single application program, or multiple application programs. Thus, based on a determination that the uncommon term is frequently observed by the user, a confidence level associated with a pair including the uncommon term and the associated mistranscribed term may be determined to be high. Similarly, based on a determination that the uncommon term is frequently observed by the user, a confidence level associated with a pair including the uncommon term and the associated mistranscribed term may be determined to be low. In general, the frequency of a particular uncommon term may be based on several factors, including term frequency, term frequency-inverse literature frequency (TF-IDF), etc.
[0062] In some embodiments, the confidence level of the pair may be based on the respective frequency of occurrence of the mistranscribed term and the uncommon term. For example, the confidence level may be based on a joint distribution of their individual frequencies. In some embodiments, the confidence level associated with the pair may vary based on the underlying application program and / or context of the uncommon term. Also, for example, the confidence level associated with the pair may vary from user to user. In some embodiments, the confidence level associated with the pair may be based on language, accent, geographic location, etc.
[0063] For example, the pairs "(haarnia, hernia)", "(narnia, hernia)", and "(hair near, hernia)" may each be associated with a high confidence level, indicating that the mistranscribed terms "haarnia", "narnia", and "hair near", respectively, have a high similarity to the non-common term "hernia". In some embodiments, the similarity may be a phonetic similarity. As another example, the pair "(hyena, hernia)" may be associated with a medium confidence level, indicating that the mistranscribed term "hyena" has a medium similarity to the non-common term "hernia". And, for example, the pair "(herein, hernia)" may be associated with a low confidence level, indicating that the mistranscribed term "herein" has a low similarity to the non-common term "hernia".
[0064] Similarly, the pairs "(haarnia, narnia)", "(hernia, narnia)", and "(hair near, narnia)" may each be associated with a high confidence level, indicating that the mistranscribed terms "haarnia", "hernia", and "hair near", respectively, have a high similarity to the non-common term "narnia". As another example, the pair "(naina, narnia)" may be associated with a medium confidence level, indicating that the mistranscribed term "naina" has a medium similarity to the non-common term "narnia". Also, for example, the pair "(aria, narnia)" may be associated with a low confidence level, indicating that the mistranscribed term "aria" has a low similarity to the non-common term "narnia".
[0065] At runtime, a user 305 may be interacting with the device 310. In some embodiments, the user 305 may be interacting with content 345. For example, the user 305 may be a doctor and the content 345 may be a live transcript of a voice interaction between the user 305 and a patient. Thus, the context-aware correction model 325 may compare the transcribed term "nurnia" 350 with the incorrectly transcribed term and determine that "nurnia" 350 matches the incorrectly transcribed term "narnia." However, based on the content 345 indicative of a medical context, the context-aware correction model 325 may determine that "hernia" 365 is a correct transcription of the transcribed term "nurnia" 350.
[0066] Some embodiments include receiving, by the computing device, a second speech input uttered by the user during a second interaction with the computing device, where a second transcription of the input includes the candidate terms. For example, during the second interaction, the user 305 may be interacting with a media playing application, and the content 345 may be a browser displaying content provided by the media playing application. For example, the user 305 may utter speech 340 to submit a voice search query to the media playing application. Such embodiments include comparing the candidate terms to one or more second incorrectly transcribed terms, where the one or more second incorrectly transcribed terms were observed by the on-device system in one or more past interactions of the user with a second application program of the computing device, and where the one or more second incorrectly transcribed terms are different from the one or more incorrectly transcribed terms. For example, the context-aware correction model 325 may compare the candidate term "nurnia" 350 to the mistranscribed terms and determine that "nurnia" 350 matches the mistranscribed term "narnia." However, based on the content 345 indicative of the media playback context, the context-aware correction model 325 may determine that "narnia" is a correct transcription of the transcribed term "nurnia" 350. Such an embodiment also includes replacing, by the on-device system, the candidate term in the transcription of the input based on the comparison with a second uncommon term, the second uncommon term paired with one of the one or more second mistranscribed terms. For example, "nurnia" 350 is replaced with the second uncommon term "Narnia."
[0067] In general, one or more of the determination of an incorrectly transcribed term, pairs containing an incorrectly transcribed term, a confidence level associated with each pair, and a threshold for accepting pairs for replacement may vary from user to user, from application program to application program, and so forth. Also, for example, such determinations may be made by a trained machine learning model. Also, for example, such determinations may vary based on the application program, the user, the user's geographic location, and / or the language used.
[0068] In some embodiments, user feedback may be incorporated into such a determination. In some embodiments, the one or more past interactions between the user and the computing device may include a voice interaction. The non-common terms in the pairs may be based on a user confirmation of the transcribed terms based on the voice interaction. For example, the ASR may transcribe speech uttered by the user and prompt the user to confirm that the transcribed terms are correct.
[0069] In some embodiments, the one or more past interactions of the user with the computing device can include interactions with a text editor. For example, the user may have entered text into an electronic message, a short message, or the like, and the computing device may prompt the user to confirm a corrected term and / or prompt the user to select a correct term from one or more candidate terms provided to the user. Thus, the non-common term in the plurality of pairs can be based on user confirmation and / or user selection of a text term in the text editor.
[0070] In some embodiments, the computing device may include a viewer interface. One or more past interactions between the user and the computing device may include text content provided by the viewer interface. For example, the user may be browsing documents on the web, browsing a library related to media content, etc. The non-common terms in the pairs may appear as text content in the document being browsed, the music library, etc.
[0071] Input recognition models are typically trained based on user preferences. However, the quality of such training depends on the quality of available logs of past user preferences. In general, maintaining such logs can require a lot of memory space and also requires continuous training of input recognition models, which can be resource intensive. However, a more efficient approach is to identify content viewed by a user, filter out uncommon terms from that content, and store these in a local repository as potentially correct terms in the context of input recognition.
[0072] Variations in user accents can be a source of errors in automatic speech recognition. A user's geographic or regional accent can be determined by the user's location information. Also, for example, a user's ethnicity may indicate ethnic dialect (e.g., ethnolect), spelling variations, etc. For example, ethnic dialect may indicate influences from the user's first language, among others. Some variability factors can also be inferred from user information. In some embodiments, user-provided preferences can be used. For example, a user is prompted to utter a sentence, and the automatic speech recognition system can infer the nuances and speech characteristics specific to the user.
[0073] In addition to the above, the system, program, or functionality described herein may provide controls to the user that allow the user to choose both whether and when the system, program, or functionality described herein may enable collection of user information (e.g., information regarding the user's ethnicity, gender, social network, social contacts, or activities, the user's preferences, or the user's current location, etc.) and whether content or communications are sent from the server to the user. Additionally, certain data may be processed in one or more ways, such as personal data being removed, protected, encrypted, etc., before being stored or used. For example, the user's identity may be processed such that the user's personally identifiable information cannot be determined, or if location information is obtained (such as at the city, zip code, or state level), the user's geographic location may be generalized such that the user's specific location cannot be determined. Thus, the user may control what information is collected about the user, how that information is used, and what information is provided to the user. In addition to user control, in embodiments where user information is used for input recognition, such user information is restricted to the user's device and is not shared with the server and / or with other devices. Also, for example, the user information may be deleted after use. For example, if a user consents to the use of such data, the data may be used to determine pairs that include incorrectly transcribed terms and uncommon terms, and the determined pairs are stored on the device, although the sources of the content associated with the incorrectly transcribed terms and uncommon terms may not be remembered. As another example, user history associated with past interactions with the device may not be stored.
[0074] FIG. 4 illustrates an exemplary training phase of a neural network 400 for a context-aware correction model for input recognition, according to an exemplary embodiment. In some embodiments, initial training of the neural network 400 may be performed without user logging. Also, domain adaptation to a particular application program may be performed, for example, by synthetically simulating errors (e.g., speech errors, context errors, etc.). For example, domain adaptation to YouTube® may be performed for the names of creators, channels, artists, videos, etc., and the base model may then be fine-tuned. For example, an initial version of the context-aware correction model 450 may be pre-trained remotely, and subsequent training may be performed on the computing device. As described herein, the input recognition system described herein may be trained and reside in a remote server, while the context-aware correction model 450 may reside in a local computing device.
[0075] As another example, a music application program may be installed on a device, and the device may provide access to a catalog of music. Thus, pronunciation variations of uncommon terms (e.g., artist names, song titles, portions of lyrics, etc.) may be generated as incorrectly transcribed terms. When a user interacts with the voice interface of the music application program, a context-aware correction model can correct the output of an automatic speech recognition system to arrive at an accurate transcription of the spoken speech.
[0076] Also, for example, an email system may be installed on the device, and the device may access contact information associated with the email system. The device may also be able to observe the user reading one or more messages, and uncommon terms may be identified from the content of such messages. Phonetic variations of such uncommon terms (e.g., names, greetings, keywords or phrases, etc.) may be generated as incorrectly transcribed terms. Thereafter, when the user interacts with the voice interface of the email system (e.g., to dictate a new message), the context-aware correction model may correct the output of the input recognition system to arrive at an accurate transcription of the spoken speech.
[0077] As another example, an application program may provide reviews of restaurants in a geographic region. Such an application may have access to restaurant names, addresses, menus, chefs' names, dish names, etc. The device may observe a user browsing through menus of particular restaurants and identify uncommon terms from such browsing activity. Pronunciation variations of such uncommon terms may be generated as incorrectly transcribed terms. The user may then interact with a voice interface of the restaurant review application program, or another application (e.g., a mapping application, a reservation application, a short messaging application, an email application, a calendar application, etc.), and the context-aware correction model may correct the output of the input recognition system to arrive at an accurate transcription of the spoken speech. For example, a user may use a mapping application to generate directions to a particular restaurant, and the context-aware correction model may enable accurate recognition of the restaurant's name, street name, etc. As another example, a user may use a short messaging application to send a text message to a contact to arrange a meeting at a particular restaurant serving a particular menu, and the context-aware correction model may enable accurate recognition of the restaurant's name, dish type, etc.
[0078] In some embodiments, training the machine learning model further includes receiving a corpus of documents. In some embodiments, training the machine learning model can be performed on a server separate from the computing device on which the context-aware correction model 450 resides. For example, the machine learning model may be pre-trained on the remote server. The document corpus 405 may be received by the pronunciation-based text API 410. The document corpus 405 may include a plurality of documents from the Internet. Also, for example, the document corpus 405 may be application-specific documents, such as, for example, content listings for a media playback application, and / or other application domain-specific documents. In some embodiments, the documents in the document corpus 405 may be associated with one or more application programs and / or concepts to which the documents relate. For example, the document corpus 405 may include documents related to music, movies, medicine, law, philosophy, history, literature, etc.
[0079] In some embodiments, training the machine learning model further includes synthetically simulating one or more errors based on a corpus of web documents. In some embodiments, synthetically simulating one or more errors is based on a text-to-speech model that utilizes a noisy channel simulator. For example, the pronunciation-based text API 410 can use a text-to-speech (TTS) model with noise and accents 415. For example, the TTS model 415 can include training data that converts text to speech under different controlled noise conditions to generate variations in the speech induced by the noise conditions. For example, different forms of noise (e.g., random noise) can be artificially injected into the speech to corrupt the speech and generate the training data. Also, for example, different forms of background noise (e.g., traffic noise, airport noise, wind, water, rain, background music, background conversation, etc.) can be artificially injected into the speech to corrupt the speech and generate the training data. The pronunciation-based text API 410 can apply the TTS model 415 to the document corpus 405 to generate a corrected corpus 430 that includes text errors in the document corpus 405, where the errors are induced by noise conditions.
[0080] Also for example, the TTS model 415 can include training data in which text is converted to speech with different accents, dialects, etc. For example, a user can be prompted to read different text documents to capture the user's variations and speech characteristics. The pronunciation-based text API 410 can apply the TTS model 415 to the document corpus 405 to generate a corrected corpus 430 that includes text errors in the document corpus 405, where the errors are induced by different accents, dialects, etc.
[0081] In some embodiments, synthetically simulating the one or more errors is based on a grapheme-to-phoneme conversion model configured to generate a pronunciation of the word based on a text version of the word. For example, the pronunciation-based text API 410 can receive the output of an automatic speech recognition (ASR) system and the output of a grapheme-to-phoneme (G2P) conversion 420. For example, the G2P conversion 420 can include training data in which a grapheme sequence including letters is converted to a phoneme sequence representing the pronunciation of the grapheme sequence. The pronunciation-based text API 410 can apply the G2P conversion 420 to the document corpus 405 to generate a corrected corpus 430 that includes errors of the text in the document corpus 405, where the errors are induced by different phoneme sequences representing the pronunciation of the grapheme sequence.
[0082] In some embodiments, synthetically simulating one or more speech errors is based on a statistical phoneme model. For example, the pronunciation-based text API 410 can receive input from a statistical phoneme model 425. For example, the statistical phoneme model 425 can utilize various phoneme characteristics provided by, for example, an ASCII transcription of the International Phonetic Alphabet (IPA). The statistical phoneme model 425 can include training data based on different articulatory modalities. For example, the source of the airflow can be considered, such as whether the airflow occurs at the lungs, whether the airflow occurs at the tongue, whether the airflow occurs at the vocal cords, etc. In some embodiments, the target of the airflow can be considered, such as whether the airflow is directed towards the mouth or towards the nose. The direction of the airflow, for example, whether the airflow is towards or away from the target, can be considered. Additional and / or alternative phoneme characteristics can be modeled by the statistical phoneme model 425. Thus, the pronunciation-based text API 410 can apply the statistical phoneme model 425 to the document corpus 405 to generate a corrected corpus 430 that includes errors in the text in the document corpus 405, where the errors are induced by different phoneme characteristics.
[0083] Text input 445 generally refers to a transcription of the input. Data from the corrected corpus 430, along with the text input 445, may be input to the context-aware correction model 450 for training purposes. In some embodiments, as part of the pre-training stage 435, labels 440 may be generated that capture the labeled training data. For example, labels 440 may be generated based on the document corpus 405 and the output of the context-aware correction model 450.
[0084] Some embodiments include training a machine learning model to generate mistranscribed terms for the non-common terms in a plurality of pairs. For example, the context-aware correction model 450 may be trained to generate mistranscribed terms for the non-common terms. For example, the non-common terms may be identified in the document corpus 405, and the context-aware correction model 450 may be trained to generate mistranscribed terms for the identified non-common terms. At runtime, once a potential mistranscribed term is identified in the transcribed audio, the potential mistranscribed term may be compared to various mistranscribed terms to identify a correct non-common term that can be used to replace the potential mistranscribed term in the transcribed audio.
[0085] Some embodiments include tuning the trained machine learning model based on an application program of the computing device. For example, an uncommon term can be associated with a first set of mistranscribed terms corresponding to a first application program and a second set of mistranscribed terms corresponding to a second application program. For example, a particular uncommon term can be associated with a set of mistranscribed terms from a music library, but the same uncommon term can be associated with a different set of mistranscribed terms based on a web browser application. In some embodiments, tuning the trained machine learning model includes generating mistranscribed terms in a plurality of pairs based on one or more errors associated with the application program. For example, in some music genres, the uncommon term can be pronounced in a particular manner that can be a phonetically incorrect version of the uncommon term. For example, the uncommon term "tomato" may be pronounced as "tometo", "tamaato", and the trained machine learning model can recognize "tometo", "tamaato" as mistranscribed terms of "tomato".
[0086] In some embodiments, the various models 415, 420, and 425 may be trained on a server and packaged with the device. Also, for example, the various models 415, 420, and 425 may be trained on a server and updates may be provided to the device to train the context-aware correction model 450.
[0087] The generated mistranscribed terms are stored in a local memory of the computing device. In some embodiments, the mistranscribed terms are not transmitted outside of the device (e.g., not shared with a server or another device). In some embodiments, the mistranscribed terms may be associated with an application program. In such embodiments, the mistranscribed terms may not be shared across application programs. For example, a first set of mistranscribed terms associated with a first application program is stored in a first memory associated with the first application program, and a second set of mistranscribed terms associated with a second application program is stored in a second memory associated with the second application program, where the first memory and the second memory are partitioned or firewalled from each other.
[0088] In general, the data generation of mistranscribed terms and pairs of mistranscribed terms and uncommon terms, as well as the training of the machine learning model, can occur on the user's device, which can provide significant computational advantages, especially in situations where millions of users may generate billions of terms per hour.
[0089] The generated mistranscribed terms extend the capabilities of existing input recognition systems. For example, the context-aware correction model 450 uses non-common terms from the user's previous interactions with the computing device, limiting the total number of mistranscribed terms that need to be considered for speech correction purposes. This approach reduces both battery power and computational power, which is especially important in resource-limited mobile devices.
[0090] In some embodiments, the generated incorrectly transcribed terms can be applied to any upstream input recognition system. Such an approach ensures that the corrected input recognition can be performed on a server or on-device according to the best available resources. For example, a hybrid approach can be applied in resource-constrained environments (e.g., mobile devices, small network bandwidth, etc.) where the context-aware correction model 450 is applied, and in less resource-constrained environments (e.g., desktop computing devices, large network bandwidth, etc.), a server-based input recognition system using deep learning can be used.
[0091] FIG. 5 is an exemplary neural network 500 for a context-aware correction model for input recognition, according to an exemplary embodiment. As shown, user characteristics 505 may be used to determine a user profile 510. In some embodiments, the user characteristics 505 may include one or more languages associated with the user, dialects associated with the user, geographic or location information, and the like. User-specific modulations 515 may be applied to layers of the neural network. For example, the user-specific modulations 515 may include a persistent personalization API ranker. Also, for example, the user may be prompted to speak a given sentence, and the user's speech characteristics may be inferred based on the sentence spoken. Such personalization allows the input recognition system to adapt to the user based on user-specific attributes.
[0092] In some embodiments, a pre-trained model with accent modulation units can be deployed on a user computing device. Biasing and / or fine-tuning processes can be performed based on the user's information, such as contacts, location, etc. Such user information is personalized and the context-aware correction model is configured to be trained on the device, and the generated data is also stored locally on the device. For example, the user may be viewing a report card, medical prescription, disease diagnosis, financial documents, or other forms of protected data, and the context-aware correction model can be configured to extract non-generic terms from such user activity and store them in the device's local memory. Thus, user information is utilized, but the data does not leave the user's computing device. This allows for a private, personalized input recognition system for the user without using user history and other forms of user interaction that the user does not provide. Minimizing the use of user history and / or logs can significantly improve the use of computational resources. In some embodiments, personal information and / or other forms of protected data can be encrypted on the device to further protect it.
[0093] The neural network 500 may be a deep neural network with an encoder-decoder architecture and several intermediate layers. In some embodiments, the neural network 500 may utilize an on-device neural machine translation (NMT) model. For example, an FNet model may be applied to a language task, in which the self-attention layer of the encoder is replaced with a Fourier layer that applies a two-dimensional (2D) Fourier transform to the input. At runtime, the input 520 may be a non-generic term. The encoder 525 may encode the input 520, which then passes through one or more layers of the neural network 500. The user-specific modulation may provide user attributes to the one or more layers. The decoder 530 may then identify mistranscribed terms that correspond to the non-generic terms provided as the input 520. The mistranscribed terms are generated as an output 535 of the neural network 500. Although the encoder 525 and the decoder 530 are shown as a single block, the neural network 500 may include multiple encoders and decoders.
[0094] Training a machine learning model to generate inferences / predictions FIG. 6 illustrates a diagram 600 showing a training phase 602 and an inference phase 604 of trained machine learning model(s) 632, according to an example embodiment. Machine learning techniques include training one or more machine learning algorithms with an input set of training data to recognize patterns in the training data and output inferences and / or predictions (regarding the patterns in the training data). The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, FIG. 6 illustrates a training phase 602 in which machine learning algorithm(s) 620 are trained with training data 610 to become trained machine learning model(s) 632. Then, during the inference phase 604, the trained machine learning model(s) 632 may receive input data 630 and one or more inference / prediction requests 640 (possibly as part of the input data 630) and, in response, provide one or more inferences and / or prediction(s) 650 as output.
[0095] Thus, the trained machine learning model(s) 632 may include one or more models of the machine learning algorithm(s) 620. The machine learning algorithm(s) 620 may include, but are not limited to, artificial neural networks (e.g., convolutional neural networks as described herein, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems). The machine learning algorithm(s) 620 may be supervised or unsupervised and may perform any suitable combination of online and offline learning.
[0096] In some examples, the machine learning algorithm(s) 620 and / or the trained machine learning model(s) 632 may be accelerated using on-device coprocessors, such as graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), and / or application specific integrated circuits (ASICs). Such on-device coprocessors may be used to accelerate the machine learning algorithm(s) 620 and / or the trained machine learning model(s) 632. In some examples, the trained machine learning model(s) 632 may be trained to provide inferences on, reside, execute, and / or otherwise perform inferences for a particular computing device.
[0097] During the training phase 602, the machine learning algorithm(s) 620 may be trained by providing at least the training data 610 as a training input using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing a portion (or all) of the training data 610 to the machine learning algorithm(s) 620, where the machine learning algorithm(s) 620 determine one or more output inferences based on the provided portion (or all) of the training data 610. Supervised learning involves providing a portion of the training data 610 to the machine learning algorithm(s) 620, where the machine learning algorithm(s) 620 determine one or more output inferences based on the provided portion of the training data 610, where the output inference(s) are either accepted or modified based on the correct results associated with the training data 610. In some examples, supervised learning of the machine learning algorithm(s) 620 may be controlled by a set of rules and / or a set of labels for the training inputs, which may be used to modify the inferences of the machine learning algorithm(s) 620.
[0098] Semi-supervised learning involves having a correct outcome for some, but not all, of the training data 610. During semi-supervised learning, supervised learning is used for the portions of the training data 610 that have the correct outcome, and unsupervised learning is used for the portions of the training data 610 that do not have the correct outcome. Reinforcement learning includes machine learning algorithm(s) 620 receiving a reward signal for a prior inference, where the reward signal can be a numerical value. During reinforcement learning, the machine learning algorithm(s) 620 can output an inference and receive a reward signal in response thereto, where the machine learning algorithm(s) 620 are configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value that represents an expected sum of the numerical values provided by the reward signal over time. In some examples, the machine learning algorithm(s) 620 and / or the trained machine learning model(s) 632 can be trained using other machine learning techniques, including, but not limited to, incremental learning and curriculum learning.
[0099] In some examples, the machine learning algorithm(s) 620 and / or the trained machine learning model(s) 632 may use transfer learning techniques. For example, transfer learning techniques may include the trained machine learning model(s) 632 being pre-trained on a set of data and additionally trained using the training data 610. More specifically, the machine learning algorithm(s) 620 may be pre-trained on data from one or more computing devices, and the resulting trained machine learning model may be provided to a particular computing device that is intended to perform the trained machine learning during the inference stage 604. Then, during the training stage 602, the pre-trained machine learning model may be additionally trained using the training data 610, which may be derived from kernel data and non-kernel data of the particular computing device. This further training of the machine learning algorithm(s) 620 and / or the pre-trained machine learning model using the training data 610 of the data of the particular computing device may be performed using either supervised learning or unsupervised learning. Once the machine learning algorithm(s) 620 and / or the pre-trained machine learning model have been trained on at least the training data 610, the training phase 602 may be completed. The resulting trained machine learning model may be utilized as at least one of the trained machine learning model(s) 632.
[0100] Specifically, once the training phase 602 is completed, the trained machine learning model(s) 632 may be provided to a computing device if not already on the computing device. The inference phase 604 may begin after the trained machine learning model(s) 632 have been provided to a particular computing device.
[0101] During the inference stage 604, the trained machine learning model(s) 632 can receive the input data 630 and generate and output one or more corresponding inferences and / or prediction(s) 650 regarding the input data 630. Thus, the input data 630 can be used as input to the trained machine learning model(s) 632 to provide the inference(s) and / or prediction(s) 650 corresponding to the kernel and non-kernel components. For example, the trained machine learning model(s) 632 can generate the inference(s) and / or prediction(s) 650 in response to one or more inference / prediction requests 640. In some examples, the trained machine learning model(s) 632 can be executed by other pieces of software. For example, the trained machine learning model(s) 632 can be executed by an inference or prediction daemon and be readily available to provide inferences and / or predictions upon request. The input data 630 may include data from the particular computing device that runs the trained machine learning model(s) 632 and / or input data from one or more computing devices other than the particular computing device.
[0102] The input data 630 can include uncommon terms. Other types of input data are possible as well. The inference(s) and / or prediction(s) 650 can include one or more mistranscribed terms that represent phonetically similar options for pronouncing the uncommon term. In some embodiments, the inference(s) and / or prediction(s) 650 can include pairs of uncommon terms and mistranscribed terms, with associated confidence levels indicating the phonetic similarity of the mistranscribed and uncommon terms. The inference(s) and / or prediction(s) 650 can include other output data generated by the trained machine learning model(s) 632 operating on the input data 630 (and training data 610). In some examples, the trained machine learning model(s) 632 can use the output inference(s) and / or prediction(s) 650 as input feedback 660. The trained machine learning model(s) 632 can also rely on past inferences as input to generate new inferences.
[0103] The convolutional neural network 500 may be an example of machine learning algorithm(s) 620. After training, a trained version of the convolutional neural network 500 may be an example of trained machine learning model(s) 632. In this approach, an example of one or more inference / prediction requests 640 may be a request to predict one or more mistranscribed terms that correspond to an uncommon term, and a corresponding example of the inference and / or prediction(s) 650 may be the output one or more mistranscribed terms.
[0104] Data Network Example 7 illustrates a distributed computing architecture 700, according to an example embodiment. The distributed computing architecture 700 includes server devices 708, 710 configured to communicate with programmable devices 704a, 704b, 704c, 704d, 704e via a network 706. The network 706 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 706 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.
[0105] Although FIG. 7 shows only five programmable devices, the distributed application architecture can serve tens, hundreds, or thousands of programmable devices. Additionally, programmable devices 704a, 704b, 704c, 704d, 704e (or any additional programmable devices) may be any type of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mountable device (HMD), a network terminal, a mobile computing device, etc. In some implementations, such as shown by programmable devices 704a, 704b, 704c, 704e, the programmable devices can be directly connected to the network 706. In other implementations, such as shown by programmable device 704d, the programmable devices can be indirectly connected to the network 706 through an associated computing device, such as programmable device 704c. In this example, programmable device 704c can function as an associated computing device to pass electronic communications between programmable device 704d and the network 706. In other examples, as shown by programmable device 704e, the computing device may be part of and / or inside a vehicle, such as a car, truck, bus, boat or ship, airplane, etc. In other examples not shown in FIG 7, the programmable device may be connected both directly and indirectly to the network 706.
[0106] The server devices 708, 710 can be configured to perform one or more services as requested by the programmable devices 704a-704e. For example, the server devices 708 and / or 710 can provide content to the programmable devices 704a-704e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content may be encrypted and / or decrypted. Other types of content are possible as well.
[0107] As another example, server devices 708 and / or 710 may provide programmable devices 704a-704e with access to software for database, search, computation, graphics, audio, video, World Wide Web / Internet utilization, and / or other functionality. Many other examples of server devices are possible as well.
[0108] Computing Device Architecture Figure 8 is a block diagram of an exemplary computing device 800, according to an exemplary embodiment. In particular, the computing device 800 shown in Figure 8 may be configured to perform at least one function of and / or associated with the context-aware correction model, neural network, method 900, and / or method 1000 described herein.
[0109] Computing device 800 may include a user interface module 801, a network communication module 802, one or more processors 803, data storage 804, one or more camera(s) 818, one or more sensors 820, and a power system 822, all of which may be linked together via a system bus, network, or other connection mechanism 805.
[0110] The user interface module 801 may be operable to transmit data to and / or receive data from external user input / output devices. For example, the user interface module 801 may be configured to transmit data to and / or receive data from user input devices such as a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, and / or other similar devices. The user interface module 801 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, either now known or hereafter developed. The user interface module 801 may also be configured to generate an audible output using devices such as a speaker, a speaker jack, an audio output port, an audio output device, an earphone, and / or other similar devices. User interface module 801 may further be configured with one or more haptic devices capable of generating haptic output, such as vibration and / or other output detectable by touch and / or physical contact with computing device 800. In some examples, user interface module 801 may be used to provide a graphical user interface (GUI) for utilizing computing device 800.
[0111] The network communication module 802 may include one or more devices providing a wireless interface 807 and / or a wired interface 808 configurable to communicate over a network. The wireless interface(s) 807 may include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth transceiver, a Zigbee transceiver, a Wi-Fi transceiver, a WiMAX transceiver, an LTE transceiver, and / or other type of wireless transceiver configurable to communicate over a wireless network. The wired interface(s) 808 may include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate over a twisted pair wire, a coaxial cable, an optical fiber link, or a similar physical connection to a wired network.
[0112] In some examples, the network communications module 802 can be configured to provide reliable, secure, and / or authenticated communications. For each communication described herein, information to facilitate reliable communications (e.g., guaranteed message delivery) may be provided, possibly as part of the message header and / or footer (e.g., packet / message ordering information, encapsulation header and / or footer, size / time information, and transmission verification information such as cyclic redundancy check (CRC) and / or parity check values). Communications may be protected (e.g., encoded or encrypted) and / or decrypted / decoded using one or more cryptographic protocols and / or algorithms, including, but not limited to, Data Encryption Standard (DES), Advanced Encryption Standard (AES), Rivest-Shamir-Adelman (RSA) algorithm, Diffie-Hellman algorithm, secure socket protocols such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms may be used to protect (and decrypt / decrypt) communications similarly or in addition to those described herein.
[0113] The one or more processors 803 may include one or more general purpose processors and / or one or more special purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application specific integrated circuits, etc.). The one or more processors 803 may be configured to execute computer readable instructions 806 contained in data storage 804 and / or other instructions described herein.
[0114] Data storage 804 may include one or more non-transitory computer-readable storage media that may be read and / or accessed by at least one of the one or more processors 803. The one or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage, which may be integrated, in whole or in part, with at least one of the one or more processors 803. In some embodiments, data storage 804 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other embodiments, data storage 804 may be implemented using two or more physical devices.
[0115] The data storage 804 can include computer readable instructions 806 and possibly additional data. In some examples, the data storage 804 can include storage necessary to execute at least some of the methods, scenarios, and techniques described herein and / or at least some of the functionality of the devices and networks described herein. In some examples, the data storage 804 can include storage for trained neural network models 812 (e.g., models of trained convolutional neural networks, such as convolutional neural network 140). In particular of these examples, the computer readable instructions 806 can include instructions that, when executed by the one or more processors 803, enable the computing device 800 to provide some or all of the functionality of the trained neural network models 812.
[0116] In some examples, computing device 800 can include camera(s) 818. Camera(s) 818 can include one or more image capture devices, such as still and / or video cameras equipped to capture light and record the captured light into one or more images. That is, camera(s) 818 can generate image(s) of the captured light. The one or more images can be one or more still images and / or one or more images utilized in video capture. Camera(s) 818 can capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as one or more other frequencies of light.
[0117] In some examples, computing device 800 may include one or more sensors 820. Sensors 820 may be configured to measure conditions within computing device 800 and / or conditions within an environment of computing device 800 and provide data regarding these conditions. For example, sensors 820 may include (i) sensors for obtaining data regarding computing device 800, such as, but not limited to, a thermometer for measuring a temperature of computing device 100, a battery sensor for measuring the power of one or more batteries of power system 822, and / or other sensors that measure conditions of computing device 800; (ii) identification sensors, such as, but not limited to, radio frequency identification (RFID) readers, proximity sensors, one-dimensional barcode readers, two-dimensional barcode (e.g., quick response (QR) code) readers, and laser trackers, that may be configured to read identifiers and / or objects configured to be read and identify at least information; and (iii) tilt sensors, gyroscopes, accelerometers, Doppler sensors, GPS devices, sonar sensors. , radar devices, laser displacement sensors, and compasses, (iv) environmental sensors that obtain data indicative of the environment of the computing device 800, including, but not limited to, infrared sensors, optical sensors, light sensors, biosensors, capacitive sensors, touch sensors, temperature sensors, wireless sensors, radio sensors, movement sensors, microphones, sound sensors, ultrasonic sensors, and / or smoke sensors, (v) force sensors that measure one or more forces (e.g., inertial forces and / or gravitational acceleration) acting about the computing device 800, such as, but not limited to, one or more sensors that measure forces including, but not limited to, one or more dimensions, torque, ground force, friction, and / or zero moment point (ZMP), and / or a ZMP sensor that identifies the location of the ZMP. Many other examples of sensors 820 are possible as well.
[0118] The power system 822 may include one or more batteries 824 and / or one or more external power interfaces 826 for providing power to the computing device 800. Each battery of the one or more batteries 824 may function as a source of stored power for the computing device 800 when electrically coupled to the computing device 800. The one or more batteries 824 of the power system 822 may be configured to be portable. Some or all of the one or more batteries 824 may be easily removable from the computing device 800. In other examples, some or all of the one or more batteries 824 may be internal to the computing device 800 and therefore may not be easily removable from the computing device 800. Some or all of the one or more batteries 824 may be rechargeable. For example, a rechargeable battery may be recharged via a wired connection between the battery and another power source, such as by one or more power sources that are external to the computing device 800 and connected to the computing device 800 via one or more external power interfaces. In other examples, some or all of the one or more batteries 824 may be non-rechargeable batteries.
[0119] The one or more external power interfaces 826 of the power system 822 can include one or more wired power interfaces, such as a USB cable and / or a power cord, that enable a wired power connection to one or more power sources external to the computing device 800. The one or more external power interfaces 826 can include one or more wireless power interfaces, such as a Qi wireless charger, that enable a wireless power connection to the one or more external power sources, such as via a Qi wireless charger. Once a power connection to an external power source is established using the one or more external power interfaces 826, the computing device 800 can draw power from the external power source via the established power connection. In some examples, the power system 822 can include associated sensors, such as a battery sensor or other type of power sensor associated with one or more batteries.
[0120] Exemplary Methods of Operation 9 is a flowchart of a method 900 according to an example embodiment. Method 900 may be performed by a computing device, such as computing device 800. Method 900 may begin at block 910, where the method includes receiving, by an on-device system operating on the computing device, input from a user during interaction with the computing device.
[0121] At block 920, the method further includes receiving a transcription of the input from the input recognition model.
[0122] At block 930, the method also includes identifying, by the on-device system, candidate terms for replacement in the transcription of the input, the candidate terms being likely to have been transcribed incorrectly.
[0123] At block 940, the method additionally includes accessing, by the on-device system, a plurality of pairs of incorrectly transcribed terms and uncommon terms based on the candidate terms, where the uncommon terms are likely to be incorrectly transcribed, the incorrectly transcribed terms are incorrect versions of the uncommon terms, and the incorrectly transcribed terms have been generated by a machine learning model.
[0124] At block 950, the method further includes replacing, by the on-device system, the candidate term with the uncommon term in the transcription of the input based on a plurality of pairs of the incorrectly transcribed term and the uncommon term.
[0125] In some embodiments, replacing the candidate term includes comparing, by the on-device system, the candidate term to one or more incorrectly transcribed terms from the plurality of pairs. Such embodiments also include determining, based on whether the candidate term matches one of the one or more incorrectly transcribed terms, whether to replace the candidate term with a corresponding uncommon term that is paired with the matching incorrectly transcribed term.
[0126] In some embodiments, the method includes identifying, by the on-device system, an additional term to replace in the transcription of the input. Such embodiments also include determining that the additional term does not match the one or more incorrectly transcribed terms. Such embodiments further include retaining the additional term in the transcription of the input.
[0127] In some embodiments, replacing the candidate term includes determining that the candidate term matches a particular mistranscribed term of a particular pair of pairs, where the particular pair is associated with a particular confidence level indicating the similarity of the particular mistranscribed term to the particular uncommon term paired with the particular mistranscribed term. The confidence level may be determined as a numerical value (e.g., between 0 and 1, a scale of 1 to 10, a percentage value, etc.). In some embodiments, the confidence level may be determined as a quality rating such as "high", "medium", or "low". Additional intermediate values may be determined as "medium-high", "medium-low", "very high", "very low", etc.
[0128] Some embodiments also include determining whether a particular confidence level exceeds a threshold. In general, the threshold can vary from user to user, from application program to application program, etc. Also, for example, the threshold can vary based on the geographic location of the user and / or the language used. For example, a user may be in a country or region that speaks a particular language or dialect, and the user may adjust their speech to accommodate local customs.
[0129] In some embodiments, replacing the candidate term further comprises determining that a particular confidence level exceeds a threshold. Such embodiments also include replacing the candidate term with a particular uncommon term.
[0130] Some embodiments include storing the pairs in a local repository on the computing device. Such embodiments also include restricting access to the contents of the local repository within the computing device.
[0131] Some embodiments include training a machine learning model to generate mistranscribed terms for the uncommon term in the plurality of pairs.
[0132] In some embodiments, training the machine learning model further includes training the machine learning model to determine a respective confidence level for each of the plurality of pairs, where the given confidence level for a given pair comprising a given mistranscribed term and a given uncommon term indicates a similarity between the given mistranscribed term and the given uncommon term.
[0133] In some embodiments, training the machine learning model further includes receiving a corpus of web documents. Such embodiments also include synthetically simulating the one or more errors based on the corpus of web documents. In some embodiments, synthetically simulating the one or more errors is based on a text-to-speech model that utilizes a noisy channel simulator. In some embodiments, synthetically simulating the one or more errors is based on a grapheme-to-phoneme conversion model configured to generate a pronunciation of the word based on a text version of the word. In some embodiments, synthetically simulating the one or more errors is based on a statistical phoneme model.
[0134] Some embodiments include adjusting the trained machine learning model based on an application program of the computing device. In some embodiments, adjusting the trained machine learning model includes generating a plurality of pairs of incorrectly transcribed terms based on one or more errors associated with the application program.
[0135] Some embodiments include receiving, by the computing device, a second input from the user during a second interaction with the computing device, where a second transcription of the second input includes a candidate term. Such embodiments include comparing the candidate term to one or more second incorrectly transcribed terms, where the one or more second incorrectly transcribed terms are different from the one or more incorrectly transcribed terms. Such embodiments also include replacing, by the on-device system, the candidate term in the transcription of the input based on the comparison with a second uncommon term, where the second uncommon term is paired with one of the one or more second incorrectly transcribed terms.
[0136] Some embodiments include synthetically simulating the uncommon terms in the plurality of pairs. For example, the uncommon terms may be generated based on aggregated statistics of commonly recognizable or known uncommon terms. Also, for example, the uncommon terms may be generated based on aggregated statistics of commonly mistranscribed terms (e.g., based on speech-to-text transcription, autocorrection processes, etc.). In some embodiments, the synthetic simulation of the uncommon terms may be based on a trained machine learning model, for example, as the machine learning model described in the context of generating mistranscribed terms.
[0137] In some embodiments, the uncommon terms in the plurality of pairs have been observed by an on-device system in one or more past interactions between the user and the computing device.
[0138] In some embodiments, the one or more past interactions of the user with the computing device include an interaction with an application program of the computing device.
[0139] In some embodiments, the one or more past interactions between the user and the computing device include voice interactions, and the non-common terms in the plurality of pairs are based on user confirmation of transcribed terms based on the voice interactions.
[0140] In some embodiments, the one or more past interactions of the user with the computing device include interactions with a text editor, and the non-common terms in the plurality of pairs are based on user identification of text terms in the text editor.
[0141] In some embodiments, the computing device includes a viewer interface. One or more past interactions between the user and the computing device include text content provided by the viewer interface. The non-common term in a plurality of pairs appears in the text content.
[0142] In some embodiments, the one or more past interactions of the user with the computing device may be one of past interactions with a search application, a text messaging application, an email messaging application, an instant messaging application, a web browsing application, a media playback application, a telephone application, a video communication application, a gaming application, a mapping application, or a navigation application.
[0143] In some embodiments, the input from the user is a keyboard-based user input. In such embodiments, the input may be received as one or more of a typing operation, a swipe operation, or a gesture operation. In some embodiments, the plurality of mistranscribed terms in the plurality of pairs of mistranscribed terms and uncommon terms may be based on one or more contextual variations of the plurality of uncommon terms.
[0144] In some embodiments, the input from the user is a voice-based user input. In some embodiments, the input recognition model may be a cloud-based automatic input recognizer.
[0145] In some embodiments, interaction with a computing device includes interaction with an application program of the computing device.
[0146] In some embodiments, the interaction with the computing device may be one of an interaction with a search application, a text messaging application, an email messaging application, an instant messaging application, a web browsing application, a media playback application, a telephone application, a video communication application, a gaming application, a mapping application, or a navigation application.
[0147] In some embodiments, the interaction with the computing device may be an interaction with a digital assistant associated with the computing device.
[0148] 10 is a flowchart of a method 1000 according to an example embodiment. Method 1000 may be performed by a computing device, such as computing device 800. Method 1000 may begin at block 1010, which includes receiving, by an on-device system operating on the computing device, input from a user during interaction with the computing device.
[0149] At block 1020, the method further includes receiving a transcription of the input from the input recognition model.
[0150] At block 1030, the method also includes identifying candidate terms for replacement in the input transcription.
[0151] At block 1040, the method additionally includes determining whether the candidate term matches the incorrectly transcribed term. For example, the method includes identifying, by the on-device system, a candidate term for replacement in the transcription of the input. The candidate term is likely to have been incorrectly transcribed. The method further includes replacing, by the on-device system, the candidate term with the uncommon term in the transcription of the input based on a plurality of pairs of the incorrectly transcribed term and the uncommon term.
[0152] In some embodiments, replacing the candidate term includes comparing, by the on-device system, the candidate term to one or more incorrectly transcribed terms from the plurality of pairs. Such embodiments also include determining, based on whether the candidate term matches one of the one or more incorrectly transcribed terms, whether to replace the candidate term with a corresponding uncommon term that is paired with the matching incorrectly transcribed term.
[0153] If it is determined that the candidate term does not match the incorrectly transcribed term, the method further includes, at block 1050, not replacing the candidate term in the transcription of the input.
[0154] If the candidate term is determined to match the mistranscribed term, the method further includes, at block 1060, determining whether a confidence level of the pair including the matched mistranscribed term and its respective paired uncommon term exceeds a threshold.
[0155] If it is determined that the confidence level of the pair including the matching incorrectly transcribed term and its respective paired uncommon term does not exceed the threshold, the method further includes, at block 1050, not replacing the candidate term in the transcription of the input.
[0156] If the confidence level of the pair including the matching incorrectly transcribed term and its respective paired uncommon term is determined to exceed a threshold, the method further includes, at block 1070, replacing the candidate term in the transcription of the input with its respective paired uncommon term.
[0157] The present disclosure should not be limited with respect to the specific embodiments described in this application, which are intended to illustrate various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from the spirit and scope thereof. In addition to those enumerated herein, functionally equivalent methods and devices within the scope of the present disclosure will be apparent to those skilled in the art from the above description. Such modifications and variations are intended to be included within the scope of the appended claims.
[0158] In the above detailed description, various features and functions of the disclosed systems, devices, and methods are described with reference to the accompanying drawings. In the figures, like symbols generally identify like components unless otherwise dictated by context. The exemplary embodiments described in the detailed description, figures, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein.
[0159] With respect to any or all of the ladder diagrams, scenarios, and flow charts in the figures and as discussed herein, each step, block, and / or communication may represent a processing of information and / or a transmission of information according to an exemplary embodiment. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, functions described as blocks, transmissions, communications, requests, responses, and / or messages may be performed out of the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functionality involved. Furthermore, more or fewer blocks and / or functions may be used in any of the ladder diagrams, scenarios, and flow charts discussed herein, and these ladder diagrams, scenarios, and flow charts may be combined with each other, either in part or in whole.
[0160] The blocks representing the processing of information may correspond to circuitry that may be configured to perform certain logical functions of the methods or techniques described herein. Alternatively or additionally, the blocks representing the processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to perform certain logical functions or operations in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device, including a disk, or hard drive, or other storage medium.
[0161] The computer readable medium may also include non-transitory computer readable media, such as register memory, processor cache, and non-transitory computer readable media that store data for a short period of time, such as random access memory (RAM). The computer readable medium may also include non-transitory computer readable media that store program code and / or data for a longer period of time, such as secondary or permanent long-term storage, such as read only memory (ROM), optical or magnetic disks, compact disk read only memory (CD-ROM). The computer readable medium may also be any other volatile or non-volatile storage system. The computer readable medium may be considered to be, for example, a computer readable storage medium, or a tangible storage device.
[0162] Additionally, blocks representing one or more information transfers may correspond to information transfers between software and / or hardware modules in the same physical device, although other information transfers may be between software and / or hardware modules in different physical devices.
[0163] For embodiments that include determining uncommon terms based on user interactions with a computing device and / or determining incorrectly transcribed terms using machine learning models or interactions between the computing device and a cloud-based server, controls may be provided to the user that allow the user to make choices about both whether and when the systems, programs, or features described herein enable collection of user information (e.g., information about the user's social network, social actions, or activities, occupation, user preferences, user demographic information, user's current location, or other personal information), as well as whether content or communications are sent from the server to the user. Additionally, certain data may be processed in one or more ways such that personally identifiable information is removed before being stored or used. For example, the user's identity may be processed such that personally identifiable information of the user cannot be determined, or if location information is obtained (such as to the city, zip code, or state level), the user's geographic location may be generalized such that the user's specific location cannot be determined. Thus, the user may control what information is collected about the user, how that information is used, and what information is provided to the user.
[0164] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those of ordinary skill in the art. The various aspects and embodiments disclosed herein are provided by way of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.
Claims
1. 1. A computer-implemented method comprising: receiving, by an on-device system operating on a computing device, input from a user during interaction with said computing device; receiving a transcription of the input from an input recognition model; identifying, by the on-device system, candidate terms for replacement in the transcription of the input, the candidate terms having a high likelihood of being mistranscribed, the method further comprising: accessing, by the on-device system, a plurality of pairs of incorrectly transcribed terms and uncommon terms based on the candidate terms, the uncommon terms having a high probability of being incorrectly transcribed, the incorrectly transcribed terms being incorrect versions of the uncommon terms, the incorrectly transcribed terms being generated by a machine learning model, the method further comprising: replacing, by the on-device system, the candidate terms with uncommon terms in the transcription of the input based on the plurality of pairs of the incorrectly transcribed terms and the uncommon terms; A computer-implemented method comprising:
2. The replacing of the candidate terms comprises: comparing, by the on-device system, the candidate term to one or more incorrectly transcribed terms from the plurality of pairs; determining whether to replace the candidate term with a non-generic term paired with the matching mistranscribed term based on whether the candidate term matches a mistranscribed term of the one or more mistranscribed terms; The computer-implemented method of claim 1 further comprising:
3. identifying, by the on-device system, additional terms to replace in the transcription of the input; determining that the additional terms do not match the one or more incorrectly transcribed terms; maintaining the additional terms in the transcription of the input; and The computer-implemented method of claim 2 further comprising:
4. The replacing of the candidate terms comprises: determining that the candidate term matches a particular mistranscribed term of a particular pair of the plurality of pairs, the particular pair being associated with a particular confidence level indicating a similarity of the particular mistranscribed term to a particular uncommon term paired with the particular mistranscribed term, the method comprising: determining whether the particular confidence level exceeds a threshold; The computer-implemented method of claim 2 further comprising:
5. The replacing of the candidate terms comprises: determining that the particular confidence level exceeds the threshold; replacing the candidate terms with the particular non-generic terms; The computer-implemented method of claim 4 further comprising:
6. storing said plurality of pairs in a local repository of said computing device; The computer-implemented method of claim 1 further comprising:
7. restricting access to content of the local repository within the computing device; The computer implemented method of claim 6 further comprising:
8. training the machine learning model to generate the mistranscribed terms for the uncommon terms in the plurality of pairs; The computer-implemented method of claim 1 further comprising:
9. Training the machine learning model includes:
10. The computer-implemented method of claim 8, further comprising training the machine learning model to determine a respective confidence level for each pair of the plurality of pairs, wherein a given confidence level for a given pair comprising a given mistranscribed term and a given uncommon term indicates a similarity of the given mistranscribed term to the given uncommon term.
10. Training the machine learning model includes: receiving a corpus of documents; synthetically simulating one or more errors based on the corpus of documents; The computer implemented method of claim 8 further comprising:
11. 11. The computer-implemented method of claim 10, wherein the synthetically simulating the one or more errors is based on a text-to-speech model that utilizes a noisy channel simulator.
12. 11. The computer-implemented method of claim 10, wherein the synthetically simulating the one or more errors is based on a grapheme-to-phoneme conversion model configured to generate a pronunciation of a word based on a text version of the word.
13. 11. The computer-implemented method of claim 10, wherein the synthetically simulating the one or more errors is based on a statistical phoneme model.
14. adjusting the trained machine learning model based on an application program of the computing device; The computer implemented method of claim 8 further comprising:
15. adjusting the trained machine learning model includes generating the mistranscribed terms in the plurality of pairs based on one or more errors associated with the application program; 15. The computer implemented method of claim 14, comprising:
16. and receiving, by the computing device, a second input from the user during a second interaction with the computing device, wherein a second transcription of the second input includes the candidate terms, and the method further includes: comparing the candidate terms to one or more second mistranscribed terms, the one or more second mistranscribed terms being different from the one or more mistranscribed terms, the method further comprising:
2. The computer-implemented method of claim 1, further comprising: replacing, by the on-device system, the candidate term in the transcription of the input with a second uncommon term based on the comparison of the candidate term to the one or more second incorrectly transcribed terms, wherein the second uncommon term is paired with a second incorrectly transcribed term of one of the one or more second incorrectly transcribed terms.
17. The computer-implemented method of claim 1 , further comprising synthetically simulating the uncommon terms in the plurality of pairs.
18. The computer-implemented method of claim 1 , wherein the uncommon terms in the plurality of pairs are observed by the on-device system in one or more past interactions between the user and the computing device.
19. 20. The computer-implemented method of claim 18, wherein the one or more past interactions of the user with the computing device include interactions with an application program of the computing device.
20. the one or more past interactions between the user and the computing device include a voice interaction; 20. The computer-implemented method of claim 18, wherein the non-common terms in the plurality of pairs are based on user confirmation of transcribed terms based on the voice interaction.
21. the one or more past interactions of the user with the computing device include an interaction with a text editor; 20. The computer-implemented method of claim 18, wherein the non-common terms in the plurality of pairs are based on user identification of text terms in the text editor.
22. the computing device includes a viewer interface; the one or more past interactions of the user with the computing device include textual content provided by the viewer interface; 20. The computer-implemented method of claim 18, wherein the uncommon terms in the plurality of pairs occur in the textual content.
23. 1. A computing device comprising: one or more processors; and data storage having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions, the functions comprising: receiving, by an on-device system operating on a computing device, input from a user during interaction with said computing device; receiving a transcription of the input from an input recognition model; and identifying, by the on-device system, candidate terms for replacement in the transcription of the input, the candidate terms being likely to be incorrectly transcribed, the function further comprising: and accessing, by the on-device system, a plurality of pairs of incorrectly transcribed terms and uncommon terms based on the candidate terms, the uncommon terms being likely to be incorrectly transcribed, the incorrectly transcribed terms being incorrect versions of the uncommon terms, the incorrectly transcribed terms being generated by a machine learning model, the function further comprising: replacing, by the on-device system, the candidate terms with uncommon terms in the transcription of the input based on the plurality of pairs of the incorrectly transcribed terms and the uncommon terms; a computing device comprising:
24. 1. An article of manufacture comprising one or more computer readable media having stored thereon computer readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions, the functions comprising: receiving, by an on-device system operating on a computing device, input from a user during interaction with said computing device; receiving a transcription of the input from an input recognition model; and identifying, by the on-device system, candidate terms for replacement in the transcription of the input, the candidate terms being likely to be incorrectly transcribed, the function further comprising: accessing, by the on-device system, a plurality of pairs of incorrectly transcribed terms and uncommon terms based on the candidate terms, the uncommon terms having a high probability of being incorrectly transcribed, the incorrectly transcribed terms being incorrect versions of the uncommon terms, the incorrectly transcribed terms being generated by a machine learning model, the method further comprising: replacing, by the on-device system, the candidate terms with uncommon terms in the transcription of the input based on the plurality of pairs of the incorrectly transcribed terms and the uncommon terms; Including the product.
Citation Information
Patent Citations
Speech recognition result output device, speech recognition result output method and speech recognition result output program
JP2012063545A
Voice interactive method, voice interactive device and voice interactive program
JP2018004976A
Misconversion dictionary creation system
JP2020184024A
Server device, communication system and information processing method
JP2021071658A
Communication system, communication method, and program
JP2021193421A