Assisted text correction and editing during voice typing
The system addresses the challenges of error correction and editing in voice typing by using natural language processing and machine learning to propose and implement changes, enhancing efficiency and usability.
Patent Information
- Application Number
- PCT/US2023/083661
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Existing voice typing systems are prone to errors, such as incorrect words, omitted words, incorrect capitalization, and incorrect punctuation, making error correction and editing cumbersome and time-consuming, especially when users' hands are occupied.
A method and system that utilize natural language processing and machine learning to receive voice input, generate structured text, identify changes to be made, and propose corrections or edits, allowing users to approve or implement these changes without manual typing.
Enables efficient and easy correction and editing of voice typing errors, allowing users to quickly produce accurate and grammatically correct text using vocal inputs, even when hands are not available.
Smart Images

Figure US2023083661_19062025_PF_FP_ABST
Abstract
Description
ASSISTED TEXT CORRECTION AND EDITING DURING VOICE TYPINGBACKGROUND
[0001] Voice typing, also called dictation, is a way to enter text in a computer system by talking into a microphone of the computing system. An automatic speech recognition (ASR) subsystem of the computing system can transcribe speech, or audio input data, into text. Despite continued efforts for improvement, ASR subsystems are still prone to errors, such as generating text with incorrect words, omitted words, incorrect capitalization, or incorrect punctuation, or having an inability to transcribe out-of-vocabulary (OOV) words. Correcting ASR errors during voice typing can be difficult, as users may be required to switch to a manual mode of operation for correction and / or to re-dictate their desired content. As a result, the correction process may be tedious and time-consuming, or in some cases, even impossible (e.g., when a user's hands are occupied, such as when driving or jogging). Additionally, aside from error correction, users may also wish to make edits after their speech is dictated. For example, users may want to change their wording or add more detailed information (e.g., date, time, or numerical values). Thus, such editing processes can be inefficient and cumbersome.SUMMARY
[0002] In general, the techniques of this disclosure are directed to methods and systems to receive natural language input data from a user, automatically generate structured text data from the natural language input data, identify one or more changes to be made to text input (entered either from natural language speech or other mechanisms for text entry such as a keyboard, keypad, touch screen, etc.), and generate one or more proposed changes to the text input. The one or more changes to be made to the text input may be identified by applying a predefined set of rules or templates to the text input (e.g., grammar rules may be applied to the text input and / or the text input may be mapped to a specific sentence structure). In some examples, a machine learning model, such as a large language model (LLM), may be applied to the structured text data, in which the LLM may parse the structured text data to identify one or more explicit and / or implicit commands indicating the one or more changes to be made to the text input (e.g., the structured text data may be representative of a natural language user input such as “Change ‘Tuesday’ to ‘Thursday’” or “Fix it”). The LLM may further generate the one or more proposed changes based on the one or more identified changes, and in some examples, the proposed changes may be output to the user (e.g.,displayed on a screen of a computing system or device) for approval. Responsive to receiving user input (e.g., approval from a user), the computing system may implement at least one of the one or more proposed changes to the text input. In some examples, the proposed changes may be implemented without requiring user input or approval. In some examples, the one or more proposed changes to the text input may be implemented to correct one or more errors in the text input. In some examples, the one or more proposed changes may be implemented to change the text input.
[0003] In one example, the disclosure is directed to a method comprising: receiving, by a computing system, and from at least one of an automatic speech recognizer and an input device providing for entry of typed characters, text input; receiving, by the computing system, natural language audio input data; automatically generating, by the computing system, structured text data based at least in part on the natural language audio input data; identifying, by a machine learning model executing on the computing system, based at least in part on the structured text data, one or more changes to be made to the text input; generating, by the machine learning model, based at least in part on the one or more changes to be made to the text input, one or more proposed changes to the text input; and outputting, by the computing system, the one or more proposed changes to the text input.
[0004] In another example, the disclosure is directed to a computing system including: a memory that stores instructions; and processor circuitry that executes the instructions to receive, from at least one of an automatic speech recognizer and an input device providing for entry of typed characters, text input; receive natural language audio input data; automatically generate structured text data based at least in part on the natural language audio input data; identify, using a machine learning model and based at least in part on the structured text data, one or more changes to be made to the text input; generate, using the machine learning model and based at least in part on the one or more changes to be made to the text input, one or more proposed changes to the text input; and output the one or more proposed changes to the text input.
[0005] In some aspects, the techniques described herein relate to a non-transitory computer-readable storage medium including instructions, that when executed by processor circuitry of a computing system, cause the processor circuitry to receive, from at least one of an automatic speech recognizer and an input device providing for entry of typed characters, text input; receive natural language audio input data; automatically generate structured text data based at least in part on the natural language audio input data; identify, using a machine learning model and based at least in part on the structured text data, one or more changes tobe made to the text input; generate, using the machine learning model and based at least in part on the one or more changes to be made to the text input, one or more proposed changes to the text input; and output the one or more proposed changes to the text input.
[0006] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 illustrates an example computing system for assisted text correction and editing, in accordance with one or more aspects of the present disclosure.
[0008] FIG. 2 illustrates further details of an example computing system for assisted text correction and editing, in accordance with one or more aspects of the present disclosure.
[0009] FIG. 3 illustrates an example display of voice typing and an explicit keyword initiating an edit mode, in accordance with one or more aspects of the present disclosure! .
[0010] FIG. 4 illustrates an example display of implementing one or more proposed changes generated from the edit mode, in accordance with one or more aspects of the present disclosure.
[0011] FIG. 5 illustrates an example display of voice typing and an implicit keyword initiating an edit mode, in accordance with one or more aspects of the present disclosure.
[0012] FIG. 6 illustrates another example display of implementing one or more proposed changes generated from the edit mode, in accordance with one or more aspects of the present disclosure.
[0013] FIG. 7 illustrates an example process for assisted text correction and editing, in accordance with one or more aspects of this disclosure.DETAILED DESCRIPTION
[0014] FIG. 1 illustrates an example computing system for artificial intelligence-assisted text correction and editing, in accordance with one or more aspects of the present disclosure. As shown in FIG. 1, computing system 100 may represent any type of computing system capable of executing speech recognition and machine learning models. Examples of computing system 100 may include a server, a disaggregated server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a portable or mobile device (e.g., a cell phone, a smart phone, other cellular handset, a tablet computer, an Internet appliance, a smart television (TV), a digital versatile disk (DVD) player, a compact disc (CD)player, a digital video recorder (DVR), a Blu-ray player, a video gaming console, a personal video recorder, a set top box, a headset (e.g., an extended reality (XR) headset (including an augmented reality (AR) headset, a virtual reality (VR) headset, a mixed reality (MR) headset, etc.)), a smart watch, or other wearable device, or any other type of computing device. In some examples, computing system 100 may represent a cloud computing system that provides one or more services via a network. That is, in some examples, computing system 100 may be a distributed computing system. In these examples, some or all of the components and / or functionality attributed to computing system 100 may be implemented or performed by a computing device, such as a remote computing device. In these examples, computing system 100 may communicate with a device via any public or private network, such as a cellular network, Wi-Fi network, a direct cell-to-satellite communication network, the Internet, or other type of network for transmitting data between computing system 100 and a computing device.
[0015] As shown in the example of FIG. 1, computing system 100 includes input components 142 and output components 146. Input components 142 may be configured to function as input devices for computing system 100 and output components 146 may be configured to function as output devices for computing system 100. Input components 142 and output components 146 may be implemented using various technologies. For instance, input components 142 may be configured to receive natural language speech 102 and / or text input 112 through tactile, audio, and / or video feedback. Examples of input components 142 include a presence-sensitive display, a presence-sensitive or touch-sensitive input device, a mouse, a keyboard, a voice responsive system, video camera, microphone or any other type of device for detecting or receiving natural language speech 102 and / or text input 112. In some examples, a presence-sensitive display includes a touch-sensitive or presence-sensitive input screen, such as a resistive touchscreen, a surface acoustic wave touchscreen, a capacitive touchscreen, a projective capacitive touchscreen, a pressure sensitive screen, an acoustic pulse recognition touch screen, or another presence-sensitive technology. That is, input components 142 of computing system 100 may include a presence-sensitive device that may receive tactile input. Input components 142 may receive indications of the tactile input by detecting one or more gestures (e.g., from a user when the user touches or points to one or more locations of input components 142 with a finger or a stylus pen).
[0016] Output components 146 may be configured to function as one or more output devices by providing output using tactile, audio, or video stimuli. Examples of output devices include a sound card, a video graphics adapter card, or any of one or more display devices,such as a liquid crystal display (LCD), dot matrix display, light emitting diode (LED) display, microLED, miniLED, organic light-emitting diode (OLED) display, e-ink, or similar monochrome or color display capable of outputting visible information. Additional examples of an output device include a speaker, a haptic device, or other device that can generate intelligible output. For instance, output components 146 may present output as text that includes one or more proposed changes 114 to text input 112. In this way, input components 142 and output components 146 may present various user interfaces of applications executing at or accessible by computing system 100 (e.g., an electronic message application, an Internet browser application, etc.). In some examples, a user may interact with a respective user interface of an application, such as a messaging application, to cause computing system 100 to perform operations related to text correction and editing.
[0017] In the example of FIG. 1, computing system 100 also includes voice typing system 103, which further includes audio input data 104, automatic speech recognizer (ASR) 106, structured text data 108, machine learning module 110, and proposed changes 114. Voice typing system 103 may perform operations described herein using hardware, software, firmware, or a mixture thereof residing in and / or executing at computing system 100. Computing system 100 may execute voice typing system 103 with one processor or with multiple processors. In some examples, computing system 100 may execute voice typing system 103 as a virtual machine executing on underlying hardware. Voice typing system 103 may execute as one or more services of an operating system or computing platform or may execute as one or more executable programs at an application layer of a computing platform.
[0018] As described herein, computing system 100 may include voice typing system 103 to provide voice typing capability to a user, such that a user can edit or correct text by providing natural language speech input. Specifically, voice typing system 103 may be configured to edit and / or correct text input 112 based on natural language speech 102. In some examples, the one or more changes to be made to text input 112 may include one or more corrections to one or more errors in text input 112. Error corrections and changes implemented by machine learning module 110 may include, for example, correcting ASR errors (based on commands such as insert, delete, replace, etc.), editing lower and upper case letters (based on commands such as upper, lower, title, etc.), post hoc word editing (based on commands such as insert, delete, replace, etc.), editing punctuation (based on commands such as insert, delete, replace, etc.), editing numbers, currencies, dates, or times, editing whitespace (based on commands such as spaces, line breaks, paragraphs, etc.), spelling out-of-vocabulary (OOV) words (including alternative spellings), and editing abbreviations (based on commands such as expand and contract, etc.).
[0019] As shown in the example of FIG. 1, input components 142 may receive natural language speech 102 and text input 112. Natural language speech 102 may comprise spoken language, such as a command from a user. For example, natural language speech 102 may be a spoken command such as “Fix it." Input components 142 may receive natural language speech 102 via, for example, a microphone or any other input device described previously. Input components 142 may also receive text input 112, which may include previously entered text and / or current text that a user intends to correct and / or edit. In some examples, text input 112 may be captured by a mode of input such as typing on a keypad or a touch sensitive screen, read from a file, etc. In some examples, text input 112 may be a typed sentence in a messaging application text entry field. For example, text input 112 may be a typed sentence such as “Let’s go on wensday.” In this example, text input 112 includes a capitalization error and a spelling error. A user may provide natural language speech 102 such as “Fix it" to indicate their desire to correct these errors, and computing system 100 may then output one or more proposed changes 114 that may include one or more edits or corrections to text input 112, such correcting “wensday" to “Wednesday.” As such, a user may not be required to perform an action of physical typing on a keyboard to edit or correct text input 112, and instead may simply provide a spoken command of “Fix it.” In this way, users may edit and / or correct text more efficiently and easily.
[0020] In some examples, text input 112 may be generated by ASR 106. Thus, computing system 100 may receive text input 112 from at least one of an input device providing for entry of typed characters (e.g., an exterior or remote device or one or more input components 142) or ASR 106. As such, a user may provide natural language speech 102, e.g., a sentence such as “Let’s go on Wednesday", which may then be converted to audio input data 104 and fed into ASR 106. In this example, ASR 106 may generate text for output such as “Let’s go on wensday.” As shown in this example, text input 112 as generated by ASR 106 may have errors (including cascading errors, where a first error causes one or more subsequent errors). A user may wish to correct these errors, but manual corrections using existing technology are often cumbersome, difficult, and inefficient. Some existing ASR subsystems recognize specific keywords spoken in a predefined syntax to indicate an intended error correction or a desired change to the text input, but these systems are often difficult to use because the user must remember the specific error correction, change command keywords, or predefined syntax and cannot perform voice typing using free form natural language speech (e.g., speechwithout predefined syntax and keywords). As such, a user may provide another free form natural language speech input such as “Fix it,” and computing system 100 may correct and / or edit the output generated by ASR 106. In this way, computing system 100 may operate iteratively, wherein the output of ASR 106 (e.g., text input 112) may be provided as input to computing system 100, and a user may provide multiple spoken commands (e.g., natural language speech 102) indicating their desired correction and / or edit until they are satisfied with a final text output.
[0021] Input components 142 may generate audio input data 104 from natural language speech 102. For example, input components 142 may convert natural language speech 102 into a digital audio signal or other form of data that can be read by ASR 106.
[0022] ASR 106 may then analyze audio input data 104 to automatically generate structured text data 108. In some examples, any suitable ASR may be used in computing system 100, including ASRs using one or more of Hidden Markov Models (HMMs), feed forward artificial neural networks, deep learning, recurrent neural networks, long short-term memory (LTSM) systems, connectionist temporal classification (CTC), transformers, attention-based models, etc.
[0023] Text input 112 may be sent from input components 142 to machine learning module 110. Machine learning module 110 may then apply one or more machine learning models to text input 112 to identify, based at least in part on structured text data 108, one or more changes to be made to text input 112. As such, machine learning module 110 may apply a machine learning model, such as a large language model, to text input 112 such as “Let’s go on wensday,” and based on structured text data 108 that is representative of natural language speech 102 such as “Fix it,” identify, for example, a spelling change to be made and a capitalization change to be made. As described above, natural language speech 102 may be in a variable and unpredictable command syntax (e.g., free form speech), which may include ambiguity, assumed linguistic knowledge, imperfect linguistic knowledge (e.g., user errors in commands), implied context awareness, repetition of words or phrases, and implicit commands and / or arguments. Machine learning module 110 may parse structured text data 108 and identify one or more explicit and / or implicit keywords in structured text data 108 that indicates commands and / or intentions to correct errors or otherwise change text input 112. As such, machine learning module 110 may be configured to separate “content” from “command” or determine whether a given utterance is meant as content or a command. For example, in a natural language utterance such as "See you tomorrow at 9 I mean 10", machine learning module 110 may segment the natural language utterance into a firstsegment "See you tomorrow at 9" that is determined as content, and a second segment "I mean 10" that is determined as a command. Machine learning module 110 may then apply the command to the content, or generate, based at least in part on the one or more changes to be made to text input 112, one or more proposed changes 114. Continuing the previous example, one or more proposed changes 114 may include a change in spelling (e.g., “wensday” to “Wednesday”) and a change in capitalization (e.g., "w” to “W"). Proposed changes 114 may be sent to output components 146, in which output components 146 may then output proposed changes 114. In some examples, proposed changes 114 may be sent to ASR 106 as feedback data that can help improve the accuracy of ASR 106. In some examples, proposed changes 114 may be sent to an input device that provided text input 112 (e.g., an exterior or remote device or one or more input components 142). In some examples, responsive to receiving user input, computing system 100 may implement at least one of one or more proposed changes 114 to text input 112. For example, a user may be presented with proposed changes 114 such as changing “wensday” to “Wednesday” and capitalizing "w” to “W”. Responsive to receiving user input that indicates a user’s approval of the changes, computing system 100 may implement proposed changes 114, such that a final text output reads “Let’s go on Wednesday.”
[0024] In this way, the techniques disclosed herein may provide users the capability to vocally correct a variety of errors and / or make specific changes to text generated by ASR systems (e.g., voice typing) or even text manually typed by a user. Furthermore, the techniques described herein may not require a user to provide an input indicating the exact change they wish to make. As such, a user may provide either an explicit or an implicit command, and given text input 112, computing system 100 may identify the one or more changes to be made to text input 112. Thus, the techniques disclosed herein may result in an improved user experience for voice typing in a computing system, as users may easily correct errors and / or make changes to any text generated from voice typing or manual typing.
[0025] FIG. 2 illustrates further details of an example computing system for artificial intelligence-assisted text correction and editing, in accordance with one or more aspects of the present disclosure. Computing system 200 of FIG. 2 may be another example of computing system 100 as illustrated in FIG. 1. FIG. 2 illustrates only one example of computing system 200, and many other examples of computing system 200 may be used in other instances and may include a subset of the components included in example computing system 200 or may include additional components not shown in example computing system 200. As shown in the example of FIG. 2, computing system 200 includes user interfacecomponents (UIC) 232, one or more processors 240 (each including processing circuitry), one or more input components) 242 (which may be similar if not substantially similar to input components 142 of FIG. 1), one or more communication unit(s) 244, one or more output components) 246 (which may be similar if not substantially similar to output components 146 of FIG. 1), and one or more storage devices 248. One or more storage devices 248 of computing system 200 also include voice typing system 203, audio input data 204, automatic speech recognizer (ASR) 206, structured text data 208, machine learning module 210, proposed changes 214, and keyword list 216. Voice typing system 203, audio input data 204, ASR 206, structured text data 208, machine learning module 210, and proposed changes 214 of FIG. 2 may be an example of voice typing system 103, audio input data 104, automatic speech recognizer 106, structured text data 108, machine learning module 110, and proposed changes 114 of FIG. 1 , respectively. As described above, some or all of the components and / or functionality attributed to computing system 200 may be implemented or performed by a computing device in communication with computing system 200.
[0026] The one or more communication units 244 of computing system 200, for example, may communicate with external devices by transmitting and / or receiving data at computing system 200, such as to and from remote computer systems or computing devices. Example communication units 244 include a network interface card (e.g., such as an Ethernet card), an optical transceiver, a radio frequency transceiver, or any other type of device that can send and / or receive information. Other examples of communication units 244 may be devices configured to transmit and receive Ultrawideband®, Bluetooth®, GPS, 3G, 4G, and Wi-Fi®, etc. that may be found in computing devices, such as mobile devices and the like.
[0027] As shown in the example of FIG. 2, communication channels 250 may interconnect each of the components as shown for inter-component communications (physically, communicatively, and / or operatively). In some examples, communication channels 250 may include a system bus, a network connection (e.g., to a wireless connection as described above), one or more inter-process communication data structures, or any other components for communicating data between hardware and / or software locally or remotely.
[0028] One or more input components 242 may receive inputs and one or more output components 246 may generate outputs. As described with respect to FIG. 1, examples of inputs are tactile, audio, kinetic, and optical input, to name only a few examples. Input components 242, in one example, may include a touchscreen, a touchpad, a mouse, a keyboard, a voice responsive system, a video camera, buttons, a control pad, a microphone orany other type of device for detecting input from a human or machine. Output components 246, in one example, may include a sound card, a video graphics adapter card, a speaker, a display, or any other type of device for generating output to a human or machine.
[0029] UIC 232 of computing system 200 may be hardware that functions as an input and / or output system for computing system 200. For example, UIC 232 may include a display component, which may be a screen at which information is displayed by UIC 232 and a presence-sensitive input component that may detect an object at and / or near the display component.
[0030] Voice typing system 203, including ASR 206, audio input data 204, structured text data 208, machine learning module 210, proposed changes 214, and keyword list 216 (hereinafter “modules 203-216”) may perform operations described herein using software, hardware, firmware, or a mixture of hardware, software, and firmware residing in and executing on computing system 200 or at one or more other computing devices (e.g., a cloudbased application - not shown). For example, some or all of modules 203-216 may be included in and executable on a local computing device. As such, the techniques described herein may all be implemented locally on a computing device.
[0031] Computing system 200 may execute one or more of modules 203-216, with one or more processors 240 or may execute any or part of one or more of modules 203-216 as or within a virtual machine executing on underlying hardware. One or more of modules 203-216 may be implemented in various ways, for example, as a downloadable or pre-installed application, remotely as a cloud application, or as part of the operating system of computing system 200. Other examples of computing system 200 that implement techniques of this disclosure may include additional components not shown in FIG. 2.
[0032] In the example of FIG. 2, one or more processors 240 may implement functionality and / or execute instructions within computing system 200. One or more processors 240 may include processor circuitry to execute instructions to perform processing. As used herein, processor circuitry is defined to include: (i) one or more special purpose electrical circuits structured to perform specific operation(s) and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors); and / or (ii) one or more general purpose semiconductor-based electrical circuits programmed with instructions to perform specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuitry include programmed microprocessors, Field Programmable Gate Arrays (FPGAs) that may instantiate instructions, Central ProcessorUnits (CPUs), Graphics Processor Units (GPUs), Digital Signal Processors (DSPs), XPUs, or microcontrollers and integrated circuits such as Application Specific Integrated Circuits (ASICs). For example, an XPU may be implemented by a heterogeneous computing system including multiple types of processor circuitry (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc., and / or a combination thereof) and application programming interface(s) (API(s)) that may assign computing task(s) to whichever one(s) of the multiple types of the processing circuitry is / are best suited to execute the computing task(s). For example, one or more processors 240 may receive and execute instructions that provide the functionality of UIC 232, communication units 244, one or more storage devices 248 and an operating system to perform one or more operations as described herein. For example, one or more processors 240 may receive and execute instructions that provide the functionality of some or all of modules 203-216 to perform one or more operations and various functions described herein. As described, the one or more processors 240 may include a central processing unit (CPU). Examples of CPUs include, but are not limited to, a digital signal processor (DSP), a general-purpose microprocessor, a tensor processing unit (TPU); a neural processing unit (NPU); a neural processing engine; a core of a CPU, VPU, GPU, TPU, NPU or another processing device, an application specific integrated circuit (ASIC), a field programmable logic array (FPGA), or other equivalent integrated or discrete logic circuitry, or other equivalent integrated or discrete logic circuitry.
[0033] One or more storage devices 248 within computing system 200 may store information, such as information retrieved from a user computing device, or other data discussed herein, for processing during the operation of computing system 200. In some examples, one or more storage devices 248 are temporary memories, meaning that a primary purpose of one or more storage devices 248 is not long-term storage. One or more storage devices 248 of computing system 200 may be configured for short-term storage of information as volatile memory and therefore may not retain stored contents if powered off. Examples of volatile memories include random access memories (RAM), dynamic randomaccess memories (DRAM), static random-access memories (SRAM), and other forms of volatile memories known in the art. Storage devices 248, in some examples, may also include one or more computer-readable storage media. Storage devices 248 may be configured to store larger amounts of information for longer terms in non-volatile memory than volatile memory. One or more storage devices 248 may further be configured for long-term storage of information as non-volatile memory space and retain information after power on / off cycles. Examples of non-volatile memories include magnetic hard disks, optical discs, floppy discs,flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. Storage devices 248 may store program instructions and / or data associated with the modules 203-216 of FIG. 2.
[0034] As described herein, computing system 200 may receive natural language speech as natural language audio input data 204, and automatically generate structured text data 208 based at least in part on natural language audio input data 204. Computing system 200, and specifically voice typing system 203, may generate structured text data 208 by applying ASR 206 to natural language audio input data 204. In some examples, ASR 206 may preprocess audio input data 204 to enhance quality and remove noise by normalizing the audio volume and filtering out any background noise. ASR 206 may then transform audio input data 204 into a more suitable format and extract features such as Mel-frequency cepstral coefficients (MFCCs), which capture information about the frequency content of the audio signal over short time intervals. ASR 206 may further perform acoustic modeling (e.g., with Hidden Markov Models (HMMs)), which may involve training a statistical model that maps the extracted audio features to phonemes. The acoustic model may learn to associate specific audio features with phonemes while taking into account the variations in pronunciation, accents, and speaking styles. ASR 206 may further implement language modeling (e.g., deep learning techniques, such as recurrent neural networks (RNNs) and transformers) to capture and predict a sequence of words or phrases while considering the context in which the words are spoken. For example, ASR 206 may use context information received by input components 242, such as, if present, the current location of a text cursor, any text that has been highlighted / selected, and / or the locations where a screen or display is tapped prior to, during, or after the natural language speech utterance. ASR 206 may further use the trained acoustic and language models to decode audio input data 204 and generate a transcription or sequence of words that best match the observed audio features. ASR 206 may further implement post-processing techniques (e.g., grammar checks, contextual analysis, spell correction, etc.) to refine the transcription and improve readability and accuracy. ASR 206 may then output the transcribed text, which may be structured text data 208, that represents audio input data 204 and can be fed into machine learning module 210 for further processing and analysis.
[0035] As described above, in some examples, the transcription output from ASR 206 may include one or more errors. Specifically, structured text data 208 may include one or more errors, and identifying the one or more changes to be made to the text input may further include inferring, by machine learning module 210 and based at least in part on structuredtext data 208 including the one or more errors, the one or more changes to be made to the text input. As such, in these examples, machine learning module 210 may still infer the correct command or intent. For example, if a user utters a command such as "Change Tuesday to Wednesday", and ASR 206 transcribes the command as "Chain Tuesday to Wednesday", machine learning module 210 (which, as described below, may include trained large language models) may still infer that the user’s actual editing intent is a word replacement.
[0036] In some examples, a user may type a command such as “Fix it” or “Change Monday to Wednesday” rather than provide natural language speech. As such, in some examples, a natural language command may be manually typed instead of uttered by voice. In some examples, computing system 200 may send, to a user computing device, instructions for generating a user interface that provides for entry of a typed input or command. For example, computing system 200 may send instructions for generating a user interface similar to graphical user interface (GUI) 324 shown in FIG. 3, wherein the GUI may be a GUI of a messaging application that may provide a user the capability to send short message service (SMS) messages and / or provide typed inputs or commands to computing system 200.
[0037] In some examples, a machine learning model included in ASR 206 may generate structured text data 208. For example, ASR 206 may convert audio input data 204 or any other input / information to an extensible Markup Language (XML), or into other structured text types, such as, but not limited to, HTML, JSON, CSV, INI Files, etc. In this way, the information and input received by voice typing system 203 can be provided to ML module 210 in a standardized format. Machine learning module 210 may further determine the type of information to include in the structured text representation. More specifically, machine learning module 210 may analyze additional information included in text input 112 of FIG. 1, such as to more accurately interpret the output of ASR 206 or a user’s command / intent.
[0038] Machine learning module 210 may further identify, based at least in part on structured text data 208, one or more changes to be made to the text input. Machine learning module 210 may then generate, based at least in part on the one or more changes to be made to the text input, one or more proposed changes 214 to the text input. As such, machine learning module 210 may assist a user in correcting or editing text that has either been manually typed by the user or outputted by ASR 206.
[0039] In some implementations, as discussed above, any information or input received by voice typing system 203 may be preprocessed. Preprocessing techniques may include extracting one or more additional features from raw data. For example, feature extractiontechniques may be applied to the user input or text input to generate one or more new, additional features.
[0040] As described herein, machine learning module 210 may employ a large language model (LLM) that can interpret structured text data 208 and identify one or more changes to be made to the text input. In some examples, machine learning module 210 may implement other machine-learned models that may be used in place of or in conjunction with LLM model and are described in more detail below. Machine learning module 210 may perform various types of natural language processing (NLP) based on structured text data 208 and text input. Structured text data 208 and text input may be referred to herein as “input data”. For example, machine learning module 210 may summarize, translate, or organize the input data. Machine learning module 210 may use recurrent neural networks (RNNs) and / or transformer models (self-attention models), such as GPT-3, BERT, and T5. In some implementations, machine learning module 210 may perform classification, summarization, name generation, regression, clustering, anomaly detection, recommendation generation, and / or other tasks.
[0041] In some implementations, machine learning module 210 may perform various types of classification based on the input data. For example, machine learning module 210 may perform binary classification or multiclass classification. In binary classification, the output data may include a classification of the input data into one of two different classes. In multiclass classification, the output data may include a classification of the input data into one (or more) of more than two classes. The classifications may be single-label or multilabel. Machine learning module 210 may perform discrete categorical classification in which the input data is simply classified into one or more classes or categories.
[0042] In cases in which machine learning module 210 performs classification, machine learning module 210 may be trained using supervised learning techniques. For example, machine learning module 210 may be trained on a training dataset that includes training examples labeled as belonging (or not belonging) to one or more classes.
[0043] In some implementations, machine learning module 210 may perform regression to provide output data in the form of a continuous numeric value. The continuous numeric value may correspond to any number of different metrics or numeric representations, including, for example, currency values, scores, or other numeric representations. In examples, machine learning module 210 may perform linear regression, polynomial regression, nonlinear regression, or nonparametric regression. In examples, machine learning module 210 may perform simple regression or multiple regression. As described above, in some implementations, a Softmax function or other function or layer may be used to squash aset of real values respectively associated with two or more possible classes to a set of real values in the range (0, 1) that sum to one.
[0044] Machine learning module 210 may perform various types of clustering. For example, machine learning module 210 may identify one or more clusters to which the input data most likely corresponds. Machine learning module 210 may identify one or more clusters within the input data. That is, in instances in which the input data includes multiple objects, documents, or other entities, machine learning module 210 may sort the multiple entities included in the input data into a number of clusters. In some implementations in which machine learning module 210 performs clustering, machine learning module 210 may be trained using unsupervised learning techniques.
[0045] Machine learning module 210 may, in some cases, act as an agent within an environment. For example, machine learning module 210 may be trained using reinforcement learning, which will be discussed in further detail below.
[0046] In some implementations, machine learning module 210 may include a parametric model while, in other implementations, machine learning module 210 may include a non- parametric model. In some implementations, machine learning module 210 may include a linear model while, in other implementations, machine learning module 210 may include a non-linear model.
[0047] As described above, machine learning module 210 may be or include one or more of various different types of machine-learned models. Examples of such different types of machine-learned models are provided below for illustration. One or more of the example models described below may be used (e.g., combined) to provide the output data in response to the input data. Additional models beyond the example models provided below may be used as well.
[0048] In some implementations, machine learning module 210 may be or include one or more classifier models such as, for example, linear classification models; quadratic classification models; etc. Machine learning module 210 may be or include one or more regression models such as, for example, simple linear regression models; multiple linear regression models; logistic regression models; stepwise regression models; multivariate adaptive regression splines; locally estimated scatterplot smoothing models; etc.
[0049] In some implementations, machine learning module 210 may be or include one or more artificial neural networks (also referred to simply as neural networks). A neural network may include a group of connected nodes, which also may be referred to as neurons or perceptrons. A neural network may be organized into one or more layers. Neural networksthat include multiple layers may be referred to as “deep” networks. A deep network may include an input layer, an output layer, and one or more hidden layers positioned between the input layer and the output layer. The nodes of the neural network may be connected or nonfolly connected.
[0050] In some examples, machine learning module 210 may be or include one or more generative networks such as, for example, generative adversarial networks. Generative networks may be used to generate new data such as artificial feedback texts.
[0051] In an example in which the input data does not include feature embeddings, one or more neural networks may be used to provide an embedding based on the input data. For example, the embedding may be a representation of knowledge abstracted from the input data into one or more learned dimensions. In some instances, embeddings may be a usefol source for identifying related entities. In some instances, embeddings may be extracted from the output of the network, while in other instances embeddings may be extracted from any hidden node or layer of the network (e.g., a close to final but not final layer of the network). Embeddings may be usefol for performing auto-suggest next video, product suggestion, entity or object recognition, etc. In some instances, embeddings are useful inputs for downstream models. For example, embeddings may be useful to generalize input data (e.g., search queries) for a downstream model or processing system.
[0052] In some implementations, machine learning module 210 may perform or be subjected to one or more reinforcement learning techniques such as Markov decision processes; dynamic programming; Q functions or Q-leaming; value function approaches; deep Q-networks; differentiable neural computers; asynchronous advantage actor-critics; deterministic policy gradient; etc.
[0053] In some implementations, machine learning module 210 may be an autoregressive model. In some instances, an autoregressive model may specify that the output data depends linearly on its own previous values and on a stochastic term. In some instances, an autoregressive model may take the form of a stochastic difference equation. One example of an autoregressive model is WaveNet, which is a generative model for raw audio.
[0054] In some implementations, machine learning module 210 may include or form part of a multiple model ensemble. As one example, bootstrap aggregating may be performed, which may also be referred to as “bagging.” In bootstrap aggregating, a training dataset is split into a number of subsets (e.g., through random sampling with replacement) and a plurality of models are respectively trained on the number of subsets. At inference time,respective outputs of the plurality of models may be combined (e.g., through averaging, voting, or other techniques) and used as the output of the ensemble.
[0055] One example ensemble is a random forest, which may also be referred to as a random decision forest. Random forests are an ensemble learning method for classification, regression, and other tasks. Random forests are generated by producing a plurality of decision trees at training time. In some instances, at inference time, the class that is the mode of the classes (classification) or the mean prediction (regression) of the individual trees may be used as the output of the forest. Random decision forests may correct for decision trees' tendency to overfit their training set.
[0056] Another example ensemble technique is stacking, which can, in some instances, be referred to as stacked generalization. Stacking includes training a combiner model to blend or otherwise combine the predictions of several other machine-learned models. Thus, a plurality of machine-learned models (e.g., of the same or different type) may be trained based on training data. In addition, a combiner model may be trained to take the predictions from the other machine-learned models as inputs and, in response, produce a final inference or prediction. In some instances, a single-layer logistic regression model may be used as the combiner model.
[0057] Another example of ensemble techniques is boosting. Boosting may include incrementally building an ensemble by iteratively training weak models and then adding to a final strong model. For example, in some instances, each new model may be trained to emphasize the training examples that previous models misinterpreted (e.g., misclassified). For example, a weight associated with each of such misinterpreted examples may be increased. One common implementation of boosting is AdaBoost, which may also be referred to as Adaptive Boosting. Other example boosting techniques include LPBoost; TotalBoost; BrownBoost; xgboost; MadaBoost, LogitBoost, gradient boosting; etc. Furthermore, any of the models described above (e.g., regression models and artificial neural networks) may be combined to form an ensemble. As an example, an ensemble may include a top-level machine-learned model or a heuristic function to combine and / or weight the outputs of the models that form the ensemble.
[0058] In some implementations, multiple machine-learned models (e.g., that form an ensemble may be linked and trained jointly (e.g., through backpropagation of errors sequentially through the model ensemble). However, in some implementations, only a subset (e.g., one) of the jointly trained models is used for inference.
[0059] In some implementations, machine learning module 210 may be used to preprocess the input data for subsequent input into another model. For example, machine learning module 210 may perform dimensionality reduction techniques and embeddings (e.g., matrix factorization, principal components analysis, singular value decomposition, word2vec / GLOVE, and / or related approaches); clustering; and even classification and regression for downstream consumption. Many of these techniques have been discussed above and will be further discussed below.
[0060] In some implementations, during training, the input data may be intentionally deformed in any number of ways to increase model robustness, generalization, or other qualities.
[0061] In response to receipt of the input data, machine learning module 210 may provide the output data. As examples, in various implementations, the output data may include content, either stored locally on the user device or in the cloud, that is relevantly shareable along with the initial content selection.
[0062] In some implementations, the output data may influence downstream processes or decision-making. As one example, in some implementations, the output data, or the second set of instructions, may be interpreted and / or acted upon by a rules-based regulator.
[0063] Machine learning module 210 described herein may be trained according to one or more of various different training types or techniques. For example, in some implementations, machine learning module 210 may be trained using supervised learning, in which machine learning module 210 is trained on a training dataset that includes instances or examples that have labels. The labels may be manually applied by experts, generated through crowdsourcing, or provided by other techniques (e.g., by physics-based or complex mathematical models). In some implementations, if the user has provided consent, the training examples may be provided by the user computing device. In some implementations, this process may be referred to as personalizing the model.
[0064] In some implementations, backward propagation of errors may be used in conjunction with an optimization technique (e.g., gradient-based techniques) to train machine learning module 210 (e.g., when the machine-learned model is a multi-layer model such as an artificial neural network). For example, an iterative cycle of propagation and model parameter (e.g., weights) update may be performed to train machine learning module 210. Example backpropagation techniques include truncated backpropagation through time, Levenberg- Marquardt backpropagation, etc.
[0065] In some implementations, machine learning module 210 described herein may be trained using unsupervised learning techniques. Unsupervised learning may include inferring a function to describe hidden structure from unlabeled data. For example, a classification or categorization may not be included in the data. Unsupervised learning techniques may be used to produce machine-learned models capable of performing clustering, anomaly detection, learning latent variable models, or other tasks.
[0066] Machine learning module 210 may be trained using semi-supervised techniques which combine aspects of supervised learning and unsupervised learning. Machine learning module 210 may be trained or otherwise generated through evolutionary techniques or genetic algorithms. In some implementations, machine learning module 210 described herein may be trained using reinforcement learning. In reinforcement learning, an agent (e.g., model) may take actions in an environment and learn to maximize rewards and / or minimize penalties that result from such actions. Reinforcement learning may differ from the supervised learning problem in that correct input / output pairs are not presented, nor sub- optimal actions explicitly corrected.
[0067] In some implementations, one or more generalization techniques may be performed during training to improve the generalization of machine learning module 210. Generalization techniques may help reduce overfitting of machine learning module 210 to the training data. Example generalization techniques include dropout techniques; weight decay techniques; batch normalization; early stopping; subset selection; stepwise selection; label smoothing; etc.
[0068] In some implementations, machine learning module 210 described herein may include or otherwise be impacted by a number of hyperparameters, such as, for example, learning rate, number of layers, number of nodes in each layer, number of leaves in a tree, number of clusters; etc. Hyperparameters may affect model performance. Hyperparameters may be hand selected or may be automatically selected through the application of techniques such as, for example, grid search; black-box optimization techniques (e.g., Bayesian optimization, random search, etc.); gradient-based optimization; etc. Example techniques and / or tools for performing automatic hyperparameter optimization include Hyperopt; Auto- WEKA; Spearmint; Metric Optimization Engine (MOE); etc.
[0069] In some implementations, various techniques may be used to optimize and / or adapt the learning rate when the model is trained. Example techniques and / or tools for performing learning rate optimization or adaptation include Adagrad; Adaptive Moment Estimation (ADAM); Adadelta; RMSprop; etc.
[0070] In some implementations, transfer learning techniques may be used to provide an initial model from which to begin training of machine learning module 210 described herein.
[0071] In some implementations, machine learning module 210 described herein may be included in different portions of computer-readable code on a computing device. In one example, machine learning module 210 may be included in a particular application or program and used (e.g., exclusively) by such particular application or program. Thus, in one example, a computing device may include a number of applications, and one or more of such applications may contain its own respective machine learning library and machine-learned model(s).
[0072] In another example, machine learning module 210 described herein may be included in an operating system of a computing device (e.g., in a central intelligence layer of an operating system) and may be called or otherwise used by one or more applications that interact with the operating system. In some implementations, each application may communicate with the central intelligence layer (and model(s) stored therein) using an application programming interface (API) (e.g., a common, public API across all applications).
[0073] In some implementations, the central intelligence layer may communicate with a central device data layer. The central device data layer may be a centralized repository of data for the computing device. The central device data layer may communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer may communicate with each device component using an API (e.g., a private API).
[0074] The technology discussed herein refers to servers, databases, software applications, and other computer-based systems, as well as actions taken, and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein may be implemented using a single device or component or multiple devices or components working in combination.
[0075] Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate sequentially or in parallel.
[0076] In addition, the machine learning techniques described herein are readily interchangeable and combinable. Although certain example techniques have been described, many others exist and may be used in conjunction with aspects of the present disclosure.
[0077] In some implementations, transfer learning (TL) may be used. Transfer learning involves reusing a model and its model parameters obtained while solving one problem and applying it to a different but related problem. Models trained on very large data sets may be retrained or fine-tuned on additional data. Often, all model designs and their parameters on a source model are copied except output layer(s). The output layers(s) are often called the head, and other layers are often called the base. The source parameters may be considered to contain the knowledge learned from the source dataset and this knowledge may also be applicable to a target dataset. Fine-tuning may include updating the head parameters with the body parameters being fixed or updated in a later step.
[0078] In some examples, computing system 200 may implement a rule-based system (either in conjunction with machine learning module 210 or separately) that may operate on predefined rules and logical conditions. For example, computing system 200 may implement a rule-based system that generates text output in which the first letter of every sentence is capitalized. In some examples, computing system 200 may implement rule-based systems such as expert systems, decision trees, or simple conditional statements in which output is generated by if-then rules. In some examples, prior to identifying the one or more changes to be made to the text input, computing system 200 may apply at least one rule from a predefined set of rules and at least one template from a predefined set of templates to the text input to generate one or more possible changes. Machine learning module 210 may then identify, based at least in part on structured text data 208 and the one or more possible changes, the one or more changes to be made to the text input.
[0079] As described, machine learning module 210 may apply a large language model, in combination with or not in combination with one or more of the machine learning techniques or rule-based systems described above, to the input data to identify one or more changes to be made to the text input and / or generate one or more proposed changes 214 to the text input. Furthermore, as described above, machine learning module 210 may apply an ensemble learning method. In some examples, machine learning module 210 can be or include one or more of a statistical language model and a large language model. In some examples, machine learning module 210 may apply a statistical language model that is trained on large datasets containing correct language usage and can probabilistically generate proposed changes 214. In some examples, ML module 210 may further assign a score to each of one or moreproposed changes 214, wherein the score of a proposed change is based at least in part on a sum of probabilities of one or more words of the change appearing in the text input. ML module 210 may then select the change with the highest score to be included in one or more proposed changes 214.
[0080] In some examples, machine learning module 210 may implement, for example, transformer-based neural networks. Transformer-based neural networks such as large language models may refer to a type of deep learning architecture specifically designed for handling sequential data, such as text or time series. As such, transformer-based neural networks like LLMs may be configured to perform natural language processing (NLP) tasks, such as question-answering, machine translation, text summarization, and sentiment analysis. Machine learning module 210 may be configured to perform tasks such as classification, sentiment analysis, entity extraction, extractive question answering, summarization, rewriting text in a different style, ad copy generation, and concept ideation.
[0081] Transformer-based neural networks may utilize a self-attention mechanism, which allows the model to weigh the importance of different elements in a given input sequence relative to each other. The self-attention mechanism may help machine learning module 210 effectively capture long-range dependencies and complex relationships between elements, such as words in a sentence.
[0082] Machine learning module 210 may include an encoder and a decoder that operate to process and generate sequential data, such as structured text data 208. Both the encoder and decoder may include one or more of self-attention mechanisms, position-wise feedforward networks, layer normalization, or residual connections. In some examples, the encoder may process an input sequence and create a representation that captures the relationships and context among the elements in the sequence. The decoder may then obtain the representation generated by the encoder and produce an output sequence. In some examples, the decoder may generate the output one element at a time (e.g., one word at a time), using a process called autoregressive decoding, where the previously generated elements are used as input to predict the next element in the sequence.
[0083] If the natural language speech or audio input data 204 is unclear or inaccurate, machine learning module 210 may be unable to determine a user’s intent with high confidence. In such instances, computing system 200 may prompt the user to clarify their input. In some examples, the prompt may include a list of options, a question, etc. For example, if a user provides an input such as “Change the day,” computing system 200 may generate a graphical user interface or text output for display that includes a clarifyingquestion such as “Which day?” In some examples, computing system 200 may output clarifying questions, suggestions, or proposed changes 214 as an audio output. As such, one or more of output components 246 or output components included in a remote computing device may present clarifying questions, suggestions, or proposed changes 214 to a user by implementing a text-to-speech technique, in which the output is read aloud to the user. In this way, a user may communicate with computing system 200 vocally, or in a conversation-like manner, such that the user may not be required to read or physically select any output generated by computing system 200. In some examples, after the user clarifies their intent or selects one or more of proposed changes 214, computing system 200 may then update any text input or output generated by computing system 200.
[0084] In some examples, an action log executed on computing system 200 and / or a remote computing device may provide a user a ledger of activity, which may show any automations or applications running in the background of computing system 200 and / or the remote computing device, as well as an accurate log of all voice typing, editing, or correction activity.
[0085] In some examples, machine learning module 210 may determine a set of information types included in the input (e.g., text or audio input data 204 or a transcription generated by ASR 206). An information type may be or otherwise include a topic, theme, point, subject, purpose, intent, keyword, etc. In some examples, machine learning module 210 may determine the information type by leveraging a self-attention mechanism to capture the relationships and dependencies between words in the input sequence. For example, machine learning module 210 may tokenize (e.g., split) a sequence of words or subwords, which machine learning module 210 may convert into vectors (e.g., numerical representations) that machine learning module 210 can process. Machine learning module 210 may use the self-attention mechanism to weigh the importance of each token in relation to the others. In this way, machine learning module 210 may identify patterns and relationships between the tokens, and in turn the words corresponding to the tokens, that indicate one or more information types.
[0086] As such, machine learning module 210 may identify one or more changes to be made to the text input based at least in part on detecting one or more keywords in structured text data 208. In some examples, the one or more keywords may include at least one word from keyword list 216, which may be a predetermined list of words, to initiate an edit mode for machine learning module 210 to generate the one or more proposed changes 214 to the text input. For example, responsive to detecting the word “Fix” in structured text data 208,machine learning module 210 may be activated or enter into an edit mode in proposed changes 214 are then generated. In some examples, the one or more keywords from keyword list 216 may indicate one or more implicit changes to be made to the text input. In these examples, machine learning module 210 may identify the one or more implicit changes to be made to the text input, and then generate, based at least in part on the one or more implicit changes to be made to the text input, the one or more proposed changes 214 to the text input.
[0087] In some examples, the one or more keywords from keyword list 216 may indicate one or more explicit changes to be made to the text input. In these examples, the one or more explicit changes may comprise at least one of an insertion of a word to the text input, a deletion of a word from the text input, a replacement of a word in the text input, an insertion of a newline character to the text input, a deletion of a newline character to the text input, an insertion of a whitespace character to the text input, a deletion of a whitespace character to the text input, an insertion of punctuation to the text input, a deletion of punctuation from the text input, a replacement of punctuation in the text input, a changing spelling of a word in the text input, a changing of a format of a number in the text input, and a changing of a text format of one or characters in the text input.
[0088] In some examples, the one or more keywords from keyword list 216 may indicate inserting a link into the text input. As such, in some examples, explicit keywords recognized by machine learning module 210 may indicate an intent by the user to include a link to a document, file, audio data, video data, web page, or other data accessible over a network or with computing system 200. For example, if structured text data 208 includes the explicit keyword(s) of “link”, “embed”, “web page”, “make a link”, “add link”, “insert link”, “linkify” or other suitable word intending a linking action, machine learning module 210 may be trained to interpret or infer the one or more words immediately before or after the explicit keyword “link” (e.g., the context) as a specification of a link to other data (such as a document, file, audio data, video data, web page, or other data accessible over a network or within a file). In these examples, responsive to detecting one or more keywords comprising a word from keyword list 216 to indicate inserting a link into the text input, computing system 200 may generate a plurality of links associated with one or more words in the text input. As such, the link, as determined by machine learning module 210, may then be inserted into proposed changes 214 in an appropriate form for the link. For example, if text input includes the phrase “Let’s go to Hawaii this fall, can you check for airline tickets?”, and structured text data 208 is generated from a natural language speech command such as “Link tickets”, machine learning module 210 may parse structured text data 208, interpret the explicitkeyword “link” as the user intending a linking action, search for a plurality of links to one or more airline websites (as machine learning module 210 may further determine context from the text input that indicates the user is requesting airline tickets), and embed a link from the plurality of links as an active “clickable” link in proposed changes 214 in place of the word “tickets”. As such, computing system 200 may be configured to receive an utterance such as “Link” or “Linkify” and determine a user’s intent for finding and embedding a link in the text input, wherein the link may be anchored by a specific word in the text input, such as the word "tickets".
[0089] In an implementation, the link may be in the form of a hyperlink (e.g., hypertext transport protocol (HTTP)) or a uniform resource locator (URL). The link may be to a site on the world wide web (WWW), the Internet (via file transfer protocol (ftp), for example), an intranet, an accessible personal or organizational storage device, etc. Computing system 200 may select a link from the plurality of links that is most associated with the one or more words in the text input, and insert the link from the plurality of links into the text input, or one or more “anchors”, as described above. As such, computing system 200 may further determine which word or words included in the text input should be the anchors for the embedded link. As described in the example above, computing system 200 may determine the anchor word to be “tickets”. In another example, computing system 200 may determine the anchor words to be “airline tickets”. In some examples, text input may include the word “link”, such as “Let’s go to Hawaii this fall, can you check this [airline name] link for tickets?”. In this example, computing system 200 may embed a link to a specific airline’s website in the words “[airline name] link” or “link”. As such, computing system 200 may not embed the link in the entirety of the text input.
[0090] In some examples, proposed changes 214 may include the link. In some examples, computing system 200 may send, to an input device, the link from the plurality of links, and receive, from the input device, an authorization to insert the link from the plurality of links into the text input. As such, the link may be output to a display of computing system 200 for the user to accept or reject the link or select a thumbnail image to represent the linked data. Computing system 200 may then insert or embed the link from the plurality of links into the text input responsive to receiving the authorization. In some examples, computing system 200 may choose a link or the plurality of links based on other data such as price, historical user activity (given the user has provided computing system 200 consent to collect or analyze such data), or any other information that may aid in ranking the plurality of links in terms of relevance.
[0091] In general, machine learning module 210 may excel at performing NLP tasks, such as generating text and other content. However, with respect to specific types of content (e.g., specific information types), machine learning module 210 may have an increased likelihood of generating false, inaccurate, or bad quality information. To address this issue, machine learning module 210 may be configured to exclude the generation of content or code relating to a set of excluded information types. For example, the set of excluded information types may include one or more of phone numbers, addresses, web addresses, etc. Thus, input information may be passed in machine learning module 210 with certain prerequisites, prompts, or “rules.” Machine learning module 210 may apply these prerequisites, prompts, or rules when generating proposed changes 214. As such, machine learning module 210 may store a plurality of text inputs or other data that further specify how proposed changes 214 should be generated by machine learning module 210. Because machine learning module 210 can interpret the rules along with the input, computing system 200 can provide more accurate instructions for generating the user’s desired edits or corrections. As such, as described, computing system 200 may be able to interpret natural language to understand user intents, and then write or generate new, robust, and / or correct text that satisfies the user’s intents.
[0092] Although primarily described herein as being a transformer-based neural network, machine learning module 210 may be or otherwise include one or more other types of neural networks. For example, machine learning module 210 may be or include an autoencoder. In some examples, the aim of an autoencoder is to learn a representation (e.g., a lowerdimensional encoding) for a set of data, typically for the purpose of dimensionality reduction. For example, in some examples, an autoencoder can seek to encode the input data and then provide output data that reconstructs the input data from the encoding. Recently, the autoencoder concept has become more widely used for learning generative models of data. In some examples, the autoencoder can include additional losses beyond reconstructing the input data. Machine learning module 210 may be or include one or more other forms of artificial neural networks such as, for example, deep Boltzmann machines, deep belief networks, stacked autoencoders, etc. Any of the neural networks described herein can be combined (e.g., stacked) to form more complex networks.
[0093] In some examples, machine learning module 210 can be or include one or more feed forward neural networks. In feed forward networks, the connections between nodes do not form a cycle. For example, each connection can connect a node from an earlier layer to a node from a later layer. In some examples, machine learning module 210 can be or includeone or more recurrent neural networks. In some examples, at least some of the nodes of a recurrent neural network can form a cycle.
[0094] Recurrent neural networks can be especially useful for processing input data that is sequential in nature. For example, a recurrent neural network can pass or retain information from a previous portion of the input data sequence to a subsequent portion of the input data sequence through the use of recunent or directed cyclical node connections. Sequential input data may include words in a sentence (e.g., for natural language processing, speech detection or processing, etc.). In some examples, sequential input data can include time-series data (e.g., sensor data versus time or imagery captured at different times). In some examples, sequential input data may include time-series data (e.g., sensor data versus time or imagery captured at different times). For example, a recurrent neural network may analyze sensor data versus time to detect or predict a swipe direction, to perform handwriting recognition, etc. Sequential input data may include words in a sentence (e.g., for natural language processing, speech detection or processing, etc.); notes in a musical composition; sequential actions taken by a user (e.g., to detect or predict sequential application usage); sequential object states; etc.
[0095] Example recurrent neural networks may include long short-term (LSTM) recurrent neural networks, gated recurrent units, bi-direction recurrent neural networks, continuous time recurrent neural networks, neural history compressors, echo state networks, Elman networks, Jordan networks, recursive neural networks, Hopfield networks, fully recurrent networks, sequence-to- sequence configurations, etc.
[0096] In some examples, machine learning module 210 can be or include one or more convolutional neural networks. In some examples, a convolutional neural network can include one or more convolutional layers that perform convolutions over input data using learned filters. Filters can also be referred to as kernels. Convolutional neural networks can be especially useful for vision problems such as when the input data includes imagery such as still images or video. However, convolutional neural networks can also be applied for natural language processing.
[0097] Machine learning module 210 may additionally train (e.g., pre-train, fine-tune, etc.) one or more machine learning models. For example, machine learning module 210 may train one or more models on a large and diverse corpus of text. This dataset may cover a wide range of topics and domains to ensure machine learning module 210 learns diverse linguistic patterns and contextual relationships. Machine learning module 210 may train models to optimize an objective function. The objective function may be or include a loss function,such as cross-entropy loss, that compares (e.g., determines a difference between) output data generated by the model from the training data and labels (e.g., ground-truth labels) associated with the training data. For example, the objective function of machine learning module 210 may be to correctly predict the next word in a sequence of words or correctly fill in missing words as much as possible.
[0098] In some examples, machine learning module 210 may continuously or periodically train machine learning models. In some examples, machine learning module 210 may fine-tune machine learning models by using feedback in the training process. In some examples, machine learning module 210 may be fine-tuned using a pre-trained Language Model for Dialog Applications (LaMDA) model. Input data to the fine-tuning may include user-generated speech-to-text data, user corrections and / or changes to the user-generated speech-to-text data, and rule-based synthetic examples generated by code or another machine learning model. Fine tuning may result in machine learning module 210 being able to learn and correct at least ASR error patterns. In some examples, prompt-tuning may be used, in which prompts or input data used to interact with a pre-trained language model, such as a Low-Rank Adaptation (LoRA) model or similar model may be refined and optimized. In these examples, the output of machine learning module 210 may be improved such that it is more specific or contextually relevant to a given query or task. In these examples, machine learning module 210 may implement the LoRA model, which may freeze pre-trained model weights and optimize rank decomposition matrices of layers of the Transformer architecture, such that the number of trainable parameters for downstream tasks is greatly reduced.
[0099] Machine learning module 210, as tuned, may also be evaluated to assess the quality of model performance in making corrections and / or changes to structured text data 208 based on free form natural language speech. In some examples, input components 242 of FIG. 2 may receive a user input via a computing device that selects feedback (e.g., thumbs up, thumbs down, etc.) relating to proposed changes 214 or any other output provided by computing system 200. In some examples, the feedback may indicate whether the output is accurate or inaccurate, correct or incorrect, high quality or low quality, etc. Input components 242 may receive this feedback and may send it to voice typing system 203. Voice typing system 203 may transmit the feedback to machine learning module 210, or in some examples, ASR 206, in which machine learning module 210 and / or ASR 206 uses the feedback for training. For example, machine learning module 210 and / or ASR 206 may convert the feedback into labeled data for supervised training. Additionally or alternatively, machine learning module 210 and / or ASR 206 may fine-tune models by monitoring the relationshipbetween the performance of machine learning module 210 and / or ASR 206 and user feedback, and iterate the fine-tuning process as necessary (e.g., to receive more positive user feedback and less negative user feedback). In this way, the techniques of this disclosure may establish a feedback loop that continuously improves the quality of the output (i.e., proposed changes 214) of machine learning module 210 and / or text generated by ASR 206.
[0100] Generally, large language models can be slow and expensive in terms of carbon, energy usage, and financial cost. Thus, in some examples, machine learning module 210 may minimize how often a large language model is invoked by caching proposed changes 214 or any other information / data relevant for generating proposed changes 214. Specifically, machine learning module 210 may be configured to perform instruction embedding in which a representation (i.e., embedding) of frequently used or critical proposed changes are stored in a cache.
[0101] In various examples, proposed changes 214 may be generated based on any data / information stored in a cache and / or any additional input, information, or updates received by computing system 200 that are not present in a cache and instead may be stored in a local memory. Machine learning module 210 may query the local memory to gather this additional input, information, or updates and use them with the cached information at runtime to generate proposed changes 214. For example, if a user provides an input such as “change the date to tomorrow,” machine learning module 210 may generate proposed changes 214 that, for example, changes “9 / 22” to “9 / 23”. If the user provides the same “change the date to tomorrow” command in the future, machine learning module 210 may use contextual information, such as data indicating the date, to generate proposed changes 214 that, for example, changes “10 / 1” to “10 / 2”. By storing frequently used or critical changes in a cache, machine learning module 210 may reuse the frequently used or critical changes without having to invoke a large language model on data other than what is included in audio input data 204 and / or the text input (e.g., machine learning module 210 may not have to request clarifying information and / or additional input to which the large language model may be applied). In some examples, machine learning module 210 may apply code caching to both compiled and interpreted languages. Machine learning module 210 may implement various types of caching, such as, for example, Just-In-Time (JIT) compilation, Ahead-Of-Time (AOT) compilation, and bytecode caching.
[0102] By leveraging machine learning module 210, computing system 200 may provide users a less time-consuming and cumbersome method of correcting and / or editing text generated through voice typing. That is, the techniques of this disclosure provide users theability to quickly and easily create accurate and grammatically correct textual information through vocal inputs.
[0103] FIG. 3 illustrates an example display of voice typing and an explicit keyword initiating an edit mode, in accordance with one or more aspects of the present disclosure. As shown in the example of FIG. 3, display 322 may be a display for showcasing output generated by the computing system described herein, which may implemented on a user computing device (e.g., display 322 may be an example input component 142 and / or output component 146 of FIG. 1). Further, as shown in the example of FIG. 3, display 322 may be a display for showcasing a graphical user interface (GUI) 324 of a messaging application that may provide a user the capability to send short message service (SMS) messages.
[0104] In the example of FIG. 3, text input 312 contains the phrase “Please help me take a look at this stock”. In some examples, text input 312 may be generated from an ASR executing on the computing system, such as ASR 206 of FIG. 2. As such, text input 312 may be an output from voice typing performed by a user. In another example, text input 312 may be entered using another mode of input, such as a keyboard, keypad, or the like. As shown in FIG. 3, text input 312 may be displayed as text on GUI 324 of the messaging application displayed by display 322.
[0105] In the example of FIG. 3, a user may utter natural language speech 302 (that is, perform voice typing with the computing system) such as “change stock to doc”, which may be received by one or more input components of the computing system, such as input components 242 of FIG. 2. As described herein, natural language speech 302 may not need to follow any predetermined command templates or syntax. As such, a user may be free to verbally express their intent to correct and / or change text input 312 using any free form speech. As described above with respect to FIG. 2, natural language speech 302 from the user may be captured by one or more input components 242 of the computing system (e.g., a microphone) and transformed into digital form as audio input data 204. As shown in the example of FIG. 3, natural language speech 302 may also be displayed as text on GUI 324 of the messaging application displayed by display 322. As described above, audio input data 204 generated from natural language speech 302 may be analyzed by ASR 206 to generate structured text data, such as structured text data 208 of FIG. 2, that can be read or parsed by a large language model executing on the computing system, such as one or more machine learning models applied by machine learning module 210 of FIG. 2.
[0106] The machine learning model may identify, based on the structured text data generated from the audio input data (or the structured text data generated from naturallanguage speech 302), one or more changes to be made to text input 312. As such, in the example of FIG. 3, the computing system may identify, based on the user’s natural language speech 302 “change stock to doc,” one or more changes to be made 320 to text input 312. In this example, one or more changes to be made 320 includes the word “stock”. The large language model may parse the structured text data generated from natural language speech 302 to identify one or more explicit and / or implicit keywords that indicate commands to change and / or correct text input 312. In this example, the explicit keyword “change” may be identified in natural language speech 302, and this keyword may initiate an edit mode for the machine learning model to generate one or more proposed changes to text input 312.
[0107] FIG. 4 illustrates an example display of implementing one or more proposed changes generated from the edit mode, in accordance with one or more aspects of the present disclosure. Display 422, text input 412, and natural language speech 402 may be similar if not substantially similar to display 322, text input 312, and natural language speech 302 of FIG.3. As described above, the machine learning model executing on the computing system (e.g., machine learning module 210 of FIG. 2) may generate, based at least in part on the one or more changes to be made to text input 412, one or more proposed changes 414 to text input 412. The computing system may further output the one or more proposed changes 414 to text input 412 (e.g., via display 422). As shown in the example of FIG. 4, one or more proposed changes 414 includes the word “doc”, which has replaced the word “stock” shown in FIG. 3. As such, because the user provided natural language speech 402 “change stock to doc”, the computing system changed the word “stock”, which was previously included in text input 412, to “doc”. In some examples, as described above, prior to implementing one or more of proposed changes 414 in text input 412, the computing system may receive a response or input from the user (e.g. via one or more input components 242 of FIG. 2) that indicates the user’s authorization or approval to implement one or more of proposed changes 414.
[0108] FIG. 5 illustrates an example display of voice typing and an implicit keyword initiating an edit mode, in accordance with one or more aspects of the present disclosure. Display 522, GUI 524, text input 512, one or more changes to be made 520, and natural language speech 502 may be similar if not substantially similar to display 322, GUI 324, text input 312, one or more changes to be made 320, and natural language speech 302 of FIG. 3. In the example of FIG. 5, text input 512 contains the phrase “They were neither excited or bored”. Additionally, as shown in FIG. 5, a user may utter natural language speech 502 “Fix it”, which may be received by one or more input components of the computing system, such as input components 242 of FIG. 2, and may be displayed as text on GUI 524 of themessaging application displayed by display 522. In this example, the user may have identified one or more changes to be made 520 to text input 512, such as an incorrect word “or”. Rather than providing explicit keywords or instructions for correcting text input 512, though, the user may simply utter natural language speech 502 “Fix it”, which contains an implicit keyword such as “Fix”, to have the computing system automatically generate one or more proposed changes that can correct text input 512. In this way, a user may not have to identify each and every error present in text input 512 in order to fully correct text input 512, and can instead utter an implicit keyword.
[0109] FIG. 6 illustrates another example display of implementing one or more proposed changes generated from the edit mode, in accordance with one or more aspects of the present disclosure. Display 622, text input 612, natural language speech 602, and one or more proposed changes 614 may be similar if not substantially similar to display 422, text input 412, natural language speech 402, and one or more proposed changes 414 of FIG. 4, respectively. As shown in the example of FIG. 6, one or more proposed changes 614 includes the word “nor”, which has corrected the word “or” shown in FIG. 5. As such, because the user provided natural language speech 602 “Fix it”, the computing system corrected the word “or”, which was previously included in text input 612, to “or”.
[0110] As shown in the examples above, one or more proposed changes may be generated based on an explicit keyword in the audio input data (e.g., “change”) to specifically replace one word with another word in the text input, or generated based on an implicit keyword in the audio input data (e.g., “fix”) to perform a general task of correcting one or more errors in the text input. However, as described above, the machine learning model may be used to generate a variety of proposed changes to the text input, such as a deletion and / or insertion of words, a deletion, insertion, or replacement of punctuation, a changing of or otherwise correction of the spelling of words (especially OOV words), and a toggling between a number format and a word format for numerical values. In some examples, the computing system may also generate one or more proposed changes including any one or more of the listed changes in response to detecting a single keyword.
[0111] Furthermore, as described above, tiie machine learning model may be trained with a set of explicit keywords and implicit keywords. For example, explicit keywords may include action words (e.g., verbs) such as change, replace, delete, modify, insert, add, also, etc., indicating a specific action to be performed on the text input. Explicit keywords to initiate an edit mode of the machine learning model may be defined in a predetermined keyword list (e.g., keyword list 216 of FIG. 2), and the machine learning model may betrained with the predetermined keyword list. Example implicit keywords may include words such as fix, correct, actually, wait, no, nope, spell check, grammar check, spell out, spell, auto, auto-correct, etc., indicating a general corrective action to be performed on the text input. Similarly, recognized implicit keywords to initiate the edit mode may be defined in a different or same predetermined keyword list (e.g., keyword list 216 of FIG. 2) and the machine learning model may be trained with the different or same predetermined keyword list.
[0112] In some examples, the machine learning model may be trained to perform automatic correction of one or more spelling, grammar, punctuation, word replacement according to a thesaurus, etc., without requiring the user to utter any specific explicit or implicit keyword. The corrections and / or changes made to the text input by the computing system may result, at least in part, from training the model on a database of natural language examples and corrections and / or changes to the natural language examples generated by users. Additionally, or alternatively, the natural language examples and corrections and / or changes to the natural language examples may comprise data generated by another machine learning model.
[0113] In some examples, a user may be required to utter a “hot word” (e.g., “hey”, “fix”, “change”, etc.) to trigger edit mode or correction mode of the machine learning model. As such, once the user utters natural language speech that is transformed into structured text data by the ASR, the machine learning model may parse the structured text data and be required to identify the “hot word” before implementing any algorithms to generate one or more proposed changes to the text input.
[0114] FIG. 7 illustrates an example process for artificial intelligence-assisted text correction and editing, in accordance with one or more aspects of this disclosure. FIG. 7 may be described with respect to one or more components of FIG. 1 and FIG. 2. In an example, computing system 100, receives, from at least one of automatic speech recognizer (ASR) 106 and an input device providing for entry of typed characters, text input 112 (750). In some examples, the input device may be one of input components 142. Computing system 100 receives natural language audio input data 104 (752). In some examples, natural language audio input data 104 may be generated from natural language speech 102 provided by a user. Computing system 100 automatically generates structured text data 108 based at least in part on natural language audio input data 104 (754).
[0115] Machine learning module 110 executing on computing system 100 identifies, based at least in part on structured text data 108, one or more changes to be made to text input112 (756). In some examples, structured text data 108 includes one or more errors, wherein identifying the one or more changes to be made to the text input 112 further comprises inferring, by machine learning module 110 and based at least in part on structured text data 108 including the one or more errors, the one or more changes to be made to text input 112. In some examples, machine learning module 110 includes one or more of a statistical language model and a large language model. In some examples, the one or more changes to be made to text input 112 include one or more corrections to one or more errors in text input 112. In some examples, machine learning module 110 identifies the one or more changes to be made to text input 112 based at least in part on detecting one or more keywords from predetermined keyword list 216 in structured text data 208. In some examples, the one or more keywords include at least one word from predetermined keyword list 216 to initiate an edit mode for machine learning module 110 to generate one or more proposed changes 114 to text input 112. In some examples, the one or more keywords from keyword list 216 indicate one or more implicit changes to be made to text input 112, in which machine learning module 110 identifies the one or more implicit changes to be made to text input 112, and generates, based at least in part on the one or more implicit changes to be made to text input 112, one or more proposed changes 114 to text input 112. In some examples, the one or more keywords from keyword list 216 indicate one or more explicit changes to be made to text input 112. In some examples, the one or more explicit changes comprise at least one of an insertion of a word to text input 112, a deletion of a word from text input 112, a replacement of a word in text input 112, an insertion of a newline character to text input 112, an insertion of punctuation to text input 112, a deletion of punctuation from text input 112, a replacement of punctuation in text input 112, a changing spelling of a word in text input 112, and a changing of a format of a number in text input 112. In some examples, the one or more keywords comprise a word from predetermined keyword list 216 to indicate inserting a link into text input 112. In some examples, prior to identifying the one or more changes to be made to text input 112, computing system 100 applies at least one rule from a predefined set of rules and at least one template from a predefined set of templates to text input 112 to generate one or more possible changes, and machine learning module 210 then identifies, based at least in part on structured text data 208 and the one or more possible changes, the one or more changes to be made to the text input 112.
[0116] Machine learning module 110 generates, based at least in part on the one or more changes to be made to text input 112, one or more proposed changes 114 to text input 112 (758). Computing system 100 outputs one or more proposed changes 114 to text input 112(760). In some examples, computing system 100 sends, to at least one of ASR 106 and the input device (e.g., input components 142), one or more proposed changes 114 to text input 112. In some examples, responsive to receiving user input, computing system 100 implements at least one of one or more proposed changes 114 to text input 112. In some examples, responsive to detecting one or more keywords comprising a word from predetermined keyword list 216 to indicate inserting a link into text input 112, computing system 100 generates a plurality of links associated with one or more words in text input 112, and selects a link from the plurality of links that is most associated with the one or more words in text input 112. In some examples, responsive to receiving user input (e.g., via input component 142), computing system 100 inserts the link from the plurality of links into text input 112.
[0117] As such, the techniques disclosed herein may improve the operation of a computing system by providing a user the capability use “voice typing” (e.g., using free form natural language speech) as a means of correcting and / or changing text input that was generated from either previous natural language speech or text manually entered via another input device (e.g., a keyboard or touchscreen). In this way, users, including users whose hands are not readily available (e.g., users who are driving or otherwise situationally impaired), may quickly and easily correct and / or change text output (e.g., output generated by ASR subsystems), as they may simply use their voice, and not be required to switch to a manual mode of operation for text correction / editing and / or be required to re-dictate their desired text content.
[0118] Example 1. A method comprising: receiving, by a computing system, and from at least one of an automatic speech recognizer and an input device providing for entry of typed characters, text input; receiving, by the computing system, natural language audio input data; automatically generating, by the computing system, structured text data based at least in part on the natural language audio input data; identifying, by a machine learning model executing on the computing system, based at least in part on the structured text data, one or more changes to be made to the text input; generating, by the machine learning model, based at least in part on the one or more changes to be made to the text input, one or more proposed changes to the text input; and outputting, by the computing system, the one or more proposed changes to the text input.
[0119] Example 2. Tire method of example 1, further comprising: sending, by the computing system and to at least one of the automatic speech recognizer and the input device, the one or more proposed changes to the text input.
[0120] Example 3. The method of any one of examples 1-2, further comprising: responsive to receiving user input, implementing at least one of the one or more proposed changes to the text input.
[0121] Example 4. The method of any one of examples 1-3, wherein the machine learning model further includes one or more of a statistical language model and a large language model.
[0122] Example 5. The method of any one of examples 1-4, wherein the one or more changes to be made to the text input include one or more corrections to one or more errors in the text input.
[0123] Example 6. The method of any one of examples 1-5, wherein identifying the one or more changes to be made to the text input is based at least in part on detecting one or more keywords in the structured text data.
[0124] Example 7. The method of example 6, wherein the one or more keywords include at least one word from a predetermined list of words to initiate an edit mode for the machine learning model to generate the one or more proposed changes to the text input.
[0125] Example 8. The method of any one of examples 6-7, wherein the one or more keywords indicate one or more implicit changes to be made to the text input, further comprising: identifying, by the machine learning model, the one or more implicit changes to be made to the text input; and generating, by the machine learning model, based at least in part on the one or more implicit changes to be made to the text input, the one or more proposed changes to the text input.
[0126] Example 9. The method of any one of examples 6-7, wherein the one or more keywords indicate one or more explicit changes to be made to the text input.
[0127] Example 10. The method of example 9, wherein the one or more explicit changes comprise at least one of an insertion of a word to the text input, a deletion of a word from the text input, a replacement of a word in the text input, an insertion of a newline character to the text input, a deletion of a newline character to the text input, an insertion of a whitespace character to the text input, a deletion of a whitespace character to the text input, an insertion of punctuation to the text input, a deletion of punctuation from the text input, a replacement of punctuation in the text input, a changing spelling of a word in the text input, a changing of a format of a number in the text input, and a changing of a text format of one or more characters in the text input.
[0128] Example 11. The method of example 6, wherein the one or more keywords comprise a word from a predetermined list of words to indicate inserting a link into the text input.
[0129] Example 12. The method of example 11 , further comprising: responsive to detecting one or more keywords comprising a word from the predetermined list of words to indicate inserting a link into the text input, generating, by the computing system, a plurality of links associated with one or more words in the text input; and selecting, by the computing system, a link from the plurality of links that is most associated with the one or more words in the text input.
[0130] Example 13. The method of example 12, further comprising: responsive to receiving user input, inserting, by the computing system, the link from the plurality of links into the text input.
[0131] Example 14. The method of example 1, further comprising: prior to the identifying the one or more changes to be made to text input, applying, by the computing system, at least one rule from a predefined set of rules and at least one template from a predefined set of templates to the text input to generate one or more possible changes; and identifying, by the machine learning model, based at least in part on the structured text data and the one or more possible changes, the one or more changes to be made to the text input.
[0132] Example 15. The method of example 1 , wherein the structured text data includes one or more errors, and wherein identifying the one or more changes to be made to the text input further comprises: inferring, by the machine learning model and based at least in part on the structured text data including the one or more errors, the one or more changes to be made to the text input.
[0133] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over, as one or more instructions or code, a computer-readable medium and executed by a hardware-based processing unit. Computer- readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media, which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors toretrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0134] By way of example, and not limitation, such computer-readable storage media can comprise random-access memory (RAM), read-only memory (ROM), EEPROM, compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage mediums and media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to nontransient, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of a computer-readable medium.
[0135] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structures or any other structures suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0136] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above,various units may be combined in a hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0137] Various embodiments have been described. These and other embodiments are within the scope of the following claims.
Claims
WHAT IS CLAIMED IS:
1. A method comprising: receiving, by a computing system, and from at least one of an automatic speech recognizer and an input device providing for entry of typed characters, text input; receiving, by the computing system, natural language audio input data; automatically generating, by the computing system, structured text data based at least in part on the natural language audio input data; identifying, by a machine learning model executing on the computing system, based at least in part on the structured text data, one or more changes to be made to the text input; generating, by the machine learning model, based at least in part on the one or more changes to be made to the text input, one or more proposed changes to the text input; and outputting, by the computing system, the one or more proposed changes to the text input.
2. The method of claim 1 , further comprising: sending, by the computing system and to at least one of the automatic speech recognizer and the input device, the one or more proposed changes to the text input.
3. The method of any one of claims 1-2, further comprising: responsive to receiving user input, implementing at least one of the one or more proposed changes to the text input.
4. The method of any one of claims 1-3, wherein the machine learning model further includes one or more of a statistical language model and a large language model.
5. The method of any one of claims 1-4, wherein the one or more changes to be made to the text input include one or more corrections to one or more errors in the text input.
6. The method of any one of claims 1-5, wherein identifying the one or more changes to be made to the text input is based at least in part on detecting one or more keywords in the structured text data.
7. The method of claim 6, wherein the one or more keywords include at least one word from a predetermined list of words to initiate an edit mode for the machine learning model to generate the one or more proposed changes to the text input.
8. The method of any one of claims 6-7, wherein the one or more keywords indicate one or more implicit changes to be made to the text input, further comprising: identifying, by the machine learning model, the one or more implicit changes to be made to the text input; and generating, by the machine learning model, based at least in part on the one or more implicit changes to be made to the text input, the one or more proposed changes to the text input.
9. The method of any one of claims 6-7, wherein the one or more keywords indicate one or more explicit changes to be made to the text input.
10. The method of claim 9, wherein the one or more explicit changes comprise at least one of an insertion of a word to the text input, a deletion of a word from the text input, a replacement of a word in the text input, an insertion of a newline character to the text input, a deletion of a newline character to the text input, an insertion of a whitespace character to the text input, a deletion of a whitespace character to the text input, an insertion of punctuation to the text input, a deletion of punctuation from the text input, a replacement of punctuation in the text input, a changing spelling of a word in the text input, a changing of a format of a number in the text input, and a changing of a text format of one or more characters in the text input.
11. The method of claim 6, wherein the one or more keywords comprise a word from a predetermined list of words to indicate inserting a link into the text input.
12. The method of claim 11, further comprising: responsive to detecting one or more keywords comprising a word from the predetermined list of words to indicate inserting a link into the text input, generating, by the computing system, a plurality of links associated with one or more words in the text input; andselecting, by the computing system, a link from the plurality of links that is most associated with the one or more words in the text input.
13. The method of claim 12, further comprising: responsive to receiving user input, inserting, by the computing system, the link from the plurality of links into the text input.
14. The method of claim 1 , further comprising: prior to identifying the one or more changes to be made to text input, applying, by the computing system, at least one rule from a predefined set of rules and at least one template from a predefined set of templates to the text input to generate one or more possible changes; and identifying, by the machine learning model, based at least in part on the structured text data and the one or more possible changes, the one or more changes to be made to the text input.
15. The method of claim 1, wherein the structured text data includes one or more errors, and wherein identifying the one or more changes to be made to tire text input further comprises: inferring, by the machine learning model and based at least in part on the structured text data including the one or more errors, the one or more changes to be made to the text input.
16. A computing system comprising: a memory that stores instructions; and processor circuitry that execu tes the instructions to perform the method of any of claims 1-15.
17. A non-transitory computer-readable storage medium comprising instructions, that when executed by processor circuitry of a computing system, cause the processor circuitry to perform the method of any of claims 1-15.
Citation Information
Patent Citations
Dictation that allows editing
US20170263248A1
Intuitive dictation
US20230306963A1
Cited By
Freezer operation state anomaly detection method and system based on voice recognition and Internet of Things
CN120808812A
Systems and processes for automated collection of data from third-party applications for predictive and adaptive utility-based personal assistance
US12511553B1