Reduce the latency caused by switching input modalities
By establishing a voice-to-text conversion session in a preemptive manner when the user switches to the high-delay input mode, the problem of delay in the input mode is solved and the user experience is improved.
Patent Information
- Application Number
- CN202011261069.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-09-09
- Filing Date
- 2016-09-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2037-03-07
AI Technical Summary
When users switch between input modes, especially when switching from low-latency input mode to high-latency input mode, they will experience significant delays, affecting the user experience.
The speech-to-text conversion session is established in a preemptive manner when the user switches from the low-latency input mode to the high-latency input mode, the processing of the speech input at the query processor is initiated, and the complete query is constructed based on the output.
This significantly reduces the latency of users when switching input modes, allowing users to start using high-latency input modes immediately or relatively quickly, improving user experience.
Smart Images

Figure CN112463938B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 201610812805.5, the application date is September 9, 2016, and the invention title is "Reducing Latency Caused by Switching Input Modes". Technical Field
[0002] This specification generally relates to various implementations that facilitate reducing and / or eliminating latency experienced by a user when switching between input modes, particularly in cases where a user switches from a low-latency input mode to a high-latency input mode. Background Art
[0003] Voice-based user interfaces are increasingly widely used for controlling computers and other electronic devices. A particularly beneficial application of voice-based user interfaces is in portable electronic devices such as mobile phones, watches, tablet computers, head-mounted devices, virtual or augmented reality devices, etc. Another beneficial application is in in-vehicle electronic systems such as automotive systems that incorporate navigation and audio capabilities. Such applications typically feature non-traditional factors that limit the utility of more traditional keyboard or touchscreen inputs and / or their use in situations where it is desirable to encourage the user to remain focused on other tasks, such as when the user is driving or walking.
[0004] For example, in terms of processor and / or memory resources, the computational resource requirements of voice-based user interfaces can be quite substantial. As a result, some conventional voice-based user interface methods employ a client-server architecture, where a relatively low-power client device receives and records voice input, transmits the recording over a network such as the Internet to an online service for speech-to-text conversion and semantic processing, and the online service generates an appropriate response and transmits it back to the client device. The online service can devote substantial computational resources to processing the voice input and can implement more complex speech recognition and semantic analysis functions than could otherwise be achieved locally within the client device. However, the client-server approach necessarily requires the client to be online (i.e., communicating with the online service) when processing voice input. Maintaining such connectivity between the client and the online service can be impractical, particularly in mobile and automotive applications where wireless signal strength is undoubtedly subject to fluctuations. Thus, when it is necessary to use the online service to convert voice input to text, a speech-to-text conversion session must be established between the client and the server. While such a session is being established, the user may experience significant latency, e.g., 1 to 2 seconds or more, which can be detrimental to the user experience. Summary of the Invention
[0005] This specification generally relates to various implementations for facilitating reduction and / or elimination of latency experienced by a user when switching between input modalities, particularly in cases where a user switches from a low-latency input modality to a high-latency input modality. For example, in some implementations, when the environment indicates that a user who provides input (e.g., text) via a low-latency input modality is likely to switch to voice input, a speech-to-text conversion session may be established in a preemptive manner.
[0006] Accordingly, in some implementations, a method may include the following operations: receiving a first input in a first modality of a multimodal interface associated with an electronic device, and in response to receiving the first input: determining that the first input meets a criterion; in response to determining that the first input meets the criterion, establishing a session in a preemptive manner between the electronic device and a query processor, the query processor being configured to process an input received in a second modality of the multimodal interface; receiving a second input in the second modality of the multimodal interface; initiating processing of at least a portion of the second input at the query processor within the session; and constructing a complete query based on an output from the query processor.
[0007] In some implementations, a method may include the following operations: receiving text input with a voice-enabled device; and in the voice-enabled device, and in response to receiving the text input: determining that the text input meets a criterion; in response to determining that the text input meets the criterion, establishing a speech-to-text conversion session in a preemptive manner between the voice-enabled device and a speech-to-text conversion processor; receiving voice input; initiating processing of at least a portion of the voice input at the speech-to-text conversion processor within the session; and constructing a complete query based on an output from the speech-to-text conversion processor.
[0008] In various implementations, the speech-to-text conversion processor may be an online speech-to-text conversion processor, and the voice-enabled device may include a mobile device configured to communicate with the online speech-to-text conversion processor when communicating with a wireless network. In various implementations, initiating the processing includes sending data associated with the text input and data associated with the voice input to the online speech-to-text conversion processor. In various implementations, sending the data may include sending at least a portion of a digital audio signal of the voice input. In various implementations, the online speech-to-text conversion processor may be configured to perform speech-to-text conversion and semantic processing of the portion of the digital audio signal based on the text input to generate an output.
[0009] In various embodiments, constructing a complete query may include combining the output with at least a portion of the text input. In various embodiments, the output from the speech-to-text conversion processor may include multiple candidate interpretations of the speech input, and constructing a complete query includes ranking the multiple candidate interpretations based at least in part on the text input. In various embodiments, preemptively initiating a speech-to-text conversion session may include activating a microphone of the voice-enabled device. In various embodiments, the method may further include providing an output to indicate that the speech-to-text conversion session is available. In various embodiments, the criteria may include that the text input satisfies a character count or word count threshold. In various embodiments, the criteria may include that the text input matches a specific language.
[0010] In addition, some embodiments include an apparatus including a memory and one or more processors operable to execute instructions stored in the memory, wherein the instructions are configured to perform any of the foregoing methods. Some embodiments also include a non-transitory computer-readable storage medium storing computer instructions, wherein the instructions are executable by one or more processors to perform any of the foregoing methods.
[0011] It should be understood that all combinations of the aforementioned concepts and additional concepts described in detail herein are contemplated as part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 An example architecture of a computer system is illustrated.
[0013] Figure 2 is a block diagram of an example distributed speech input processing environment.
[0014] Figure 3 It is used as a graphic Figure 2 Flowchart of an example method for processing speech input in an environment.
[0015] Figure 4 An example communication exchange that may occur between various entities configured with selected aspects of the present disclosure is illustrated according to various implementations.
[0016] Figure 5 is a flow chart illustrating an example method of establishing a speech-to-text session in a preemptive manner in accordance with various implementations. DETAILED DESCRIPTION
[0017] In the embodiments discussed below, an application executing on a resource-constrained electronic device such as a mobile computing device (e.g., a smart phone or a smart watch) may provide a so-called "multimodal" interface that supports a variety of different input modalities. These input modalities may include low-latency inputs such as text that respond to user input with substantially no latency, and high-latency inputs such as speech recognition that incur a relatively high latency due to the various latency-causing routines they require, such as establishing a session with a translation processor configured to convert the input received via the high-latency modality into a form that matches the lower-latency input modality. To reduce latency (or at least the perceived latency) when a user switches from a first input (e.g., text input) that provides low latency to a second input (e.g., speech) that has higher latency, the electronic device may establish a session with the translation processor in a preemptive manner, e.g., in response to a determination that the first input meets one or more criteria. The electronic device is thereby able to immediately initiate processing of the second input by the translation processor rather than being required to first establish a session, which significantly reduces the latency experienced by the user when switching input modalities.
[0018] Further details regarding the selected embodiments are discussed below. However, it should be understood that other embodiments are also contemplated, so the embodiments disclosed herein are not exclusive.
[0019] Turning now to the drawings, where like numerals in several views represent like parts, Figure 1 is a block diagram of the electronic components in an example computer system 10. The system 10 generally includes at least one processor 12 that communicates with a plurality of peripheral devices via a bus subsystem 14. These peripheral devices may include a storage subsystem 16, e.g., including a memory subsystem 18 and a file storage subsystem 20, a user interface input device 22, a user interface output device 24, and a network interface subsystem 26. The input and output devices allow a user to interact with the system 10. The network interface subsystem 26 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0020] In some embodiments, the user interface input device 22 may include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into a display, an audio input device such as a speech recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways for inputting information into the computing system 10 or onto a communication network.
[0021] The user interface output device 24 may include a display subsystem, a printer, a fax machine, or a non-visual display, such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways for outputting information from the computing system 10 to a user or another machine or computer system.
[0022] The storage subsystem 16 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 16 may include logic for performing selected aspects of the methods disclosed hereinafter.
[0023] These software modules are typically executed by the processor 12 independently or in combination with other processors. The memory subsystem 18 in the storage subsystem 16 may include multiple memories, including a main random access memory (RAM) 28 for storing instructions and data during program execution and a read-only memory (ROM) 30 for storing fixed instructions therein. The file storage subsystem 20 may provide permanent storage for program and data files and may include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical disk drive, or a removable media cartridge. The modules implementing the functionality of certain embodiments may be stored by the file storage subsystem 20 in the storage subsystem 16 or stored in other machines accessible by the processor 12.
[0024] The bus subsystem 14 provides a mechanism for allowing the various components and subsystems of the system 10 to communicate with each other as expected. Although the bus subsystem 14 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0025] The system 10 may have different types, including mobile devices, portable electronic devices, embedded devices, desktop computers, laptop computers, tablet computers, wearable devices, workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Additionally, the functionality implemented by the system 10 may be distributed among multiple systems interconnected with each other via one or more networks, e.g., in a client-server, peer-to-peer, or other network arrangements. Due to the ever-changing nature of computers and networks, the description of the system 10 depicted in Figure 1 is only intended as a specific example for illustrating some embodiments. Many other configurations of the system 10 may have more or fewer components than the computer system depicted in Figure 1
[0026] The embodiments discussed below may include one or more methods of various combinations for implementing the functions described herein. Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform a method such as one or more of the methods described herein. Still other embodiments may include an apparatus including a memory and one or more processors operable to execute instructions stored in the memory to perform a method such as one or more of the methods described herein.
[0027] The various program codes described below may be identified based on an application, within which the various program codes are implemented in a particular embodiment. However, it should be understood that any specific program terms are used for convenience only. Additionally, in view of the myriad ways in which computer programs can be organized into routines, procedures, methods, modules, objects, etc. and the various ways in which program functions can be allocated among the various software layers (e.g., operating systems, libraries, API applications, applets, etc.) residing within a typical computer, it should be understood that some embodiments may not be limited to the specific combinations and allocations of program functions described herein.
[0028] Furthermore, it should be understood that the various operations described herein that may be performed by any program code or within any routine, workflow, etc. may be combined, split, reordered, omitted, performed successively or in parallel, and / or supplemented with other techniques, and thus certain embodiments are not limited to the specific order of operations described herein.
[0029] Figure 2An example distributed speech input processing environment 50 is illustrated, for example, using a speech-enabled device 52 that communicates with one or more online services such as an online search service 54. In the embodiments discussed below, for example, the speech-enabled device 52 is described as a mobile device, such as a cellular phone or a tablet computer. While other embodiments may utilize a variety of other speech-enabled devices, reference to a mobile device in the following is only for the purpose of simplifying the following discussion. Countless other types of speech-enabled devices can use the functionality described herein, for example, including laptops, watches, head-mounted devices, virtual or augmented reality devices, other wearable devices, audio / video systems, navigation systems, cars and other vehicle systems, etc. In addition, many of such speech-enabled devices may be considered resource-constrained because the memory and / or processing power of such devices may be constrained based on technical, economic or other factors, especially when compared to online or cloud-based services that can devote almost unlimited computing resources to independent tasks. Some such devices may also be considered offline devices, in which respect such devices are able to operate "offline" at least part of the time and are not connected to online services, for example, based on the expectation that such devices may experience temporary network connectivity interruptions from time to time under ordinary use.
[0030] The voice-enabled device 52 can be operated to communicate with a variety of online services. A non-limiting example is an online search service 54. In some embodiments, the online search service 54 can be implemented as a cloud-based service using a cloud infrastructure, for example, using a server farm or a cluster of high-performance computers running software suitable for processing a large number of requests from multiple users. In the illustrated embodiment, the online search service 54 can query one or more databases to locate the requested information, for example, to provide a list of websites including the requested information. The online search service 54 may not be limited to voice-based searches, and may also be able to handle other types of searches, for example, text-based searches, image-based searches, etc.
[0031] The voice-enabled device 52 can also communicate with other online systems (not shown), and these other online systems are not necessarily required to process searches. For example, some online systems can process voice-based requests for non-search actions, such as setting an alarm or reminder, managing a list, initiating communication with other users via phone, text, email, etc., or performing other actions that can be initiated via voice input. For the purposes of this disclosure, voice-based requests and other forms of voice input can be collectively referred to as voice-based queries, regardless of whether the voice-based query seeks to initiate a search, ask a question, issue a command, dictate an email or text message, etc. Thus, generally speaking, within the context of the illustrated embodiments, any voice input, such as including one or more words or phrases, can be considered a voice-based query.
[0032] In Figure 2 an embodiment, the voice input received by the voice-enabled device 52 is processed by a voice-enabled search application (or “app”) 56. In other embodiments, the voice input can be processed within the operating system or firmware of the voice-enabled device. The application 56 in the illustrated embodiment provides a multimodal interface that includes a text action module 58, a voice action module 60, and an online interface module 62. Although not shown in Figure 2 the application 56 can also be configured to accept input using input modalities other than text and voice, such as motion (e.g., gestures made with the phone), biometrics (e.g., retina input, fingerprint, etc.).
[0033] The text action module 58 receives text input directed to the application 56 and performs various actions, such as filling one or more presented input fields of the application 56 with the provided text. The voice activity module 60 receives voice input directed to the application 56 and coordinates the analysis of the voice input. The voice input can be analyzed locally (e.g., by components 64 to 72 as described below) or remotely (e.g., by a separate online speech-to-text conversion processor 78 or a voice-based query processor 80 as described below). The online interface module 62 provides an interface to the online search service 54 and to the separate online speech-to-text conversion processor 78 and voice-based query processor 80.
[0034] If the voice-enabled device 52 is offline, or if its wireless network signal is too weak and / or insufficient to delegate the analysis of the voice input to an online speech-to-text conversion processor (e.g., 78, 80), then the application 56 can rely on a local speech-to-text conversion processor to process the voice input. The local speech-to-text conversion processor can include various middleware, frameworks, operating systems, and / or firmware modules. For example, in Figure 2In this case, the local speech-to-text conversion processor includes a streaming speech-to-text module 64 and a semantic processor module 66 equipped with a parsing module 70.
[0035] The streaming speech-to-text module 64 receives an audio recording of a speech input, such as in the form of digital audio data, and converts the digital audio data into one or more text words or phrases (also referred to herein as tokens). In the illustrated embodiment, module 64 takes the form of a streaming module so that the speech input is converted to text on a token-by-token basis in real time or near real time, such that tokens can be efficiently output from module 64 while the user is speaking and thus before the user has articulated a complete dictated request. Module 64 may rely on one or more locally stored offline acoustic and / or language modules 68, which together model the relationship between an audio signal and the phonetic units in a language along with the word order in that language. In some embodiments, a single module 68 may be used, while in other embodiments, multiple modules may be supported, e.g., to support multiple languages, multiple speakers, etc.
[0036] Given that module 64 converts speech to text, the semantic processor module 66 attempts to discern the semantics or meaning of the text output by module 64 in order to formulate an appropriate response. For example, the parsing module 70 relies on one or more offline grammar modules 72 to map the interpreted text to various constructs, such as statements, questions, etc. As shown, the parsing module 70 may provide the parsed text to the application 56 so that the application 56 can, for example, populate input fields and / or provide the text to the online interface module 62. In some embodiments, a single module 72 may be used, while in other embodiments, multiple modules may be supported. It should be understood that in some embodiments, modules 68 and 72 may be combined into fewer modules or divided into additional modules, as may be the functionality of modules 64 and 66. Additionally, when the device 52 is not communicating with the online search service 54, modules 68 and 72 are locally stored on the speech-enabled device 52 and are thus accessible offline, so these modules are referred to herein as offline modules.
[0037] On the other hand, if the voice-enabled device 52 is online, or if its wireless network signal is strong enough and / or sufficient to delegate the analysis of the voice input to an online speech-to-text conversion processor (e.g., 78, 80), the application 56 can rely on remote capabilities to process the voice input. The remote capabilities can be provided by various sources, such as a standalone online speech-to-text conversion processor 78 and / or a voice-based query processor 80 associated with the online search service 54, either of which can rely on various acoustic / language, grammar, and / or action modules 82. It should be understood that in some embodiments, particularly when the voice-enabled device 52 is a resource-constrained device, the online speech-to-text conversion processor 78 and / or the voice-based query processor 80 and the modules 82 used thereby can implement more complex and computationally resource-intensive voice processing functions than locally to the voice-enabled device 52. However, in other embodiments, no complementary online capabilities can be used.
[0038] In some embodiments, both online and offline capabilities can be supported, e.g., such that the online capabilities are used whenever the device communicates with an online service, and the offline capabilities are used when no connection exists. In other embodiments, the online capabilities are used only when the offline capabilities fail to adequately process a particular voice input.
[0039] For example, Figure 3 Illustrated is a voice processing routine 100 that can be executed by the voice-enabled device 52 to process a voice input. The routine 100 begins at block 102 by receiving a voice input in the form of, for example, a digital audio signal. At block 104, a preliminary attempt is made to forward the voice input to the online search service. If unsuccessful, e.g., due to lack of connectivity or lack of a response from the online speech-to-text conversion processor 78, block 106 passes control to block 108 to convert the voice input into text tokens (e.g., using Figure 2 module 64), and to parse the text tokens (block 110, e.g., using Figure 2 module 70), and the processing of the voice input is complete.
[0040] Returning to block 106, if the attempt to forward the voice input to the online search service is successful, block 106 bypasses blocks 108 to 110 and passes control directly to block 112 to perform client-side rendering and synchronization. Subsequently, the processing of the voice input is complete. It should be understood that in other embodiments, offline processing can be attempted before online processing, e.g., to avoid unnecessary data communication when the voice input can be processed locally.
[0041] As described in the background art, a user may experience latency when switching input modalities, particularly when the user switches from a low-latency input modality such as text to a high-latency input modality such as voice. For example, assume that a user wishes to submit a search query to an online search service 54. The user may type text into the text input of a voice-enabled device 52, but may decide that typing is too cumbersome or may become distracted (e.g., due to driving) such that the user can no longer efficiently type text. In existing electronic devices such as smart phones, the user is required to press a button or touch screen icon to activate the microphone and initiate the establishment of a session with a voice-to-text conversion processor implemented locally on the voice-enabled device 52 or online at a remote computing system (e.g., 78 or 80). Establishing such a session can take time, which can degrade the user experience. For example, establishing a session with an online voice-to-text conversion processor 78 or an online voice-based query processor 80 may take on the order of one to two seconds or more, depending on the strength and / or reliability of the available wireless signal.
[0042] To reduce or avoid such latency and using the techniques described herein, for example, while the user is still typing the first part of her query using the keypad, the voice-enabled device 52 can establish a session with the voice-to-text conversion processor in a preemptive manner. By the time the user decides to switch to voice, the session may already be established or at least the establishment of the session may be in progress. Either way, the user can start speaking immediately or at least relatively quickly. The voice-enabled device 52 can respond with little or no perceivable latency.
[0043] Figure 4 Illustrates an example of the communications that may be exchanged between an electronic device such as a voice-enabled device 52 and a voice-to-text conversion processor such as a voice-based query processor 80 according to various embodiments. This particular example illustrates the scenario of establishing a session between the voice-enabled device 52 and an online voice-based query processor 80. However, this is not limiting. Similar communications may be exchanged between the voice-enabled device 52 and a stand-alone online voice-to-text conversion processor 78. Additionally or alternatively, similar communications may be exchanged between internal modules of a suitably equipped voice-enabled device 52. For example, when the voice-enabled device 52 is offline (and Figure 3 the operations of blocks 108 to 112 are performed), various internal components of the voice-enabled device 52 such as one or more of the streaming voice-to-text module 64 and / or the semantic processor module 66 can collectively perform operations similar to those performed by Figure 4Tasks performed by the online voice-based query processor 80 in (but certain aspects such as the depicted handshake protocol may be simplified or omitted). Similarly, a user 400 of the voice-enabled device 52 is schematically depicted.
[0044] At 402, text input can be received from the user 400 at the voice-enabled device 52. For example, the user 400 can start a search by typing text on a physical keypad or a graphical keypad presented on a touch screen. At 404, the voice-enabled device 52 can evaluate the text input and / or the current context of the voice-enabled device 52 to determine whether various criteria are met. If the criteria are met, the voice-enabled device 52 can establish a voice-to-text conversion session with the voice-based query processor 80. In Figure 4 it, the process is indicated as a three-way handshake at 406 to 410. However, other handshake protocols or session establishment routines can be used instead. At 412, the voice-enabled device 52 can provide some output indicating that the session has been established so that the user 400 will know that he or she can start speaking instead of typing.
[0045] Various criteria can be used to evaluate the text input received by the voice-enabled device at 402. For example, length-based criteria such as the character or word count of the text input received up to that point can be compared with a length-based threshold (e.g., a character or word count threshold). Satisfaction of the character / word count threshold can suggest that the user may be getting tired of typing and will switch to voice input. Additionally or alternatively, the text input can be compared with various grammars to determine the matching language of the text input (e.g., German, Spanish, Japanese, etc.). Some languages may include long words that the user is more likely to switch the input modality (e.g., text-to-speech) to complete. Additionally or alternatively, it can be determined whether the text input matches one or more patterns, such as regular expressions or other similar mechanisms.
[0046] In some embodiments, as a supplement or alternative to evaluating the text input against various criteria, the context of the voice-enabled device 52 can be evaluated. If the context of the voice-enabled device 52 is "driving", it is very likely that the user will want to switch from text input to voice input. The "context" of the voice-enabled device 52 can be determined based on a variety of signals, including but not limited to sensor signals, user preferences, search history, etc. Examples of sensors that can be used to determine the context include but are not limited to position coordinate sensors (e.g., Global Positioning System or "GPS"), accelerometers, thermometers, gyroscopes, light sensors, etc. User preferences and / or search history can indicate environments in which the user prefers and / or tends to switch the input modality when providing input.
[0047] ReviewFigure 4 , sometimes after indicating to the user at 412 that a session has been established, at 414, the voice-enabled device 52 can receive voice input from the user 400. For example, the user can stop typing text input and can start speaking into the microphone and / or mouthpiece of the voice-enabled device 52. The voice-enabled device 52 can then initiate online processing of at least a portion of the voice input at the online voice-based query processor 80 within the session established at 406 to 410. For example, at 416, the voice-enabled device 52 can send at least a portion of the digital audio signal of the voice input to the online voice-based query processor 80. In some embodiments, at 418, the voice-enabled device 52 can also send data associated with the text input received at 402 to the online voice-based query processor 80.
[0048] At 420, the online voice-based query processor 80 can perform speech-to-text conversion and / or semantic processing on the portion of the digital audio signal to generate output text. In some embodiments, the online voice-based query processor 80 can further generate an output based on the text input it receives at 418. For example, the online voice-based query processor 80 can create a bias through the text input it receives at 418. Suppose the user says the word "socks" into the microphone of the voice-enabled device 52. Without any other information, the user's dictated voice input might simply be interpreted by the online voice-based query processor 80 as "socks". However, if the online voice-based query processor 80 takes into account the text input of "red" for the continuing voice input, the online voice-based query processor 80 might bias the interpretation of the dictated word "socks" as "Sox" (as in "Boston Red Sox").
[0049] As another example, the language of the text input can bias the online voice-based query processor 80 towards a particular interpretation. For example, certain languages such as German have relatively long words. If the online voice-based query processor 80 determines that the text input is German, the online voice-based query processor 80 is more likely to link the text interpreted from the voice input with the text input, rather than separating them into separate words / tokens.
[0050] In addition to the text input, the online voice-based query processor 80 can consider other signals, such as the user's context (e.g., a user located in New England is more likely to mention the Red Sox than a user in Japan), the user's accent (e.g., a Boston accent can significantly increase the likelihood of interpreting "socks" as "Sox"), the user's search history, and so on.
[0051] Review Figure 4 At 422, the online voice-enabled query processor 80 can provide output text to the voice-enabled device 52. The output can have various forms. In embodiments where the text input and / or the context of the voice-enabled device 52 is provided to the voice-based query processor 80, the voice-based query processor 80 can return a "best" guess about the text corresponding to the voice input received by the voice-enabled device 52 at 414. In other embodiments, the online voice-based query processor 80 can output or return multiple candidate interpretations of the voice input.
[0052] Regardless of the form of output provided by the online voice-based query processor 80 to the voice-enabled device 52, at 424, the voice-enabled device 52 can use the output to construct a complete query that can be submitted to, for example, the online search service 52. For example, in embodiments where the online voice-based query processor 80 provides a single best guess, the voice-enabled device 52 can incorporate the best guess as one token in a multi-token query that also includes the original text input. Or, if the text output appears to be the first part of a relatively long word (especially in languages such as German), the voice-enabled device 52 can directly concatenate the best guess of the online voice-based query processor 80 with the text input to form a single word. In embodiments where the online voice-based query processor 80 provides multiple candidate interpretations, the voice support device 52 can rank the candidate interpretations based on various signals such as one or more attributes of the text input received at 402 (e.g., character count, word count, language, etc.), the context of the voice-enabled device 52, etc., so that the voice-enabled device 52 can select the "best" candidate interpretation.
[0053] Although the examples described herein mainly relate to a user switching from text input to voice input, this is not limiting. In various embodiments, when a user switches between any input modalities, especially when a user switches from a low-latency input modality to a high-latency input modality, the techniques described herein can be employed. For example, an electronic device can provide a multimodal interface, which can be an interface capable of accepting various different types of input, such as a web interface or an application interface (e.g., a text messaging application, a web search application, a social networking application, etc.). Assume that a first input is received in a first, low-latency modality of the multimodal interface provided by the electronic device. The electronic device can be configured to establish a session in a preemptive manner between the electronic device and a conversion processor (e.g., online or local), which is configured to process inputs received in a second, high-latency modality of the multimodal interface. This process can be performed, for example, in response to determining that the first input meets the criteria. Then, when a second input is received in the second modality of the multimodal interface, the electronic device can be prepared to immediately or very quickly initiate processing of at least a portion of the second input at the conversion processor within the session. This can reduce or eliminate the latency experienced by the user when switching from the first input modality to the second input modality.
[0054] Figure 5 FIG. illustrates a routine 500 that can be executed by a voice-enabled device 52 to establish a voice-to-text conversion session in a preemptive manner with a (online or local) voice-to-text conversion processor, according to various embodiments. The routine 500 begins at block 502 by receiving text input. At block 504, the text input can be analyzed for one or more criteria to determine whether to establish a voice-to-text conversion session in a preemptive manner.
[0055] After determining that one or more criteria are met, at block 508, the voice-enabled device 52 can establish the aforementioned voice-to-text conversion session with a voice-to-text conversion processor (e.g., 64 to 72) that includes components local to the voice-enabled device or an online voice-to-text conversion processor (such as 78 or 80). At block 510, voice input can be received, for example, at a microphone of the voice-enabled device 52. At block 512, the voice-enabled device 52 can initiate processing of the voice input received at block 510 within the session established at 508. At block 514, a complete query can be constructed based at least on the output provided by a voice-based query processor with which a session was established at block 508. Thereafter, the complete query can be used, however the user desires, for example, as a search query submitted to an online search service 54 or as part of a text communication (e.g., a text message, an email, a social media post) to be sent by the user.
[0056] Although several implementations have been described and illustrated in the text, various other devices and / or structures may be utilized for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each such variation and / or modification is to be regarded as within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application for which the teachings are used. Those skilled in the art will recognize or be able to ascertain, by routine experimentation, many equivalents to the specific implementations described herein. It will thus be understood that the foregoing implementations are presented by way of example only, and that the implementations may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. The implementations of the present disclosure are directed to each separate feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
Claims
1. A method, comprising: receiving a first input in a first modality of a multimodal interface associated with an electronic device; determining that a context of the electronic device satisfies criteria, wherein determining that the context of the electronic device satisfies the criteria comprises: determining the context based on one or more signals from sensors of the electronic device, wherein the sensors are at least one of a position coordinate sensor, an accelerometer, a thermometer, a gyroscope, and a light sensor; and in the electronic device, in response to receiving the first input and determining that the context of the electronic device satisfies the criteria: establishing a session in a preemptive manner between the electronic device and a query processor, the query processor being configured to process an input received in a second modality of the multimodal interface; receiving a second input in the second modality of the multimodal interface; initiating, within the session, processing of at least a portion of the second input at the query processor; and constructing a complete query based on an output from the query processor.
2. The method according to claim 1, wherein, the first input comprises a text input, and wherein constructing the complete query comprises: combining the output with at least a portion of the text input.
3. The method according to claim 1, wherein the output from the query processor comprises a plurality of candidate interpretations of the second input, and wherein constructing the complete query comprises: ranking the plurality of candidate interpretations at least partially based on the first input.
4. The method according to claim 1, wherein, the query processor is an online query processor, and wherein the electronic device comprises a mobile device configured to communicate with the online query processor when communicating with a wireless network.
5. The method according to claim 1, wherein, initiating the processing comprises: sending data associated with the first input and data associated with the second input to the query processor.
6. The method according to claim 1, further comprising: providing an output to indicate that the session between the electronic device and the query processor is available.
7. The method according to claim 1, further comprising: providing the complete query to a search service.
8. The method according to claim 1, further comprising: using the query for one or more non-search actions.
9. The method according to claim 8, wherein, the one or more non-search actions comprise initiating communication with another user.
10. The method according to claim 8, wherein, the one or more non-search actions comprise setting a reminder.
11. The method according to any one of claims 1-10, wherein, the second modality is a voice modality, and establishing a session in a preemptive manner between the electronic device and the query processor comprises: if the electronic device is associated with an online service, establishing an online session in a preemptive manner between the electronic device and an online speech-to-text conversion processor; and If the electronic device is not associated with the online service, a local session is established preemptively between the electronic device and a local speech-to-text conversion processor.
12. An electronic device, comprising: a sensor; a memory storing instructions; one or more processors configured to execute the instructions for: receiving a first input in a first modality of a multimodal interface associated with the electronic device; determining that a context meets criteria based on one or more signals from the sensor, wherein determining that the context meets the criteria includes: determining the context based on the one or more signals from the sensor, wherein the sensor is at least one of a position coordinate sensor, an accelerometer, a thermometer, a gyroscope, and a light sensor; and in response to receiving the first input and determining that the context meets the criteria: establishing a session preemptively with a query processor configured to process an input received in a second modality of the multimodal interface; receiving a second input in the second modality of the multimodal interface; initiating, within the session, processing of at least a portion of the second input at the query processor; and constructing a complete query based on an output from the query processor.
13. A method, comprising: determining that a context of an electronic device meets criteria when a multimodal interface of the electronic device is in a first modality; wherein determining that the context of the electronic device meets the criteria is based on one or more signals from a sensor of the electronic device, and wherein the sensor is at least one of a position coordinate sensor, an accelerometer, a thermometer, a gyroscope, and a light sensor and is a sensor other than a microphone of the electronic device; and in response to determining that the context of the electronic device meets the criteria, by the electronic device: establishing a session preemptively between the electronic device and a query processor configured to process an input received in a second modality of the multimodal interface; receiving a second modality input in the second modality of the multimodal interface; initiating, within the session, processing of at least a portion of the second modality input at the query processor; and constructing a complete query based on an output from the query processor.
14. The method according to claim 13, further comprising: receiving a first modality input in the first modality of the multimodal interface, and wherein constructing the complete query includes: combining the output with at least a portion of the first modality input.
15. The method according to claim 13, wherein, the query processor is an online query processor, and wherein the electronic device includes a mobile device configured to communicate with the online query processor when communicating with a wireless network.
16. The method according to any one of claims 13-15, further comprising: providing an output to indicate that the session between the electronic device and the query processor is available.
17. An electronic device, comprising: a sensor; a microphone; a memory storing instructions; One or more processors configured to execute instructions for: Determining that a context meets criteria based on one or more signals from the sensors and when a multimodal interface of an electronic device is in a text mode, wherein determining that the context meets the criteria includes: Determining the context based on one or more signals from the sensors, wherein the sensors are at least one of a position coordinate sensor, an accelerometer, a thermometer, a gyroscope, and a light sensor; and In response to determining that the context meets the criteria: Establishing a speech-to-text conversion session in a preemptive manner between the electronic device and a speech-to-text conversion processor, the speech-to-text conversion session for processing speech input received via a microphone in a speech mode of the multimodal interface; Providing an output to indicate that the speech-to-text conversion session is available; Receiving speech input; Initiating, within the session, processing of at least a portion of the speech input at the speech-to-text conversion processor; and Constructing a complete query based on an output from the speech-to-text conversion processor.
18. The apparatus according to claim 17, wherein, When constructing the complete query based on an output from the speech-to-text processor, the one or more processors further combine the output with text input received in the text mode.
19. A method, comprising: Receiving a first input in a first mode of a multimodal interface associated with an electronic device; In response to receiving the first input, determining that criteria are met, wherein the criteria include that the first input matches a specific language; In response to determining that the criteria are met, establishing a session in a preemptive manner between the electronic device and a query processor configured to process input received in a second mode of the multimodal interface; Receiving a second input in the second mode of the multimodal interface; Initiating, within the session, processing of at least a portion of the second input at the query processor; and Constructing a complete query based on an output from the query processor, wherein the output from the query processor includes multiple candidate interpretations of the second input, and wherein constructing the complete query includes: Ranking the multiple candidate interpretations at least partially based on the first input.
20. The method according to claim 19, wherein determining that the criteria are met includes: Determining that the first input meets the criteria.
21. The method according to claim 20, wherein, The criteria further include: the first input meets a character count threshold.
22. The method according to claim 20, wherein, The criteria further include: the first input meets a word count threshold.
23. The method according to claim 19, wherein, The query processor is an online query processor, and wherein the electronic device includes a mobile device configured to communicate with the online query processor when communicating with a wireless network.
24. The method according to claim 19, wherein, Initiating the processing includes: Send data associated with the first input and data associated with the second input to the query processor.
25. The method according to any one of claims 19-24, further comprising: Providing an output to indicate that the session between the electronic device and the query processor is available.
Citation Information
Patent Citations
Method and apparatus for improved text input
WO2008032169A2