Multimodal disambiguation of speech-supported input
Patent Information
- Application Number
- DE102016109521
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-06-17
- Filing Date
- 2016-05-24
- Publication Date
- 2026-08-27
- Estimated Expiration
- 2036-05-24
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND Electronic devices (e.g., tablets, smartphones, smartwatches, laptops, PCs, etc.) allow users to provide voice input, such as voice commands, spoken text, and the like. Traditionally, when a voice input is entered into a currently active application (e.g., a virtual assistant, a speech-to-text application, etc.), the input is processed to identify words for text input or commands, depending on the currently active application. In some cases, multiple voice-enabled tasks occur simultaneously, and it is difficult for the system to determine a suitable target application for the voice input. Such voice inputs are traditionally handled sequentially; for example, a user must manually switch from one currently active voice-enabled application to another. From US patent application 2014 / 0184550A1, a system and a method are disclosed that use gaze direction information to improve interactions with objects in an environment. From US patent application 2012 / 0016678A1, an intelligent automated assistance system is disclosed that interacts with a user using natural language and, when necessary, calls upon external services to obtain information or perform actions. SUMMARY The object of the present invention is to provide an improved voice input for a voice-enabled application. This problem is solved by the subject matter of main claim 1 and dependent claims 10 and 18, which define the present invention. Preferred embodiments of the present invention are the subject of the dependent claims. In short, one aspect provides a procedure that includes the following steps: receiving a speech input at an audio receiver of a device; selecting, using a processor of a device, an active speech-enabled target application for the speech input from a variety of active speech-enabled target applications that are capable of processing speech input; and providing, using a processor of the device, the speech input to the selected active speech-enabled target application. Another aspect provides an electronic device comprising: an audio receiver; a display device; a processor operationally coupled with the audio receiver and the display device; and a memory that stores instructions executable by the processor to: receive speech input at the audio receiver; select an active speech-enabled target application for speech input from a variety of active speech-enabled target applications capable of processing speech input; and provide the speech input to the selected active speech-enabled target application. Another aspect provides a product that includes: a storage device that stores code executable by a processor, wherein the code includes: code that receives speech input at an audio receiver of a device; code that selects an active speech-enabled target application for speech input from a variety of active speech-enabled target applications capable of processing speech input; and code that provides the speech input for the selected active speech-enabled target application. The foregoing is a summary and may therefore include simplifications, generalizations and missing details; consequently, the person skilled in the art will understand that the summary is purely explanatory and is in no way intended to be limiting. For a better understanding of the embodiments, along with their other features and advantages, reference is made to the following description in conjunction with the accompanying drawings. The scope of the invention is defined in the accompanying claims. BRIEF DESCRIPTION OF THE DIFFERENT VIEWS OF THE DRAWINGS Figure 1 shows an example of the circuitry of an electronic device. Figure 2 shows another example of the circuitry of an electronic device. Figure 3 shows an exemplary method for the multimodal disambiguation of speech-based text input. DETAILED DESCRIPTION It is self-evident that the components of the embodiments, as generally described herein and illustrated in the figures, can be arranged and designed in many different configurations in addition to the described exemplary embodiments. Therefore, the following more detailed description of the exemplary embodiments, as shown in the figures, is not intended to limit the scope of the claimed embodiments, but is merely representative of exemplary embodiments. Any reference in this entire description to "an embodiment" (or similar expressions) means that a particular feature, structure, or distinguishing characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the occurrence of phrases such as "in an embodiment" and the like at various points in this entire description does not necessarily always refer to the same embodiment. Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to facilitate a thorough understanding of the embodiments. However, those skilled in the art will recognize that the various embodiments can be implemented in practice without one or more of the specific details or with different methods, components, materials, and so on. In other cases, well-known structures, materials, or processes are not shown or described in detail for the sake of clarity. As described here, if multiple voice-activated tasks occur simultaneously, it is difficult for the system to determine a suitable target for voice input. Voice-activated tasks are conventionally handled sequentially rather than in parallel. This is primarily because voice-enabled applications are unable to coordinate voice input, requiring the user to manually switch between available voice-enabled applications. Accordingly, one embodiment allows multiple speech-enabled applications or programs to operate in parallel. In this embodiment, the use of multimodal inputs is employed to achieve coordinated handling of speech inputs, ensuring that different speech inputs are directed to a suitable application or task. One embodiment can select an active target resource, such as an application, a device, etc. That is, one embodiment can select an active target device, an active target application, a specific subsystem, or a combination thereof, to route the speech input. For example, if a user is typing text into a speech-to-text application and a messaging application is active at the same time, and a notification from the messaging application appears, the user might want to speak their response. Traditionally, the user is forced to explicitly close the speech-to-text application and open the messaging application to speak a response to the notification or message. However, through several mechanisms, one implementation allows the user to select the desired speech-enabled application for the spoken text. Exemplary mechanisms that allow the user to target speech input to the correct application include, but are not limited to, eye tracking, gestures, situational and contextual data, or even touch input. The use of trigger words or phrases can also be employed. However, trigger words, touch input, or other similar mechanisms do not force a switch between active speech-enabled applications. Rather, these mechanisms allow the speech input to be routed without requiring the user to turn off the active application, explicitly activate another, etc. Instead, the initial application, e.g.,In the example above, the voice input application does not provide voice input to respond to the speech-to-text messaging application. Instead, voice input is redirected to the second application, such as the speech-to-text messaging application, and the first application remains active, listening for additional audio to input into the voice input application. For example, the user might shift their gaze from an input window or area of the speech-to-text application to the notification window or area provided in the dialog box by the speech-to-text messaging application, respond to the notification, and then return their attention to the input window or area. Using an eye-tracking subsystem, one embodiment switches the speech input audio from the user to the speech-to-text messaging application and redirects this input so that it is not sent to the speech-to-text application. Thus, the original application remains active, and the user can temporarily switch their speech input to the alternative application with minimal effort. Alternatively, the user can shift their gaze from the input area or window to the notification area or window, speak a response, and then return their attention to the writing window. One implementation can thus use more than one additional input (in this example, touch input and eye-tracking input) to increase the accuracy of forwarding the speech input. One embodiment may use other types or modes of input to route the voice input to the appropriate application or task. For example, a user may provide a gesture input, captured by a camera, such as raising a hand to signal "stop" to the input window or area of the voice-enabled text input application, and then shift their gaze to the notification window or area to speak a response. The user may then return their gaze to the input window, optionally reinforcing this movement with another gesture. Situation and context perception can be implemented using data that specifies a predetermined situation or context for speech input. For example, discontinuities in a flow of words or actions can serve as a predetermined input pattern linked to a context change, such as switching the target application to which the speech input data should be directed. Returning to the previously mentioned example, the user might interrupt dictating a message mid-sentence, for example, with completely different words or phrases (such as "Close!"), the target of which should be a different speech-enabled program, application, or task (e.g., a speech-enabled application that displays a notification dialog).Discontinuities could be subtle indicators, such as hesitation or sudden pausing, that increase the likelihood of a focus shift. This, in turn, can be coupled with additional data that enhances situational awareness, such as knowing what is in focus on the screen when the user speaks. By monitoring this data, an implementation can increase the probability or accuracy of correctly identifying a target for the spoken input. The illustrated embodiments are best understood with reference to the figures. The following description is intended to be purely exemplary and depicts only certain embodiments. Although various other circuits, circuitry, or components can be used in information handling devices, with regard to the circuitry 100 of a smartphone and / or tablet, an example shown in Fig. 1 comprises a system-on-a-chip design, which is found, for example, in tablets or other mobile computing platforms. The software and the processor(s) are combined in a single chip 110. The processors include internal arithmetic units, registers, buffers, buses, I / O ports, etc., as is well known in engineering. Internal buses and the like depend on different manufacturers, but essentially all peripheral devices (120) can be attached to a single chip 110. The circuitry 100 combines the processor, memory control, and I / O control node all together in a single chip 110. Such systems 100 also typically do not use SATA, PCI, or LPC.Common interfaces include, for example, SDIO and I2C. There is one or more power management chips 130, e.g., a battery management unit (BMU), which manages the power, such as that supplied by a rechargeable battery 140, which can be recharged by connecting to a power source (not shown). In at least one form factor, a single chip, such as 110, is used to provide BIOS-like functionality and DRAM memory. System 100 typically includes one or more WWAN transceivers 150 and WLAN transceivers 160 for connecting to various networks, such as telecommunications networks and wireless internet devices, e.g., access points. It also usually includes devices 120, e.g., an audio receiver, such as a microphone, which works with a speech processing system to provide speech input and related data, as described further herein; a camera that captures image data and works with a gesture recognition machine; and so on. System 100 often includes a touchscreen 170 for data input and display / playback. System 100 also typically includes various storage devices, e.g., flash memory 180 and SDRAM 190. Figure 2 shows a block diagram of another example of the circuits, circuitry, or components of an information handling device. The example shown in Figure 2 may correspond to computer systems, such as the ThinkPad series of PCs sold by Lenovo (US) Inc. of Morrisville, NC, or to other devices. As can be seen from the present description, embodiments may include other features or only some of the features of the example shown in Figure 2. The example in Fig. 2 includes a so-called chipset 210 (a group of integrated circuits or chips that work together, chipsets) with an architecture that may vary depending on the manufacturer (e.g., Intel, AMD, ARM, etc.). Intel is a registered trademark of Intel Corporation in the United States and other countries. AMD is a registered trademark of Advanced Micro Devices, Inc. in the United States and other countries. ARM is an unregistered trademark of ARM Holdings plc in the United States and other countries. The architecture of the chipset 210 includes a core and memory control group 220 and an I / O control node 250, which exchanges information (e.g., data, signals, instructions, etc.) via a Direct Management Interface (DMI) 242 or a Link Controller 244. In Fig.2. The DMI 242 is a chip-to-chip interface (occasionally also referred to as a link between a "northbridge" and a "southbridge"). The core and memory control group 220 comprises one or more processors 222 (for example, single-core or multi-core) and a memory control node 226, which exchange information via a front-side bus (FSB) 224; it should be noted that the components of the group 220 can be integrated into a single chip, replacing the conventional "northbridge" architecture. One or more processors 222 comprise internal arithmetic units, registers, buffers, buses, I / O ports, etc., as is well known in engineering. In Fig. 2, the memory control node 226 provides an interface with the memory 240 (for example, to provide support for a type of RAM that can be referred to as "system memory" or "memory"). The memory control node 226 also includes a low-voltage differential signaling (LVDS) interface 232 for a display device 292 (e.g., a CRT, a flat panel display, a touchscreen, etc.). A block 238 includes certain technology that can be supported via the LVDS interface 232 (e.g., serial digital video, HDMI / DVI, display connector). The memory control node 226 also includes a PCI Express (PCI-E) interface 234, which can support discrete graphics 236. In Fig. 2, the I / O control node 250 includes a SATA interface 251 (for example, for HDDs, SSDs 280, etc.), a PCI-E interface 252 (for example, for wireless connections 282), a USB interface 253 (for example, for devices 284, such as a digitizer, keyboard, mice, cameras, telephones, microphones, storage media, other connected devices, etc.), a network interface 254 (for example, LAN), a GPIO interface 255, an LPC interface 270 (for ASICs 271, a TPM 272, a Super I / O 273, a Firmware Hub 274, BIOS support 275, and various types of memory 276, such as a ROM 277, a Flash 278, and an NVRAM 279), a power management interface 261, a clock interface 262, a Audio interface 263 (for example, for loudspeakers 294), a TCO interface 264, a system management bus interface 265 and SPI flash 266, which may include a BIOS 268 and boot code 290.The I / O control node 250 can include Gigabit Ethernet support. Upon power-up, the system may be configured to execute the boot code 290 for the BIOS 268, which is stored in the SPI flash 266, and subsequently processes data under the control of one or more operating systems and application software (such as those stored in system memory 240). An operating system may be stored in any of several locations and may be accessible, for example, according to the instructions of the BIOS 268. As described herein, a device may include fewer or more features than those shown in the system in Fig. 2. Information handling device circuits, such as those shown in Fig. 1 or Fig. 2, can be used in devices such as tablets, smartphones, laptop computers, or other personal computing devices in general, and / or other electronic devices to which users can provide voice input for various voice-enabled applications. Examples of voice-enabled applications include text-to-speech applications or voice-enabled applications in general, with specific examples being applications supported by voice input, such as note-taking or word processing applications; applications activated by voice commands, such as virtual assistants or navigation applications; or applications that support voice input as provided by other applications. Referring generally to Fig. 3, one embodiment assists a user by implementing a mechanism that directs a voice input to a suitable voice-enabled application. For example, when receiving a voice input at 301, e.g., at an audio receiver such as the microphone of a smartphone or a tablet computer, one embodiment at 302 specifies that more than one active target resource, e.g., a voice-enabled application, could receive and respond to the voice input. If no other voice-enabled application is active or available, the voice input can, of course, be directed to the only possible destination. Otherwise, one embodiment may require the user to select the correct voice-enabled application to which the voice input is directed. For example, when a notification of an incoming text message is received, the user might provide voice input to a voice-enabled note-taking application. To reply, the user might speak an input and intend for the input to be provided to the messaging application rather than the note-taking application. Thus, one embodiment of 302 specifies that a voice input, e.g., "Close!", received at 301, while both the note-taking application and the messaging application are capable of using voice input, is directed to one of the voice-enabled applications. With a conventional device, a user would be forced to disable the note-taking application, activate the messaging application, and then provide the "close" voice input to prevent the "close" command from being written into the note being created by the note-taking application. To streamline this process, one embodiment includes a processing capability that uses multimodal inputs to disambiguate the possible targets and select an appropriate target application. For example, an embodiment of 303 can select an active target speech-enabled application for speech input using various additional data sources or processing techniques that enable situation or context perception. For example, an additional data source, such as input data from an eye-tracking system, can be used to determine where the user is looking on the display while providing speech input. If this area of the display device coincides with a displayed notification from the messaging application, an embodiment can select this application as the appropriate target application. As another example, an additional data source, such as input from a camera or a gesture recognition system, can be used to determine whether the user is performing a predetermined gesture while providing voice input. For instance, a predetermined hand gesture can be used to practically direct the voice input to one application or another. As another example, data from a touchscreen display can be used to determine where the user touched the display screen when providing voice input. This additional data input via alternative (i.e., non-speech) channels can be used by an embodiment at 303 when selecting a target application. One embodiment can also apply additional processing, for example, to the speech input itself, to select a suitable target application in 303. For instance, one embodiment can analyze the word(s) used in the speech output to determine whether the note-taking application or the messaging application is the appropriate target application for the speech input. It is evident from the above that more than one technique can be used to improve the accuracy of the selection in 303. Once a speech-enabled target application has been selected at 303, an embodiment at 304 provides the speech input to the selected active speech-enabled target application. Thus, the speech input is sent to one of the possible speech-enabled target applications. In one embodiment, the speech data can be temporarily buffered while the selection step at 303 is completed and then sent to the appropriate application at 304, preventing the speech input from reaching the wrong application. As a specific example, the selection at 303 can include obtaining an additional input, such as eye-tracking input. This eye-tracking input can be obtained while the user is providing the voice input, allowing an embodiment to determine that the voice input, e.g., "Close," is provided while the user is focused on a notification displayed by the messaging application. Thus, an embodiment at 303 can select the messaging application and provide the buffered voice data (or text or other output derived from the voice data or its processing) to the messaging application. The note-taking application can continue listening for the additional voice input; that is, it does not need to be closed or disabled to send the voice input data to the messaging application. As described herein, many additional data sources, such as gesture input, touch input, and / or speech input, can be used alone or in combination to enable selection at 303. In one embodiment, selection at 303 involves analyzing one or more words of speech input with an additional analysis of previous speech input, for example, to determine whether the words of speech input appear in the note-taking application without any context related to previously entered words. This can be combined with another input mode, such as eye-tracking data or gesture data, to increase the accuracy of the analysis that determines that the words of speech input are unrelated to the note-taking application.Therefore, if the speech input contains words that are determined with a lower probability or accuracy to be unrelated to an active speech input application, this probability or accuracy can be increased by analyzing eye-tracking or gesture data, e.g., whether the user's gaze is focused on a notification from a messaging service application when the unrelated speech input is received. One embodiment therefore represents a technical improvement in providing mechanisms by which received speech input can be routed to speech-enabled applications. In addition to reducing the time and complexity of conventional techniques, one embodiment is suitable for enabling the acceptance of speech input as a form of interface with the device, because users are no longer burdened with selecting which application should be considered active and receive the speech input. As those skilled in the art will understand, various aspects can be designed as a system, method, or device program product. Accordingly, the aspects can take the form of an embodiment consisting entirely of hardware or an embodiment with software, which is generally referred to here as a "circuit," "module," or "system." Furthermore, the aspects can take the form of a device program product, which is designed as one or more device-readable media containing device-readable program code. It should be noted that the various functions described here can be implemented using instructions stored on a device-readable storage medium, such as a non-signal storage device, which are executed by a processor. A storage device can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any suitable combination thereof. More specific examples of a storage medium would include: a portable computer disk, a hard disk, random access memory (RAM), programmable memory (ROM), erasable programmable memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only storage device (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.In the context of this publication, a storage device is not a signal, and the term ‘non-transient’ includes all media except signal media. Program code stored on a storage medium can be transferred using any suitable medium, including without limitation wireless, wired, fiber optic, RF, etc., or any suitable combination thereof. Program code for performing operations can be written in any combination of one or more programming languages. The program code can run entirely on a single device, partially on a single device, as a standalone software package, partially on one device and partially on another, or entirely on another device. In some cases, the devices can be connected via any type of connection or network, including a local area network (LAN) or a wide area network (WAN), or the connection can be established via other devices (for example, via the internet using an internet service provider), via wireless connections such as near-field communication (NFC), or via a wired connection such as a USB connection. This document describes exemplary embodiments with reference to the figures, which depict exemplary methods, devices, and program products according to various embodiments. It is understood that the actions and functionality can be implemented, at least partially, by program instructions. These program instructions can be provided to a processor of a device, a special information handling device, or another programmable data processing device to form a machine such that the instructions, executed by a processor of the device, implement the specified functions / actions. It should be noted that although specific blocks are used in the figures and a particular order of blocks is depicted, these are not restrictive examples. In certain contexts, two or more blocks can be combined, a block can be divided into two or more blocks, or certain blocks can be rearranged or repositioned as needed, since the examples explicitly shown are used for descriptive purposes only and are not to be interpreted as restrictive. As it is used here, the singular form ‘ein’ can be interpreted as encompassing the plural form ‘ein oder mehr’’ unless clearly indicated otherwise.
Claims
Method comprising the following steps: Receiving (301) a speech input at an audio receiver of a device; Selecting (303), using a processor of a device, an active speech-enabled target application for the speech input from a variety of active speech-enabled target applications that are capable of processing speech inputs; and Providing (304), using a processor of the device, the received speech input to the selected active speech-enabled target application. Method according to claim 1, wherein the selection (303) comprises analyzing one or more words of the speech input. The method of claim 1, further comprising receiving an additional input, wherein the selection includes using the additional input. Method according to claim 3, wherein the additional input comprises an eye-tracking input. Method according to claim 4, wherein the eye tracking input is linked to a predetermined area. Method according to claim 5, wherein the predetermined area comprises a display area occupied by one of the plurality of active speech-enabled target applications. Method according to claim 6, wherein the display area comprises a notification issued by one of the plurality of active speech-enabled target applications. Method according to claim 3, wherein the additional input is selected from the group consisting of a gesture input, a touch input and a voice input. Method according to claim 2, wherein the analysis comprises the analysis of a previous speech input. Electronic device comprising: an audio receiver; a display device; a processor operationally coupled with the audio receiver and the display device; and a memory storing instructions executable by the processor for: receiving (301) a speech input at the audio receiver; selecting (303) an active speech-enabled target application for speech input from a plurality of active speech-enabled target applications capable of processing speech input; and providing (304) the speech input to the selected active speech-enabled target application. Electronic device according to claim 10, wherein the selection (303) comprises analyzing one or more words of the speech input. Electronic device according to claim 10, wherein the instructions are further executable by the processor to obtain an additional input, wherein selecting (303) an active speech-enabled target application includes using the additional input. Electronic device according to claim 12, wherein the additional input comprises an eye-tracking input. Electronic device according to claim 13, wherein the eye tracking input is linked to a predetermined area. Electronic device according to claim 14, wherein the predetermined area comprises a display area of the display device occupied by one of the plurality of active speech-enabled target applications. Electronic device according to claim 15, wherein the display area comprises a notification issued by one of the plurality of active speech-enabled target applications. Electronic device according to claim 12, wherein the additional input is selected from the group consisting of a gesture input, a touch input and a voice input. Product comprising: a storage device that stores code executable by a processor, the code comprising: code that receives speech input at an audio receiver of a device (301); code that selects an active speech-enabled target application for speech input from a plurality of active speech-enabled target applications capable of processing speech input (303); and code that provides the speech input to the selected speech-enabled target application (304).
Citation Information
Patent Citations
Speech-Enabled Content Navigation And Control Of A Distributed Multimodal Browser
US20080255851A1
Intelligent Automated Assistant
US20120016678A1
Interactive spoken dialogue interface for collection of structured data
US20130339030A1
System and Method for Using Eye Gaze Information to Enhance Interactions
US20140184550A1
Displaying speech command input state information in a multimodal browser
US20140208210A1