Windshield Multimodal Interface for Precise Gaze and Speech Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automobile user interfaces pose a risk of distraction and collision due to the need for drivers to take their eyes off the road and operate physical buttons, and speech interfaces are cumbersome for precise input, especially when describing vague objects.
Innovation Solution
A multimodal user interface that accepts input via speech and geometric modalities like gaze detection and gesture recognition, allowing users to indicate positions and tasks without physical interaction, using models to determine precise inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If physical buttons and conventional interfaces are used, then precise input can be obtained, but driver distraction and collision risk increase
Solution Approach 1:
The patent replaces mechanical button interfaces with gaze detection and gesture recognition systems. The gaze detection facility uses cameras to track eye position and determine fixation points on the windshield display, while gesture recognition uses computer vision to interpret hand movements. This substitution eliminates the need for physical interaction, allowing drivers to input commands without taking their eyes off the road, thus resolving the contradiction between input precision and safety.
2Reliability
If speech interfaces are used, then driver eyes remain on the road, but precise input becomes cumbersome
Solution Approach 1:
The patent merges multiple input modalities (gaze detection, gesture recognition, and speech recognition) into a unified multimodal interface. The system can process speech commands while simultaneously using gaze and gesture data to disambiguate vague references. For example, when a driver says 'move that information here,' the system combines the speech command with gaze fixation location and hand gesture direction to precisely determine which information to move and where to place it, thereby achieving both safety and precision.
3Measurement precision
If multiple input modalities are integrated, then input precision improves, but system complexity increases
Solution Approach 1:
The patent segments the multimodal input system into distinct functional modules: a gaze detection facility with camera and tracking algorithms, a gesture recognition facility with computer vision processing, a speech recognition facility, and a task facility that integrates all inputs. Each module operates independently but communicates through standardized interfaces. This modular segmentation allows the system to achieve high input precision through multiple modalities while managing complexity through clear separation of concerns and independent module development.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Some embodiments described herein relate to a multimodal user interface for use in an automobile. The multimodal user interface may display information on a windshield of the automobile, such as by projecting information on the windshield, and may accept input from a user via multiple modalities, which may include a speech interface as well as other interfaces. The other interfaces may include interfaces allowing a user to provide geometric input by indicating an angle. In some embodiments, a user may define a task to be performed using multiple different input modalities. For example, the user may provide via the speech interface speech input describing a task that the user is requesting be performed, and may provide via one or more other interfaces geometric parameters regarding the task. The multimodal user interface may determine the task and the geometric parameters from the inputs.