Multimodal Text Input Module for Mobile Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mobile devices face cumbersome text input limitations, especially with foreign languages and special characters, as existing methods require tedious keyboard input or have limitations in speech recognition and camera-based OCR, restricting their applicability to specific applications.
Innovation Solution
A multimodal text input method and module that allows text input via both keyboard and camera modes, using optical character recognition (OCR) to recognize text from images, with a graphical user interface that adapts between modes, enabling seamless integration with existing applications and potentially replacing the standard keyboard module.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If keyboard input is used for text input on mobile devices, then text input functionality is provided, but text input becomes cumbersome especially for foreign languages and special characters
Solution Approach 1:
The keyboard module is enhanced to provide multiple input modes (keyboard input, camera-based OCR input, speech recognition input) within a single integrated module. This allows the same module to handle different input methods seamlessly, making it universal for various text input needs including foreign languages and special characters without requiring separate applications or modules.
2Adaptability or versatility
If camera-based OCR is implemented for text input, then text input for foreign languages is improved, but the module cannot replace the standard keyboard module due to application compatibility requirements
Solution Approach 1:
The camera-based OCR functionality and keyboard input functionality are merged into a single integrated keyboard module. The module detects whether an application is using it as a standard keyboard module and automatically adapts its behavior, providing camera-based input when appropriate while maintaining compatibility with applications that expect standard keyboard module behavior.
Solution Approach 2:
The keyboard module dynamically adapts its behavior based on the application context. It detects whether it is being used as a standard keyboard module and adjusts its functionality accordingly, enabling camera-based text input when the application supports it while maintaining standard keyboard functionality when required for compatibility.
3Adaptability or versatility
If multiple text input modes are integrated in one module, then text input versatility is improved, but the module size and complexity increase
Solution Approach 1:
The integrated keyboard module is segmented into distinct functional components: keyboard input processing, camera-based OCR input processing, and speech recognition input processing. Each component operates independently but can be activated based on user selection or application requirements, allowing the module to provide multiple input modes while maintaining a manageable internal structure.
Data Source
Figure 1a~1b
Figure 2a~2b
Figure 3a~3b
AI summary
The present invention relates to a method and a module for a multimodal text input in a mobile device (1) via a keyboard or in a camera mode by holding the camera of the mobile device (1) over a written text, such that an image is taken of the written text and the written text is recognized, wherein the input text is output to an application requesting the input text, the method comprising the following steps: a) activating a keyboard mode; b) providing an A-Z-keyboard in a first field (4) for text input; c) activating the camera mode; d) capturing the image of the written text and displaying the captured image with the written text in a second field (5) of a display (2) of the mobile device (1); e) converting the captured image to character text by optical character recognition (OCR) and displaying the recognized character text on the display (2); outputting a selected character as the input text to the application upon a selection of the character on the A-Z-keyboard, or outputting a selected part of the recognized character text as the input text to the application upon a selection of the part of the recognized character text; wherein the respective selection takes place by a single keypress or control command, or by a single gesture.