OCR Voice Command Navigation for Cloud EHR Text Fields
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems struggle to access and control multiple text fields in target applications, especially in cloud-hosted EHR environments, where not all applications expose necessary APIs for voice commands, limiting functionality and accessibility.
Innovation Solution
Implementing optical character recognition (OCR) to identify and map text fields within a user interface, allowing voice recognition systems to capture images, perform OCR, and navigate cursors to execute voice commands across multiple text fields, even when only one is active.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional voice recognition APIs are used to control text fields, then voice commands can be executed in applications that expose API functionality, but voice commands cannot access or control text fields in applications that do not provide necessary APIs (such as cloud-hosted EHR applications)
Solution Approach 1:
The patent introduces an intermediary system that acts as a bridge between voice recognition commands and the target application interface. This intermediary captures screen images, performs OCR to extract text and coordinate information, and translates voice commands into simulated input actions (mouse clicks, keyboard input) that work with any application displaying its interface visually, regardless of API availability
Solution Approach 2:
The patent replaces the traditional API-based mechanical system with an optical-acoustic-mechanical system. Instead of using application programming interfaces, the system uses optical character recognition on screen displays combined with acoustic voice recognition to control applications through simulated physical input actions, making the control mechanism independent of the target application's internal architecture
2Productivity
If only the active text field is accessed through traditional methods, then current voice commands can operate on the active field, but other text fields in the application remain inaccessible until manually activated
Solution Approach 1:
The patent adds a spatial dimension to text field access by using screen coordinate mapping. Instead of relying on application-program-defined field access sequences, the system captures the visual layout of all text fields on screen, maps their coordinates, and allows voice commands to navigate directly to any field's location, enabling parallel access to multiple fields simultaneously rather than sequential activation
3Adaptability or versatility
If cloud-hosted EHR applications are used to improve accessibility and deployment, then applications can be delivered through virtualization platforms, but traditional API access methods become unavailable or limited
Solution Approach 1:
The patent creates a visual copy of the application interface through screen capture, then performs OCR on this copy to extract text and coordinate information. This allows the system to interact with cloud-hosted applications through their visual representation rather than requiring direct API access to the underlying application logic, effectively bypassing the limitations of virtualized deployment environments
Data Source
AI summary
Systems and methods for using optical character recognition (OCR) with voice recognition commands are provided. Some embodiments include receiving a user interface that includes a text field, capturing an image of at least a portion of the user interface, and performing OCR on the image to identify the text field and a word in the user interface. Some embodiments include mapping a coordinate of the word in the text field, receiving a voice command that includes the word, and navigating, by the computing device, a cursor to the coordinate to execute the voice command.


