OCR Voice Command Navigation for Cloud EHR Text Fields

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems struggle to access and control multiple text fields in target applications, especially in cloud-hosted EHR environments, where not all applications expose necessary APIs for voice commands, limiting functionality and accessibility.

Innovation Solution

Implementing optical character recognition (OCR) to identify and map text fields within a user interface, allowing voice recognition systems to capture images, perform OCR, and navigate cursors to execute voice commands across multiple text fields, even when only one is active.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional voice recognition APIs are used to control text fields, then voice commands can be executed in applications that expose API functionality, but voice commands cannot access or control text fields in applications that do not provide necessary APIs (such as cloud-hosted EHR applications)

Engineering Contradiction:
Improvevoice command compatibilityVSAvoidsystem architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system that acts as a bridge between voice recognition commands and the target application interface. This intermediary captures screen images, performs OCR to extract text and coordinate information, and translates voice commands into simulated input actions (mouse clicks, keyboard input) that work with any application displaying its interface visually, regardless of API availability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the traditional API-based mechanical system with an optical-acoustic-mechanical system. Instead of using application programming interfaces, the system uses optical character recognition on screen displays combined with acoustic voice recognition to control applications through simulated physical input actions, making the control mechanism independent of the target application's internal architecture

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If only the active text field is accessed through traditional methods, then current voice commands can operate on the active field, but other text fields in the application remain inaccessible until manually activated

Engineering Contradiction:
Improvetext field accessibilityVSAvoidfield navigation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent adds a spatial dimension to text field access by using screen coordinate mapping. Instead of relying on application-program-defined field access sequences, the system captures the visual layout of all text fields on screen, maps their coordinates, and allows voice commands to navigate directly to any field's location, enabling parallel access to multiple fields simultaneously rather than sequential activation

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If cloud-hosted EHR applications are used to improve accessibility and deployment, then applications can be delivered through virtualization platforms, but traditional API access methods become unavailable or limited

Engineering Contradiction:
Improvedeployment flexibilityVSAvoidAPI accessibility
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent creates a visual copy of the application interface through screen capture, then performs OCR on this copy to extract text and coordinate information. This allows the system to interact with cloud-hosted applications through their visual representation rather than requiring direct API access to the underlying application logic, effectively bypassing the limitations of virtualized deployment environments

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12026456B2Systems and methods for using optical character recognition with voice recognition commands
Publication Date: 2024.07.02 DOLBEY
  • US12026456B2 patent drawing
  • US12026456B2 patent drawing
  • US12026456B2 patent drawing

AI summary

Systems and methods for using optical character recognition (OCR) with voice recognition commands are provided. Some embodiments include receiving a user interface that includes a text field, capturing an image of at least a portion of the user interface, and performing OCR on the image to identify the text field and a word in the user interface. Some embodiments include mapping a coordinate of the word in the text field, receiving a voice command that includes the word, and navigating, by the computing device, a cursor to the coordinate to execute the voice command.