Voice Command Image Search Automation for API-Less Applications
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition applications lack the ability to automate tasks in target applications that do not expose necessary APIs, particularly in cloud-hosted environments like EHR systems, limiting their functionality.
Innovation Solution
Systems and methods that utilize image searching with voice recognition commands, allowing users to select icons or areas on a computer screen via voice commands by associating voice commands with search images, enabling automation of tasks without relying on traditional APIs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional API-based voice command interfaces are used, then automation functionality is achieved in applications that expose APIs, but functionality is lost in applications without API access such as cloud-hosted EHR systems
Solution Approach 1:
The patent introduces an intermediary component that captures screen images and performs image recognition to identify UI elements. This mediator layer sits between the voice recognition system and the target application, translating voice commands into coordinate-based actions without requiring direct API access to the application. The intermediary handles the complexity of cross-application automation by providing a universal interface that works with any application visible on the screen.
Solution Approach 2:
The patent replaces the traditional mechanical/API-based interaction system with an optical recognition system. Instead of using application programming interfaces to communicate with target applications, the system uses image capture and pattern recognition to identify UI elements and simulate user actions. This substitution allows voice commands to control applications that would otherwise be inaccessible to automation tools.
2Ease of operation
If cloud-hosted virtual applications are used, then accessibility and deployment are improved, but direct API access for automation is limited or unavailable
Solution Approach 1:
The patent creates a visual copy or representation of the application interface by capturing screen images. Instead of interacting with the actual application through its native APIs, the system works with image copies of the UI that can be analyzed and manipulated. This copying approach allows automation to function on cloud-hosted applications by treating their visual representation as the interface to control, bypassing the need for direct application programming access.
3Adaptability or versatility
If image searching with voice commands is implemented, then automation capability is extended to applications without APIs, but system complexity and processing requirements increase
Solution Approach 1:
The patent performs preliminary actions by pre-processing and analyzing screen images to identify and catalog UI elements before voice commands are issued. The system captures screenshots, identifies clickable elements, and stores their visual characteristics and coordinates in advance. When a voice command is received, the system queries this pre-processed data rather than performing full image analysis in real-time, significantly reducing computational resources required during actual voice interaction.
Data Source
AI summary
Embodiments described herein include systems and methods for using image searching with voice recognition commands. Embodiments of a method may include providing a user interface via a target application and receiving a user selection of an area on the user interface by a user, the area including a search image. Embodiments may also include receiving an associated voice command and associating, by the computing device, the associated voice command with the search image.


