Screen Element Analysis for Third-Party App Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in navigating and controlling applications without using the application programming interface (API), particularly for third-party applications, which requires cumbersome remote control operations due to the inability to identify and execute commands directly.
Innovation Solution
An electronic device method that receives user input, determines user intent, identifies UI elements on the screen, and executes commands by performing tasks corresponding to sub-goals, allowing navigation and control of applications without relying on API, using automatic speech recognition and natural-language understanding for voice inputs and dynamic sub-goal determination based on screen changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If API-based control methods are used for third-party applications, then control precision is improved, but device complexity increases due to requiring multiple APIs and integration efforts
Solution Approach 1:
The patent introduces an intermediary system that captures screen images and uses AI models to identify UI elements and infer their functions. This mediator translates visual screen information into control commands, eliminating the need for direct API integration with third-party applications while maintaining control precision.
Solution Approach 2:
The patent replaces the mechanical/API-based control system with a vision-based system. Instead of using application programming interfaces to communicate with third-party apps, the system uses screen capture, image processing, and AI inference to identify and control UI elements, simplifying the control architecture.
2Ease of operation
If remote control operations are simplified, then ease of operation is improved, but control reliability deteriorates due to inability to execute commands directly
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously captures screen images before and after executing control commands. By comparing the screen states and validating whether the intended UI element changes occur, the system ensures reliable command execution while maintaining simplified remote control operations.
Solution Approach 2:
The patent performs preliminary actions by capturing the current screen state before executing a control command and capturing the screen state after execution. This preliminary and subsequent screen capture validates that the control command was executed correctly, ensuring reliability without complicating the user interface.
3Adaptability or versatility
If screen analysis is performed to identify UI elements, then adaptability is improved for third-party applications, but processing time increases
Solution Approach 1:
The patent segments the screen into multiple regions and processes each region independently to identify UI elements. By dividing the screen analysis task into smaller regional segments, the system reduces the overall processing time while maintaining comprehensive UI element detection across the entire screen, thus improving efficiency without sacrificing adaptability.
Data Source
AI summary
Provided are an electronic device for navigating an application screen, and an operating method thereof. The method may include receiving a user input; determining, based on the user input, a user intent for controlling the electronic device; determining a command for performing a control operation corresponding to the user intent as a goal; identifying elements of a user interface on the screen of the application; determining, based on the user intent and the elements of the user interface, at least one sub-goal for executing the command; and executing the command by performing at least one task corresponding to the at least one sub-goal, wherein the at least one sub-goal is changeable based on a validation of an operation of navigating the application for executing the command, and the at least one task includes units of action for navigating the application.


