Voice UI Control via Knowledge Graph Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices struggle to efficiently allow users to access UI elements using voice-based interactions, particularly due to the need for specific voice commands and limitations in accessing sub-functionality or sub-pages not displayed on the screen.

Innovation Solution

The method involves generating a knowledge graph based on characteristics of UI elements, such as position, function, and appearance, to predict natural language utterances. This allows users to access UI elements using natural language voice inputs, enabling access to both displayed and non-displayed application features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If specific voice commands are used to access UI elements, then the electronic device can be controlled remotely, but the user must memorize unnatural phrases and remain close to the device to read UI element names or numbers

Engineering Contradiction:
Improveease of accessing UI elementsVSAvoidcomplexity of voice commands
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical interaction methods (touchscreen, keyboard, remote controller) with voice-based interaction. The system captures voice inputs, converts them to text, and uses natural language processing to interpret user intent, thereby substituting physical interaction mechanisms with acoustic field-based interaction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary layer consisting of the knowledge graph and natural language processing system. This intermediary translates natural language voice inputs into actionable commands by matching them with predicted utterances derived from UI element characteristics, bridging the gap between user speech and device control without requiring users to memorize specific commands.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If traditional remote control or touchscreen is used, then the interface can be accessed, but the method is unusual and lacks fine-grained navigation or selection capability

Engineering Contradiction:
Improveversatility of interaction methodVSAvoidease of navigation and selection
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements dynamic adaptation of the interaction system by generating predicted natural language utterances based on current UI element characteristics. The system dynamically updates the knowledge graph as UI elements change, allowing the voice recognition system to adapt to different screens, applications, and contexts without requiring reconfiguration or memorization of fixed commands.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter space of interaction from fixed numeric codes or predefined commands to continuous natural language expressions. By deriving predicted utterances from UI element characteristics such as position, type, and function, the system allows users to navigate and select elements using flexible, context-appropriate language rather than rigid parameter-based commands.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If voice recognition technology with specific phrases is implemented, then remote control capabilities are improved, but non-technical users find it difficult to use

Engineering Contradiction:
Improveautomation of device controlVSAvoidease of use for non-technical users
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent implements a self-service mechanism where the system automatically generates predicted natural language utterances based on UI element characteristics. The knowledge graph autonomously derives expected user phrases from element properties such as position, type, and function, eliminating the need for users to learn or memorize commands. The system serves itself by continuously adapting its expected input based on the current interface state.

Inventive Principle:
Principle #25Self-service

4Measurement precision

If users must remain close to the device to read UI element names, then accurate selection is possible, but this is not always practicable

Engineering Contradiction:
Improveprecision of UI element identificationVSAvoidpracticability of interaction
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces visual reading and physical proximity requirements with acoustic field-based interaction. By capturing voice inputs and converting them to text through speech recognition, the system allows users to identify and access UI elements from a distance without needing to visually read element names or numbers on the screen.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12230261B2Electronic apparatus and method for controlling user interface elements by voice
Publication Date: 2025.02.18 SAMSUNG ELECTRONICS CO LTD
  • US12230261B2 patent drawing
  • US12230261B2 patent drawing
  • US12230261B2 patent drawing

AI summary

A method for controlling an electronic device is provided. The method includes identifying one or more user interface (UI) elements displayed on a screen of an electronic device, determining a characteristic(s) of one or more identified UI elements, generating a data base based on the characteristic of one or more identified UI elements, where the database comprises to predict NL utterances of one or more identified UI elements, where the NL utterances are predicted based on the at least one characteristic of one or more identified UI elements, receiving a voice input of a user of the electronic device, where the voice input comprises an utterance indicative of the at least one characteristic of one or more identified UI elements presented in the database, and automatically accessing UI element(s) of one or more UI elements in response to determining that the utterances of the received voice input from the user matches with the predicted NL utterances of one or more identified UI elements.