Display Voice Control with UI Text Fallback for Third-Party Apps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Smart voice systems struggle to recognize and execute voice commands for third-party applications and menu options due to lack of semantic configuration, leading to incomplete or incorrect actions in complex scenarios.

Innovation Solution

A display apparatus equipped with a detector to acquire voice information, a controller to extract keywords, and a method to traverse a configuration library for matching actions, or recognize text and layout information to generate control instructions, enabling execution of user commands even for unconfigured applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the smart voice system uses a configuration library with pre-defined semantic functions, then recognition accuracy for built-in applications is improved, but adaptability to third-party applications and menu options deteriorates

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidcompatibility with third-party applications
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary component that captures text information and layout information from the display interface. This intermediary layer bridges the gap between the voice input system and the application control system, enabling the voice system to adapt to third-party applications without requiring pre-configuration of semantic functions for each application.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by pre-obtaining text information and layout information of the display interface before voice recognition processing. This allows the system to prepare the necessary contextual information in advance, enabling faster and more accurate matching of voice commands to interface elements without requiring real-time configuration.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the system configures semantic functions for all possible applications and menu options, then voice command recognition accuracy is improved, but device complexity and configuration workload increase

Engineering Contradiction:
Improvevoice command recognition accuracyVSAvoidconfiguration workload
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system implements self-service by automatically obtaining text information and layout information from the display interface without requiring manual configuration. The voice recognition system autonomously adapts to the current interface state by retrieving necessary information dynamically, eliminating the need for developers to pre-configure semantic functions for each application or menu option.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transitions from a static configuration approach to a dynamic one where text information and layout information are obtained in real-time based on the current display state. This dynamic adaptation allows the voice recognition system to adjust to different applications and interface configurations without requiring complex pre-setup procedures.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If the system retrieves text information and layout information from the display interface, then adaptability to complex interfaces is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveinterface adaptation capabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary retrieval of text information and layout information from the display interface before voice command processing. By obtaining this contextual information in advance, the system reduces the processing time required during actual voice command execution, as the necessary interface data is already available for rapid matching and recognition.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12437758B2Display apparatus and a voice control method
Publication Date: 2025.10.07 HISENSE VISUAL TECH CO LTD
  • US12437758B2 patent drawing
  • US12437758B2 patent drawing
  • US12437758B2 patent drawing

AI summary

Some embodiments of the present application disclose a display apparatus and a voice control method for the display apparatus. The display apparatus comprises a display, a detector and a controller. The display is configured to present a user interface, and the detector is configured to acquire user voice information; and the controller is configured to cause the display apparatus to perform: acquiring voice information inputted from a user; in response to the voice information, extracting at least one keyword from the voice information; traversing action items in a configuration library; in response to determining that no action item in the configuration library matches the at least one keyword, obtaining text information of the user interface on the display to in order to determine an control instruction according to the text information.