Natural human-computer interaction for virtual personal assistant systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual personal assistant systems often require conventional human-computer interactions and do not fully model natural human interaction, leading to interference with other applications and inefficient speech recognition.

Innovation Solution

Implementing audio distortion techniques to generate multiple semantically distinct variations of user input, combined with eye tracking and engagement modeling, to enhance speech recognition accuracy and facilitate more natural interactions with virtual personal assistants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a humanlike avatar is displayed to facilitate natural interaction, then the ease of operation is improved, but the avatar may interfere with use of other applications and occupy significant display space

Engineering Contradiction:
Improveease of interactionVSAvoiddisplay space
Core Design Contradiction:
Ease of operationVSArea of stationary object

Solution Approach 1:

The avatar's display properties are dynamically adjusted based on user engagement detection. When engagement is detected, the avatar becomes more prominent; when not engaged, the avatar reduces visibility or moves to minimize display space occupation, allowing other applications to use the display area.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Different regions of the display are allocated differently based on avatar state. The avatar occupies significant space only in specific local areas when needed for interaction, while other regions remain available for applications, creating a dynamic spatial allocation strategy.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If conventional speech recognition systems are used, then the device complexity is reduced, but the speech recognition accuracy deteriorates in noisy environments

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by capturing audio input and generating distorted variations before the actual speech recognition occurs. Multiple versions of the audio signal are prepared in advance through controlled distortion to handle noise and improve recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The audio signal parameters are changed by generating multiple distorted variations of the audio input. These distortions modify characteristics like timing, amplitude, or frequency to create alternative representations that can improve speech recognition accuracy in challenging acoustic environments.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If the avatar is made prominent to improve interaction, then the ease of operation is improved, but the reliability deteriorates when user engagement is not detected

Engineering Contradiction:
Improveease of interactionVSAvoidinteraction reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system continuously monitors user engagement through eye tracking and provides feedback by adjusting the avatar's visibility and interaction state. When engagement is detected, the avatar becomes interactive; when not detected, it reduces prominence, preventing spurious interactions and improving overall system reliability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12399560B2Natural human-computer interaction for virtual personal assistant systems
Publication Date: 2025.08.26 INTEL CORP
  • US12399560B2 patent drawing
  • US12399560B2 patent drawing
  • US12399560B2 patent drawing

AI summary

Technologies for natural language interactions with virtual personal assistant systems include a computing device configured to capture audio input, distort the audio input to produce a number of distorted audio variations, and perform speech recognition on the audio input and the distorted audio variants. The computing device selects a result from a large number of potential speech recognition results based on contextual information. The computing device may measure a user's engagement level by using an eye tracking sensor to determine whether the user is visually focused on an avatar rendered by the virtual personal assistant. The avatar may be rendered in a disengaged state, a ready state, or an engaged state based on the user engagement level. The avatar may be rendered as semitransparent in the disengaged state, and the transparency may be reduced in the ready state or the engaged state. Other embodiments are described and claimed.