Natural human-computer interaction for virtual personal assistant systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual personal assistant systems often require conventional human-computer interactions and do not fully model natural human interaction, leading to interference with other applications and inefficient speech recognition.
Innovation Solution
Implementing audio distortion techniques to generate multiple semantically distinct variations of user input, combined with eye tracking and engagement modeling, to enhance speech recognition accuracy and facilitate more natural interactions with virtual personal assistants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a humanlike avatar is displayed to facilitate natural interaction, then the ease of operation is improved, but the avatar may interfere with use of other applications and occupy significant display space
Solution Approach 1:
The avatar's display properties are dynamically adjusted based on user engagement detection. When engagement is detected, the avatar becomes more prominent; when not engaged, the avatar reduces visibility or moves to minimize display space occupation, allowing other applications to use the display area.
Solution Approach 2:
Different regions of the display are allocated differently based on avatar state. The avatar occupies significant space only in specific local areas when needed for interaction, while other regions remain available for applications, creating a dynamic spatial allocation strategy.
2Measurement precision
If conventional speech recognition systems are used, then the device complexity is reduced, but the speech recognition accuracy deteriorates in noisy environments
Solution Approach 1:
The system performs preliminary actions by capturing audio input and generating distorted variations before the actual speech recognition occurs. Multiple versions of the audio signal are prepared in advance through controlled distortion to handle noise and improve recognition accuracy.
Solution Approach 2:
The audio signal parameters are changed by generating multiple distorted variations of the audio input. These distortions modify characteristics like timing, amplitude, or frequency to create alternative representations that can improve speech recognition accuracy in challenging acoustic environments.
3Ease of operation
If the avatar is made prominent to improve interaction, then the ease of operation is improved, but the reliability deteriorates when user engagement is not detected
Solution Approach 1:
The system continuously monitors user engagement through eye tracking and provides feedback by adjusting the avatar's visibility and interaction state. When engagement is detected, the avatar becomes interactive; when not detected, it reduces prominence, preventing spurious interactions and improving overall system reliability.
Data Source
AI summary
Technologies for natural language interactions with virtual personal assistant systems include a computing device configured to capture audio input, distort the audio input to produce a number of distorted audio variations, and perform speech recognition on the audio input and the distorted audio variants. The computing device selects a result from a large number of potential speech recognition results based on contextual information. The computing device may measure a user's engagement level by using an eye tracking sensor to determine whether the user is visually focused on an avatar rendered by the virtual personal assistant. The avatar may be rendered in a disengaged state, a ready state, or an engaged state based on the user engagement level. The avatar may be rendered as semitransparent in the disengaged state, and the transparency may be reduced in the ready state or the engaged state. Other embodiments are described and claimed.


