Display-Dependent Response Generation for Lower Voice Assistant Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Excessive network transmissions and inefficient bandwidth utilization occur when computing devices process network traffic data, particularly in voice-activated environments where both audio and visual responses are generated, leading to resource consumption and delayed responses.
Innovation Solution
A data processing system that generates responses based on client device configuration, such as the state of the display, to reduce resource consumption by omitting visual components if the display is OFF, thereby optimizing processor, battery, and bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If both audio and video responses are generated for voice-activated requests, then the response is more comprehensive and informative, but network bandwidth utilization increases and processing time is extended
Solution Approach 1:
The system dynamically changes the response format parameters based on the display state. When the display is OFF, the response is generated as audio-only; when the display is ON, both audio and video responses are generated. This parameter adaptation resolves the contradiction by matching the response comprehensiveness to the actual user reception capability.
Solution Approach 2:
The response generation system transitions from a static approach (always generating both audio and video) to a dynamic approach where the response format is adjusted in real-time based on display state. This dynamic adaptation allows the system to optimize bandwidth utilization while maintaining response completeness when needed.
2Loss of information
If both audio and video responses are generated and transmitted, then the user receives more information, but processor utilization and response time are negatively impacted
Solution Approach 1:
The system changes the response generation parameters based on display state. When display is OFF, only audio responses are generated, reducing processor workload and improving response efficiency. When display is ON, both audio and video are generated to maximize information delivery. This parameter adaptation resolves the contradiction between information delivery and response efficiency.
3Reliability
If audio and video responses are generated regardless of display state, then the system maintains consistent response capability, but resource consumption increases unnecessarily
Solution Approach 1:
The system transitions from a static response generation approach (always generating both audio and video) to a dynamic approach that adapts to display state. This dynamic adjustment maintains reliability by ensuring appropriate responses are always generated while reducing energy consumption by avoiding unnecessary video generation when the display is OFF.
Solution Approach 2:
The system extracts only the necessary response components based on display state. When the display is OFF, the video component is extracted/removed from the response generation process, leaving only audio. This extraction approach maintains response capability consistency for the available output modes while reducing processor and battery consumption.
4Loss of information
If video data is included in all responses, then visual information is always available, but network bandwidth is wasted when display is OFF
Solution Approach 1:
The system changes the response composition parameters based on display state. When display is OFF, the video data parameter is set to excluded, transmitting only audio data. When display is ON, video data is included. This parameter change resolves the contradiction by ensuring visual information is available when needed while preventing bandwidth waste when the display is unavailable.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure is generally related to a data processing system to process data packets in a voice activated computer network environment. The data processing system can improve the efficiency of the network by generating non-video data responses to voice commands received from a client device if a display associated with a client device is in an OFF state. A digital assistant application executed on the client device can send to the data processing system client device configuration data, which includes the state of the display device, among status data of other components of the client device. The data processing system can receive a current volume of speakers associated with the client device, and set a volume level for the client device based on the current volume level and a minimum response volume level at the client device.