Wearable Intent Inference for Display-Free Voice Assistance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing devices, such as desktop computers, lack the ability to capture relevant environmental images and audio due to their fixed positioning, limiting the quality and type of computer-implemented services they can provide to users.
Innovation Solution
A display-free body wearable computing device is worn by a user to capture audio and image data, inferring user intent through a large language model and sensors to provide computer-implemented services, including image capture and command execution based on user speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a desktop computer is used, then the device structure is simple and stable, but the ability to capture relevant environmental images and audio is limited due to fixed positioning
Solution Approach 1:
The patent transitions from a static desktop computer to a dynamic wearable device that moves with the user. The computing device is worn on the user's body, enabling it to dynamically capture environmental images and audio from various positions and perspectives, thereby significantly improving adaptability to different environments while the wearable form factor keeps structural complexity manageable.
2Reliability
If a wearable device is used to capture environmental data, then the quality and relevance of services improve, but the device complexity increases
Solution Approach 1:
The patent divides the computing system into multiple functional components: wearable sensors for data collection, a processing unit for analyzing sensor data and inferring user intent, and a service execution component for providing computer-implemented services. This segmentation allows the system to achieve high service quality through coordinated specialized components while managing overall device complexity by distributing functions across modules.
Solution Approach 2:
The wearable computing device integrates multiple functions into a single platform: environmental sensing, audio capture, image capture, user intent inference through large language models, and execution of various computer-implemented services. This multi-functionality improves service reliability and quality while the integrated design prevents excessive complexity by consolidating diverse functions in one universal device.
3Measurement precision
If the device infers user intent using large language models and sensors, then the accuracy of command execution improves, but the processing time and energy consumption increase
Solution Approach 1:
The system performs preliminary actions by continuously collecting and pre-processing sensor data in the background, maintaining a ready state for rapid intent analysis. Environmental images and audio data are captured and pre-analyzed by sensors before user interaction occurs, so when the user speaks or interacts, the system can quickly process the intent using large language models without significant delay, thus improving measurement precision while minimizing time loss.
Data Source
AI summary
Methods and systems for providing assistance to users of display free body wearable computing devices are disclosed. The method may include identifying that a user of a display free body wearable computing device is speaking. The method may also include inferring whether at least one other person is in a detection range of the user. In an instance where no other persons are inferred as being in the detection range, a large language model may be prompted using an intention analysis prompt and a transcription of the speaking by the user to obtain an assistance request outcome. The assistance request outcome may indicate that the speaking may include a question and/or command directed to the display free body wearable computing device. The display free body wearable computing device may subsequently provide computer-implemented services to the user based at least in part on the transcription of the speaking by the user.


