Wearable Multimedia Device for Object Recognition and Health Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies lack efficient and user-friendly methods for personal assistance and health monitoring using wearable multimedia devices, particularly in providing information about items of interest, summarizing communications and events, and tracking nutritional intake.

Innovation Solution

A wearable multimedia device equipped with image and depth sensors, machine learning models, and a laser projector to provide personalized assistance by identifying items of interest, summarizing communications and events, and monitoring nutritional intake, using machine learning models to generate and present relevant data through a user interface on a surface or audio output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If wearable multimedia devices are equipped with multiple sensors and machine learning models to provide comprehensive personal assistance and health monitoring, then the functionality and accuracy of the device is improved, but the device complexity increases

Engineering Contradiction:
ImprovefunctionalityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The wearable multimedia device integrates multiple functions including object recognition, voice recognition, health monitoring, nutritional tracking, and communication summarization into a single device. The device uses image sensors, depth sensors, microphones, and machine learning models to perform diverse tasks such as identifying items of interest, detecting gestures, monitoring health metrics, and providing personalized assistance, thereby achieving multi-functionality that reduces the need for multiple separate devices

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines various sensing capabilities (image sensors, depth sensors, microphones) and processing functions (object recognition, voice recognition, health monitoring) into a unified wearable multimedia device. The device merges these components to work together seamlessly, with data from multiple sensors being processed by machine learning models to generate comprehensive outputs for personal assistance and health monitoring

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If the device processes image data and voice inputs through machine learning models to provide accurate information, then the measurement precision and reliability are improved, but the processing time and energy consumption increase

Engineering Contradiction:
ImproveaccuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The device performs preliminary processing of image data and voice inputs locally using on-device machine learning models to quickly identify objects, detect gestures, and extract key information before transmitting data to remote servers for more comprehensive analysis. This preliminary action reduces the time required for initial processing and enables real-time responses while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses an intermediary processing layer that combines local machine learning models with remote cloud computing resources. The device first processes data locally using embedded models to get immediate results, then uses cloud-based models for more complex analysis when needed. This intermediary approach balances processing speed and accuracy by leveraging both on-device and remote computational capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If the device continuously monitors health metrics and tracks nutritional intake using sensors and machine learning, then the reliability of health monitoring is improved, but the energy consumption increases

Engineering Contradiction:
Improvehealth monitoring reliabilityVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The device uses periodic action by monitoring health metrics at scheduled intervals rather than continuously. The image sensor captures images at specific times when the user is likely to be consuming food or engaging in activities relevant to health tracking. The depth sensor and microphones are activated periodically to detect gestures and voice commands related to health monitoring, thereby reducing energy consumption while maintaining reliable health tracking through strategic sampling

Inventive Principle:
Principle #19Periodic action

4Loss of information

If the device provides comprehensive information about items of interest, communications, and events through multiple output channels, then the information completeness is improved, but the device complexity and data processing requirements increase

Engineering Contradiction:
Improveinformation completenessVSAvoiddata processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments information processing into distinct modules: object recognition for identifying items of interest, voice recognition for processing commands and communications, health monitoring for tracking nutritional intake, and communication summarization for processing messages. Each module processes specific types of data independently using dedicated machine learning models, then integrates the results to provide comprehensive information outputs through multiple channels including visual projections and audio feedback

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient and user-friendly personal assistance and health monitoring by providing accurate information on items of interest, summarizing communications and events, and tracking nutritional intake, enhancing user experience through intuitive data presentation.

Implementation Method 1

accessing, by a wearable multimedia device worn by a user, image data regarding an environment of the user, where the image data is generated using one or more image sensors of the wearable multimedia device

Methodology Applied
Scientific EffectLight reflection: Reflection

Implementation Method 2

Three-dimensional (3D) depth sensors (e.g., a time of flight (TOF) camera) can be used to detect user gestures that are interacting with one or more VI elements projected on the surface

Methodology Applied
Scientific EffectTime of flight: Time of Flight

Implementation Method 3

a laser projected virtual interface (VI) can be projected onto the palm of a user's hand or other surface

Methodology Applied
Scientific EffectLaser: Laser

Implementation Method 4

receiving, by the wearable multimedia device, a first user input from the user, where the first user input includes a first command with respect to the item of interest, and where the first user input is received using one or more microphones of the wearable multimedia device

Methodology Applied
Scientific EffectSound wave detection: Sound

Data Source

PatentUS12411542B2Electronic devices using object recognition and/or voice recognition to provide personal and health assistance to users
Publication Date: 2025.09.09 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US12411542B2 patent drawing
  • US12411542B2 patent drawing
  • US12411542B2 patent drawing

AI summary

In an example method, a wearable multimedia device is worn by a user. Further, the device receives one or more communications during a first period of time; receives information regarding one or more events during the first period of time; receives a first spoken command from the user during a second period of time, where the second period of item is subsequent to the first period of time, and where the first spoken command includes a request to summarize the one or more communications and the one or more events; generates, using one or more of machine learning models, a summary of the one or more communications and the one or more events; and presents at least a portion of the summary to the user.