In-Vehicle Virtual Assistant Using Multi-Sensor Proactive Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional virtual personal assistants are passive and lack access to sensor data, limiting their functionality and ability to provide proactive insights and information to users.

Innovation Solution

A two-way virtual personal assistant system that continuously monitors sensor data from various sources, enabling proactive notifications and generating accurate outputs based on analyzed data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional virtual personal assistant is used, then device simplicity is maintained, but functionality and proactive capability are limited

Engineering Contradiction:
ImprovefunctionalityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple sensors (microphones, cameras, accelerometers, gyroscopes, barometers, magnetometers, GPS) and their data processing capabilities into a unified virtual personal assistant system. This merging enables the system to access diverse data sources including audio, video, motion, environmental, and location information, thereby significantly enhancing functionality and proactive notification capabilities while managing system complexity through integrated architecture.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If passive virtual personal assistant is used, then ease of operation is maintained, but user engagement and information provision are reduced

Engineering Contradiction:
Improveuser interaction simplicityVSAvoidinformation provision to user
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system continuously monitors and analyzes sensor data in the background before user initiation, proactively identifying conditions that warrant notification. By performing preliminary analysis of audio, video, motion, environmental, and location data, the system can alert users to relevant events, safety concerns, or information needs without requiring them to actively query the system, thus maintaining ease of operation while significantly reducing information loss.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If access to sensor data is limited, then system complexity is reduced, but measurement precision and insight capability are diminished

Engineering Contradiction:
Improvedata analysis accuracyVSAvoidsensor integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The virtual personal assistant system is designed with multi-functionality to process and analyze diverse sensor data types including audio from microphones, visual data from cameras, motion data from accelerometers and gyroscopes, environmental data from barometers and magnetometers, and location data from GPS. This universal processing capability enables precise measurement and analysis across multiple domains while managing complexity through a unified data processing architecture that handles various sensor inputs through common analysis pipelines.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12512099B2Two-way in-vehicle virtual personal assistant
Publication Date: 2025.12.30 HARMAN INT IND INC
  • US12512099B2 patent drawing
  • US12512099B2 patent drawing
  • US12512099B2 patent drawing

AI summary

Embodiments include a computer-implemented method for interacting with a user. The method comprises utilizing a processor associated with a virtual personal assistance system. The processor is configured to perform the steps of obtaining first sensor data from a first sensor included in a plurality of sensors. The first sensor data comprises audio speech input. The steps further comprise analyzing the audio speech input using a first deep learning model to generate an intended meaning, selecting second sensor data to obtain based on the intended meaning, obtaining the second sensor data from a second sensor included in the plurality of sensors, analyzing the second sensor data and the intended meaning to generate a prediction value, applying a second deep learning model to the prediction value to generate a text segment, and converting the text segment into a natural language audio output; and outputting the natural language audio output to the user.