In-Vehicle AI Interaction Using Surrounding Scene Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing vehicle-mounted technologies lack the capability to interact with drivers based on the surrounding environment of the vehicle, which could help reduce fatigue and improve concentration.
Innovation Solution
A vehicle-mounted apparatus equipped with a camera, processor, and output interface that captures environment images, generates descriptive text using a composite model, and produces response texts through a language model to execute interactive operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If vehicle-mounted technology is equipped with environment recognition and interaction capabilities, then driver fatigue reduction and concentration improvement are achieved, but device complexity increases
Solution Approach 1:
The system is divided into distinct functional modules: environment image acquisition module, composite model processing module, language model processing module, and output execution module. Each module handles a specific task in the interaction pipeline, making the overall complex system manageable and maintainable through clear separation of concerns
Solution Approach 2:
The patent introduces intermediate processing layers (composite model and language model) that mediate between the raw environment images and the final interaction outputs. These intermediaries transform and structure the data flow, enabling complex AI processing while maintaining system organization and reducing direct complexity between components
2Measurement precision
If multiple processing models are integrated for environment understanding, then interaction accuracy improves, but computing resource consumption increases
Solution Approach 1:
The processing pipeline is segmented into specialized models: composite model for environment understanding and language model for interaction generation. This segmentation allows each model to be optimized for its specific function, improving overall accuracy while enabling selective execution to manage computational resources
Solution Approach 2:
The system processes environment images through multiple modeling stages (composite model then language model), which may seem excessive but ensures high accuracy in environment understanding and interaction generation. The multi-stage processing is optimized to balance thoroughness with computational efficiency
Data Source
AI summary
A vehicle-mounted apparatus and vehicle-mounted apparatus control method are provided. The vehicle-mounted apparatus is configured on a vehicle in which a user rides. The apparatus captures an environmental image around the vehicle. The apparatus inputs the environmental image into a composite model to generate a corresponding environmental text, wherein the environmental text is configured to describe the environmental image. The apparatus inputs the environmental text into a language model to generate a response text corresponding to the vehicle. The apparatus executes an interactive operation corresponding to the user based on the response text.


