Surroundings Information Actuation Using VQA and Position-Annotated Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face limitations in classifying objects for surroundings information, leading to information loss and dependence on predefined object definitions, restricting the scope of shared information.

Innovation Solution

A system captures image data, annotates position data using a GNSS module, and generates a text description through a visual question answering system, actuating components based on detected triggers and sharing surroundings information via V2X communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a restricted number of objects are classified according to predefined standards, then the system complexity is reduced, but the information completeness and adaptability of surroundings information is lost

Engineering Contradiction:
Improveadaptability of object classificationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

A visual question answering system acts as an intermediary between the image data and the V2X communication system. This intermediary generates natural language descriptions of detected objects and situations, enabling flexible and adaptable surroundings information sharing without requiring complex predefined classification schemes. The VQA system translates visual data into contextualized descriptions that can be directly used for navigation and safety applications.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If predefined object definitions are used for classification, then the ease of operation is improved, but the measurement precision and information accuracy deteriorate

Engineering Contradiction:
Improveease of object classificationVSAvoidobject classification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

Instead of forcing detected objects into predefined classification categories, the system inverts the approach by generating natural language descriptions that capture the essential characteristics and context of detected objects. This inversion allows the system to maintain ease of operation through automated processing while achieving higher precision by describing objects in their actual contextual form rather than constrained categories.

Inventive Principle:
Principle #13The other way round (Inversion)

3Reliability

If comprehensive surroundings information is generated and shared, then the traffic safety and navigation accuracy are improved, but the data transmission volume and processing requirements increase

Engineering Contradiction:
Improvetraffic safetyVSAvoiddata transmission volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system extracts only the most relevant and critical surroundings information needed for navigation and safety, rather than transmitting all detected data. The visual question answering system generates concise natural language descriptions that capture essential contextual information about obstacles, traffic conditions, and navigation-relevant features, reducing data volume while maintaining high reliability for safety-critical applications.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250376184A1Method for actuating at least one system component of a system according to surroundings information generated by the system
Publication Date: 2025.12.11 AUDI AG
  • US20250376184A1 patent drawing

AI summary

A method and system configured to actuate at least one system component of a system according to surroundings information generated by the system. Image data are captured by a capture device of the system. Upon detection of a trigger, position data of the system are annotated on the image data. Subsequently, a text description message for the image data is prepared by a visual question answering system of the system (24), so that the surroundings information is generated. At least one system component is then actuated according to the surroundings information. An actuated system component may include a warning message and/or a navigation instruction and/or providing the surroundings information as a message to at least one further system is effectuated if the further system is located at a distance to the system and/or on a trajectory corresponding at least in some sections to that of the system is traveled.