Driving Risk Localization Using Visual and Optical Flow Reasoning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current advanced driving and advanced driving assistance systems face challenges in identifying and communicating important objects and risks in driving scenes effectively, limiting high-level automation and situational awareness in intelligent vehicles.

Innovation Solution

A computer-implemented method and system that uses an encoder-decoder neural network to extract and concatenate visual and optical flow features from images, identifying important traffic agents and infrastructure, and controlling vehicle systems to respond to potential risks, utilizing a pre-trained driving risk assessment mechanism and dataset for risk localization and reasoning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current advanced driving systems are used, then basic driving functions are maintained, but the ability to identify and communicate important objects and risks is insufficient

Engineering Contradiction:
Improverisk identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex driving scene analysis into distinct functional modules: an encoder that extracts visual and optical flow features, a concatenation layer that integrates multi-source features, and a decoder that localizes and reasons about risks. This segmentation allows each module to specialize in specific tasks, improving overall risk identification accuracy while managing system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from traditional 2D image processing to 4D spatiotemporal feature extraction by incorporating optical flow information across multiple time frames. This dimensional expansion enables the model to capture motion dynamics and temporal relationships, significantly improving risk identification accuracy by analyzing not just spatial patterns but also temporal evolution of driving scenes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If traditional object detection methods are used, then processing speed is maintained, but situational awareness and risk communication capabilities are limited

Engineering Contradiction:
Improvesituational awareness capabilityVSAvoidnetwork architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The encoder-decoder network is designed as a universal framework that simultaneously performs multiple functions: feature extraction, risk localization, and risk reasoning. The same network architecture handles both visual and optical flow inputs, and can identify various types of traffic agents and infrastructure elements, providing comprehensive situational awareness through a single multi-functional system rather than separate specialized detectors.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The concatenated features serve as an intermediary representation that bridges the gap between raw multi-source inputs (images and optical flow) and the final risk assessment output. This intermediate feature fusion layer integrates information from multiple modalities and time frames, enabling the decoder to perform comprehensive risk localization and reasoning while managing the complexity of processing diverse input types.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If simple risk assessment is used, then system responsiveness is maintained, but the ability to communicate risks to the driver is insufficient

Engineering Contradiction:
Improverisk information completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary feature extraction and concatenation during the encoding phase, preparing integrated spatiotemporal representations before the decoding and risk localization stage. This preliminary processing organizes and pre-integrates information from multiple sources, reducing the computational burden during final risk assessment and enabling faster, more complete risk communication to the driver.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12065167B2System and method for completing joint risk localization and reasoning in driving scenarios
Publication Date: 2024.08.20 HONDA MOTOR CO LTD
  • US12065167B2 patent drawing
  • US12065167B2 patent drawing
  • US12065167B2 patent drawing

AI summary

A system and method for completing joint risk localization and reasoning in driving scenarios that include receiving a plurality of images associated with a driving scene of an ego agent. The system and method also include inputting image data associated with the plurality of images to an encoder and inputting concatenated features to a decoder that identifies at least one of: an important traffic agent and an important traffic infrastructure that is located within the driving scene of the ego agent. The system and method further include controlling at least one system of the ego agent to provide a response to account for the at least one of: the important traffic agent and the important traffic infrastructure that is located within the driving scene of the ego agent.