Multimodal Neural Network for Simultaneous Depth and Semantic Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Advanced Driver Assistance Systems (ADAS) face challenges in simultaneously performing semantic segmentation and depth estimation due to high processing complexity and limited accuracy, as existing solutions require separate neural networks that increase processing demands and capacity requirements.

Innovation Solution

A multimodal neural network architecture that combines an encoder with both a depth decoder and a semantic segmentation decoder, utilizing inverted residual blocks and skip connections to reduce processing complexity while improving accuracy, allowing for simultaneous depth estimation and semantic segmentation from a single image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If two separate deep neural networks are used for semantic segmentation and depth estimation, then both functions can be performed accurately, but processing complexity and capacity requirements increase significantly

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines two separate deep neural networks (one for semantic segmentation and one for depth estimation) into a single unified multimodal neural network. This unified network processes a single input image to simultaneously generate both semantic segmentation maps and depth maps, thereby reducing processing complexity and capacity requirements while maintaining the accuracy of both functions.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If two separate deep neural networks are used for semantic segmentation and depth estimation, then both functions can be performed accurately, but processing capacity and battery consumption increase

Engineering Contradiction:
Improvefunctional completenessVSAvoidbattery capacity
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent merges two separate neural network processing pipelines into a single multimodal network that handles both semantic segmentation and depth estimation tasks. By consolidating the computational workload into one network, the system reduces overall processing capacity requirements and consequently lowers battery consumption while maintaining complete functional capability.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If multiple label types are assigned in semantic segmentation, then object classification accuracy improves, but processing complexity increases

Engineering Contradiction:
Improvelabel assignment accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal encoder in the multimodal neural network that serves multiple functions: it extracts features for both semantic segmentation with multiple label types and depth estimation. This shared encoder enables the system to handle complex multi-class classification accurately while avoiding the processing complexity increase that would result from separate specialized networks for each task.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240338938A1Multimodal method and apparatus for segmentation and depth estimation
Publication Date: 2024.10.10 HARMAN INT IND INC
  • US20240338938A1 patent drawing
  • US20240338938A1 patent drawing
  • US20240338938A1 patent drawing

AI summary

A multimodal neural network model for combined depth estimation and semantic segmentation of images and a method of training the multimodal neural network model. The multimodal neural network comprising a single encoder, a depth decoder to estimate the depth of the image and a semantic segmentation decoder to determine semantic labels from the image. The method for training the multimodal neural network model comprising receiving a plurality of images at a single encoder, after encoding the images providing them to a depth estimation decoder and a semantic segmentation decoder to estimate the depth of the images and semantic labels to the images. The method further comprising comparing the estimated depth with the actual depth of the images and comparing the calculated semantic labels with the actual labels of the images to determine a depth loss and a semantic segmentation loss, respectively.