Neural Network Image Processing for Distortion-Free Semantic Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing methods for intelligent vehicles struggle to accurately convert target information from a vehicle's forward view into a real-world coordinate system without causing distortion, especially due to camera vibrations affecting extrinsic parameters.
Innovation Solution
An end-to-end neural network model is trained to learn the correspondence between pixels in the forward view and top-down view, using camera extrinsic parameters, including variations caused by vibrations, to directly output semantic segmentation results in the real-world coordinate system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Inverse Perspective Mapping (IPM) is used to transform target information from forward view to top-down view, then target information can be obtained in real-world coordinate system, but distortion occurs which affects traffic safety
Solution Approach 1:
The patent replaces the traditional geometric transformation method (IPM) with a neural network-based learning approach. The neural network learns the mapping relationship between forward view images and top-down view semantic segmentation results through training, substituting the mechanical geometric transformation with an intelligent learning system that can adapt to various camera parameters and vibration conditions without causing distortion.
Solution Approach 2:
The patent incorporates camera extrinsic parameters (including those affected by vibration) as input variables to the neural network. By dynamically adjusting the network's input parameters based on actual camera conditions, the system maintains accurate mapping between forward view and top-down view coordinate systems while compensating for vibration-induced parameter changes, thereby eliminating distortion.
2Measurement precision
If camera extrinsic parameters are used to project image features from forward view to top-down view, then target information can be obtained, but vibration affects the extrinsic parameters causing inaccuracy
Solution Approach 1:
The patent uses camera extrinsic parameters (which include vibration information) as feedback input to the neural network. The network learns to process these varying parameters and adjust its transformation accordingly, turning the vibration-induced parameter changes from harmful factors into useful information that helps the system adapt to actual camera conditions and maintain accuracy.
Solution Approach 2:
The system transitions from using fixed geometric transformation rules to a dynamic neural network model that can adapt its transformation behavior based on input camera parameters. The neural network dynamically adjusts its mapping based on the actual extrinsic parameters including vibration effects, making the system flexible and accurate under varying conditions.
3Device complexity
If traditional coordinate transformation methods are used, then the process is simple, but the results are distorted and less accurate
Solution Approach 1:
The patent replaces simple geometric transformation algorithms with a neural network-based intelligent processing system. Although this increases computational complexity, it dramatically improves accuracy by learning complex non-linear mappings and adapting to various camera conditions, providing precise semantic segmentation results in real-world coordinate system.
Data Source
AI summary
An image processing method includes (i) obtaining a forward view of a vehicle, wherein the forward view at least shows an area in front of the vehicle, (ii) obtaining parameters of a camera used to capture the forward view, the parameters including extrinsic parameters of the camera, and (iii) inputting the forward view and the camera parameters into a neural network model to obtain a semantically segmented top-down view. Thus, an end-to-end method is provided, which avoids distortions that may occur during the transformation of features from the forward view to features in the top-down view.


