Multi-scale Deep Learning Object Detection via Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current convolutional neural networks (CNNs) face high computational expense and difficulty in identifying both large and small objects within images, making them impractical for efficient classification, especially with high-resolution 3D images.
Innovation Solution
The multi-scale deep learning (MSDL) system employs segmentation techniques to identify candidate segments of varying sizes and shapes, resamples these segments into fixed window sizes, and processes them using a multi-scale neural network with feature extracting CNNs and classifiers to classify objects efficiently, reducing computational load and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a CNN processes entire high-resolution 3D images, then classification accuracy is improved, but computational expense and storage requirements increase significantly
Solution Approach 1:
The patent divides the 3D image into multiple segments of varying sizes and shapes, processing each segment independently through the CNN. This segmentation approach reduces the computational load compared to processing the entire image while maintaining classification accuracy by preserving contextual information around objects of interest.
Solution Approach 2:
The patent applies different processing strategies to different regions of the image based on their characteristics. By identifying segments with objects of interest and processing those specifically, while using coarser processing for background regions, the system achieves efficient computation without sacrificing accuracy in critical areas.
2Adaptability or versatility
If a CNN uses large convolution windows to detect large objects, then detection capability for large objects is improved, but ability to detect small objects deteriorates
Solution Approach 1:
The patent segments the image into regions of different sizes and processes each through the CNN independently. This allows the system to use appropriate convolution window sizes for each segment, enabling effective detection of both large and small objects without the compromise that would result from using a single fixed window size for the entire image.
Solution Approach 2:
The patent dynamically adjusts the processing approach based on the size and characteristics of detected segments. By adapting the convolution window size and processing parameters to match the scale of objects within each segment, the system achieves versatile detection capability across different object sizes.
Data Source
AI summary
A system for identifying objects in an image is provided. The system identifies segments of an image that may contain objects. For each segment, the system generates a segment score by inputting to a multi-scale neural network windows of multiple scales that include the segment that have been resampled to a fixed window size. A multi-scale neural network includes a feature extracting convolutional neural network (“feCNN”) for each scale and a classifier that inputs each feature of each feCNN. The segment score indicates whether the segment contains an object. The system generates a pixel score for pixels of the image. The pixel score for a pixel indicates that that pixel is within an object based on the segment scores of segments that contain that pixel. The system then identifies the object based on the pixel scores of neighboring pixels.


