Multi-scale Deep Learning Object Detection via Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current convolutional neural networks (CNNs) face high computational expense and difficulty in identifying both large and small objects within images, making them impractical for efficient classification, especially with high-resolution 3D images.

Innovation Solution

The multi-scale deep learning (MSDL) system employs segmentation techniques to identify candidate segments of varying sizes and shapes, resamples these segments into fixed window sizes, and processes them using a multi-scale neural network with feature extracting CNNs and classifiers to classify objects efficiently, reducing computational load and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a CNN processes entire high-resolution 3D images, then classification accuracy is improved, but computational expense and storage requirements increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational expense
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent divides the 3D image into multiple segments of varying sizes and shapes, processing each segment independently through the CNN. This segmentation approach reduces the computational load compared to processing the entire image while maintaining classification accuracy by preserving contextual information around objects of interest.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing strategies to different regions of the image based on their characteristics. By identifying segments with objects of interest and processing those specifically, while using coarser processing for background regions, the system achieves efficient computation without sacrificing accuracy in critical areas.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If a CNN uses large convolution windows to detect large objects, then detection capability for large objects is improved, but ability to detect small objects deteriorates

Engineering Contradiction:
Improvedetection capability for large objectsVSAvoiddetection capability for small objects
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the image into regions of different sizes and processes each through the CNN independently. This allows the system to use appropriate convolution window sizes for each segment, enabling effective detection of both large and small objects without the compromise that would result from using a single fixed window size for the entire image.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the processing approach based on the size and characteristics of detected segments. By adapting the convolution window size and processing parameters to match the scale of objects within each segment, the system achieves versatile detection capability across different object sizes.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10521699B2Multi-scale deep learning system
Publication Date: 2019.12.31 LAWRENCE LIVERMORE NAT SECURITY LLC
  • US10521699B2 patent drawing
  • US10521699B2 patent drawing
  • US10521699B2 patent drawing

AI summary

A system for identifying objects in an image is provided. The system identifies segments of an image that may contain objects. For each segment, the system generates a segment score by inputting to a multi-scale neural network windows of multiple scales that include the segment that have been resampled to a fixed window size. A multi-scale neural network includes a feature extracting convolutional neural network (“feCNN”) for each scale and a classifier that inputs each feature of each feCNN. The segment score indicates whether the segment contains an object. The system generates a pixel score for pixels of the image. The pixel score for a pixel indicates that that pixel is within an object based on the segment scores of segments that contain that pixel. The system then identifies the object based on the pixel scores of neighboring pixels.