Universal Scene Descriptor for Visual Metadata Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision systems are limited in extracting multi-level semantic metadata from complex visual environments, as they are typically modeled after human vision constraints and are not robust to changes in viewing conditions, lacking scalability and universality in scene understanding.

Innovation Solution

A universal scene descriptor framework that emulates the human visual cortex and higher-level cognitive processes by decomposing scenes into salient regions of interest, extracting hierarchical biologically-inspired visual features, and classifying them to provide multi-level semantic information, integrating these into a universal feature vector for metadata extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If computer vision systems are modeled after human vision constraints, then the system complexity is reduced, but the robustness to changes in viewing conditions deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidrobustness to changes in viewing conditions
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the scene into multiple regions of interest (ROIs) based on saliency detection, allowing the system to process only relevant portions of the image. This segmentation approach reduces computational complexity while maintaining robustness by focusing resources on salient regions that are more likely to contain important semantic information across varying viewing conditions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts features at multiple hierarchical levels (from low-level edge detection to high-level semantic concepts) and combines them into a multi-level feature vector. This dimensional expansion from single-level to multi-level feature representation enables the system to maintain robustness across different viewing conditions while managing complexity through hierarchical organization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If supervised segmentation is used to identify image segments, then the precision of region identification is improved, but the ease of operation deteriorates due to requiring manual annotation

Engineering Contradiction:
Improveprecision of region identificationVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements unsupervised saliency-based segmentation that automatically identifies regions of interest without requiring manual annotation or supervised training data. The system performs self-service by detecting salient regions through computational algorithms that analyze image properties such as color, texture, and motion, thereby achieving both automated operation and acceptable precision for metadata extraction.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If application-specific scene descriptors are used, then the precision for specific applications is improved, but the adaptability to different applications deteriorates

Engineering Contradiction:
Improveprecision for specific applicationsVSAvoidadaptability to different applications
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal multi-level feature vector that captures semantic information at multiple hierarchical levels (edges, textures, shapes, objects, and scene concepts). This universal descriptor can be adapted to different applications by selecting appropriate levels of the hierarchy, providing both precision for specific applications and versatility across different domains without requiring application-specific customization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8494259B2Biologically-inspired metadata extraction (BIME) of visual data using a multi-level universal scene descriptor (USD)
Publication Date: 2013.07.23 TELEDYNE SCIENTIFIC & IMAGING LLC
  • US8494259B2 patent drawing
  • US8494259B2 patent drawing
  • US8494259B2 patent drawing

AI summary

A computer vision system provides a universal scene descriptor (USD) framework and methodology for using the USD framework to extract multi-level semantic metadata from scenes. The computer vision system adopts the human vision system principles of saliency, hierarchical feature extraction and hierarchical classification to systematically extract scene information at multiple semantic levels.