Enhanced Max Margin Learning for Scalable Image Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multimodal data mining in multimedia databases faces challenges such as the semantic gap and inefficiencies in learning relationships between image and text data, particularly in large-scale databases, where existing methods struggle with scalability and convergence rates.

Innovation Solution

The Enhanced Max Margin Learning (EMML) approach is introduced, which formulates the image annotation problem as a structured prediction task using a new max margin learning framework, allowing for faster convergence and scalability by reducing the number of constraints and optimizing the Lagrange dual problem, enabling efficient learning of relationships between images and text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing max margin learning methods are used for image annotation, then the relationship between images and text can be learned, but the convergence rate is slow and the method does not scale well to large databases

Engineering Contradiction:
Improveconvergence rateVSAvoidlearning time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the learning problem by formulating image annotation as a structured prediction task with decomposable constraints. The Lagrange dual problem is broken down into smaller subproblems that can be solved more efficiently, enabling faster convergence while maintaining learning accuracy for image-text relationships

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the optimization parameters by transforming the primal problem into a Lagrange dual formulation. This parameter transformation allows the use of more efficient optimization algorithms that converge faster, directly addressing the slow convergence issue of existing max margin learning methods

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If existing multimodal data mining methods are applied to large-scale databases, then comprehensive analysis can be performed, but the query response time increases with database scale

Engineering Contradiction:
Improvedatabase scaleVSAvoidquery response time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent segments the large-scale database processing into manageable structured prediction tasks. By decomposing the annotation problem and using efficient Lagrange dual optimization, the system can handle larger databases without proportionally increasing query response time, achieving scalability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional brute-force optimization mechanisms with an optimized Lagrange dual formulation. This substitution of the optimization mechanism enables the system to scale to large databases while maintaining efficient query response times

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10007679B2Enhanced max margin learning on multimodal data mining in a multimedia database
Publication Date: 2018.06.26 THE RES FOUNDATION FOR THE STATE UNIV OF NEW YORK
  • US10007679B2 patent drawing
  • US10007679B2 patent drawing
  • US10007679B2 patent drawing

AI summary

Multimodal data mining in a multimedia database is addressed as a structured prediction problem, wherein mapping from input to the structured and interdependent output variables is learned. A system and method for multimodal data mining is provided, comprising defining a multimodal data set comprising image information; representing image information of a data object as a set of feature vectors in a feature space; clustering in the feature space to group similar features; associating a non-image representation with a respective image data object based on the clustering; determining a joint feature representation of a respective data object as a mathematical weighted combination of a set of components of the joint feature representation; optimizing a weighting for a plurality of components of the mathematical weighted combination with respect to a prediction error between a predicted classification and a training classification; and employing the mathematical weighted combination for automatically classifying a new data object.