Content Identification via High-Order Conjunction Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search systems cannot identify and classify content, such as images, without manual association of descriptive text, limiting their ability to retrieve content based on search queries.

Innovation Solution

A method is developed to generate a high-order conjunction that predicts labels associated with content, allowing for automatic tagging and classification of content, enabling search queries to be satisfied by descriptive labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional search systems rely on manual association of descriptive text with content, then content can be identified when descriptive text is present, but content cannot be automatically identified or classified without manual entry of descriptive text

Engineering Contradiction:
Improveautomatic content identificationVSAvoidmanual text entry time
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The system enables content to identify and classify itself automatically through machine learning algorithms that analyze content features and generate descriptive labels without human intervention. The classifier processes content independently, performing the identification task that would otherwise require manual descriptive text entry.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary classification of content by generating descriptive labels in advance, before any search query is executed. This pre-computation of content metadata allows for rapid retrieval during search operations without requiring manual text association at the time of content ingestion.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If conventional search systems use simple text matching, then search queries can be processed quickly, but the systems cannot identify content based on visual or structural features

Engineering Contradiction:
Improvecontent identification capabilityVSAvoidclassification system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The content classification system is segmented into distinct functional modules: feature extraction components that analyze different aspects of content (visual, textual, structural), clustering algorithms that group similar features, and label generation components that produce descriptive tags. This modular architecture manages complexity while enabling versatile content identification across multiple content types.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary classifier layer between raw content and search queries. This classifier acts as a mediator that transforms diverse content types into standardized descriptive labels, enabling the search system to handle complex content identification tasks without requiring direct complex analysis for each search operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If manual descriptive text is associated with content, then content can be searched using that text, but new content cannot be automatically tagged without user intervention

Engineering Contradiction:
Improvecontent tagging efficiencyVSAvoiduser intervention requirement
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The content tagging system operates autonomously by automatically analyzing new content, extracting relevant features, and generating appropriate descriptive labels without requiring user intervention. The machine learning classifier performs the entire tagging process self-service style, dramatically improving productivity while maintaining ease of operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where classification results are continuously refined based on performance metrics and user interactions. This feedback loop enables the system to improve tagging accuracy over time while maintaining automatic operation, balancing productivity gains with operational simplicity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8788503B1Content identification
Publication Date: 2014.07.22 GOOGLE LLC
  • US8788503B1 patent drawing
  • US8788503B1 patent drawing
  • US8788503B1 patent drawing

AI summary

Systems, computer program products, and methods can identify a training set of content, and generate one or more clusters from the training set of content, where each of the one or more clusters represent similar features of the training set of content. The one or more clusters can be used to generate a classifier. New content is identified and the classifier is used to associate at least one label with the new content.