Visual Search System Segmenting Images for Task-Focused Results

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image search technologies often return unsatisfactory results due to vague textual annotations and lack of visual-centric relevance, failing to efficiently assist users in completing tasks related to images, such as purchasing items or finding recipes, as they do not effectively utilize user intent and segment-specific information.

Innovation Solution

A visual search system that identifies user intent and segments associated with an image, providing relevant search results by accessing a visual content store and incorporating multi-modal inputs like mobile device data, to offer an efficient and interactive visual-centric search experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If general search-by-image feature is used, then search functionality is provided, but search relevance and annotation coverage are not satisfactory

Engineering Contradiction:
Improvesearch relevanceVSAvoidannotation coverage
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the image into multiple regions and identifies different segments (e.g., food item, dining setting) to generate targeted annotations for each segment. This allows the system to provide comprehensive annotation coverage while maintaining high search relevance by focusing on specific regions rather than treating the image as a whole.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes parameters by identifying multiple user intents (e.g., recipe search, restaurant search, ingredient search) associated with different segments of the image. By dynamically adjusting the search parameters based on identified user intents and segments, the system achieves both high relevance and broad annotation coverage.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If human crowdsourcing is used for annotations, then text annotations are provided, but annotations are too vague and results are not visual centric

Engineering Contradiction:
Improveannotation detailVSAvoidvisual centric search
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent replaces manual human crowdsourcing with an automated computer vision system that performs image segmentation and annotation. This substitution eliminates the vagueness inherent in human annotations while maintaining visual-centricity, as the system directly processes visual features to generate precise, actionable annotations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service annotation by automatically analyzing the image content, identifying segments, determining user intents, and generating relevant annotations without human intervention. This automated approach provides detailed, precise annotations that are inherently visual-centric since they derive directly from image analysis.

Inventive Principle:
Principle #25Self-service

3Reliability

If multiple searches are performed to complete a task, then comprehensive results are obtained, but search efficiency is reduced

Engineering Contradiction:
Improvecomprehensive resultsVSAvoidsearch efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary actions by automatically segmenting the image, identifying multiple user intents, and generating targeted search queries for each intent and segment combination in advance. This preliminary processing enables the system to retrieve comprehensive results from multiple specialized searches simultaneously, rather than requiring sequential user-initiated searches.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system merges multiple search operations into a single unified process by combining results from different user intents and segments. By integrating the search functionality across multiple dimensions (intents × segments) and presenting unified comprehensive results, the system achieves thorough coverage while maintaining high efficiency through automated parallel processing.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10664515B2Task-focused search by image
Publication Date: 2020.05.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10664515B2 patent drawing
  • US10664515B2 patent drawing
  • US10664515B2 patent drawing

AI summary

Systems, computing devices, and methods for performing an image search are presented. A search query including an image is received from a user. A segment associated with the image is identified. A user intent associated with the image and the segment is identified. Search results associated with the identified segment and user intent are generated, and presented to the user.