Attention-Fused Embeddings for Partial-Image Object Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image search technologies struggle with partial or distorted object depictions, unconventional lighting, and duplicate object records, failing to consider sensory characteristics like smell, taste, and sound, and often yield inaccurate results due to focusing solely on pixel color.

Innovation Solution

A system utilizing multi-feature and multi-modal data, including image, text, sound, odor, taste, and tactility, with attention-based fusion of embeddings, GAN-autoencoder for image enhancement, and RDBMS for data management, to accurately identify and categorize objects, and integrate inventory and supplier databases for contextual search.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional image search methods focusing only on pixel color values are used, then the search process is simple and fast, but the identification accuracy is low and results lack meaningful resemblance

Engineering Contradiction:
Improveidentification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the identification process into multiple independent modules: image processing module extracts visual features, text processing module extracts semantic features, and sensor data processing module extracts sensory characteristics. Each module operates independently and contributes to the final identification, allowing the system to achieve high accuracy without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional 2D image pixel analysis to multi-dimensional feature space by incorporating text descriptions, sensory characteristics (smell, taste, texture, sound), and contextual information. This dimensional expansion enables more accurate identification by considering objects from multiple perspectives simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If multi-modal data and attention-based fusion are used to improve identification accuracy, then search results become more precise, but processing time and computational resources increase

Engineering Contradiction:
Improvesearch accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-processing and storing sensory characteristics, text descriptions, and image features in structured databases before actual search queries. During search operations, the system retrieves pre-processed data and performs attention-based fusion, significantly reducing real-time processing requirements while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs dynamic attention mechanisms that adaptively weight different modalities based on query context and data availability. The system dynamically adjusts which features receive more attention during fusion, optimizing processing efficiency by focusing computational resources on the most relevant features for each specific search scenario.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If the system accommodates partial or distorted object depictions with multiple features, then it handles diverse query conditions better, but data processing complexity increases

Engineering Contradiction:
Improvehandling diverse queriesVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal identification framework that handles multiple query types (complete objects, partial objects, distorted views, different lighting conditions) through a single multi-modal system. The same architecture processes images, text, and sensor data regardless of query complexity, achieving versatility without proportionally increasing processing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an attention-based fusion mechanism as an intermediary layer between raw multi-modal data and final identification results. This fusion layer integrates features from different modalities and handles partial or distorted information by weighing available evidence, simplifying the processing of diverse query conditions through a unified intermediate representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If sensory characteristics like smell, taste, and sound are incorporated, then identification accuracy for challenging objects improves, but system complexity and data collection requirements increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges sensory characteristic data (smell, taste, texture, sound) with traditional visual and text data into a unified multi-modal feature space. By combining these diverse data types through attention-based fusion, the system achieves improved identification accuracy for challenging objects while managing complexity through integrated processing rather than separate systems.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250308223A1Linking different variations of multi-feature and multi-modal information to a unique object in a data space, using attention-basesd fused embeddings and rdbms, identifying a unique entity from partial or incomplete image query data, and displaying location and acquisition time using artificial intelligence
Publication Date: 2025.10.02 LIM CO LTD
  • US20250308223A1 patent drawing
  • US20250308223A1 patent drawing
  • US20250308223A1 patent drawing

AI summary

The present disclosure describes methods, systems, apparatus, and media for object identification and classification, utilizing multi-feature and multi-modal data. This includes shape, material, brand, price, odor, taste, tactility, and sound. The system integrates a server space for data processing, a querying device for iterative searches, and a data interface module for refining results. It features AI-driven image optimization, feature extraction, and pattern recognition, employing novel techniques for fusing multi-feature and multi-modal embeddings utilizing multi-head attention. Additionally, a linker module powered by two active learning with feedback loops AI models consolidates scattered data into a unified object information database. The system also employs novel AI algorithms for isolating the object of interest through a saliency map and semantic analysis, as well as for enhancing raw images with a GAN-autoencoder.