Multimodal Attribute Extraction for Catalog Error Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for attribute error detection in multi-modal item data, such as product names and images, are inadequate, often relying on manual verification which is time-consuming and prone to human errors.

Innovation Solution

A system that employs multi-modal machine learning models to extract attributes from various data sources, including text and visual data, and cross-checks these attributes to ensure accuracy, thereby reducing manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual verification methods are used for attribute error detection, then accuracy can be maintained through human judgment, but time consumption and labor intensity increase significantly

Engineering Contradiction:
Improveattribute error detection accuracyVSAvoidverification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical verification with an automated machine learning system that processes multi-modal item data. The system uses trained models to extract and verify attributes from text, images, and other data sources, substituting human labor with automated computational processes while maintaining detection accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces machine learning models as intermediaries between the raw item data and the final attribute extraction. These models act as mediators that process multi-modal data sources and produce verified attributes, eliminating the need for direct manual verification while preserving accuracy through automated cross-checking.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If traditional text-based extraction methods are used, then processing simplicity is maintained, but accuracy deteriorates due to inability to handle visual format attributes

Engineering Contradiction:
Improveextraction method simplicityVSAvoidattribute extraction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent implements a multi-modal extraction system that handles multiple data types (text, images, and other formats) through unified machine learning models. The system performs both text-based and visual attribute extraction using the same framework, making the extraction process universally applicable to different data formats while improving accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines multiple data sources and modalities (text data, image data, and other attributes) into a composite extraction approach. By integrating different types of data sources and processing them through coordinated machine learning models, the system achieves superior extraction accuracy compared to single-modality methods.

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If multi-modal machine learning models are deployed for attribute extraction, then extraction accuracy improves through cross-checking, but system complexity increases

Engineering Contradiction:
Improveattribute extraction accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the attribute extraction system into separate specialized models for different data modalities (text processing models, image processing models, etc.). Each model is optimized for its specific data type, and their results are cross-checked and integrated. This segmentation allows high accuracy through specialization while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250291808A1Attribute extraction and error detection using multi-modal data sources and machine-learning models
Publication Date: 2025.09.18 MAPLEBEAR INC
  • US20250291808A1 patent drawing
  • US20250291808A1 patent drawing
  • US20250291808A1 patent drawing

AI summary

An online system enhances the accuracy and completeness of item attribute data in a catalog database by extracting attribute values from multiple data sources with different data modalities. The system applies machine-learning models to information sources such as text descriptions, images, third-party databases, and user engagement data. Extracted attributes are verified before being stored in the catalog. Contradictory attribute values are identified through cross-checking and flagged for audit. The system ranks data sources to prioritize high-confidence attribute extractions. A client interface enables users to specify desired attributes and extraction criteria, which guide multi-modal machine-learning models in retrieving relevant attributes. The system supports iterative refinement of attribute extraction processes based on user feedback and evaluation results. Additionally, extracted attributes can be used to filter catalog items, enhancing search and selection functionality.