Multi-Modal Embedding Correction for Accurate Attribute Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing search engines struggle to accurately provide search results that align with user intent when handling multi-modal inputs such as images and text, as they fail to effectively combine and manipulate attributes due to inherent non-linearity in deep learning models.
Innovation Solution
A method that decomposes multi-modal embedding into a delta space with non-linearity and a vector space allowing linear expression, using a correction function to perform vector operations, enabling accurate and efficient search results by transforming multi-modal inputs into a virtual attribute vector space for linear computation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If deep learning models are used for multi-modal attribute combination, then search capability is enhanced, but non-linearity prevents accurate vector operations between attributes
Solution Approach 1:
The patent segments the embedding space into two distinct components: a delta space that captures non-linear characteristics and a linear space that enables accurate vector operations. This segmentation allows the system to preserve the strengths of deep learning models while eliminating their limitations for attribute manipulation.
Solution Approach 2:
The patent introduces a correction function as an intermediary that maps attributes from the non-linear delta space to a linear space where vector operations can be performed accurately. This intermediary layer enables precise attribute manipulation while maintaining the benefits of multi-modal deep learning representations.
2Measurement precision
If multi-modal embedding is used to capture complex attributes, then search quality improves, but computational complexity increases
Solution Approach 1:
By dividing the embedding space into delta and linear components, the patent enables simplified computations in the linear space while preserving complex multi-modal relationships in the delta space representation, reducing overall computational burden.
Solution Approach 2:
The patent transforms the problem from operating in a complex non-linear space to operating in a simplified linear space through the correction function, changing the computational parameters from complex non-linear operations to simple linear vector operations.
3Ease of operation
If traditional vector operations are performed in non-linear embedding space, then attribute manipulation is attempted, but results are inaccurate
Solution Approach 1:
The correction function serves as an intermediary transformation that converts attributes from the non-linear delta space to a linear representation, enabling accurate vector operations while maintaining the full expressive power of multi-modal embeddings.
Solution Approach 2:
The patent effectively adds a transformation dimension by mapping attributes between the delta space and linear space through the correction function, allowing vector operations to be performed in the linear dimension while preserving the original non-linear characteristics.
Data Source
AI summary
A method of providing search results based on multi-modal features includes performing a vector operation between attributes according to a user query on a multi-modal embedding space; and providing search results corresponding to the user query based on an embedding vector acquired through the vector operation.


