Stochastic Similarity Search Refinement via Binary Hashing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale similarity searches in machine learning and artificial intelligence are compute and memory intensive due to the sheer volume and richness of data, and existing hashing methods provide imperfect data conversions, leading to degraded search accuracy.

Innovation Solution

The implementation of a compute device with column-read enabled memory that performs stochastic associative searches by refining initial result sets using Euclidean distances and binary hashing, leveraging a three-dimensional cross-point architecture for efficient data access and processing, and utilizing random sparse lifting and Procrustean orthogonal sparse hashing to generate distance-preserving sparse binary hash codes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If hashing methods are used to perform stochastic associative searches, then search speed is improved, but search accuracy is degraded

Engineering Contradiction:
Improvesearch speedVSAvoidsearch accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The search process is divided into two distinct stages: a hashing stage for rapid initial filtering and a refinement stage for accurate result verification. The hashing stage uses binary hash codes for fast stochastic associative search, while the refinement stage uses original floating-point vectors to correct inaccuracies and eliminate false positives, thereby resolving the accuracy-speed tradeoff.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Binary hash codes serve as an intermediary representation that enables fast search operations. These hash codes are derived from original floating-point vectors but provide a compressed, discrete representation that can be efficiently compared using Hamming distances, facilitating rapid initial search results before final accuracy verification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If large-scale datasets are processed, then search comprehensiveness is improved, but computational overhead and memory requirements increase

Engineering Contradiction:
Improvedata volumeVSAvoidcomputational overhead
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system extracts and stores only the essential information needed for search operations in a compressed format. Original floating-point vectors are converted to binary hash codes, extracting only the critical similarity information while discarding redundant precision, thereby reducing memory requirements and computational overhead for large-scale datasets.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary hashing of data before actual search operations. By pre-computing binary hash codes from floating-point vectors and storing them in a database, the system prepares data in an optimized format that enables rapid search operations without repeatedly processing the full complexity of original datasets.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11829376B2Technologies for refining stochastic similarity search candidates
Publication Date: 2023.11.28 SK HYNIX NAND PRODUCT SOLUTIONS CORP
  • US11829376B2 patent drawing
  • US11829376B2 patent drawing
  • US11829376B2 patent drawing

AI summary

Technologies for refining stochastic similarity search candidates include a device having a memory that is column addressable and circuitry connected to the memory. The circuitry is configured to add a set of input data vectors to the memory as a set of binary dimensionally expanded vectors, including multiplying each input data vector with a projection matrix. The circuitry is also configured to produce a search hash code from a search data vector, including multiplying the search data vector with the projection matrix. Additionally, the circuitry is configured to identify a result set of the binary dimensionally expanded vectors as a function of a Hamming distance of each binary dimensionally expanded vector from the search hash code and determine, from the result set, a refined result set as a function of a similarity measure in an original input space of the input data vectors.