In-Storage Graph Traversal for Approximate Nearest Neighbor Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for approximate nearest neighbor search (ANNS) are computationally extensive and inefficient in terms of resource utilization, particularly in terms of storage and processing power.

Innovation Solution

The proposed solution involves a near-data processing (NDP) approach that leverages in-storage computing architectures and logic units (LUs) inside solid-state drive (SSD) devices to accelerate graph traversal in ANNS. This includes a two-level scheduling process to exploit spatial and temporal locality, and a speculative searching mechanism to further enhance performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If graph traversal is performed using existing techniques (KD-tree, LSH), then nearest neighbor search can be conducted, but computational cost is excessively high and resource utilization is inefficient

Engineering Contradiction:
ImproveANNS search speedVSAvoidcomputational cost
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent segments the graph traversal process into multiple stages: candidate generation, candidate filtering, and final selection. It divides the dataset into multiple partitions and processes them in parallel using multiple processing units. This segmentation reduces the computational burden on each unit and improves overall search efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-computing and storing graph structures (such as HNSW graphs) in advance. The graph information including vertex connections and distance metrics is prepared beforehand and stored in memory, allowing the search process to directly traverse pre-built structures rather than computing distances from scratch during query time.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If traditional ANNS techniques are used, then search functionality is provided, but storage resources are not utilized efficiently

Engineering Contradiction:
Improvestorage resource utilizationVSAvoidsystem architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the storage system and computing system into an integrated architecture where the storage device (SSD) directly performs graph traversal computations. The storage controller is enhanced with computing capabilities, allowing data to be processed in-place during storage operations without requiring separate computation hardware, thus improving storage resource utilization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a storage controller as an intermediary between the host processor and storage medium. This controller acts as a mediator that receives search queries, performs graph traversal computations using stored graph structures, and returns results. It bridges the gap between storage and computation functions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Speed

If parallel computation is implemented using multiple logic units, then processing speed increases, but query allocation and management complexity increases

Engineering Contradiction:
Improveparallel processing speedVSAvoidquery allocation complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent implements dynamic query allocation where the storage controller adaptively distributes queries to available logic units based on current system state and workload. The allocation strategy adjusts in real-time to balance the load across multiple processing units, optimizing parallel processing efficiency while managing complexity through adaptive control.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250165468A1Systems and methods for graph traversal for approximate nearest neighbor search
Publication Date: 2025.05.22 SAMSUNG ELECTRONICS CO LTD
  • US20250165468A1 patent drawing
  • US20250165468A1 patent drawing
  • US20250165468A1 patent drawing

AI summary

A system and a method for approximate nearest neighbor search are disclosed. A query storage circuit stores query information related to at least one query from a host in a query property table. A generator and allocator circuit is configured to generate graph information using a batch of at least one vertex corresponding to the at least one queries from the query property table and to allocate the at least one queries to at least one logic unit (LU) based on the graph information. A search circuit has the at least one LU and is configured to compute at least one distance, using the graph information, between the at least one vertex and at least one candidate neighbor of the at least one vertex to generate at least one distance result. The query property table is modified based on the at least one distance result.