Multi-Chiplet Homomorphic Encryption Accelerator for High Yield

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing size of central processing units (CPUs), graphics processing units (GPUs), and neural processing units (NPUs) leads to yield limitations in micro-processor production and high manufacturing costs for monolithic accelerators for homomorphic encryption (HE), which are also resource-intensive and slow due to the large size and complexity of HE operations.

Innovation Solution

A multi-chiplet architecture is proposed, where the chip is broken into smaller, interconnected chiplets forming a ring topology, each performing number-theoretic transform (NTT) operations independently, with each chiplet connected to a unique memory chiplet, allowing for efficient data processing and reduced manufacturing costs through scalable and flexible design.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If monolithic accelerator architecture is used for homomorphic encryption operations, then processing capability is improved, but chip size increases and manufacturing yield decreases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidmanufacturing yield
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent divides the monolithic accelerator into multiple independent chiplets, each capable of performing homomorphic encryption operations. This segmentation allows smaller, more manufacturable units to be produced with higher yield, while maintaining overall processing capability through parallel operation of multiple chiplets.

Inventive Principle:
Principle #1Segmentation

2Productivity

If chip size is increased to enhance processing capability, then performance is improved, but manufacturing cost increases and yield decreases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidmanufacturing cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

By segmenting the accelerator into standard-sized chiplets, the patent enables mass production at smaller scales with lower individual costs. The modular design allows reuse of identical chiplet units, further reducing manufacturing costs while achieving enhanced processing capability through parallel processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The chiplets are designed with universal interfaces and standardized architectures that can be used across different configurations. This multi-functionality allows the same chiplet design to serve various homomorphic encryption workloads, reducing development and manufacturing costs through economies of scale.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If more operators are added to enhance HE performance, then processing capability is improved, but chip size increases and yield decreases

Engineering Contradiction:
ImproveHE performanceVSAvoidchip size
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The patent distributes operators across multiple chiplets rather than concentrating them in a single large chip. Each chiplet contains a manageable set of operators, keeping individual chip sizes small while achieving high overall HE performance through parallel execution across the chiplet array.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250023707A1Method and device with homomorphic encryption operation
Publication Date: 2025.01.16 SAMSUNG ELECTRONICS CO LTD
  • US20250023707A1 patent drawing
  • US20250023707A1 patent drawing
  • US20250023707A1 patent drawing

AI summary

An operation method includes obtaining an input matrix including a coefficient of a polynomial, based on a preprocessing unit (PU), performing a preprocessing operation on the coefficient, based on a first number-theoretic transform (NTT) architecture, performing a first NTT operation on a column element of the input matrix for which the preprocessing operation is completed, performing a Hadamard product operation between a result of the first NTT operation and a twiddle factor, and based on a second NTT architecture, performing a second NTT operation on a row element of the input matrix for which the Hadamard product operation is completed.