FeFET In-Memory Computing for Cosine Distance Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current von Neumann computer architectures for cosine-based operations in artificial intelligence applications face significant energy consumption and delay issues, and existing in-memory computing solutions for cosine distance calculations are scarce and not applicable to wider applications like binary neural networks.
Innovation Solution
An in-memory computing architecture utilizing two FeFET-based storage arrays, Translinear circuits, and a WTA circuit, where each storage array stores different vectors and outputs inner products and sum of squares, allowing for efficient cosine distance calculations with reduced energy consumption and delay through the use of FeFET and resistor structures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If traditional von Neumann computer architecture is used for cosine-based operations, then computation can be performed, but energy consumption and delay are significant
Solution Approach 1:
The patent merges storage and computation functions into a unified in-memory computing architecture. FeFET-based storage cells simultaneously store vector data and perform multiplication operations, while Translinear circuits compute inner products and sum of squares directly in memory, eliminating the need to transfer data between CPU and memory and thus reducing energy consumption and computational delay
Solution Approach 2:
The patent replaces the traditional von Neumann architecture with an in-memory computing system that uses FeFET-based storage cells and Translinear circuits to perform computations directly in memory. This substitution of the computational mechanism enables parallel processing of cosine similarity calculations for multiple vectors simultaneously, dramatically improving productivity while reducing energy consumption
2Adaptability or versatility
If in-memory computing cell is designed for Hamming code calculation, then delay and energy consumption are solved, but applicability to wider scenarios like binary neural network is limited
Solution Approach 1:
The patent designs an in-memory computing architecture that is universally applicable to multiple scenarios including binary neural networks and hyperdimensional computing. The FeFET-based storage cells and Translinear circuits can compute both Hamming distances and cosine similarities, making the system versatile for different AI applications while maintaining calculation accuracy through precise analog computation
Solution Approach 2:
The patent achieves versatility by changing the computation parameter from Hamming distance to cosine similarity while using the same in-memory computing hardware. By adjusting the mathematical operations performed by Translinear circuits (computing inner products and sum of squares instead of bit differences), the system adapts to different application requirements without sacrificing reliability
3Productivity
If existing in-memory computing cell is used for approximate cosine similarity calculation, then some computation is enabled, but implementation is not applicable to wider applications
Solution Approach 1:
The patent implements continuous analog computation throughout the in-memory computing process. Translinear circuits continuously compute inner products and sum of squares for all storage vectors simultaneously, and the WTA circuit continuously identifies the maximum value, enabling efficient nearest neighbor search while maintaining the capability to handle various application scenarios through parameter adjustment
Data Source
AI summary
Disclosed are an in-memory computing architecture for a nearest neighbor search of a cosine distance and an operating method thereof. The in-memory computing architecture comprises two FeFET-based storage arrays, Translinear circuits and a WTA circuit, and the two storage arrays are a first storage array and a second storage array, respectively; wherein each of the storage cells comprises a FeFET and a resistor which are electrically connected; an input vector is inputted into the first storage array for outputting the inner product X of the input vector multiplied by all the storage vectors in the first storage array; the second storage array outputs the sum of squares Y of all vector elements in the storage vectors; the output values of the first storage array and the second storage array are respectively inputted into the Translinear circuits through current mirrors; and the Translinear circuits output X2/Y to the WTA circuit.


