Vector database based on three-dimensional storage calculation
By performing storage and computing fusion and three-dimensional integration in the vector database, the problems of low accuracy, long search time and high cost in the existing technology are solved, and accurate and fast search of high-dimensional giant vector databases are achieved and system costs are reduced.
Patent Information
- Application Number
- CN202411919272.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-12-24
- Publication Date
- 2025-06-27
AI Technical Summary
When processing high-dimensional giant vector data, existing vector databases have problems such as low accuracy, long search time and high cost, especially in dynamic data and dynamic algorithm scenarios.
By fusing processing circuits into storage circuits, specifically, sinking vector distance calculation circuits (VDCCs) into storage arrays, adopting large-scale parallel computing, implementing memory fusion, and supporting three-dimensional integration to improve integration density and reduce costs.
It realizes accurate and fast search of high-dimensional giant vector databases, reduces the relationship between search time and data scale, improves search accuracy, and reduces system costs, so that the vector database can support dynamic data and dynamic algorithms.
Smart Images

Figure CN120215816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of databases, and more particularly, to vector databases. Background Art
[0002] A vector database is a database specifically designed to store and query high-dimensional vectors. It is widely used in fields such as artificial intelligence (AI). The spread of vector databases benefits from two trends: one trend is the explosive growth of unstructured data (such as images, videos, audio, text, etc.); the other trend is the development of vector embeddings, which can effectively convert unstructured data into feature vectors. Although vector databases originate from traditional databases, they have made significant technological improvements to traditional databases for AI. Given the wide application of AI and the importance of vector databases to AI, vector databases have become an independent field.
[0003] For traditional databases (storing scalar data), retrieval is to find an exact match of keywords. For vector databases (storing vector data), the search is similarity-based, that is, to find the data vector that is most similar to the query vector. Vector databases can use many different algorithms, all of which are based on approximate nearest neighbor (ANN) search and allow for fast and efficient searching of a large amount of relevant information.
[0004] Figure 1A represents a typical vector database 70. It stores a series of data vectors V0, V1, V2… V i … The vector is a mathematical representation of a data point in a multi-dimensional space, where each dimension corresponds to a specific feature or attribute. The dimension of the data vectors in the vector database 70 ranges from dozens to thousands. The input 76 of the vector database 70 includes the query vector V q , and the output 77 includes the result vector V r . V r is the data vector that is most similar to V q .
[0005] Figure 1B represents a process of querying the vector database 70. The two-dimensional vector space d1-d2 is a simplified vector space. Each data vector V i is represented by a point in the two-dimensional vector space, and the query vector V q is also represented by a point in the vector space. Their distance D i is the vector distance between the query vector V q and the data vector V i . The retrieval process of the vector database 70 is actually to find the one among numerous V i that is most similar to V qThe one or more data vectors with the closest distance (since the retrieval of the vector database is called search based on similarity). In theory, all Vs in the vector database 70 i need to be "traversed", that is, for each V i the distance D between it and V q needs to be calculated, and then comparison and sorting are performed. In the i example of Figure 1B , the data vector V m and V q have the minimum distance D m , so V m is the nearest neighbor of V q ; the data vector V m-1 and V q have the second minimum distance D m-1 , so V m-1 is the second nearest neighbor of V q ; and so on. Generally speaking, the data vectors with k (k≥1) minimum distances (V m , V m-1 …) are the first k nearest neighbors of V q , usually called the first k (top-k) nearest neighbors (KNN). The above process of traversal retrieval is called "brute-force search" (or, "exhaustive search").
[0006] Brute-force search is not practical for traditional hardware (especially the von Neumann architecture). Existing storage technologies can easily meet the storage requirements of the vector database. Existing storage technologies only store passively and do not have supporting search capabilities. Since the processing circuit and the storage circuit are separated in traditional hardware, each processing circuit needs to process at least the data (Tb level) of one storage chip, and the brute-force retrieval takes hours, which is much longer than the time acceptable to humans.
[0007] To solve the above problems, existing technologies usually use software methods - especially indexing methods - to accelerate the search of the vector database. Common indexing methods include quantization-based indexing, graph-based indexing, and tree-based indexing. Figure 1C shows quantization-based indexing. It applies the quantizer Z to map the vectors V1 - V 10 to three buckets 10 to 30, and their centroids are c1 to c3 respectively. Therefore, Z(V1), Z(V2), Z(V3), Z(V4) = c1; Z(V5), Z(V6), Z(V7) = c2; and, Z(V8), Z(V9), Z(V 10 ) = c3. During the query process, the query vector V q is first mapped to bucket 20 by the quantizer Z, that is, Z(V q) = c2, and then search for its nearest neighbor within the bucket 20, finding it to be V7. By restricting the search to a single bucket, the quantization-based index can accelerate queries.
[0008] The accuracy of the index is relatively low. As shown in Figure 1C, for the query vector V q `, since it is quantized into the bucket 20, i.e., Z(V q `) = c2, its nearest neighbor is determined to be V5. In fact, since V q ` is closer to V2 in bucket 10, its nearest neighbor should be V2. When the query vector is close to the boundary of the bucket, this type of error is inevitable. For high-dimensional vectors (dimension > 1,000), the above errors are more likely to occur. Another problem encountered by the index is that it requires pre-computation, i.e., the index file of the entire vector database needs to be calculated before the query. For a giant vector database (vector scale Z > 10 9 ), the index calculation is computationally intensive and time-consuming. In addition, the index is not suitable for dynamic data or dynamic algorithms. Once the index is calculated, if there are any changes in the vector data or the vector distance algorithm, the entire index needs to be recalculated. Finally, during the query process, the index file needs to be loaded into the main memory, which requires a large amount of main memory, resulting in a high system cost.
[0009] Although the index is suitable for traditional databases (storing scalar-like data), it is not suitable for vector databases (storing vector-like data). Index-based vector databases are usually limited to small dimensions (thousands of dimensions) and small scales (billions of 10 9 -level vectors). This may be suitable for small knowledge bases, but it is far from enough for a general knowledge base (world knowledge base) that stores all human knowledge. The world knowledge base contains trillions (10 12 )-level to quadrillions (10 15 )-level vectors. For a giant vector database of this scale, it is difficult to handle solely by software, and hardware improvement is required. Summary of the Invention
[0010] The main objective of the present invention is to provide a technical path to achieve Generative AI and large language models (LLMs) by replacing "big computing power" with "big storage power". "Big storage power" not only requires a huge (passive) storage space but also a high-speed (active) search ability.
[0011] Another objective of the present invention is to provide a high-dimensional giant vector database with high precision, high performance, and low cost.
[0012] Another objective of the present invention is to support a vector database for dynamic data.
[0013] Another object of the present invention is a vector database that supports a dynamic vector distance algorithm.
[0014] Another object of the present invention is to overcome the deficiencies of the indexing method.
[0015] To achieve these and other objects, the present invention improves the existing vector database. Different from traditional hardware (such as the von Neumann architecture, where the processing circuit and the storage circuit are separated, i.e., storage and computing are separated), the present invention integrates the processing circuit into the storage circuit, i.e., storage and computing are integrated. Specifically, it sinks the processing circuit - the vector distance calculation circuit (VDCC) - to the bottom layer of the storage circuit - the storage array. Since the storage array is the smallest building block of the storage circuit, the granularity of the present invention (the amount of data to be processed by each VDCC) is the smallest (array Mb level), much smaller than the granularity of traditional hardware (chip Tb level). By significantly reducing the granularity and using large-scale parallel computing, the present invention can achieve accurate and fast brute-force search for high-dimensional giant vector databases.
[0016] The vector database of the present invention is stored in a storage and computing body, which includes multiple storage and computing cores (SPCs), and each storage and computing core includes multiple storage and computing units (SPUs) and a minimum distance search circuit (MDSC). Here, the storage and computing core is the basic building block of the storage and computing body at the physical level, and the storage and computing unit is the basic building block of the storage and computing core at the circuit level. Each storage and computing unit includes at least one storage array and a vector distance calculation circuit (VDCC). The MDSC searches for the minimum distance D from the distances D calculated by the VDCCs of all storage and computing units. m Brute-force search will scan all data vectors in the vector database, calculate their distances (calculated by the VDCC) from the query vector V in sequence, and then search for D (performed by the MDSC) by comparing them with each other to find the nearest neighbor V of V. q m (performed by the MDSC) and find the nearest neighbor V of V. q m
[0017] Accordingly, the present invention proposes a vector database based on storage and computing fusion, comprising an input (110) having at least one query vector (115) and at least one storage and computing core (100) coupled to the input (110), characterized in that the storage and computing core (100) contains: a plurality of storage and computing units (100ij...), each storage and computing unit (100ij) containing at least one storage array (170ij) and a vector distance calculation circuit VDCC (185), the storage array (170ij) storing at least one data vector (175) in the vector database, the VDCC (185) calculating the vector distance between the query vector (115) and the data vector (175); a minimum distance search circuit MDSC (190), the MDSC (190) searching for at least one minimum distance D from the vector distances calculated by all VDCCs of the plurality of storage and computing units (100ij...). m .
[0018] Since brute force search does not use any approximation methods (such as indexing), the present invention can obtain accurate results: the nearest neighbors found by MDSC are the true nearest neighbors. This solves the problem of low accuracy of high-dimensional vector databases in the prior art. In addition, due to the use of large-scale parallel computing, the brute force search time is basically independent of the size Z of the vector database: whether Z is a trillion (10 12 ) level, quadrillion (10 15 ) level or even higher, and the brute force search time is roughly the same (10ms level). This is far superior to the prior art. In addition, the present invention also supports dynamic data (i.e., data vectors in the vector database can be changed, added and / or deleted) and dynamic algorithms (i.e., VDCC uses different vector distance algorithms in different scenarios).
[0019] The present invention further supports the top-k nearest neighbors (k is an integer > 1). MDSC is suitable for searching V q The minimum distance D m To fully utilize MDSC, find and output D m After that, the reset circuit is used to store D m The register is reset to a predetermined value (usually the maximum value, such as each bit is reset to 1). m Excluded from the next round of minimum distance search process. Therefore, the minimum distance found in the next round is actually the second minimum distance D m-1 . Repeat the above process to find all the first k nearest neighbors.
[0020] Traditional memory - computing integration adopts two - dimensional integration, that is, the memory circuit and the logic circuit are integrated on the same substrate in a side - by - side manner. This leads to two problems: one is that two - dimensional integration increases the chip area (the memory array becomes sparser), increasing the cost; the other is that the manufacturing processes of the memory circuit and the logic circuit are incompatible. Blindly integrating the two on the same substrate results in a complex process flow, further increasing the cost. To overcome the above problems, the present invention proposes to use three - dimensional integration to achieve memory - computing integration. In the memory - computing unit based on three - dimensional integration (three - dimensional memory - computing unit), the memory circuit and the logic circuit are in different physical planes: the memory array is located on the first plane, and the VDCC and MDSC are located on the second plane. The first and second planes are stacked on top of each other and are coupled through inter - plane connections. One function of these inter - plane connections is to provide a high - bandwidth electrical connection between the two planes; another function is to provide a high - strength mechanical connection between the two planes.
[0021] Accordingly, the present invention proposes a vector database based on three - dimensional integration, including an input (110) having at least one query vector (115) and at least one memory - computing core (100) coupled to the input (110), characterized in that the memory - computing core (100) comprises: a plurality of memory - computing units (100ij…), each memory - computing unit (100ij) comprising at least one memory array (170ij) and a vector distance calculation circuit VDCC (185), the memory array (170ij) storing at least one data vector (175) in the vector database, the VDCC (185) calculating the vector distance between the query vector (115) and the data vector (175); at least one semiconductor substrate (0), the VDCC (185) being located within the semiconductor substrate (0); the memory array (170ij) being stacked on top of the VDCC (185); the projection of the VDCC (185) on the semiconductor substrate (0) at least partially overlapping with the memory array (170ij); and the VDCC (185) being electrically coupled to the memory array (170ij).
[0022] There are two forms of the memory - in - computing core (3D memory - in - computing core) based on 3D integration: one is a single - chip form, and the other is a chip - stack form. The single - chip form adopts 3D monolithic integration, where its memory circuit (including the 3D - M memory array) is stacked on top of its logic circuit (including VDCC, MDSC, and the peripheral circuits of the 3D - M array), and they are all located in the same chip. The inter - layer connection is a contact via hole, which serves as the intra - chip connection within the single - chip. On the other hand, the chip - stack form adopts 3D bonding integration, which contains multiple bonded chips (such as memory chips and logic chips). Its memory chip (including the memory array) and logic chip (including VDCC, MDSC, and part of the peripheral circuits of the memory array) are bonded after stacking, and the inter - layer connection is an inter - chip connection. In the chip - stack, the memory circuit can be a traditional 2D memory (such as RAM, ROM, or NVM), or it can be 3D - M. An important type of chip - stack is a pair of two chips (chip pair) using face - to - face hybrid bonding.
[0023] Similar to flash memory and solid - state drives, the memory - in - computing core can be encapsulated into a memory - in - computing card; the memory - in - computing card can be encapsulated into a memory - in - computing disk; the memory - in - computing disks can further form a memory - in - computing cluster. The memory - in - computing core, memory - in - computing card, memory - in - computing disk, and memory - in - computing cluster constitute a memory - in - computing hierarchy, and each layer can have its own vector computing circuits (such as VDCC, screening circuits, and sorting circuits). The vector computing circuits at the lower layer are usually simpler than those at the higher layer, and the screening conditions at the lower layer are more relaxed. Since the data in the vector database will be screened out layer by layer, the data output from the lower - layer computing circuits to the higher - layer computing circuits becomes less and less, which can greatly reduce the bandwidth pressure of data transmission. Brief Description of the Drawings
[0024] Figure 1A Represents a vector database and its data structure; Figure 1B Represents a simplified two - dimensional vector space and the corresponding query schematic diagram; Figure 1C represents a quantization - based indexing algorithm (prior art).
[0025] Figure 2A Is the architecture of the memory - in - computing core (SPC) in a vector database; Figure 2B Is the circuit block diagram of the memory - in - computing unit (SPU) in this memory - in - computing core.
[0026] Figure 3 Is a perspective view of the physical structure of a 3D memory - in - computing unit.
[0027] Figures 4A - 4B Are cross - sectional views of two memory - in - computing single - chips.
[0028] Figures 5A - 5B Are cross - sectional views of two memory - in - computing chip - stacks.
[0029] Figures 6A - 6CIt is a circuit block diagram of three types of 3D memory - computing units.
[0030] Figures 7A - 7C It is a layout diagram of three types of 3D memory - computing units.
[0031] Figures 8A - 8B It is a circuit diagram of two types of vector distance calculation circuits (VDCCs).
[0032] Figure 9 It is a block diagram of a reconfigurable memory - computing unit.
[0033] Figure 10 It is a circuit diagram of a reconfigurable VDCC.
[0034] Figure 11 It is a circuit diagram of a minimum distance search circuit (MDSC).
[0035] Figure 12 It is a circuit diagram of another MDSC.
[0036] Figure 13 It is a flowchart of a top - k nearest neighbor search algorithm.
[0037] Figure 14 It is a block diagram of a memory - computing hierarchy.
[0038] Figure 15 It is a basic block diagram of a hybrid VDCC.
[0039] Figures 16A - 16D It is a block diagram of four embodiments of hybrid VDCCs.
[0040] Note that these figures are only schematic diagrams and are not drawn to scale. For visibility and convenience, some dimensions and structures in the figures may be enlarged or reduced. In different embodiments, the letter suffixes or superscripts / subscripts after numbers represent different instances of the same type of structure; the same number prefixes represent the same or similar structures. In this article, " / " represents the relationship of "and" or "or".
[0041] In this article, "vector database" can refer to a vector database in the mathematical sense (i.e., an organized set of vectors), or it can refer to the hardware / software for implementing a vector database (hardware includes processors, storage, and related architectures; software includes index files and related indexing algorithms). This article does not make a distinction here.
[0042] In this article, "storage body" refers to any semiconductor information storage device that can store information permanently or temporarily, such as random access memory (RAM), read-only memory (ROM), non-volatile memory (NVM), or three-dimensional memory (3D-M). "Storage array" is the basic building block of a memory chip, which is a collection of storage elements that share at least one address line. "Storage-computing body" refers to any storage body with internal computing power (the ability to process the data it stores), such as storage-computing core, storage-computing card, storage-computing disk, or storage-computing cluster.
[0043] In this article, "core" refers to a physical entity with mechanical strength. It has sufficient mechanical strength in the direction perpendicular to the substrate, and external mechanical force cannot pry it open before destroying the substrate. "Memory and computing core" is the basic building block of the memory and computing body at the physical level. There are two forms of memory and computing core: one is a single chip (containing a single chip), and the other is a chip stack (containing multiple bonded chips, such as a chip pair).
[0044] In this article, "in the substrate" means that at least part of the active components (such as transistors, memory cells) are located in the substrate, even if the interconnections between them are located above the substrate. "Above the substrate" means that the active components (such as transistors, memory cells) are all located above the substrate, not in the substrate. Unless otherwise specified, "upper" and "lower" in this article refer to the relative positions when the chip is upright (substrate at the bottom). Even if the chip is flipped, the meaning of "upper" and "lower" still follows the meaning when the chip is upright.
[0045] In this article, "plane" refers to a plane in an abstract sense, which may include a physical structure with a finite thickness. For example, a plane may include a memory array with multiple memory layers, or a circuit with multiple interconnect layers. "A first plane stacked on a second plane" means that a first physical structure (such as a memory array) with a first finite thickness is stacked on a second physical structure (such as a substrate circuit) with a second finite thickness. DETAILED DESCRIPTION
[0046] The present invention improves the existing vector database. Unlike traditional hardware (such as the von Neumann architecture, where the processing circuit is separated from the storage circuit, i.e., storage and computation are separated), the present invention integrates the processing circuit into the storage circuit, i.e., storage and computation are integrated. Specifically, it sinks the processing circuit - the vector distance calculation circuit (VDCC) - to the bottom layer of the storage circuit - the storage array. Since the storage array is the smallest building block of the storage circuit, the granularity (the amount of data required to be processed by each VDCC) of the present invention is the smallest (array Mb level), which is much smaller than the granularity of traditional hardware (chip Tb level). By significantly reducing the granularity and using large-scale parallel computing, the present invention can achieve accurate and fast brute force search of high-dimensional giant vector databases.
[0047] The vector database of the present invention is stored in a memory - computing unit, and the memory - computing unit includes multiple memory - computing cores (SPCs), and each memory - computing core includes multiple memory - computing units (SPUs) and a minimum - distance search circuit (MDSC). Figure 2A Disclosed is a memory - computing core (SPC) 100 based on memory - computing fusion. The memory - computing core 100 is the basic building unit of the memory - computing unit at the physical level, and it is an entity with mechanical strength. The memory - computing core 100 can not only store data (vector database) but also process data (calculate vector distance), and these data are stored locally and are very close to the processing circuit. The input 110 of the memory - computing core 100 includes a query vector V q and the output 120 includes a result vector V r . The memory - computing core 100 is located on a semiconductor substrate 0 and contains m x n memory - computing units (100aa…100ij…100mn) and a minimum - distance search circuit (MDSC) 190.
[0048] Figure 2B Disclosed is a memory - computing unit 100ij. The memory - computing unit 100ij is the building unit of the memory - computing core 100 at the circuit level, and it includes a storage circuit 170 and a logic circuit 180. The storage circuit 170 includes a storage array 170ij which stores at least one data vector V i 175 of the vector database. The logic circuit 180 includes a vector - distance calculation circuit (VDCC) 185. The inputs 110, 175 of the VDCC 185 include the query vector V q and the data vector V i , and the output 150 ( Figure 2A the 150aa…150mn in) includes the vector distance D q between V i and V i . The MDSC 190 compares the vector distances D i from all memory - computing units (100aa…100mn), and searches for the minimum distance D m and its corresponding minimum - distance data vector V m . V m is the result vector V r .
[0049] Since brute-force search does not use any approximation methods (such as indexing), the present invention can obtain accurate results: the nearest neighbors found by MDSC190 are the true nearest neighbors. This solves the problem of low accuracy in high-dimensional vector databases in the prior art. In addition, since the data reading and distance calculation of vectors are executed in parallel in all memory-computation units 100ij, and the minimum vector search time (the only quantity related to the scale Z of the vector database) has a very weak relationship with Z, and its value is much smaller than the reading time of vectors, the brute-force search time of the present invention is basically independent of Z: regardless of whether Z is in the trillions (10 12 ), quadrillions (10 15 ), or even higher, the brute-force search time is approximately the same (on the order of 10 ms), which is far better than the prior art.
[0050] Traditional memory-computation integration adopts two-dimensional integration, that is, the memory circuit 170 and the logic circuit 180 are integrated on the same plane side by side. This causes two problems: one is that two-dimensional integration increases the chip area (the memory array becomes sparser), increasing the cost; the other is that the manufacturing processes of the memory circuit 170 and the logic circuit 180 are incompatible, and blindly integrating the two on the same plane results in a complex process flow, further increasing the cost. To overcome the above problems, the present invention proposes to use three-dimensional integration to achieve memory-computation integration.
[0051] As Figure 3 shown, in the memory-computation unit based on three-dimensional integration (three-dimensional memory-computation unit) 100ij, the memory circuit 170 includes a memory array 170ij and is located on the first plane 0L1, and the logic circuit 180 includes a VDCC 185 and an MDSC 190 and is located on the second plane 0L2. The first and second planes are stacked on top of each other and are coupled through inter-plane connections 160. One function of these inter-plane connections 160 is to provide a high-bandwidth electrical connection between the two planes 0L1 and 0L2, and the other function is to provide a high-strength mechanical connection.
[0052] The three-dimensional memory-computation core 100 has two forms: one is a single chip ( Figure 4A and Figure 4B ), and the other is a chip stack ( Figure 5A and Figure 5B). The single chip adopts three-dimensional monolithic integration. Among them, the storage circuit 170 (the first plane 0L1) includes a three-dimensional memory (3D-M) array, and the logic circuit 180 (the second plane 0L2) includes VDCC and MDSC. The peripheral circuits of VDCC, MDSC, and the 3D-M array are located in the semiconductor substrate 0. The 3D-M array 170 is stacked on top of the logic circuit 180, and they are all located in the same chip. The inter-plane connection 160 is a contact via hole, which is an intra-chip connection within the single chip. On the other hand, the chip stack adopts three-dimensional bonding integration. It contains a bonded memory chip 100a (the first plane 0L1) and a logic chip 100b (the second plane 0L2). Among them, the memory chip 100a includes a memory array, and the logic chip 100b includes VDCC, MDSC, and part of the peripheral circuits of the memory array. The memory chip 100a and the logic chip 100b are stacked and bonded to each other, and the inter-plane connection 160 is an inter-chip connection between the memory chip 100a and the logic chip 100b. In the chip stack, the storage circuit 170 can be a traditional two-dimensional memory (such as RAM, ROM, or NVM), or it can be 3D-M. For example, when the storage circuit 170 is an internal memory DRAM, the present invention is an internal memory with a built-in minimum distance search circuit; when the storage circuit 170 is an external memory flash, the present invention is an external memory with a built-in minimum distance search circuit; when the storage circuit 170 is a medium memory 3D-XPoint (the speed of 3D-XPoint is between that of the internal memory DRAM and the external memory flash, and is referred to as "medium memory" in this article), the present invention is a medium memory with a built-in minimum distance search circuit.
[0053] The single chip 100 stacks the storage circuit 170 on top of the logic circuit 180. Therefore, the integration of the storage circuit 170 and the logic circuit 180 does not increase the chip area, which is a progress for two-dimensional integration. Figure 4A and Figure 4B Two types of memory-computing single chips 100 are disclosed. They are integrated 3D-M, that is, the 3D-M array 170ij and the logic circuit 180 are three-dimensionally monolithically integrated. Different from traditional memories (two-dimensional memories, that is, memory cells are located on a two-dimensional plane), the memory cells of the 3D-M array 170ij are located in three-dimensional space. Note that the underlying substrate circuit 0K includes VDCC 185, MDSC 190, and the peripheral circuit of the 3D-M array 170ij.
[0054] According to its physical structure, 3D-M is divided into three-dimensional horizontal memory (3D-M H ) and three-dimensional vertical memory (3D-M V ). In 3D-M H , all address lines are horizontal, and its memory cells form multiple horizontal memory layers, and these horizontal memory layers are stacked on top of each other on the substrate circuit. In 3D-M HA typical example is 3D-XPoint. 3D-M V In V , at least one set of address lines is vertical, and its memory cells form multiple vertical memory strings along the vertical address lines, and the memory strings are arranged side by side on the substrate circuit. 3D-M V A typical example is 3D-NAND. 3D-M H Faster, 3D-M V Higher storage density.
[0055] According to its function, 3D-M is divided into 3D-RAM (three-dimensional random access memory) and 3D-ROM (three-dimensional read-only memory). 3D-RAM is used for caching; 3D-ROM can store information permanently, and it is also called non-volatile memory (NVM). Memory types applicable to 3D-M include memristors, resistive random access memories (RRAM or ReRAM), phase change memories (PCM), programmable metallization cell (PMC) memories, conductive bridge memories (CBRAM), etc.
[0056] In Figure 4A Figure 4A , the logic circuit 180 (including VDCC 185) is part of the substrate circuit 0K (formed by transistor 0t and substrate interconnections 0m1, 0m2). The memory circuit 170 includes a 3D-M H array. In this embodiment, the memory circuit 170 includes four address line layers 0a1-0a4. Each address line layer (e.g., 0a1) includes multiple address lines (e.g., 1a). These address line layers 0a1-0a4 form two memory layers 16A, 16B, where the memory layer 16A is stacked on the substrate circuit 0K, and the memory layer 16B is stacked on the memory layer 16A. Memory cells (e.g., 7aa) are located at the intersections between two address lines (e.g., 1a, 2a). Contact via holes 1av, 3av implement the on-chip (inter-plane) connection 160 between the 3D-M H array 170 and the logic circuit 180.
[0057] In Figure 4A Figure 4A , the memory cell 7aa includes a programmable layer 5 and a diode layer 6. The programmable layer 5 can be an antifuse layer (programmable once) or a resistive RAM (RRAM) layer (programmable repeatedly). A diode generally refers to a device whose resistance is much greater than the resistance under the read voltage when the value of the applied voltage is less than the read voltage or the polarity is opposite. For those skilled in the art, a diode is also called a selector, a steering element, a unidirectional switch, etc.
[0058] Figure 4B Similar to Figure 4A , the only difference is that the 3D-M array is 3D-M VArray: The memory circuit 170 includes eight stacked horizontal address line layers 0a1 - 0a8. Each horizontal address line layer (such as 0a5) includes multiple horizontal address lines (such as 15). The memory circuit 170 also includes a set of vertical address lines 19 perpendicular to the surface of the substrate 0. For simplicity, 3D-M is not marked in this figure. V The on-chip (inter-plane) connections 160 between the array 170 and the substrate circuit 0K are well-known to those skilled in the art.
[0059] Figure 4B 3D-M in V The array is based on vertical transistors or transistor-like devices and includes multiple memory strings 16X, 16Y placed side by side. Each memory string (such as 16Y) includes multiple stacked memory cells (such as 18ay - 18hy). Each memory cell (such as 18fy) contains a vertical transistor that includes a gate (horizontal address line) 15, a memory layer 17, and a vertical channel (vertical address line) 19. The memory layer 17 can include an oxide-nitride-oxide layer, an oxide-polysilicon-oxide layer, etc. Additionally, 3D-M V The array can also be based on diodes or diode-like devices.
[0060] Different from a single chip, the chip stack 100 contains multiple bonded chips. The chip stack 100 stacks the memory chip 100a (containing a semiconductor substrate 0`) on top of the logic chip 100b (containing a semiconductor substrate 0), so the integration of the memory circuit 170 and the logic circuit 180 does not increase the chip area. More importantly, since the memory circuit 170 and the logic circuit 180 are formed on the memory chip 100a and the logic chip 100b respectively, different manufacturing processes can be used to optimize the memory circuit 170 and the logic circuit 180 respectively, which will simplify the design of the process flow and further reduce costs.
[0061] Figure 5A The chip stack 100 in uses face-to-back bonding, that is, the surface of the second chip 100b is bonded to the back of the first chip 100a. During the manufacturing process, the through-silicon vias (TSVs) 161 are etched through the first chip 100a. The first and second chips 100a, 100b are both facing up (+z direction) during bonding. The first chip 100a is coupled to the second chip 100b through micro-bumps 162. The inter-chip (inter-plane) connections 160 are realized through the TSVs 161 and the micro-bumps 162.
[0062] Figure 5BThe chip stack 100 in [[ ]] uses face-to-face bonding, that is, the first chip 100a is flipped and facing down (i.e., along the -z direction), while the second chip 100b is facing up (i.e., along the +z direction). Two (a pair of) chips using face-to-face hybrid bonding are called a "chip pair". The two chips 100a and 100b are bonded by metal-metal (such as copper-copper) bonding. During the manufacturing process, a first dielectric film 166a is deposited on the top of the first chip 100a, and a first via 164a is etched in the first dielectric film 166a. Then, a second dielectric film 166b is deposited on the top of the second chip 100b, and a second via 164b is etched in the second dielectric film 166b. After that, the first chip 100a is flipped, the first via 164a and the second via 164b are aligned, and the first chip 100a and the second chip 100b are bonded. The first via 164a and the second via 164b that form electrical contacts achieve inter-chip (plane) connection 160. Since when hybrid bonding, the first via 164a and the second via 164b do not need to pass through any substrate, they can be made small in size and numerous in quantity. Therefore, the chip pair can obtain a large bandwidth and provide good mechanical strength. As is well known to those skilled in the art, in addition to metal-metal bonding, dielectric-semiconductor bonding, dielectric-dielectric bonding, etc. can also be used.
[0063] Figure 5A and Figure 5B Embodiments of [[ ]] can all adopt wafer-scale bonding, and after wafer bonding, they are cut uniformly. Therefore, the first and second chips 100a and 100b have the same chip size, and all their edges are aligned. For example, the left edge of the first chip 100a is aligned with the left edge of the second chip 100b; the right edge of the first chip 100a is aligned with the right edge of the second chip 100b.
[0064] Figures 6A - Figures 7C How to enhance the function of the logic circuit 180 is illustrated by comparing three embodiments. In these three-dimensional memory and computing units 100ij, the VDCC 185 serves different numbers of memory arrays 170ij (calculating vector distance).
[0065] In [[ ]] Figure 6A , the VDCC 185 serves a single memory array 170ij. In [[ ]] Figure 6B , the VDCC 185 serves four memory arrays 170ijA - 170ijD. In [[ ]] Figure 6C , the VDCC 185 serves nine memory arrays 170ijA - 170ijI. As [[ ]] Figures 7A - 7C shows, the more memory arrays the VDCC 185 serves, the larger its area and the more functions it has. In [[ ]] Figures 6A - Figures 7CAmong them, since the storage array 170ij and the VDCC 180ij are located in different physical layers 0L1, 0L2 (see Figure 3 ), the storage array 170ij is represented by a dashed line.
[0066] Figures 7A - Figures 7C Shows the projection of the storage array 170ij (physically located on the first plane 0L1 drawn with a dashed line) on the second plane 0L2 and their relative positions with respect to the VDCC 185. Figure 7A The embodiment corresponding to Figure 6A has an embodiment in which the VDCC 185 is covered by the storage array 170ij. In this embodiment, the pitch of the VDCC is equal to the pitch of the storage array. Since its area is smaller than the footprint of the storage array, Figure 7A the VDCC 185 in Figure 8A has limited functionality and can be used to implement a simple distance algorithm (such as
[0067] Figures 7B - 7C Shows two complex VDCCs 185, which can be used to implement complex vector distance algorithms (such as Figure 8B and Figures 9 - 10 ). Figure 7B The embodiment corresponding to Figure 6B has an embodiment in which the VDCC 185 is covered by four storage arrays 170ijA - 170ijD. Below the four storage arrays 170ijA - 170ijD, the VDCC 185 can be freely laid out. Since the VDCC pitch is twice the pitch of the storage array and the chip area of the VDCC 185 is four times that of each storage array, it can have complex functionality.
[0068] Figure 7C The embodiment corresponding to Figure 6C has an embodiment in which the VDCC 185 is covered by nine storage arrays (170ijA…170ijI). Below the nine storage arrays (170ijA…170ijI), the VDCC 185 can be freely laid out. Since the VDCC pitch is three times the pitch of the storage array and the chip area of the VDCC 185 is nine times that of each storage array, it can have more complex functionality.
[0069] Figures 8A - Figures 10 Discloses a variety of VDCCs 185. Figure 8A and Figure 8B Show two VDCCs 185 that implement the Manhattan distance and the Euclidean distance respectively. Figure 8A The first VDCC 185 in i [m] calculates the Manhattan distance. The subtract - absolute value (sub - abs) unit 282 calculates the difference Di [m] -V q [m] | (i.e., the absolute value of the subtraction), where V i [m] is the data vector V i the value of the m-th dimension of 175 (from the storage circuit 170); V q [m] is the query vector V q the value of the m-th dimension of 115 (from the input 110). Then, the accumulator 288 (including the adder 284 and the register 286) sums D i [m] for all m dimensions. Since only subtraction (and taking the absolute value) needs to be performed, this VDCC 185 can be implemented by a relatively simple circuit, such as Figure 7A the VDCC 185 in
[0070] Figure 8B The second VDCC 185 in i [m] calculates the Euclidean distance. The subtraction-squaring (sub-sqr) unit 283 calculates the difference D i [m] = (V q [m] ) 2 (i.e., the square of the subtraction). Then, the accumulator 288 sums D i [m] for all m dimensions, and then performs a square root operation. Since it involves square and square root operations, the second VDCC 185 requires a relatively complex circuit to implement, such as Figure 7B the VDCC 185 in
[0071] As is well known to those skilled in the art, in addition to the Manhattan distance and the Euclidean distance, other distance algorithms can also be used, including: Chebyshev distance, Minkowski distance, cosine similarity, Haversine distance, Pearson correlation coefficient, Earth mover’s distance (i.e., EMD), Jaccard similarity, Sorensen-Dice index, Hamming distance, inner product, etc.
[0072] Generally speaking, different scenarios require the use of different distance algorithms (i.e., dynamic algorithms). To support dynamic algorithms, the present invention also discloses a reconfigurable vector database, whose vector distance algorithm can be reconfigured during use. As Figure 9As shown, the memory - computing unit is a reconfigurable memory - computing unit 100ij, and its VDCC is a reconfigurable VDCC 185. By adjusting the configuration parameter P c 118, the vector distance algorithm executed by the VDCC 185 can be reset. The input 110 also includes the configuration parameter P c 118. As Figure 2A shown, the input 110 sends P c 118 to each memory - computing unit 100ij, and they are configured simultaneously. For example, at time t1, all memory - computing units 100ij are configured with the first distance algorithm; at time t2, all memory - computing units 100ij are configured with the second distance algorithm; and so on.
[0073] Figure 10 Disclosed is a reconfigurable VDCC 185, which implements the reconfigurable memory - computing unit 100ij by using a weighted distance algorithm. In this embodiment, the configuration parameter P c 118 is the weight W 。 The subtractor 285 calculates the difference D q [m] between the m - th dimensional value of the query vector V d [m] and the data vector V i [m] =V i [m] -V q [m] . Then, the multiplier 287 multiplies the difference D i [m] by the value of the m - th dimension in the weight W [m] (i.e., weighted multiplication), and the vector distance D is the sum of all m - dimensional D i [m] ×W [m] . By adjusting the weight W [m] , different distance algorithms can be used, which is useful for optimizing the search algorithm of the vector database.
[0074] Figures 11 - Figures 13 Multiple MDSCs 190 are disclosed. Figure 11 The first MDSC 190 in uses a binary - tree comparison circuit to search for the minimum distance. This embodiment has 16 memory - computing units, and its outputs 150aa - 150dd are vector distances D1 - D 16 . At the comparator 130 11 -130 18 , the vector distances D1 - D 16 are compared pairwise. Each comparator (such as 130 11 ) has four inputs (such as D1, A1; D2, A2) and two outputs (such as D 11, A 11 ). Among them, D1 is the output from the first memory - computing unit, A1 is the address of the first memory - computing unit; D2 is the output from the second memory - computing unit, A2 is the address of the second memory - computing unit; D 11 is the smaller one of D1 and D2, and A 11 is the corresponding address to it (the smaller value). After pairwise comparison at the comparator 130 11 - 130 18 , the result D 11 - D 18 is sent to a higher - level comparator 130 21 - 130 24 for another round of comparison. This process is repeated until the top - level comparator 130 41 finds the minimum vector distance D m and its A m . Then, this A m is sent to the storage circuit 170 to find the minimum - distance vector V m , which is the nearest neighbor of V q (i.e., the result vector V r 120).
[0075] Figure 11 is a simplified case where the storage circuit 170 in each memory - computing unit 100ij stores only one data vector. Generally speaking, the storage circuit 170 in each memory - computing unit 100ij stores X (X is a positive integer) data vectors. At this time, it takes X cycles to read out all the data vectors. For example, in the first cycle, the first set of data vectors is read out from all memory - computing units 100ij, and the D m1 of the first cycle is found by MDSC 190 and stored in the D m1 register; in the second cycle, the second set of data vectors is read out from all memory - computing units 100ij, and the D m2 of the second cycle is found by MDSC 190 and stored in the D m2 register; and so on. After all X cycles, compare D m1 , D m2 ... to find the D m of the entire vector database.
[0076] The present invention also supports the top - k (k is an integer > 1) nearest neighbors. MDSC is only suitable for searching for the minimum distance D q of V m . To make full use of MDSC, a reset circuit is added. After finding D m , the storage of D m is reset using the reset circuit.The distance register is reset to a predetermined value (usually the maximum distance, for example, each bit is reset to 1). This excludes D m from the next round of minimum distance search process. Therefore, the minimum distance found in the next round is actually the second minimum distance D m-1 . By repeating the above process, all the first k nearest neighbors can be found.
[0077] Figures 12 - Figures 13 Disclosed are an MDSC 190* for finding the first k nearest neighbors and its method. For this MDSC190* ( Figure 12 ), the output of each VDCC 185 (such as 150aa or vector distance D1) is stored in a distance register (such as REG D1 140aa). The MDSC 190* also includes a reset circuit 135, whose input is the address A m corresponding to the minimum distance D m , and the outputs R1, R2,... are respectively connected to the reset terminals of the distance registers REG D1 140aa, REG D2 140ab..., so that the distance register corresponding to A m can be reset to a predetermined value (usually the maximum distance, for example, each bit is reset to 1).
[0078] V m search loop ( Figure 13 ) is similar to Figure 11 : The VDCC 185 calculates the vector distances D of all memory and computing units (step 610); the MDSC 190 searches for the minimum distance D m and its corresponding address A m (step 620); based on A m , V m is found from the storage circuit 170 (step 630). After finding V m , the next step is to search for V m-1 (step 640 - 650). The key step is to reset the distance register storing A m to a predetermined value (such as all 1s). For example, if D2 is the minimum distance and A m points to 150ab, the reset circuit 135 is a decoder using A m as the input, which sets the signal on line R2 to high level, and this will reset the distance register REG D2 140ab. Since V m is excluded, the second search loop (steps 620 - 630) will obtain V m-1 . After excluding V m-1 , the third search loop (steps 620 - 630) will obtain Vm-2 Continue the search loop until the top k nearest neighbors are found and the resulting vector V stored in the result register file 145 is output r (Step 660).
[0079] Figures 14 - Figures 16D A memory - computing hierarchy 1000 and related hybrid VDCCs are disclosed. As Figure 14 shown, the memory - computing core 100 can be encapsulated into the memory - computing card 200; the memory - computing card 200 can be encapsulated into the memory - computing board 300; the memory - computing board 300 can be encapsulated into the memory - computing cluster 400, and they form the memory - computing hierarchy 1000. Each layer of the memory - computing hierarchy 1000 has its own vector computing circuits (such as VDCCs, screening circuits, and sorting circuits). The vector computing circuits of the lower layer are generally simpler than those of the higher layer, and the screening conditions of the lower layer are looser. Accordingly, the present invention proposes a hybrid VDCC. As Figure 15 shown, the hybrid VDCC includes the in - array VDCC 185 and the out - of - array VDCC 285. The in - array VDCC 185 is located within / under the storage array, while the out - of - array VDCC 285 is located outside the storage array. The complexity of the in - array VDCC 185 is generally lower than that of the out - of - array VDCC 285.
[0080] Figures 16A - Figures 16D Embodiments representing four multi - layer memory - computing bodies are shown, and they can all apply Figure 15 the hybrid VDCC. In Figure 16A the embodiment, the out - of - array VDCC 285A and the in - array VDCC 185 are on the same semiconductor substrate 0, but it is located outside all the memory - computing arrays 100aa - 100mn, and it serves all the storage - computing units 100aa - 100mn within the same memory - computing core 100. In Figure 16B the embodiment, the out - of - array VDCC 285B and multiple memory - computing cores 100A, 100B are on different chips, they overlap with each other, and are encapsulated in the same memory - computing card 200. The out - of - array VDCC 285B serves the memory - computing cores 100A, 100B. In Figure 16C the embodiment, the out - of - array VDCC 285C and multiple memory - computing cards 200A, 200B are in the same memory - computing board 300, and it serves the memory - computing cards 200A, 200B. In Figure 16D the embodiment, the out - of - array VDCC 285D serves the memory - computing boards 300A, 300B.
[0081] In Figures 16A - Figures 16D the embodiment, as the level gets higher ( Figure 16A → Figure 16D ), the function of the out - of - array computing circuit 285 becomes more powerful, and the screening conditions for data are more stringent. For example, in the memory - computing core 100 (Figure 16A ), within the array, the VDCC 185 uses a simple vector distance algorithm such as Manhattan distance (using subtraction without multiplication); outside the array, the VDCC 285A uses a more complex vector distance algorithm such as Euclidean distance (using square and square root). In the memory-computation card 200 ( Figure 16B ), outside the array, the VDCC 285B uses a complex vector distance algorithm such as standard Euclidean distance (using square, division and square root). In the memory-computation board 300 and the memory-computation cluster 400 ( Figure 16C and Figure 16D ), outside the array, the VDCCs 285C and 285D can use even more complex vector distance algorithms such as cosine similarity (using square, division and square root). In addition, the VDCC can also be a reconfigurable VDCC. Since the data in the vector database is screened layer by layer, the data output from the lower-layer computing circuit to the upper-layer computing circuit becomes less and less, which can greatly reduce the bandwidth pressure of data transmission.
[0082] An important advantage of the present invention is that its brute-force search time is basically independent of the scale Z of the vector database. Brute-force search includes the following three operations: vector reading, distance calculation, and minimum distance search. Since both vector reading and distance calculation are performed in parallel in all memory-computation units, the total brute-force search time includes: the data reading time (t wa ) of all vectors in a single memory-computation unit, the vector distance calculation time (t vd ) of all vectors in a single memory-computation unit, and the time (T ms ) to search for the minimum distance among all vector distances in the entire vector database. Among them, t wa and t vd are independent of Z and are only proportional to the storage capacity z of a single memory-computation unit; only T ms is related to Z, and its relationship is T ms ∝log2(Z) (see Figure 11 ), and this relationship is much weaker than the relationship in the prior art where the search time is proportional to Z. Taking 3D-XPoint as an example, t wa + t vd ~10ms; for a giant vector library, Z = 10 12 , T ms ~20us; Z = 10 15 , T ms ~30us. Obviously, T ms << t wa + t vd . Therefore, the brute-force search time of the present invention is basically independent of Z: regardless of whether Z is in the trillions (10 12 ) level or the quadrillions (10 15Level 0 or even higher, and the brute-force search time is roughly the same (at the 10 ms level). Compared with the prior art where the search time increases as Z increases, this consistency in search time of the present invention can greatly simplify the system design. On the other hand, for considerations such as energy consumption, it is not necessary to perform brute-force search on the entire vector database. It only needs to be performed on the vector database in the selected partition (such as the relevant field).
[0083] Finally, the significance of the present invention is discussed. AI is an important driver of the information revolution. Modeling plays a key role in AI. In large language models (LLMs), the computational cost (C) of modeling is proportional to the product of the parameter scale (P) and the data scale (D) (i.e., C≈6PD). Since P is roughly proportional to D, the modeling cost is proportional to the square of the data scale (i.e., C∝D 2 ), which has serious consequences for "AI scaling".
[0084] The purpose of "AI scaling" is to expand the data scale that large language models can handle. It is expected that the data scale will increase by a factor of a thousand (1,000) in the future. According to the above square relationship (C∝D 2 ), the modeling cost (including chip cost and energy cost) will increase by a factor of one million (1,000,000). For such huge chip and energy demands, the earth's resources will not be able to support. Therefore, AI based on modeling is not sustainable.
[0085] In fact, modeling is not the only means of numerical analysis. The two fundamental tools of numerical analysis are "fitting (modeling)" and "interpolation". Interpolation estimates the value of an unknown point by interpolating a continuous function from neighboring data points, and can also achieve the same effect as fitting. Fitting can reduce the storage requirements (the parameter scale needs to be much smaller than the data scale), but requires modeling and has high requirements for computing power; interpolation does not require modeling, has low requirements for computing power (mainly searching for neighboring data points), but requires storing the entire dataset and has high requirements for storage and its retrieval power.
[0086] In the past, people were accustomed to fitting, and computing power was relatively abundant, so large models adopted fitting. When the computing power was insufficient to meet the fitting requirements, interpolation methods began to gain prominence in the field of AI. Vector databases and Retrieval-Augmented Generation (RAG) are typical representatives of interpolation methods. The top-k nearest neighbors in a vector database is exactly a process of searching for neighboring data points. Although many improvements have been made in the software of vector databases in the existing technologies, vector databases are still built on traditional hardware, which limits their accuracy, speed, dimension size, and data scale, resulting in the existing vector databases being only applicable to small knowledge bases. The present invention improves the existing hardware to achieve accurate and fast search for high-dimensional giant vector databases and extends their application to general human knowledge bases (world knowledge bases). Particularly importantly, its cost (storage cost and computing cost) is only proportional to the data scale (C∝D 1 ), compared with large models (the computing cost is proportional to the square of the data scale C∝D 2 ), the AI based on vector databases has greater potential for sustainable development. Therefore, "high computing power" is not the only path to achieve large models (world knowledge bases), and "high storage power" (not only having storage capacity, but also having accurate and fast search power) is also a path to achieve large models (world knowledge bases).
[0087] The present invention can also be applied to various scientific calculations and engineering measurements. For example, in the field of metrology, it is often necessary to compare the measured vectors (measurement vectors) with a pre-computed vector (budget vector) library. By finding the closest budget vector, the corresponding measurement parameters can be found. Using brute-force search, the initial value of the measurement parameters can be found faster. This method can directly find the global minimum, thus avoiding falling into the local minimum, thereby reducing the computing power required for regression calculations.
[0088] It should be understood that without departing from the spirit and scope of the present invention, changes can be made to the form and details of the present invention, which does not prevent them from applying the spirit of the present invention. Therefore, the present invention should not be limited by anything other than the spirit of the appended claims.
Claims
1. A vector database, comprising an input (110) having at least one query vector (115) and at least one storage and computing core (100) coupled to the input (110), characterized in that: The memory computing core (100) comprises: A plurality of storage and calculation units (100ij...), each storage and calculation unit (100ij) comprising at least one storage array (170ij) and a vector distance calculation circuit VDCC (185), the storage array (170ij) storing at least one data vector (175) in the vector database, the VDCC (185) calculating the vector distance between the query vector (115) and the data vector (175); At least one semiconductor substrate (0), the VDCC (185) being located in the semiconductor substrate (0); the memory array (170ij) being stacked on the VDCC (185); the projections of the VDCC (185) and the memory array (170ij) on the semiconductor substrate (0) at least partially overlapping; the VDCC (185) and the memory array (170ij) being electrically coupled.
2. The vector database according to claim 1, further characterized in that The invention comprises: a minimum distance search circuit MDSC (190), wherein the MDSC (190) searches for at least one minimum distance D from the vector distances calculated by all VDCCs of the plurality of storage units (100ij...) m .
3. A vector database comprising an input (110) having at least one query vector (115) and at least one storage and computing core (100) coupled to the input (110), characterized in that: The memory computing core (100) comprises: A plurality of storage and calculation units (100ij...), each storage and calculation unit (100ij) comprising a vector distance calculation circuit VDCC (185) and at least one storage array (170ij), wherein the storage array (170ij) stores at least one data vector (175) in the vector database, and the VDCC (185) calculates the vector distance between the query vector (115) and the data vector (175); A minimum distance search circuit MDSC (190) searches for at least one minimum distance D from the vector distances calculated by all VDCCs of the plurality of storage units (100ij...) m .
4. The vector database according to claim 3, further characterized in that The invention comprises: at least one semiconductor substrate (0), wherein the VDCC (185) is located in the semiconductor substrate (0); the memory array (170ij) is stacked on the VDCC (185); the projections of the VDCC (185) and the memory array (170ij) on the semiconductor substrate (0) at least partially overlap; and the VDCC (185) and the memory array (170ij) are electrically coupled.
5. The vector database according to claim 1 or 4, further characterized in that: The memory computing core (100) is a single chip, and the storage array (170ij) and the VDCC (185) are both located within the single chip.
6. The vector database according to claim 1 or 4, further characterized in that: The storage and computing core (100) is a chip stack and comprises a first chip (100a) and a second chip (100b), wherein the first chip (100a) and the second chip (100b) are stacked and bonded; The VDCCs of the plurality of storage and computing units (100ij...) are located on the first chip (100a); The storage array of the multiple storage and computing units (100ij...) is located in the second chip (100b).
7. The vector database according to claim 6, further characterized in that: The first chip (100a) and the second chip (100b) are bonded face to face, and the first chip (100a) and the second chip (100b) are bonded face to face to form a chip pair.
8. The vector database according to any one of claims 1 to 4, further characterized in that: The input (110) has at least one setting parameter (118), and the VDCC (185) is a reconfigurable VDCC and sets a vector distance algorithm based on the setting parameter (118).
9. The vector database according to any one of claims 1 to 4, further characterized in that contain: A plurality of distance registers (140aa...), wherein the distance registers (140aa...) store the vector distances calculated by the VDCC in the plurality of storage and calculation units (100ij...); A reset circuit (135) searches for the minimum distance D in the MDSC (190). m And the D m After output, the reset circuit (135) will store the D m The distance register is reset to a preset value.
10. The vector database according to any one of claims 1 to 4, further characterized in that Contains: an out-of-array VDCC (285); the VDCC (185) and the out-of-array VDCC (285) use different vector distance algorithms.