In-memory computing device and data processing method thereof
By utilizing in-memory computing devices and methods, and employing memory arrays and computing circuits to select appropriate sorting patterns, the performance bottleneck of the traditional top-k function in large-scale databases with stringent latency requirements is resolved, achieving efficient data processing and memory performance optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA GROUP HOLDING LTD
- Filing Date
- 2020-09-07
- Publication Date
- 2026-04-24
AI Technical Summary
Traditional software implementations of top-k functions cannot process a large number of elements within a reasonable timeframe, resulting in performance bottlenecks in applications with strict latency requirements. In particular, storage performance and data transfer become limiting factors in large-scale databases.
The method employs in-memory computing devices and methods, utilizing memory arrays and computing circuits for data processing. By selecting different sorting modes (first sorting mode and second sorting mode), the first K data elements are output. In the first mode, K is greater than a threshold, and in the second mode, K is less than or equal to the threshold, thereby reducing unnecessary data movement.
It improves the computational efficiency and performance of large-scale data processing tasks, reduces memory performance bottlenecks, meets stringent latency requirements, and is suitable for data processing in large databases and cloud systems.
Smart Images

Figure CN115836346B_ABST
Abstract
Description
Background Technology
[0001] Similarity search has been widely applied in various computing fields, including multimedia databases, data mining, and machine learning. The top-k function can be used in similarity search tasks to find the K elements with the highest or lowest similarity among a given set of elements (e.g., N elements). For example, the top-k function is used in fast region-convolution neural networks (RCNNs). Traditionally, the top-k function is implemented in software.
[0002] However, traditional software implementations of top-k functions cannot process a large number of elements within a reasonable timeframe, making them unsuitable for applications with stringent latency requirements. With the rapid growth of database sizes, the large amount of data transfer between processing units and storage devices due to limited storage performance becomes a performance bottleneck for top-k functions. Summary of the Invention
[0003] Embodiments of this disclosure provide a processing-in-memory (PIM) device. The PIM device includes a memory array and computing circuitry. The memory array is configured to store data. The computing circuitry is configured to execute an instruction set to cause the PIM device to perform the following steps: selecting, based on a configuration from a host, a computing mode including a first sorting mode and a second sorting mode, wherein the host is communicatively coupled to the PIM device; accessing data elements in the memory array of the PIM device; and outputting the first K data elements from the data elements to the memory array or the host in either the first or second sorting mode. K is an integer greater than a threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0004] Embodiments of this disclosure also provide a data processing method. The data processing method includes: selecting from multiple computing modes configured in an in-memory computing device, the multiple computing modes including a first sorting mode and a second sorting mode; accessing multiple data elements in a memory array of the in-memory computing device; and, in the first sorting mode or the second sorting mode, outputting the first K data elements from the multiple data elements to the memory array or to a host communicatively coupled to the in-memory computing device, wherein K is an integer greater than a threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0005] Embodiments of this disclosure also provide a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores a set of instructions executable by one or more computing circuits of the device to cause the device to begin executing a data processing method. The data processing method includes: selecting among a plurality of computing modes based on a configuration, the plurality of computing modes including a first sorting mode and a second sorting mode; accessing a plurality of data elements in a memory array of the device; and, in the first sorting mode or the second sorting mode, outputting the first K data elements of the plurality of data elements to the memory array or to a host communicatively connected to the device. K is an integer greater than a threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0006] Embodiments of this disclosure also provide a data processing system. The data processing system includes a host and a plurality of in-memory computing devices communicatively coupled to the host. Each of the plurality of in-memory computing devices includes a memory array and computing circuitry. The memory array is configured to store data, and the computing circuitry is configured to execute an instruction set to cause the in-memory computing device to perform the following steps: selecting among a plurality of computing modes based on a configuration from the host, the plurality of computing modes including a first sorting mode and a second sorting mode; accessing a plurality of data elements in the memory array of the in-memory computing device; and, in the first sorting mode or the second sorting mode, outputting the first K data elements of the plurality of data elements to the host. K is an integer greater than a threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0007] Additional features and advantages of the disclosed embodiments will be set forth in part in the description which follows, and in part will be apparent from the description, or may be learned by practice of the embodiments. The features and advantages of the disclosed embodiments may be realized and obtained by the elements and combinations set forth in the claims.
[0008] It should be understood that, as stated, the above general description and the following detailed description are exemplary and explanatory only, and do not limit the disclosed embodiments. Attached Figure Description
[0009] Figure 1 The structure of an exemplary in-memory computing block consistent with some embodiments of this disclosure is shown;
[0010] Figure 2A An exemplary neural network accelerator architecture consistent with some embodiments of this disclosure is shown;
[0011] Figure 2B A schematic diagram of an exemplary cloud system including a neural network accelerator architecture, consistent with some embodiments of this disclosure, is shown;
[0012] Figure 3 An exemplary memory chip structure consistent with some embodiments of this disclosure is shown;
[0013] Figure 4 An exemplary in-memory computing processing unit consistent with some embodiments of this disclosure is shown;
[0014] Figure 5 Exemplary examples of some embodiments of this disclosure are shown. Figure 4 The in-memory computation processing unit performs operations for a top-k sorting method;
[0015] Figure 6 Exemplary embodiments consistent with some of the embodiments shown are provided by [the present disclosure]. Figure 4 The in-memory computation processing unit performs operations for another top-k sorting method;
[0016] Figure 7A and Figure 7B An exemplary in-memory computing processing unit consistent with some embodiments of this disclosure is shown;
[0017] Figure 8 An exemplary in-memory computing-based accelerator architecture consistent with some embodiments of this disclosure is shown;
[0018] Figure 9 Exemplary embodiments consistent with some embodiments of this disclosure are shown. Figure 7A and Figure 7B The in-memory computation processing unit performs operations for similarity searches;
[0019] Figure 10 Exemplary examples of some embodiments of this disclosure are shown. Figure 7A and Figure 7B The in-memory computation processing unit performs operations for similarity searches;
[0020] Figure 11 and Figure 12 Exemplary embodiments consistent with some embodiments of this disclosure are shown. Figure 7A and Figure 7B The in-memory computation processing unit performs operations for k-means clustering calculations;
[0021] Figure 13 A flowchart illustrating an exemplary method for performing data processing consistent with some embodiments of this disclosure is shown;
[0022] Figure 14 A flowchart illustrating an exemplary data processing method consistent with some embodiments of the present disclosure is shown. Detailed Implementation
[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, and unless otherwise stated, the same numbers in different figures represent the same or similar elements. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with this application. Rather, they are merely examples of apparatuses, systems, and methods consistent with the aspects related to this disclosure as set forth in the appended claims. In the event of any conflict with terms or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0024] Unless otherwise expressly stated, the term "or" covers all possible combinations except those that are impractical. For example, if a component is declared to include A or B, then unless otherwise expressly stated or impractical, the component may include A, or B, or A and B. As a second example, if a component is declared to include A, B, or C, then unless otherwise expressly stated or impractical, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C. The term "exemplary" is used in the sense of "example" rather than "best practice."
[0025] Today, the scale of databases and data processing tasks is growing significantly and rapidly across various applications. Furthermore, many applications are designed to meet stringent latency requirements in order to provide a satisfactory user experience. Currently, many simple computations with high parallelism, such as similarity search and k-means calculation, are limited by the bandwidth and capacity of the system's memory components, which has become one of the major performance bottlenecks.
[0026] The embodiments of this disclosure alleviate the aforementioned problems by providing devices and methods for data processing to perform top-k ranking, k-means clustering, or other similarity search calculations. By utilizing in-memory computing technology and the high bandwidth of Dynamic Random Access Memory (DRAM), unnecessary data movement is reduced, enabling efficient parallel computing. Therefore, memory performance bottlenecks in similarity search and k-means calculations can be significantly reduced. With the devices and methods disclosed in the embodiments, computation time remains within an acceptable range despite increased data volume, and the overall performance and efficiency of various calculations are improved. The proposed devices and methods for data processing can be applied to various applications with large databases and numerous data processing tasks, including various cloud systems utilizing artificial intelligence (AI) computing.
[0027] Specifically, the embodiments disclosed herein can be used in various applications or environments, such as artificial intelligence training and inference, database and big data analytics acceleration, etc. Artificial intelligence-related applications involve neural network-based machine learning (ML) or deep learning (DL). For example, some embodiments can be used with neural network architectures such as deep neural networks (DNN), convolutional neural networks (CNN), recurrent neural networks (RNN), etc. Furthermore, some embodiments are configured to support various processing architectures, such as data processing units (DPU), neural network processing units (NPU), graphics processing units (GPU), field-programmable gate arrays (FPGA), tensor processing units (TPU), application-specific integrated circuits (ASIC), any other type of heterogeneous accelerator processing unit (HAPU), etc.
[0028] As used herein, the term "accelerator" refers to hardware used to accelerate certain computations. For example, an accelerator may be configured to accelerate top-k ranking computations, k-means clustering computations, or other computations performed in similarity searches. In some embodiments, an accelerator may be configured to accelerate workloads (e.g., neural network computation tasks) in any artificial intelligence-related application. Accelerators with Dynamic Random-Access Memory (DRAM) or Embedded Dynamic Random-Access Memory (eDRAM) are referred to as DRAM-based or eDRAM-based accelerators.
[0029] Figure 1An exemplary in-memory computing block structure consistent with some embodiments of this disclosure is shown. The in-memory computing block 100 includes a memory cell array 110, a block controller 120, a block row driver 131, and a block column driver 132. Although dynamic random-access memory is used as an example to illustrate some embodiments, it should be understood that the in-memory computing block 100 according to embodiments of this disclosure can be implemented based on various memory technologies, including static random-access memory (SRAM), resistive random-access memory (ReRAM), etc. The memory cell array 110 includes r1 to r... m m rows and c1 to c n n columns. For example Figure 1 As shown, storage unit 111 is connected to r1 to r2. m Each of the m rows and c1 to c n Between each of the n columns. In some embodiments, the data may be stored in a cross memory such as a multi-bit memristor.
[0030] Block row driver 131 and block column driver 132 can direct traffic from r1 to r m m rows and c1 to c n The n columns provide signals, such as voltage signals, for processing corresponding operations. In some embodiments, the block row driver 131 and the block column driver 132 are configured to allow analog signals to pass through the storage cell 111. In some embodiments, the analog signals are converted from digital input data.
[0031] Block controller 120 may include an instruction register for storing instructions. In some embodiments, the instructions include instructions on when the block row driver 131 or block column driver 132 provides signals to the corresponding column or row, which signals to provide, etc. Block controller 120 may decode the instructions stored in the register into signals used by the block row driver 131 or block column driver 132.
[0032] The in-memory computation block 100 may further include a row sense amplifier 141 or a column sense amplifier 142 for reading data from a memory cell or storing data into a memory cell. In some embodiments, the row sense amplifier 141 and the column sense amplifier 142 store buffered data. In some embodiments, the in-memory computation block 100 also includes a digital-to-analog converter (DAC) 151 or an analog-to-digital converter (ADC) 152 to convert input signals or output data between the analog and digital domains. In some embodiments of this disclosure, the row sense amplifier 141 or the column sense amplifier 142 is omitted because computations in the in-memory computation block 100 can be performed directly on the values stored in the memory cells without reading these values or using any sense amplifier.
[0033] According to embodiments of this disclosure, the in-memory computation block 100 implements parallel computation by using memory as multiple single instruction multiple data (SIMD) processing units. The in-memory computation block 100 can support computational operations including bitwise operations on integer and floating-point values, addition, subtraction, multiplication, and division. For example, in Figure 1 In the storage cell array 110, the first column c1 and the second column c2 store the first vector A and the second vector B, respectively. By applying formatting signals to the rows corresponding to the lengths of the first to third columns c1-c3 and vectors A, B, and C, the vector operation result C of adding vectors A and B can be stored in the third column c3. Similarly, Figure 1 The memory cell array 110 can also support vector multiplication and addition operations. For example, the calculation C = aA + bB can be performed by applying a voltage signal corresponding to multiplier a to the first column c1 and a voltage signal corresponding to multiplier b to the second column c2, and by applying a formatting signal to the corresponding column and row to perform addition and storing the result C in the third column c3.
[0034] In some embodiments, a vector is stored in multiple columns to represent the n-bit values of its elements. For example, a vector in which elements have 2-bit values can be stored in two columns of storage cells. In some embodiments, when the length of the vector exceeds the number of rows in the storage cell array 110 constituting the storage block, the vector can be stored in multiple storage blocks. Multiple storage blocks can be configured to compute different vector segments in parallel. Although the in-memory computing architecture in the embodiments does not use arithmetic logic outside the storage cells to perform computational operations, this disclosure can also be applied to in-memory computing architectures that include arithmetic logic in order to perform arithmetic logic in the in-memory computing architecture. As mentioned above, computational operations such as addition and multiplication can also be performed as column-by-column vector computations in the in-memory computing architecture. The disclosed embodiments provide an in-memory computing accelerator architecture capable of performing efficient top-k operations, k-means clustering, or similarity search in large databases. In some embodiments, the top-k operation, i.e., finding the k largest or smallest elements from a set, can be widely used in predictive modeling in information retrieval, machine learning, and data mining.
[0035] Figure 2A An exemplary accelerator architecture 200 consistent with some embodiments of this disclosure is shown. In some embodiments, the accelerator architecture 200 is referred to as a neural network processing unit architecture. In the context of this disclosure, a neural network accelerator is also referred to as a machine learning accelerator or a deep learning accelerator. In various embodiments, the accelerator architecture 200 can also be applied to in-memory computing accelerators with various functionalities, such as accelerators for parallel graphics processing, database queries, or other computing tasks. Figure 2A As shown, the accelerator architecture 200 includes an in-memory computing accelerator 210, an interface 212, etc. It should be understood that the in-memory computing accelerator 210 performs algorithmic operations based on the transmitted data.
[0036] The in-memory computing accelerator 210 includes one or more memory chips 2024. In some embodiments, the memory chip 2024 includes a plurality of memory blocks for data storage and computation. The memory blocks are configured to perform one or more operations (e.g., multiplication, addition, multiplication-accumulation, etc.) on the transmitted data. In some embodiments, each memory block included in the memory chip 2024 has a... Figure 1 The in-memory compute block 100 shown has the same configuration. Due to the hierarchical design of the in-memory compute accelerator 210, the in-memory compute accelerator 210 can provide versatility and scalability. The in-memory compute accelerator 210 can include any number of memory slabs 2024 and each memory slab 2024 can have any number of memory blocks.
[0037] Interface 212 (e.g., a Peripheral Component Interconnect Express (PCIe) interface) can be used as an inter-chip bus to provide communication between the in-memory compute accelerator 210 and the host unit 222. The inter-chip bus connects the in-memory compute accelerator 210 to other devices (e.g., off-chip memory or peripheral devices). In some embodiments, the accelerator architecture 200 also includes a Direct Memory Access (DMA) unit. The DMA unit can be considered part of interface 212 or a separate component (not shown) within the in-memory compute accelerator 210, facilitating data transfer between host memory 224 and the in-memory compute accelerator 210. Furthermore, the DMA unit can assist in data transfer between multiple accelerators. The DMA unit allows off-chip devices to access on-chip and off-chip memory without interrupting the host's central processing unit (CPU). Therefore, the DMA unit can also generate memory addresses and initiate memory read or write cycles. The DMA unit may also contain multiple hardware registers that can be read and written by one or more processors. Multiple hardware registers, including memory address registers, byte count registers, one or more control registers, and other types of registers, can specify a combination of source, destination, transfer direction (reading from or writing to an input / output device), transfer unit size, or number of bytes transferred in a burst. It should be understood that accelerator architecture 200 may include a second direct memory access unit for transferring data between other accelerator architectures to allow direct communication between multiple accelerator architectures without involving the host central processing unit.
[0038] Although Figure 2A The accelerator architecture 200 is interpreted as including storage blocks (e.g., Figure 1 The in-memory computing accelerator 210 is an in-memory computing block 100, but it should be understood that the disclosed embodiments can be applied to any type of memory block that supports arithmetic operations to accelerate certain applications (e.g., deep learning).
[0039] Accelerator architecture 200 can also communicate with host unit 222. Host unit 222 can be one or more processing units (e.g., x86 central processing units). In some embodiments, in-memory computing accelerator 210 is considered a coprocessor of host unit 222.
[0040] like Figure 2AAs shown, host unit 222 may be associated with host memory 224. In some embodiments, host memory 224 is integrated memory or external memory associated with host unit 222. Host memory 224 may be local or global memory. In some embodiments, host memory 224 includes a host disk. The host disk is configured as external memory providing additional storage for host unit 222. Host memory 224 may be Double Data Rate Synchronous Dynamic Random-Access Memory (DDRSDRAM), etc. Compared to the on-chip memory of in-memory computing accelerator 210, host memory 224 is configured to store large amounts of data at slower access speeds and acts as a higher-level cache. Data stored in host memory 224 can be transferred to in-memory computing accelerator 210 for various computing tasks or to execute neural network models.
[0041] In some embodiments, the host system 220, having a host unit 222 and a host memory 224, includes a compiler (not shown). A compiler is a program or computer software that translates computer code written in a programming language into instructions to create an executable program. In machine learning applications, a compiler can perform various operations, such as preprocessing, lexical analysis, syntax parsing, semantic analysis, converting input programs into intermediate representations, code optimization, and code generation, or combinations thereof.
[0042] In some embodiments, the compiler pushes one or more commands to the host unit 222. Based on these commands, the host unit 222 can assign any number of tasks to one or more memory chips (e.g., memory chip 2024) or processing elements. Some commands may instruct the direct memory access unit to retrieve instructions and data from host memory (e.g., ...). Figure 2A The host memory 224) is loaded into the accelerator (e.g., Figure 2A The in-memory computing accelerator 210). Instructions can be loaded into each memory slice (e.g., the memory slice assigned the corresponding task) Figure 2A (The memory chip 2024), and one or more memory chips can process these instructions.
[0043] It should be understood that the first few instructions may instruct data to be loaded / stored from host memory 224 into one or more local memories of the memory slice. Each memory slice may then initiate an instruction pipeline, which involves fetching instructions from local memory (e.g., via a fetch unit), decoding the instructions (e.g., via an instruction decoder), generating local memory addresses (e.g., corresponding to operands), reading source data, performing or loading / store operations, and then writing back the results.
[0044] Figure 2B A schematic diagram of an exemplary cloud system including a neural network accelerator architecture, consistent with some embodiments of this disclosure, is shown. Figure 2B As shown, cloud system 230 provides cloud services with artificial intelligence capabilities and includes multiple computing servers (e.g., computing servers 232 and 234). In some embodiments, computing server 232 may, exemplarily, be incorporated into... Figure 2A The accelerator architecture is 200. For simplicity and clarity, in Figure 2B The accelerator architecture 200 is shown in a simplified manner.
[0045] With the assistance of accelerator architecture 200, cloud system 230 is able to provide extended data processing capabilities. For example, in some embodiments, cloud system 230 can provide artificial intelligence functions such as image recognition, facial recognition, translation, and 3D modeling. It should be understood that accelerator architecture 200 can be deployed to computing devices in other forms. For example, accelerator architecture 200 can also be integrated into computing devices such as smartphones, tablets, and wearable devices.
[0046] Figure 3 An exemplary memory chip structure consistent with some embodiments of this disclosure is shown. Memory chip 300 includes a memory block assembly 310, a controller 320, a row driver 331, a column driver 332, a global buffer 340, an instruction memory 350, a data transfer table 360, and a block table 370. According to some embodiments of this disclosure, the memory block assembly 310 includes a plurality of memory blocks arranged in a two-dimensional grid.
[0047] Controller 320 provides commands to each storage block in storage block assembly 310 via row driver 331, column driver 332, and global buffer 340. Row driver 331 is connected to each row of storage blocks in storage block assembly 310, and column driver 332 is connected to each column of storage blocks in storage block assembly 310. In some embodiments, a block controller (e.g., in each storage block) is included. Figure 1 The block controller 120 is configured to receive commands from the controller 320 via the row driver 331 or the column driver 332 and to send commands to the block row driver (e.g., Figure 1 Block row driver 131) and block column driver (e.g., Figure 1 The block column driver 132 in the memory sends a signal to perform a corresponding operation in the memory. According to embodiments of the present disclosure, by using the block controller in the memory block of the memory block assembly 310, the memory block can independently perform different operations, thereby enabling block-level parallel processing, and data can be efficiently transferred between memory cells arranged in rows and columns in the respective memory block through the block controller.
[0048] In some embodiments, the global buffer 340 is used to transfer data between storage blocks in the storage block assembly 310. For example, the controller 320 uses the global buffer 340 when transferring data from one storage block in the storage block assembly 310 to another. According to some embodiments of this disclosure, the global buffer 340 is shared by all storage blocks in the storage block assembly 310. The global buffer 340 can be configured to store commands for each storage block to process tasks assigned in processing a neural network model. In some embodiments, the controller 320 is configured to send commands stored in the global buffer 340 to the corresponding storage blocks via row driver 331 and column driver 332. In some embodiments, these commands are received from a host unit (e.g., Figure 2A The data is transmitted from the host unit 222. The global buffer 340 can be configured to store data for processing assigned tasks and send the data to a storage block. In some embodiments, data stored in and sent from the global buffer 340 originates from the host unit (e.g., host unit 222). Figure 2A The data is transmitted from the host unit 222 or other storage blocks in the storage block assembly 310. In some embodiments, the controller 320 is configured to store data from storage blocks in the storage block assembly 310 into a global buffer 340. In some embodiments, the controller 320 receives an entire row of data from a storage block in the storage block assembly 310 in one cycle and stores it into the global buffer 340. Similarly, the controller 320 sends an entire row of data from the global buffer 340 to another storage block in one cycle.
[0049] In some embodiments, Figure 3 The storage chip 300 includes an instruction memory 350 configured to pipeline instructions for executing neural network models within the storage block assembly 310. The instruction memory 350 may store computation instructions or instructions for data movement between storage blocks of the storage block assembly 310. A controller 320 may be configured to access the instruction memory 350 to retrieve instructions stored therein. The instruction memory 350 may be configured to have a separate instruction segment assigned to each storage block. In some embodiments, the storage chip 300 includes a data transfer table 360 for recording data transfers within the storage chip 300. The data transfer table 360 may be configured to record data transfers between storage blocks. In some embodiments, the data transfer table 360 may be configured to record pending data transfers. In some embodiments, the storage chip 300 may include a block table 370 for recording the state of storage blocks. The block table 370 may have a state field storing the current state of a corresponding storage block. According to some embodiments of this disclosure, during computation, the storage block assembly 310 may have one of several states, including, for example, an idle state, a computation state, and a ready state.
[0050] Figure 4 An exemplary in-memory computing processing unit 400 consistent with some embodiments of this disclosure is shown. In some embodiments, the in-memory computing processing unit 400 is applied to... Figure 2A The accelerator architecture shown in Figure 200 is the same as or similar to the architecture shown in Figure 200. Figure 3 The storage chip configuration shown is (e.g., storage chip 300). In some embodiments, the in-memory computing processing unit 400 is referred to as a processing-in-memory data processing unit (PIM-DPU). The in-memory computing processing unit 400 includes a memory array 410, a memory interface 420, computing circuitry 430, a host interface 440, a configuration register 450, and a controller 460, which may be integrated into the same chip or the same die or embedded in the same package. For example, in some embodiments, the in-memory computing processing unit 400 is on a die of dynamic random access memory (DRAM), wherein the storage device is DRAM with memory array 410 or embedded DRAM, the memory array 410 including storage cells arranged in rows and columns. Furthermore, the memory array 410 may be divided into multiple logical blocks or partitions, also referred to as “trunks” for storing data, each trunk including one or more rows of the memory array 410. For example, the in-memory computing processing unit 400 includes 4 Gbits of DRAM, but this disclosure is not limited thereto. The in-memory computing processing unit 400 may further include dynamic random access memory (DRAM) units of various capacities. In some embodiments, the in-memory computing processing unit 400 having DRAM or embedded DRAM units is referred to as a DRAM-based or embedded DRAM-based accelerator.
[0051] External agents, such as a host, can communicate via a peripheral interface (e.g., host interface 440) and exchange commands, instructions, or data with each other to communicate with the in-memory computing processing unit 400 and the program configuration register 450, thereby configuring various parameters for performing computations. For example, host interface 440 may be a Fast Peripheral Component Interconnect interface, but this disclosure is not limited thereto. Configuration register 450 can store configurations including parameters, such as the K value for top-k sorting calculations, the partition block size of memory blocks in the memory array 410 of dynamic random access memory, etc.
[0052] The controller 460 can communicate with the configuration register 450 to access stored parameters and accordingly instruct the computing circuit 430 to perform a series of operations to perform various calculations, such as top-k ranking calculations, k-means clustering calculations, or other calculations to accelerate similarity search methods on large datasets. For example, in a top-k ranking calculation, the computing circuit 430 calculates and outputs the first to the Kth maximum or minimum values in the dataset, where K can be any integer. In a k-means clustering calculation, the computing circuit 430 can divide N data points (or "observations") into K sets (or "clusters") to minimize the within-cluster sum of squares (WCSS) or variance. K can be any integer greater than 1, and N can be any integer greater than K. In some applications, the value of K for top-k ranking or k-means calculation is a number between 64 and 1500, but this disclosure is not limited thereto. In some embodiments, the number of vectors in the dataset used for top-k ranking or k-means calculation is 10 8 Up to 10 10 The scope is approximately 1000 to 1000, but this disclosure is not limited to this.
[0053] In some embodiments, instructions from controller 460 are decoded and then executed by computing circuitry 430. Computing circuitry 430 includes storage components and computing components for performing computations. For example, the storage components in computing circuitry 430 include one or more vector registers 432 and one or more vector registers 436, and the computing components in computing circuitry 430 include one or more reducers 434 and a scalar arithmetic-logic unit (ALU) 438. Computing circuitry 430 can read data from or write data to storage block 410 via memory interface 420. For example, memory interface 420 can be a wide input / output interface connecting storage block 410 and computing circuitry 430, providing 1024 bits of read / write per cycle (e.g., 2 nanoseconds).
[0054] For reference Figure 5 This illustrates exemplary operations performed by the in-memory computing processing unit 400 for a top-k sorting method, consistent with some embodiments of this disclosure. Figure 5 In the illustrated embodiment, the in-memory computing processing unit 400 performs a top-k sorting method, where the K value is greater than the number of elements the vector can hold. For example, if the data is stored in 32-bit format, a 1024-bit vector can hold 32 data elements. In response to a K value greater than 32 programmed in the configuration register 450, the controller 460 can instruct the computing circuitry 430 to perform... Figure 5 The method shown.
[0055] In this scenario, the memory array 410 includes multiple logic blocks 412, 414, and 416 for storing data elements. The computing circuit 430 receives data elements from logic blocks 412, 414, and 416 and further calculates the maximum or minimum element in each of the logic blocks 412, 414, and 416.
[0056] For example, computation circuit 430 reads a vector (e.g., a 1024-bit vector storing multiple data elements) and compares the minimum value stored in the vector with the scalar value of the current minimum value in the current logic block. By repeating the above process and reading each vector in the current logic block, computation circuit 430 can obtain the minimum element in the current logic block, as well as the block identifier (ID) associated with the minimum element. Computation circuit 430 can also perform a similar operation to obtain the maximum element in the current logic block and the block identifier associated with the maximum element.
[0057] The computation circuit 430 can store the minimum (or maximum) element of each of logic blocks 412, 414, and 416 in an entry of one of the multiple vector registers 432. The computation circuit 430 can then use a reducer 434 to determine the global minimum (or maximum) element based on the block minimum (or maximum) elements of logic blocks 412, 414, and 416. For example, a minimum reducer is used to determine the global minimum element, and a maximum reducer is used to determine the global maximum element.
[0058] After storing the global minimum (or maximum) element as one of the first K data elements, the computing circuit 430 disables the stored global minimum (or maximum) element and repeats the above operation to obtain a new block minimum (or maximum) element of the logic block associated with the stored, disabled global minimum (or maximum) element.
[0059] After obtaining the new block minimum (or maximum) element of the logic blocks, the calculation circuit 430 can again use the reducer 434 to determine the second global minimum (or maximum) element based on the block minimum (or maximum) elements of logic blocks 412, 414, and 416. Therefore, in order to determine the first K data elements (which may be the largest K data elements or the smallest K data elements), the calculation circuit 430 can repeat the above operation for K cycles to obtain the first to the Kth global minimum (or maximum) elements.
[0060] Now for reference Figure 6 This illustrates an exemplary operation performed by an in-memory computation processing unit for another top-k sorting method, consistent with some embodiments of this disclosure. Figure 5 Compared to the embodiments shown, in Figure 6In the illustrated embodiment, the in-memory computing processing unit 400 performs a top-k sorting method, where the K value is less than or equal to the number of elements the vector can hold. For example, if data is stored in a 1024-bit vector as 32 bits, the controller 460 can instruct the computing circuitry 430 to perform a top-k sorting method in response to a K value programmed in the configuration register 450 being less than or equal to 32. Figure 6 The method shown.
[0061] In this scenario, vector register 432 is configured to store the current minimum K values. Calculation circuitry 430 uses reducer 434 (e.g., a maximum reducer) to obtain the maximum value of the current minimum K values and stores that maximum value in scalar register 436.
[0062] When computation circuitry 430 reads a vector (e.g., a 1024-bit vector storing multiple data elements) from memory array 410, it stores one or more minimum values from the vector in scalar register 436. Scalar arithmetic logic unit 438 can communicate with scalar register 436 and compare one or more minimum values from the vector with the maximum value of the current K minimum values. In response to one or more minimum values from the vector being less than the maximum value of the current K minimum values, computation circuitry 430 can replace the maximum value of the current K minimum values in vector register 432 with one or more minimum values from the vector, and then recalculate the new maximum value of the current K minimum values in vector register 432.
[0063] The computation circuit 430 can repeat the above operation until all data elements have been read and processed. Therefore, the K values retained in the vector register 432 after this iterative process are the K smallest data elements.
[0064] By storing the current K largest values in vector register 432, and by comparing one or more maximum values of the vector read from memory array 410 with the minimum value of the current K largest values stored in scalar register 436 through scalar arithmetic logic unit 438, and updating vector register 432 based on the comparison result, computation circuit 430 can perform a similar operation to store the K largest data elements in vector register 432. Therefore, computation circuit 430 can determine the first K data elements, which can be either the K largest data elements or the K smallest data elements.
[0065] Now for reference Figure 7A and Figure 7B It illustrates an exemplary in-memory computing processing unit 700 consistent with some embodiments of this disclosure. Similar to... Figure 4 In some embodiments, the in-memory computing processing unit 400 can be applied to the same memory computing processing unit 700 as the memory computing processing unit 400. Figure 2AThe accelerator architecture shown in Figure 200 is the same as or similar to the architecture shown in Figure 200. Figure 3 The memory chip configuration shown is (e.g., memory chip 300). With Figure 4 Compared to the in-memory computing processing unit 400, in some embodiments, the in-memory computing processing unit 700 includes more memory components and computing components for performing computations. For example, the computing circuitry 430 in the in-memory computing processing unit 700 may also include a static random access memory 732, a decoder 734, and a single instruction multiple data processor 736, the single instruction multiple data processor 736 including one or more adders, subtractors, multipliers, multiply-accumulators, or any combination thereof.
[0066] Figure 7B This illustrates how the storage and computing components in the in-memory computing processing unit 700 communicate and collaborate to perform various computing tasks. In some embodiments, the in-memory computing processing unit 700 uses static random access memory 732 and decoder 734 to perform a product quantization (PQ) compression method to compress or reconstruct data received from memory array 410 or controller 460 for later data processing or manipulation in computing circuitry 430.
[0067] In some embodiments, computation circuitry 430 performs a similarity search or k-means algorithm and uses a single instruction multiple data processor 736 and a reducer 434 to compute the distance between two vectors in a highly parallel and scalable manner. After computation, computation circuitry 430 can store the computed distance value in vector register 432 or scalar register 436. Computation circuitry 430 can perform a top-k sorting operation based on heap sort or any other suitable algorithm using its registers (e.g., vector register 432 or scalar register 436) and the maximum reducer in reducer 434. Detailed operation will be discussed further in the following paragraphs.
[0068] Figure 8 An accelerator architecture 800 based on in-memory computing, consistent with some embodiments of this disclosure, is shown. For example... Figure 8 As shown, in some embodiments, the in-memory computing processing units 810a-810n are composed of... Figure 4 In-memory computing processing unit 400 or Figure 7A It is implemented using an in-memory computing unit 700, and provides high scalability and high capacity. Figure 8In the accelerator architecture 800 based on in-memory computing shown, the in-memory computing system may include multiple in-memory computing processing units 810a-810n, and each of the in-memory computing processing units 810a-810n communicates with the host 820 through a handshake protocol, a double data rate (DDR) protocol, or any other suitable protocol.
[0069] In some embodiments, the in-memory computing processing units 810a-810n do not have direct communication capabilities. Alternatively, the in-memory computing processing units 810a-810n only communicate with the host 820, but this disclosure is not limited thereto. In some other embodiments, some or all of the in-memory computing processing units 810a-810n may also communicate directly with one or more other in-memory computing processing units 810a-810n via appropriate protocols. In some embodiments, the in-memory computing-based accelerator architecture 800 includes hundreds or thousands of in-memory computing processing units 810a-810n depending on the different capacity requirements of various applications. Generally, Figure 8 The in-memory computing processing units 810a-810n can manipulate and process many different types of highly parallel computing and send the final computing results to the host 820. Therefore, data communication between the in-memory computing chip and the host 820 is reduced.
[0070] Now for reference Figure 9 This illustrates exemplary asynchronous computation operations performed by an in-memory computing processing unit for performing similarity searches, consistent with some embodiments of this disclosure. For example... Figure 9 As shown, the memory array 410 may include four dynamic random access memory blocks. In some embodiments, during similarity search, the in-memory computation processing unit 700 calculates the distance values between vectors stored in the dynamic random access memory blocks and performs a top-k sorting method to sort the calculated distance values.
[0071] As described above, the computing circuit 430 can compute the distance values between vectors in a highly parallel computing manner using the single instruction multiple data processor 736 and the summator in the reducer 434. The parallel computing output from the single instruction multiple data processor 736 can be stored or accumulated in the vector accumulator 738.
[0072] for Figure 9The asynchronous computation in the memory array 410 accumulates or stores the distance value in the vector accumulator 738, which can then be written back to a dynamic random access memory block in the memory array 410. The computation circuit 430 can then access the distance values in the memory array 410 and perform a top-k sort operation based on a merge sort algorithm or any other suitable algorithm through its registers (e.g., vector register 432 or scalar register 436) and reducer 434 (e.g., a minimum reducer). The minimum value of the dynamic random access memory block and its corresponding tag can be stored in the registers of the computation circuit 430. Therefore, the computation circuit 430 can first use the minimum reducer in the reducer 434 to find the minimum value stored in the register, and then continue searching for the corresponding minimum value in the dynamic random access memory block storing the distance values and output the minimum value to the host. That is, in Figure 9 In asynchronous computation, the calculation of distance values and the top-k sorting operation are performed in different time periods.
[0073] Now for reference Figure 10 This illustrates exemplary synchronous computation operations performed by the in-memory computing processing unit 700 for performing similarity searches, consistent with some embodiments of this disclosure. Figure 9 Compared to asynchronous computing in [the context of], for Figure 10 The synchronous calculation shown can store the distance value in a register and not write it back to memory array 410.
[0074] The computation circuit 430 performs a top-k sorting operation based on a heap sort algorithm using its registers (e.g., vector register 432 or scalar register 436) and a reducer 434 (e.g., a maximum reducer). Specifically, the computation circuit 430 can use the maximum reducer in the reducer 434 to maintain a minimum top-k heap and uses registers to store the heap. Therefore, in Figure 10 In the synchronous calculation, the distance value calculation and the top-k sorting operation are performed simultaneously within the calculation circuit 430.
[0075] Now for reference Figure 11 and Figure 12 This illustrates an exemplary operation performed by the in-memory computing processing unit 700 for k-means clustering computation, consistent with embodiments of this disclosure. k-means clustering is a vector quantization method designed to divide n observations into k sets (e.g., clusters). Each observation is a vector and belongs to the cluster with the nearest mean. That is, each observation is assigned to the cluster with the nearest cluster centroid (i.e., the cluster centroid).
[0076] Given an initial set of k-means values, k-means clustering computation is performed by alternating between assignment and update steps. Figure 11An exemplary operation of the allocation step is shown, wherein the in-memory computation processing unit 700 assigns each vector to a cluster with the nearest mean (e.g., a cluster with the minimum squared Euclidean distance). Figure 12 An exemplary operation of the update step is illustrated, wherein the in-memory computation processing unit 700 recalculates the mean or centroid (i.e., the hypothetical or actual data point at which the cluster center is located) for the vectors assigned to each cluster. For example, in some embodiments, the centroid of a cluster can be calculated and defined based on the following formula:
[0077]
[0078] Where, x1 to x n Let m1 be the n vectors to be clustered. (t) to m k (t) Representing k clusters S1 respectively (t) To S k (t) The centroid in the t-th iteration.
[0079] like Figure 11 As shown, memory array 410 stores vectors in its memory blocks 1010 and stores the current centroid (mean) of a cluster in one of its row buffers 1020. In some embodiments, when row buffer 1020 is allocated to store the current centroid, the associated memory block is not allocated to store vectors, so the vectors stored in memory array 410 can be read accordingly through their respective row buffers. Computation circuitry 430 can read the eigenvectors in memory array 410 and all centroids in row buffer 1020 into registers 432 or 436. Therefore, computation circuitry 430 can compute the distance between the eigenvectors and centroids in a highly parallel computing manner using single instruction multiple data processor 736 and the sum reducer in reducer 434, and store the calculated distance between the eigenvectors and centroids in static random access memory 732.
[0080] Then, the computation circuit 430 can use the minimum reducer in the reducer 434 to find the minimum value stored in the static random access memory 732 to assign the eigenvector to the cluster with the nearest mean. Thus, the computation circuit 430 can mark the eigenvector with an identifier indicating the cluster whose centroid is closest to it, and write the eigenvector with the cluster identifier back to the memory array 410.
[0081] like Figure 12As shown, in the update step, the computation circuit 430 reads the vectors labeled with cluster identifiers from the memory array 410 into registers 432 or 436, and calculates the updated centroids using the single instruction multiple data processor 736 based on one or more vectors labeled with corresponding cluster identifiers (e.g., the same cluster identifier). The updated centroids can then be written back to the row buffer 1020. In the update step, random access can be reduced by changing the access order of vectors and centroids, which reduces random memory access in k-means clustering calculations.
[0082] By repeating the assignment and update steps, the computation circuit 430 can cluster the vectors stored in the memory array 410 until convergence is achieved. In response to the centroids not changing after the update step (e.g., no vectors were assigned to different clusters in the assignment step), the computation circuit 430 can output the clustering results to the host or store the clustering results in the memory array 410.
[0083] Figure 13 A flowchart illustrating an exemplary data processing method 1300 performed on an in-memory computing-based accelerator architecture consistent with some embodiments of this disclosure is shown. According to some embodiments of this disclosure, an in-memory computing-based accelerator architecture (e.g., Figure 2A Accelerator architecture 200 in China Figure 3 Storage chip 300 and Figure 8 The in-memory computing-based accelerator architecture 800 is used to perform top-k ranking, k-means clustering, or similarity search. Specifically, any in-memory computing processing unit in the in-memory computing-based accelerator architecture (e.g., Figure 8 The in-memory computing processing unit (810a-810n) can be based on a host (e.g., a host that communicates with the in-memory computing processing unit) Figure 8 The host (820) is configured to select between computation modes of the in-memory computing processing unit. The computation mode may include one or more of a first top-k sorting mode, a second top-k sorting mode, and a k-means clustering mode. Depending on the selected mode, the in-memory computing processing unit can access data elements in the memory array to perform corresponding operations.
[0084] Figure 13 The data processing method 1300 describes the top-k sorting computation operation performed by an in-memory computing-based accelerator architecture when the top-k sorting mode is selected. In step 1310, the in-memory computing processing unit (e.g., in-memory computing-based accelerator architecture) in the in-memory computing-based accelerator architecture... Figure 8 The in-memory computing processing units 810a-810n are communicatively connected to the host computer of the in-memory computing processing unit (e.g., Figure 8The host 820 receives the configuration. The in-memory computing processing unit can select between computing modes according to the configuration. In the data processing method 1300, the computing modes include a first top-k sorting mode and a second top-k sorting mode.
[0085] In step 1320, the in-memory computing processing unit determines whether to operate using a first top-k sorting mode or a second top-k sorting mode. Specifically, the in-memory computing processing unit compares the configured K value with a threshold. If the configured K value is greater than the threshold (yes in step 1320), the in-memory computing processing unit selects the first top-k sorting mode and executes steps 1331-1338. If the configured K value is less than or equal to the threshold (no in step 1320), the in-memory computing processing unit selects the second top-k sorting mode and executes steps 1341-1346.
[0086] In the first top-k sorting mode, in step 1331, the computing circuit of the in-memory computing processing unit receives data elements from multiple logic blocks in the memory array.
[0087] In step 1332, the computation circuit calculates the maximum (or minimum) element of each logic block. For example, if a top-k sorting pattern is used to determine the maximum K data, the computation circuit calculates the maximum element of the block. On the other hand, if a top-k sorting pattern is used to determine the minimum K data, the computation circuit calculates the minimum element of the block.
[0088] In step 1333, the computation circuit stores the largest (or smallest) element of the logic block in one or more vector registers within the computation circuit. Then, the computation circuit repeats steps 1334-1338 until the first K data elements have been determined.
[0089] In step 1334, the computation circuit determines the global maximum (or minimum) element based on the block maximum (or minimum) element of the logic block. In step 1335, the computation circuit stores the determined global maximum (or minimum) element as one of the first K data elements. After storing the global maximum (or minimum) element, in step 1336, the computation circuit disables the global maximum (or minimum) element in the associated logic block. In step 1337, the computation circuit obtains the next block maximum (or minimum) element of the logic block associated with the stored, disabled global maximum or minimum element.
[0090] In other words, through steps 1336 and 1337, when the global maximum (or minimum) element is determined and stored, the computing circuit only needs to recalculate and update a new block maximum (or minimum) element for the corresponding logic block, and the block maximum (or minimum) elements of other logic blocks can be reused in the next iteration to determine the next global maximum (or minimum) element.
[0091] In step 1338, the computing circuit determines whether all the first K data elements have been identified and obtained. If all the first K data elements have been identified and stored (yes in step 1338), the computing circuit executes step 1350 and outputs the first K data elements to the host or memory array. Otherwise (no in step 1338), steps 1334-1338 are repeated to identify and store the first K data elements sequentially.
[0092] When the second top-k sorting mode is selected based on the determination in step 1320, in step 1341, the computation circuit of the in-memory computation processing unit first stores the K initial data elements from the memory array into the first register. Then, the computation circuit repeatedly updates the first register by repeating steps 1342-1346 until all data elements from the memory array have been received and processed.
[0093] In step 1342, the computation circuit selects the largest (or smallest) element in the first register as the target element. For example, if a top-k sorting pattern is used to determine the K largest data items, the computation circuit selects the smallest element in the first register as the target element. Conversely, if a top-k sorting pattern is used to determine the K smallest data items, the computation circuit calculates the largest element in the first register as the target element. That is, the target element is an element that can be evicted from the first register and replaced by another data element during a subsequent update process.
[0094] In step 1343, the computational circuitry determines the top K candidate elements from one or more remaining data elements received from the memory array. For example, the remaining data elements may be new data read from an unprocessed vector of a dynamic random access memory (DRAM) data array. The computational circuitry may receive a vector from the memory array and select the smallest (or largest) element in the vector as the top K candidate elements.
[0095] In step 1344, the computation circuit compares the top K candidate elements with the target element. If a top-k sorting pattern is used to determine the maximum K data, the computation circuit determines whether the top K candidate elements are greater than the target element currently stored in the first register. If a top-k sorting pattern is used to determine the minimum K data, the computation circuit determines whether the top K candidate elements are less than the target element currently stored in the first register. Therefore, the computation circuit can determine whether to replace the target element in the first register with the top K candidate elements based on the comparison results. Specifically, in some embodiments, the computation circuit stores the target element and the top K candidate elements in a second register within the computation circuit, and compares the top K candidate elements with the target element using a scalar arithmetic logic unit to obtain the comparison result.
[0096] If the computational circuit determines that the target element in the first register should be replaced (yes in step 1344), the computational circuit executes step 1345 to replace the target element with the top K candidate elements. Otherwise (no in step 1344), step 1345 is skipped and the data stored in the first register remains unchanged.
[0097] In step 1346, the computation circuitry determines whether all data elements of interest in the memory array have been processed. If there are still remaining data elements (e.g., data in the vector that have not yet been processed) that need to be processed (step 1346 is no), steps 1342-1346 are repeated to update the current first K data elements stored in the first register. If all data elements of interest in the memory array have been processed (step 1346 is yes), the computation circuitry executes step 1350 and outputs the first K data elements stored in the first register to the host or the memory array.
[0098] Therefore, through the above operations, Figure 13 The data processing method 1300 can perform top-k sorting calculations to output K largest (or smallest) data elements, where the value of K can be greater than or less than the number of elements that a vector in the vector register can hold. Specifically, if the value of K is an integer greater than a threshold, the first top-k sorting mode is selected; if the value of K is an integer less than or equal to the threshold, the second top-k sorting mode is selected.
[0099] Figure 14 A flowchart illustrating an exemplary data processing method 1400 performed on an accelerator architecture based on in-memory computing, consistent with some embodiments of this disclosure, is shown. Figure 14 The data processing method 1400 in the example illustrates a k-means clustering pattern using an accelerator architecture based on in-memory computation (e.g., Figure 2A Accelerator architecture 200 in China Figure 3 Storage chip 300 and Figure 8 The k-means clustering operation is performed by the in-memory computing-based accelerator architecture 800.
[0100] In the k-means clustering pattern, in step 1410, at least one in-memory computation processing unit (e.g., Figure 8The in-memory computing processing units 810a-810n initialize the centroids of the clusters. For example, to have K clusters, the in-memory computing processing units can provide K initial centroids and store them in the row buffer of the memory array. Then, the in-memory computing processing units cluster multiple vectors stored in the memory array by repeating the allocation step 1420 and the update step 1430. In the allocation step 1420, the in-memory computing processing units assign each vector to one of the current clusters based on the distance between the vector point and the centroid of the cluster. In the update step 1430, the in-memory computing processing units update the centroid of each cluster based on the corresponding vector. Each centroid is the average of the vectors assigned to the same cluster in the allocation step 1420.
[0101] Specifically, allocation step 1420 includes sub-steps 1421-1425. In step 1421, the computation circuit receives centroids from the row buffer. In step 1422, the computation circuit receives eigenvectors selected from the vectors of the memory array. In step 1423, the computation circuit assigns cluster identifiers to the eigenvectors, where the cluster identifiers indicate the centroids closest to the eigenvectors. In step 1424, the computation circuit writes the eigenvectors with cluster identifiers back to the memory array.
[0102] In step 1425, the computational circuitry determines whether all vectors of interest in the memory array have been assigned to their associated clusters. If there are remaining vectors to process (no in step 1425), steps 1421-1425 are repeated to assign the remaining vectors. If all vectors of interest in the memory array have been assigned (yes in step 1425), the computational circuitry proceeds to update step 1430.
[0103] Update step 1430 also includes sub-steps 1431 and 1432. In step 1431, the computational circuit calculates the updated centroids based on one or more vectors labeled with corresponding cluster identifiers. In step 1432, the computational circuit determines whether to update the centroids of all cluster identifiers based on the latest allocation result obtained in step 1420. If there are remaining centroids to update (if step 1432 is negative), steps 1431 and 1432 are repeated.
[0104] When all centroids have been updated (yes in step 1432), in step 1440, the computation circuit checks whether the centroids have remained unchanged since step 1430 in the current cycle. If one or more centroids have changed (no in step 1440), the in-memory computation processing unit repeats step 1420 and update step 1430 until convergence is achieved.
[0105] If the centroids remain unchanged after the update operation (step 1440 is yes), the in-memory computation processing unit executes step 1450 and outputs the clustering results to the host or stores the clustering results in the memory array. Therefore, through the above operations, Figure 14 The data processing method 1400 in the middle can realize k-means clustering calculation and output the clustering results.
[0106] In summary, as presented in the various embodiments of this disclosure, the proposed devices and methods can utilize the high bandwidth of dynamic random access memory (DRAM) and guarantee efficient, parallel, and fast computation. The storage performance bottlenecks in similarity search and k-means calculation are significantly reduced through the wide input / output (e.g., 1024-bit read / write per cycle) between the memory array (e.g., DRAM data array) and the computation circuitry.
[0107] Furthermore, by using the proposed apparatus and method to perform k-means clustering computation, unnecessary data movement between the memory array and computational circuitry is reduced and minimized. Additionally, random memory access is reduced by storing centroids in a row buffer and changing the access order of vectors and centroids during the update step of the k-means clustering computation. Therefore, the overall efficiency of k-means clustering is improved. In some embodiments, the efficiency of similarity search can depend solely on the ratio of bandwidth to memory capacity. Even with increasing data volume, the computation time for similarity search can still be kept within tens of milliseconds.
[0108] The embodiments of this disclosure can be applied to many products, environments, and scenarios. For example, some embodiments of this disclosure can be applied to in-memory computing, such as in-memory computing for artificial intelligence, including processing units based on dynamic random access memory. Some embodiments of this disclosure can also be applied to tensor processing units, data processing units, neural network processing units, etc.
[0109] Embodiments of this disclosure also provide a computer program product. This computer program product includes a non-transitory computer-readable storage medium having computer-readable program instructions on it for causing a processor to perform the methods described above.
[0110] A computer-readable storage medium can be a tangible device capable of storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A more specific, exemplary, and non-exhaustive list of computer-readable storage media includes the following devices: portable computer floppy disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), static random access memory, compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanical encoding devices (e.g., punched cards or raised structures in recesses for recording instructions thereon), and any suitable combination thereof.
[0111] The computer-readable program instructions used to perform the above methods can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages (including object-oriented programming languages and traditional procedural programming languages). The computer-readable program instructions can be executed entirely on the computer system as a standalone software package, or partially on the first computer and partially on a second computer located remotely from the first computer. In the latter case, the remote second computer can be connected to the first computer via any type of network, including local area networks (LANs) or wide area networks (WANs).
[0112] Computer-readable program instructions can be provided to a computer processor or other programmable data processing device to form a machine, such that the instructions, which are executed by the computer processor or other programmable data processing device, create means for implementing the above-described methods.
[0113] The flowcharts and diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this specification. In this regard, blocks in flowcharts or diagrams may represent software programs, a section or portion of code, which includes one or more executable instructions for implementing a specific function. It should also be noted that in some alternative implementations, the functions marked in the boxes may not appear in the order indicated in the figures. For example, two blocks shown consecutively may actually execute substantially simultaneously, or these blocks may sometimes execute in reverse order, depending on the functions involved. It will also be noted that each block in a diagram or flowchart, and combinations of blocks in diagrams and flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified function or action.
[0114] It should be understood that, for clarity, certain features of this disclosure described in the context of different embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, various features of this disclosure described in the context of a single embodiment may also be provided individually, or in any suitable sub-combination, or as appropriate in any other described embodiment of this disclosure. Certain features described in the context of various embodiments should not be considered as essential features of those embodiments unless the embodiments cannot be practiced without these elements.
[0115] The embodiments may be further described using the following claims:
[0116] 1. An in-memory computing device, comprising:
[0117] Memory arrays, configured to store data; and
[0118] The computing circuitry is configured to execute a set of instructions to cause the in-memory computing device to perform the following steps:
[0119] Based on configuration from a host, a selection is made among multiple computing modes, the host being communicatively coupled to the in-memory computing device, wherein the multiple computing modes include a first sorting mode and a second sorting mode;
[0120] Accessing multiple data elements in the memory array of the in-memory computing device; and
[0121] In either the first or second sorting mode, the first K data elements among the plurality of data elements are output to the memory array or the host.
[0122] Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0123] 2. The in-memory computing device of claim 1, wherein the computing circuitry further comprises a vector register, wherein, in the first sorting mode, the computing circuitry receives the plurality of data elements from a plurality of logic blocks in the memory array, and wherein the computing circuitry is further configured to execute the instruction set to enable the in-memory computing device to determine the first K data elements by means of the following steps:
[0124] For each of the plurality of logical blocks, the maximum or minimum element is;
[0125] The maximum or minimum element of the plurality of logical blocks is stored in the vector register; and
[0126] Repeat the following steps until the first K data elements are determined:
[0127] The global maximum or minimum element is determined based on the maximum or minimum element of the multiple logical blocks;
[0128] Store the global maximum or minimum element as one of the first K data elements;
[0129] Disable the global maximum or minimum element in its logical block; and
[0130] Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element.
[0131] 3. The in-memory computing device according to claim 1 or 2, wherein the computing circuitry includes a first register, and in the second sorting mode, the computing circuitry is further configured to execute the instructions to cause the in-memory computing device to determine the first K data elements by means of the following steps:
[0132] Storing multiple initial data elements from the memory array into the first register; and
[0133] The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed:
[0134] Select the maximum or minimum element in the first register as the target element;
[0135] Determining candidate elements from one or more remaining data elements received from the memory array; and
[0136] Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.
[0137] 4. The in-memory computing device of claim 3, wherein the computing circuitry includes a second register and a scalar arithmetic logic unit, wherein the computing circuitry is further configured to execute the instructions to cause the in-memory computing device to determine whether to replace the target element in the first register with the candidate element:
[0138] Store the target element and the candidate element in the second register; and
[0139] The candidate element is compared with the target element by the scalar arithmetic logic unit.
[0140] 5. The in-memory computing device according to any one of claims 1-4, wherein the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to perform the following steps:
[0141] In response to the K value in the configuration being greater than the threshold, the first sorting method is selected; and
[0142] In response to a K value in the configuration being less than or equal to the threshold, the second sorting method is selected.
[0143] 6. The in-memory computing device according to any one of claims 1-5, wherein the plurality of computing modes includes a k-means clustering mode, wherein in the k-means clustering mode, the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to perform the following steps:
[0144] Multiple vectors stored in the memory array are clustered by repeating the following steps:
[0145] Assign each of the plurality of vectors to one of the plurality of clusters; and
[0146] Update multiple centroids of the plurality of clusters, wherein each of the plurality of centroids is the average of one or more corresponding vectors assigned to the same cluster.
[0147] 7. The in-memory computing device of claim 6, wherein, in the k-means clustering mode, the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to allocate each of the plurality of vectors through the following steps:
[0148] Receive the plurality of centroids from the row buffer of the memory array;
[0149] Receive a feature vector selected from the plurality of vectors from the memory array;
[0150] A cluster identifier is assigned to the feature vector, the cluster identifier indicating the nearest centroid to the feature vector; and
[0151] Write the feature vector with the cluster identifier back to the memory array.
[0152] 8. The in-memory computing device of claim 7, wherein, in the k-means clustering mode, the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to update each of the plurality of centroids through the following steps:
[0153] The updated centroid is calculated based on one or more vectors with corresponding cluster identifiers in the memory array.
[0154] 9. The in-memory computing device according to any one of claims 6-8, wherein, in the k-means clustering mode, the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to:
[0155] In response to the fact that the plurality of centroids remain unchanged after updating each of them, the clustering result is output to the host or the memory array.
[0156] 10. The in-memory computing device according to any one of claims 1-9, further comprising:
[0157] A host interface is configured to enable the in-memory computing device to communicate with a host.
[0158] A configuration register is configured to store the configuration from the host; and
[0159] The controller is configured to send the instruction set according to the configuration.
[0160] 11. The in-memory computing device according to any one of claims 1-10, wherein the computing circuitry further comprises:
[0161] One or more registers are configured to store data used for computation;
[0162] Storage device, configured to store the instruction set; and
[0163] The decoder is configured to decode the instruction set.
[0164] 12. The in-memory computing device according to any one of claims 1-11, wherein the computing circuit further comprises one or more single instruction multiple data units, one or more reducer units, one or more arithmetic logic units, or any combination thereof.
[0165] 13. The in-memory computing device according to any one of claims 1-12, wherein the memory array includes a dynamic random access memory array, and the in-memory computing device further includes an input / output interface for communication between the dynamic random access memory array and the computing circuitry.
[0166] 14. A data processing method, comprising:
[0167] Based on the selection among multiple computing modes configured in the in-memory computing device, wherein the multiple computing modes include a first sorting mode and a second sorting mode;
[0168] Accessing multiple data elements in the memory array of the in-memory computing device; and
[0169] In the first sorting mode or the second sorting mode, the first K data elements among the plurality of data elements are output to the memory array or a host that is communicatively coupled to the in-memory computing device;
[0170] Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0171] 15. The data processing method according to claim 14, further comprising:
[0172] When the in-memory computing device is configured to operate in the first sorting mode:
[0173] Receive the plurality of data elements from the plurality of logical blocks in the memory array; and
[0174] The first K data elements are determined through the following steps:
[0175] For each of the plurality of logical blocks, the maximum or minimum element is; and
[0176] Repeat the following steps until the first K data elements are determined:
[0177] Determine the global maximum or minimum element based on the maximum or minimum element of each of the multiple logical blocks:
[0178] Store the global maximum or minimum element as one of the first K data elements:
[0179] Disable the global maximum or minimum element in its logical block; and
[0180] Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element.
[0181] 16. The data processing method according to claim 14 or 15, further comprising:
[0182] With the in-memory computing device configured to operate in a second sorting mode, the first K data elements are determined through the following steps:
[0183] Storing multiple initial data elements from the memory array into a first register; and
[0184] The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed:
[0185] Select the maximum or minimum element in the first register as the target element;
[0186] Determine candidate elements from one or more remaining data elements received from the memory array; and
[0187] Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.
[0188] 17. The data processing method according to claim 16, wherein determining whether to replace the target element in the first register with the candidate element includes:
[0189] The target element and the candidate element are stored in the second register; and
[0190] The candidate element is compared with the target element.
[0191] 18. The data processing method according to any one of claims 14-17, wherein the plurality of calculation modes further includes a k-means clustering mode, and the data processing method further includes:
[0192] When the in-memory computing device is configured to operate in k-means clustering mode, multiple vectors stored in the memory array are clustered by repeating the following steps:
[0193] Assign each of the plurality of vectors to one of the plurality of clusters; and
[0194] Update multiple centroids of the plurality of clusters, wherein each of the plurality of centroids is the average of one or more corresponding vectors assigned to the same cluster.
[0195] 19. The data processing method of claim 18, wherein allocating each of the plurality of vectors comprises:
[0196] Receive the plurality of centroids;
[0197] Receive a feature vector selected from the plurality of vectors;
[0198] A cluster identifier is assigned to the feature vector, the cluster identifier indicating the nearest centroid to the feature vector; and
[0199] Write the feature vector with the cluster identifier back to the memory array.
[0200] 20. The data processing method according to claim 19, further comprising:
[0201] When the in-memory computing device is configured to operate in k-means clustering mode, each of the plurality of centroids is updated by computing an updated centroid based on one or more vectors in the memory array that have corresponding cluster identifiers.
[0202] 21. The data processing method according to any one of claims 18-20, further comprising:
[0203] When the in-memory computing device is configured to operate in k-means clustering mode, in response to the fact that the plurality of centroids remain unchanged after updating each of them, the clustering result is output to the host or the memory array.
[0204] 22. A non-transitory computer-readable medium storing a set of instructions executed by one or more computing circuits of a device to cause the device to begin implementing a data processing method, the data processing method comprising:
[0205] Based on the configuration, selection is made among multiple computing modes, wherein the multiple computing modes include a first sorting mode and a second sorting mode;
[0206] Accessing multiple data elements in the memory array of the device; and
[0207] In either the first or second sorting mode, the first K data elements among the plurality of data elements are output to the memory array or a host connected in communication with the device.
[0208] Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0209] 23. The non-transitory computer-readable medium of claim 22, wherein the instruction set is executed by the one or more computing circuits of the device to cause the device to still execute in the first sorting mode:
[0210] Receive the plurality of data elements from the plurality of logical blocks in the memory array, and
[0211] The first K data elements are determined through the following steps:
[0212] For each of the plurality of logical blocks, the maximum or minimum element is; and
[0213] Repeat the following steps until the first K data elements are determined:
[0214] The global maximum or minimum element is determined based on the maximum or minimum element of the block of the plurality of logical blocks;
[0215] Store the global maximum or minimum element as one of the first K data elements;
[0216] Disable the global maximum or minimum element in its logical block; and
[0217] Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element.
[0218] 24. The non-transitory computer-readable medium of claim 22 or 23, wherein the instruction set is executed by the one or more computing circuits of the device to cause the device to still execute in the second sorting mode:
[0219] The first K data elements are determined through the following steps:
[0220] Store multiple initial data elements from the memory array into the first register; and
[0221] The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed:
[0222] Select the maximum or minimum element in the first register as the target element;
[0223] Determine candidate elements from one or more remaining data elements received from the memory array; and
[0224] Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.
[0225] 25. The non-transitory computer-readable medium of claim 24, wherein the instruction set is executed by the one or more computing circuits of the device to cause the device to determine whether to replace the target element in the first register with the candidate element by means of the following steps:
[0226] Store the target element and the candidate element in a second register; and
[0227] The candidate element is compared with the target element.
[0228] 26. The non-transitory computer-readable medium according to any one of claims 22-25, wherein the plurality of computing modes includes a k-means clustering mode, and the instruction set is executed by the one or more computing circuits of the device to cause the device to still execute in the k-means clustering mode:
[0229] Multiple vectors stored in the memory array are clustered by repeating the following steps:
[0230] Assign each of the plurality of vectors to one of the plurality of clusters; and
[0231] Update multiple centroids of the plurality of clusters, wherein each of the plurality of centroids is the average of one or more corresponding vectors assigned to the same cluster.
[0232] 27. The non-transitory computer-readable medium of claim 26, wherein the instruction set is executed by the one or more computing circuits of the device to cause the device to assign each of the plurality of vectors by means of the following steps:
[0233] Receive the plurality of centroids from the row buffer of the memory array;
[0234] Receive a feature vector selected from the plurality of vectors from the memory array;
[0235] A cluster identifier is assigned to the feature vector, the cluster identifier indicating the nearest centroid to the feature vector; and
[0236] Write the feature vector with the cluster identifier back to the memory array.
[0237] 28. The non-transitory computer-readable medium of claim 27, wherein the instruction set is executed by the one or more computing circuits of the device to cause the device to update each of the plurality of centroids by the following steps:
[0238] The updated centroid is calculated based on one or more vectors with corresponding cluster identifiers in the memory array.
[0239] 29. The non-transitory computer-readable medium according to any one of claims 26-28, wherein the instruction set is executed by the one or more computing circuits of the device to cause the device to further perform the following in a k-means clustering mode:
[0240] In response to the fact that the plurality of centroids remain unchanged after updating each of them, the clustering result is output to the host or the memory array.
[0241] 30. A data processing system, comprising:
[0242] Host; and
[0243] A plurality of in-memory computing devices are communicatively coupled to the host, wherein any one of the plurality of in-memory computing devices includes a memory array and computing circuitry, the memory array being configured to store data, and the computing circuitry being configured to execute an instruction set to cause the in-memory computing device to perform the following steps:
[0244] Based on the configuration from the host, a selection is made among multiple computing modes, including a first sorting mode and a second sorting mode;
[0245] Accessing multiple data elements in the memory array of the in-memory computing device; and
[0246] Under either the first or the second sorting method, the first K data elements among the plurality of data elements are output to the host.
[0247] Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected.
[0248] In the foregoing description, embodiments have been described with reference to numerous specific details, which may vary with different implementations. Certain adjustments and modifications can be made to the described embodiments. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice of this disclosure herein. The specification and embodiments are intended to be illustrative only. The order of steps shown in the figures is also intended for illustrative purposes only and is not intended to limit to any particular order of steps. Therefore, those skilled in the art will understand that these steps can be performed in different orders to implement the same method. Exemplary embodiments have been disclosed in the drawings and specification. However, many variations and modifications can be made to these embodiments. Therefore, although specific terminology is used, it is used only in a general and descriptive sense and not for limiting purposes. The scope of the embodiments is defined by the appended claims.
Claims
1. An in-memory computing device, comprising: A memory array, configured to store data; and The computing circuitry includes a vector register or a first register; the computing circuitry is configured to execute an instruction set to cause the in-memory computing device to perform the following steps: The selection of multiple computing modes is based on configuration from the host, the host communication being coupled to the in-memory computing device, wherein the multiple computing modes include a first sorting mode and a second sorting mode. Accessing multiple data elements in the memory array of the in-memory computing device; and In either the first or second sorting mode, the first K data elements among the plurality of data elements are output to the memory array or the host. Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected; In the first sorting mode, the computing circuit receives the plurality of data elements from the plurality of logic blocks in the memory array; For each of the plurality of logical blocks, the maximum or minimum element is; Store the largest or smallest element of the plurality of logical blocks in the vector register; and repeat the following operation until the first K data elements are determined: The global maximum or minimum element is determined based on the maximum or minimum element of the multiple logical blocks; Store the global maximum or minimum element as one of the first K data elements; Disable the global maximum or minimum element in its logical block; and Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element; In the second sorting mode, multiple initial data elements from the memory array are stored in the first register; and The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed: Select the maximum or minimum element in the first register as the target element; Determine candidate elements from one or more remaining data elements received from the memory array; and Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.
2. The in-memory computing device according to claim 1, wherein the computing circuit includes a second register and a scalar arithmetic logic unit, wherein, The computing circuitry is also configured to execute the instruction set to enable the in-memory computing device to determine whether to replace the target element in the first register with the candidate element: The target element and the candidate element are stored in the second register; as well as The candidate element is compared with the target element by the scalar arithmetic logic unit.
3. The in-memory computing device according to claim 1, wherein, The plurality of computing modes includes a k-means clustering mode, wherein the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to perform the following steps: Multiple vectors stored in the memory array are clustered by repeating the following steps: Assign each of the plurality of vectors to one of the plurality of clusters; and Update multiple centroids of the plurality of clusters, wherein each of the plurality of centroids is the average of one or more corresponding vectors assigned to the same cluster.
4. The in-memory computing device according to claim 3, wherein, In the k-means clustering pattern, the computing circuitry is further configured to execute the instruction set to cause the in-memory computing device to allocate each of the plurality of vectors through the following steps: Receive the plurality of centroids from the row buffer of the memory array; Receive a feature vector selected from the plurality of vectors from the memory array; A cluster identifier is assigned to the feature vector, the cluster identifier indicating the nearest centroid to the feature vector; and Write the feature vector with the cluster identifier back to the memory array.
5. The in-memory computing device according to claim 1, further comprising: A host interface is configured to enable the in-memory computing device to communicate with a host. A configuration register is configured to store the configuration from the host; and The controller is configured to send the instruction set according to the configuration.
6. The in-memory computing device according to claim 1, wherein, The computing circuit also includes: One or more registers are configured to store data used for computation; Storage device, configured to store the instruction set; and The decoder is configured to decode the instruction set.
7. The in-memory computing device according to claim 1, wherein the computing circuit further comprises one or more single instruction multiple data units, one or more reduction units, one or more arithmetic logic units, or any combination thereof.
8. A data processing method, comprising: The selection is based on multiple computing modes configured in the in-memory computing device, wherein the multiple computing modes include a first sorting mode and a second sorting mode. Accessing multiple data elements in the memory array of the in-memory computing device; and In the first sorting mode or the second sorting mode, the first K data elements among the plurality of data elements are output to the memory array or output to a host that is communicatively coupled to the in-memory computing device; Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected; When the in-memory computing device is configured to operate in the first sorting mode: Receive the plurality of data elements from the plurality of logical blocks in the memory array; and The first K data elements are determined through the following steps: For each of the plurality of logical blocks, the maximum or minimum element is; and Repeat the following steps until the first K data elements are determined: Determine the global maximum or minimum element based on the maximum or minimum element of each of the multiple logical blocks: Store the global maximum or minimum element as one of the first K data elements: Disable the global maximum or minimum element in its logical block; and Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element; With the in-memory computing device configured to operate in a second sorting mode, the first K data elements are determined through the following steps: Storing multiple initial data elements from the memory array into a first register; and The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed: Select the maximum or minimum element in the first register as the target element; Determine candidate elements from one or more remaining data elements received from the memory array; and Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.
9. The data processing method according to claim 8, wherein, Determining whether to replace the target element in the first register with the candidate element includes: The target element and the candidate element are stored in the second register; and The candidate element is compared with the target element.
10. The data processing method according to claim 8, wherein, The plurality of calculation modes also includes k-means clustering mode, and the data processing method further includes: With the in-memory computing device configured in k-means clustering mode, multiple vectors stored in the memory array are clustered by repeating the following steps: Assign each of the plurality of vectors to one of the plurality of clusters; and Update multiple centroids of the plurality of clusters, wherein each of the plurality of centroids is the average of one or more corresponding vectors assigned to the same cluster.
11. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executed by one or more computing circuits of a device to cause the device to begin implementing a data processing method, the data processing method comprising: Based on the configuration, a selection is made among multiple computing modes, wherein the multiple computing modes include a first sorting mode and a second sorting mode; Accessing multiple data elements in the memory array of the device; and In the first sorting mode or the second sorting mode, the first K data elements among the plurality of data elements are output to the memory array or to a host that is communicatively connected to the device; Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected; Execute in the first sorting mode: Receive the plurality of data elements from the plurality of logical blocks in the memory array, and The first K data elements are determined through the following steps: For each of the plurality of logical blocks, the maximum or minimum element is; and Repeat the following steps until the first K data elements are determined: The global maximum or minimum element is determined based on the maximum or minimum element of the block of the plurality of logical blocks; Store the global maximum or minimum element as one of the first K data elements; Disable the global maximum or minimum element in its logical block; and Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element; Execute in the second sorting mode: The first K data elements are determined through the following steps: Store multiple initial data elements from the memory array into the first register; and The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed: Select the maximum or minimum element in the first register as the target element; Determine candidate elements from one or more remaining data elements received from the memory array; and Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.
12. The non-transitory computer-readable medium according to claim 11, wherein, The instruction set is executed by one or more computing circuits of the device to enable the device to determine whether to replace the target element in the first register with the candidate element by the following steps: Store the target element and the candidate element in a second register; and The candidate element is compared with the target element.
13. The non-transitory computer-readable medium according to claim 11, wherein, The plurality of computation modes includes a k-means clustering mode, and the instruction set is executed by one or more computation circuits of the device to enable the device to still operate in the k-means clustering mode: Multiple vectors stored in the memory array are clustered by repeating the following steps: Assign each of the plurality of vectors to one of the plurality of clusters; and Update multiple centroids of the plurality of clusters, wherein each of the plurality of centroids is the average of one or more corresponding vectors assigned to the same cluster.
14. The non-transitory computer-readable medium according to claim 13, wherein, The instruction set is executed by the one or more computing circuits of the device to cause the device to allocate each of the plurality of vectors through the following steps: Receive the plurality of centroids from the row buffer of the memory array; Receive a feature vector selected from the plurality of vectors from the memory array; A cluster identifier is assigned to the feature vector, the cluster identifier indicating the nearest centroid to the feature vector; and Write the feature vector with the cluster identifier back to the memory array.
15. A data processing system, comprising: Host; and A plurality of in-memory computing devices are communicatively coupled to the host, wherein any one of the plurality of in-memory computing devices includes a memory array and computing circuitry, the memory array being configured to store data, and the computing circuitry being configured to execute an instruction set to cause the in-memory computing device to perform the following steps: Based on the configuration from the host, a selection is made among multiple computing modes, including a first sorting mode and a second sorting mode; Accessing multiple data elements in the memory array of the in-memory computing device; and In either the first or second sorting mode, the first K data elements among the plurality of data elements are output to the host. Wherein, K is an integer greater than the threshold when the first sorting mode is selected, and an integer less than or equal to the threshold when the second sorting mode is selected; In the first sorting mode, the computing circuit receives the plurality of data elements from the plurality of logic blocks in the memory array; For each of the plurality of logical blocks, the maximum or minimum element is; Store the largest or smallest element of the plurality of logical blocks in a vector register; and repeat the following operation until the first K data elements are determined: The global maximum or minimum element is determined based on the maximum or minimum element of the multiple logical blocks; Store the global maximum or minimum element as one of the first K data elements; Disable the global maximum or minimum element in its logical block; and Get the next block maximum or minimum element of the logical block associated with the disabled global maximum or minimum element; In the second sorting mode, multiple initial data elements from the memory array are stored into a first register; and The first register is updated by repeating the following operation until the plurality of data elements from the memory array are received and processed: Select the maximum or minimum element in the first register as the target element; Determine candidate elements from one or more remaining data elements received from the memory array; and Based on the comparison result between the candidate element and the target element, determine whether to use the candidate element to replace the target element in the first register.