Memory device for performing in-memory prefetching and system including the memory device

By using memory cell arrays, information registers and prefetching circuits in memory devices, and using indirect memory access information for prefetching, the problem of low indirect memory access efficiency in the prior art is solved, and more efficient in-memory prefetching and processing is achieved.

CN111009268BActive Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201910742115.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-08
Filing Date
2019-08-12
Publication Date
2025-05-13
Estimated Expiration
2039-08-12

Smart Images

  • Figure CN111009268B_ABST
    Figure CN111009268B_ABST
Patent Text Reader

Abstract

A memory device includes a memory cell array, an information register, and a prefetch circuit. The memory cell array stores a valid data array, a base array, and a target data array, wherein the valid data array includes valid elements among elements of first data, the base array includes position elements indicating position values ​​corresponding to the valid elements, and the target data array includes target elements corresponding to the position values ​​of second data. The information register stores indirect memory access information including a start address of the target data array and a unit size of the target element. The prefetch circuit prefetches the target element corresponding to the position element read from the memory cell array based on the indirect memory access information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the priority of Korean Patent Application No. 10-2018-0119527 filed on October 8, 2018 in the Korean Intellectual Property Office (KIPO), the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] Example embodiments relate generally to semiconductor integrated circuits, and more particularly, to a memory device that performs in-memory prefetching and a system including the memory device. Background Art

[0004] Prefetching is a technique used by computer processors to improve execution performance by prefetching instructions or data from original storage in slower memory to faster local memory before they are actually needed. In the case of irregular data access such as indirect memory access, it is difficult to improve the operation speed when prefetching based on the location of the data. Although software and hardware prefetching methods that consider indirect memory access patterns have been proposed, excessive bandwidth may be consumed when detecting and predicting indirect memory access, or bandwidth may be consumed when unnecessary data is read due to incorrect detection and prediction. Summary of the invention

[0005] At least one exemplary embodiment of the inventive concept may provide a memory device capable of improving indirect memory access efficiency and a system including the memory device.

[0006] At least one exemplary embodiment of the inventive concept may provide a memory device and a system including the memory device capable of improving efficiency of a Process-In-Memory (PIM) architecture through indirect memory access.

[0007] According to an exemplary embodiment of the present invention, a memory device includes: a memory cell array, an information register and a prefetch circuit. The memory cell array stores a valid data array, a base array and a target data array, wherein the valid data array sequentially includes valid elements among elements of first data, the base array sequentially includes position elements indicating position values ​​corresponding to the valid elements, and the target data array sequentially includes target elements of second data corresponding to the position values. The information register stores indirect memory access information including a starting address of the target data array and a unit size of the target element. The prefetch circuit prefetches the target element corresponding to the position element read from the memory cell array based on the indirect memory access information.

[0008] According to an exemplary embodiment of the present invention, a memory device includes: a plurality of memory semiconductor dies stacked in a vertical direction, a memory cell array is formed in the plurality of memory semiconductor dies; a plurality of through-silicon vias electrically connecting the plurality of memory semiconductor dies; an information register; and a pre-fetch circuit. The memory cell array stores a valid data array, a base array, and a target data array. The valid data array sequentially includes valid elements among elements of first data. The base array sequentially includes position elements indicating position values ​​corresponding to the valid elements. The target data array sequentially includes target elements of second data corresponding to the position values. The information register is configured to store indirect memory access information including a starting address of the target data array and a unit size of the target element. The pre-fetch circuit is configured to pre-fetch the target element corresponding to the position element read from the memory cell array based on the indirect memory access information.

[0009] According to an exemplary embodiment of the present invention, a system includes: a memory device and a host device, the host device including a memory controller configured to control access to the memory device. The memory device includes: a memory cell array configured to store a valid data array, a base array, and a target data array: an information register; and a prefetch circuit. The valid data array sequentially includes valid elements among elements of first data. The base array sequentially includes position elements indicating position values ​​corresponding to the valid elements. The target data array sequentially includes target elements of second data corresponding to the position values. The information register is configured to store indirect memory access information including a starting address of the target data array and a unit size of the target element. The prefetch circuit is configured to prefetch the target element corresponding to the position element read from the memory cell array based on the indirect memory access information.

[0010] According to an exemplary embodiment of the present invention, a memory device includes: a memory cell array configured to store a valid data array; a RAM; and a controller. The valid data array sequentially includes valid elements among elements of first data. The base array sequentially includes position elements indicating position values ​​corresponding to the valid elements. The target data array sequentially includes target elements of second data corresponding to the position values. The controller is configured to receive indirect memory access information including a starting address of the target data array and a unit size of the target element from an external device. The controller performs a prefetch operation in response to receiving the indirect memory access information, wherein the prefetch operation reads the target element using the starting address and the unit size, and stores the read target element into the RAM.

[0011] A memory device and system according to at least one exemplary embodiment of the inventive concept may enhance the accuracy and efficiency of in-memory prefetching by performing indirect memory access based on indirect memory access information provided from a memory controller.

[0012] Furthermore, a memory device and system according to at least one exemplary embodiment of the inventive concept may reduce latency of indirect memory access and increase speed of in-memory prefetching by parallelizing data arrangement for indirect memory access and utilizing bandwidth within the memory device.

[0013] Furthermore, a memory device and a system according to at least one exemplary embodiment of the inventive concept may efficiently perform sparse data operations by performing in-memory processing operations using in-memory prefetching. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Exemplary embodiments of the present disclosure will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.

[0015] Figure 1 is a block diagram illustrating a memory system including a memory device according to an exemplary embodiment of the inventive concept.

[0016] Figure 2 is a flowchart illustrating a method of controlling a memory device according to an exemplary embodiment of the inventive concept.

[0017] Figure 3 is a diagram illustrating an example of sparse matrix-vector multiplication (SpMV).

[0018] Figure 4A and Figure 4B It is shown that Figure 3 A plot of the SpMV corresponding to the data array.

[0019] Figure 5 is a block diagram illustrating a memory device according to an exemplary embodiment of the inventive concept.

[0020] Fig. 6A , Figure 6B and Figure 6C is a diagram illustrating an exemplary embodiment of indirect memory access information for in-memory prefetch according to an exemplary embodiment of the inventive concept.

[0021] Figure 7 is a diagram illustrating an example of data parallel disposition for indirect memory access according to an exemplary embodiment of the inventive concept.

[0022] Figure 8is a diagram illustrating an exemplary embodiment of a prefetch circuit included in a memory device according to an exemplary embodiment of the inventive concept.

[0023] Fig. 9 is a block diagram illustrating a system including a memory device according to an exemplary embodiment of the inventive concept.

[0024] Fig.10 is a diagram illustrating an exemplary embodiment of an operation unit for in-memory processing included in a memory device according to an exemplary embodiment of the inventive concept.

[0025] Fig.11 and Fig.12 is an exploded perspective view of a system including a stacked memory device according to an exemplary embodiment of the inventive concept.

[0026] Fig.13 and Fig.14 is a diagram illustrating a package structure of a stacked memory device according to an exemplary embodiment of the inventive concept.

[0027] Fig.15 is a diagram illustrating an example structure of a stacked memory device according to an exemplary embodiment of the inventive concept.

[0028] Fig.16 is a block diagram illustrating a computing system according to an exemplary embodiment of the inventive concept. DETAILED DESCRIPTION

[0029] The inventive concept will be described more fully below with reference to the accompanying drawings, in which some exemplary embodiments are shown. In the accompanying drawings, the same reference numerals always represent the same elements. Repeated descriptions may be omitted.

[0030] Figure 1 is a block diagram illustrating a memory system including a memory device according to an exemplary embodiment of the inventive concept.

[0031] refer to Figure 1 , the memory system 10 includes a memory controller 20 and a semiconductor memory device 30 .

[0032] The memory controller 20 may control the overall operation of the memory system 10. The memory controller 20 may control the entire data exchange between the external host and the semiconductor memory device 30. For example, the memory controller 20 may write data to or read data from the semiconductor memory device 30 in response to a request from the host. In addition, the memory controller 20 may issue an operation command (e.g., a read command, a write command, etc.) to the semiconductor memory device 30 to control the semiconductor memory device 30. In an exemplary embodiment, the memory controller 20 is located within the host device.

[0033] In some exemplary embodiments, the semiconductor memory device 30 may be a memory device including a dynamic memory cell, such as a dynamic random access memory (DRAM), a double data rate 4 (DDR4) synchronous DRAM (SDRAM), a low power DDR4 (LPDDR4) SDRAM, or a LPDDR5 SDRAM.

[0034] The memory controller 20 may transmit a clock signal CLK, a command CMD, and an address (signal) ADDR to the semiconductor memory device 30, and exchange data DQ with the semiconductor memory device 30. In addition, the memory controller 20 may provide indirect memory access information IMAI to the semiconductor memory device 30 according to a request of the host. The indirect memory access information IMAI may be provided to the semiconductor memory device 30 as an additional control signal, or provided through a mode register write command for setting a mode register included in the semiconductor memory device 30.

[0035] The semiconductor memory device 30 may include a memory cell array MC40, an information register IREG 100, and a prefetch circuit PFC 200. In an embodiment, a controller (e.g., a control circuit) includes the information register IREG 100 and the prefetch circuit PFC 200, and the controller performs prefetching in response to receiving memory access information IMAI from a host device or a memory controller.

[0036] The memory cell array 40 may include a plurality of memory cells to store data. The memory cells may be grouped into a plurality of banks, and each bank may include a plurality of data blocks.

[0037] According to the control of the memory controller 20, the valid data array, the base array, and the target data array may be stored in the memory cell array 40. The valid data array may sequentially include valid elements among the elements of the first data. The base array may sequentially include position elements indicating position values ​​corresponding to the valid elements. The target data array may sequentially include target elements corresponding to the position values ​​of the second data. Figure 3 , Figure 4A and Figure 4B The valid data array, base array, and target data array are further described.

[0038] The memory 100 may store the indirect memory access information IMAI provided from the memory controller 20. In an embodiment, the indirect memory access information IMAI includes at least the starting address of the target data array and the unit size of the target element. Fig. 6A , Figure 6B , Figure 6C and Figure 7 The indirect memory access information IMAI provided from the memory controller is further described.

[0039] In an embodiment, the prefetch circuit 200 prefetches a target element corresponding to a position element read from the memory cell array 40 based on the indirect memory access information IMAI.

[0040] A method for accelerating memory access is prefetching, and conventional prefetching adopts a streaming or stride scheme that reads data in adjacent memory areas simultaneously. Conventional prefetching assumes spatial locality of data to read adjacent data simultaneously. However, conventional prefetching may consume memory bandwidth unnecessarily.

[0041] A new prefetch technique according to an exemplary embodiment of the inventive concept includes detecting an indirect memory access, predicting an address of the indirect memory access, and reading data from the predicted address.

[0042] Prefetch costs must be minimized to prevent delays to other memory accesses caused by prefetching. It is very important to accurately detect and predict indirect memory accesses and optimize prefetch operations. If incorrect detection and prediction consume bandwidth or prefetching itself consumes considerable bandwidth, the performance of the entire system may be degraded due to delays incurred by other workloads.

[0043] In an embodiment, the prefetching technology for reducing the latency of indirect memory access is divided into software prefetching and hardware prefetching.

[0044] Based on the fact that software access has an accuracy advantage over hardware access, software prefetching focuses on the accuracy of detection and prediction of indirect memory access. In other words, compared with hardware that cannot grasp the content of the application workload executed by the host device, software can accurately predict when indirect memory access occurs and which part of the memory device must be accessed at a lower cost through source code inspection. For example, when the compiler captures an indirect memory access code when scanning the source code, the compiler can insert a prefetch command before the indirect memory access code. Delays in indirect memory access can be prevented by prefetching data based on the inserted prefetch command and using the prefetched data in subsequent operations.

[0045] However, software prefetching may not sufficiently reduce the prefetching cost. Although software prefetching can be performed more easily than hardware prefetching, indirect memory accesses may not be perfectly detected by automatic code analysis, and false detection or detection oversight may occur. In addition, command insertion for each detected code may significantly increase the number of prefetch commands.

[0046] Hardware prefetching can solve these problems of software prefetching. For example, future cache misses can be predicted by monitoring the cache misses that occur immediately after the prefetch table access in order to obtain information about indirect memory accesses. Prefetching can be performed by calculating the address of the data to be prefetched through such monitoring and prediction.

[0047] However, hardware prefetching may not be able to avoid the accuracy and cost issues. Hardware cannot capture the execution context of each workload, so in hardware prefetching, the cache miss immediately following the prefetch table access is considered an indirect memory access. The insufficient information may not be enough to improve the accuracy of detection and prediction, so unnecessary prefetching may occur due to false detection of indirect memory access.

[0048] In contrast, a memory device and system according to exemplary embodiments of the inventive concept may enhance the accuracy and efficiency of in-memory prefetching by performing indirect memory access based on indirect memory access information provided from a memory controller.

[0049] Figure 2 is a flowchart illustrating a method of controlling a memory device according to an exemplary embodiment of the inventive concept.

[0050] refer to Figure 1 and Figure 2, the valid data array, the base array, and the target data array provided from the memory controller 20 are stored in the memory cell array 40 included in the memory device 30 (S100). In an embodiment, the valid data array sequentially includes valid elements among the elements of the first data. In an embodiment, the base array sequentially includes position elements indicating position values ​​corresponding to the valid elements. In an embodiment, the target data array sequentially includes target elements corresponding to the position values ​​of the second data.

[0051] The indirect memory access information IMAI is provided from the memory controller 20 to the memory device 30 (S200), and the memory device 30 stores the indirect memory access information IMAI in the information register 100 included in the memory device 30 (S300). The indirect memory access information IMAI includes at least the start address of the target data array and the unit size of the target element.

[0052] Using the pre-fetch circuit 200 included in the memory device 30 and based on the indirect memory access information IMAI, a target element corresponding to the position element read from the memory cell array 40 is pre-fetched (S400). The pre-fetch circuit 200 can be used to efficiently perform in-memory pre-fetching based on the indirect memory access information IMAI. In an embodiment, before performing an operation that requires the use of the read target element or data derived from the read target element, the read target element is temporarily stored in a fast memory (e.g., DRAM, SRAM, register, etc.) of the memory 30 as part of the pre-fetching.

[0053] In the following, reference Figure 3 , Figure 4A and Figure 4B An example sparse matrix vector multiplication (SpMV) is described as an example to which prefetching of indirect memory access according to an exemplary embodiment can be applied. For example, if a process of a host device predicts that the host device will soon need a data result of the SpMV, the processing host device can output indirect memory access information IMAI to trigger prefetching and execution of the SpMV. For example, the process can predict that the host device will soon need the data result by parsing the source code of an executable file executed by the host device to obtain instructions corresponding to the SpMV.

[0054] Figure 3 is a diagram showing an example of SpMV, and Figure 4A and Figure 4B It is shown that Figure 3 A plot of the SpMV corresponding to the data array.

[0055] Figure 3 A sparse matrix SM and a column vector CV are shown. Figure 4A Shown with Figure 3 The sparse matrix SM corresponds to the valid data array AA and the basis array BA, Figure 4B Shown with Figure 3 The column vector CV corresponds to the target data array. Figure 3 , Figure 4A and Figure 4B An example multiplication of a sparse matrix SM and a column vector CV is described, and it will be understood that the same description can be applied to various multiplications, such as multiplication of a row vector and a sparse matrix, multiplication of a sparse matrix and another matrix, and the like.

[0056] One of the common workloads of high-performance computing, machine learning, deep learning, and graph analysis is sparse data operations. For example, sparse data operations are basic operations in finite element simulations, and are widely used in engineering analysis such as mechanics, thermodynamics, and fluid dynamics, recursive neural networks (RNNs) for processing time-varying speech data, and page ranking algorithms that assign weight values ​​to web documents based on their relative importance. Here, as Figure 3 The sparse matrix SM of , sparse data indicates that a small portion of all data elements are valid elements having valid values ​​(eg, values ​​other than zero).

[0057] For sparse data, only valid elements (e.g., non-zero) are stored for storage space efficiency. Figure 3 In the case of the sparse matrix SM, seven of the fifty elements are valid elements. Only seven valid elements and corresponding information of all elements are stored to save storage space of the memory device. Figure 4A As shown, a valid data array AA sequentially including only valid elements A(1) to A(7) among fifty elements and a base array BA sequentially including position elements B(1) to B(7) indicating position values ​​corresponding to the valid elements A(1) to A(7) are stored.

[0058] The main operation of sparse matrices is multiplication with vectors, such as Figure 3 The SpMV shown. The position of the vector element (ie, the target element) to be multiplied with the matrix element (ie, the effective element) is determined by the column position of the matrix element, which can be called indirect memory access.

[0059] Considering Figure 3 The vector elements of the column vector CV in the SpMV are accessed. For the first row of the sparse matrix SM, the calculation 7*1+(-2.5)*5 is performed, so the first element and the eighth element of the column vector CV must be read. For the second row of the sparse matrix SM, the calculation (-5)*(-5) is performed, so the fourth element of the column vector CV must be read.

[0060] In summary, only valid elements of the sparse matrix SM are stored, and the column positions of the stored valid elements are irregular, such as 1, 8, 4, 2, 1, 3, and 8. The positions of vector elements or target elements are determined by the irregular column positions of the valid elements, so the vector elements are accessed irregularly.

[0061] To summarize the above, the array including the valid elements of the sparse matrix can be called the valid data array AA, the array including the position or column information corresponding to the valid elements can be called the base array BA, and the array including the vector elements can be called the target data array TA. Figure 3 , Figure 4A and Figure 4B In the case of , AA={A(1), A(2), A(3), A(4), A(5), A(6), A(7)}={7, -2.5, -5, 3, -6, 37, 9}, BA={B(1), B(2), B(3), B(4), B(5), B(6), B(7)}={1, 8, 4, 2, 1, 3, 8}, TA={T(1), T(2), T(3), T(4), T(5), T(6), T(7), T(8), T(9, T(10)}={1, 32, 4, -5, 8, 7, 13, 5, 43, -7}. SpMV can be represented by A(i)*T[B(i)], where i=1~7 and T[B(i)] is the operation that causes indirect memory access.

[0062] A representative method for accelerating memory access is prefetching, and conventional prefetching adopts a streaming or stride scheme that simultaneously reads data in adjacent memory areas to prevent duplicate reading. When performing the operation A(i)*T[B(i)] of SpMV, A(i) corresponds to sequential access, and conventional prefetching can be effective. However, when only valid elements of a sparse matrix are stored, T[B(i)] becomes irregular, and thus the address of T[B(i)] becomes irregular.

[0063] exist Figure 3 , Figure 4A and Figure 4B In the example of , BA = {1, 8, 4, 2, 1, 3, 8}, and therefore the necessary vector elements as target elements are T[1], T[8], T[4], T[2], T[1], T[3], and T[8], which results in irregular access. For such irregular access, conventional prefetching has little effect because the target elements are read separately when needed. Therefore, a significant delay in access to the memory device may be caused, and the execution of SpMV may be degraded.

[0064] According to an exemplary embodiment, the execution of SpMV may be enhanced by efficiently prefetching target elements within a memory device upon indirect memory access.

[0065] Figure 5 is a block diagram illustrating a memory device according to an exemplary embodiment of the inventive concept.

[0066] Although the reference Figure 5 DRAM is described as an example of a memory device, but the memory device according to an exemplary embodiment can be any of a variety of memory cell architectures, including but not limited to volatile memory architectures such as DRAM, (thyristor RAM) TRAM, and SRAM, or non-volatile memory architectures such as read-only memory (ROM), flash memory, ferroelectric RAM (FRAM), magnetic RAM (MRAM), etc.

[0067] refer to Figure 5 , the memory device 400 includes a control logic 410 (e.g., a logic circuit), an address register 420, a memory bank control logic 430 (e.g., a logic circuit), a row address multiplexer 440, a column address latch 450, a row decoder 460 (e.g., a decoding circuit), a column decoder 470 (e.g., a decoding circuit), a memory cell array 480, a sense amplifier unit 485, an input / output (I / O) selection circuit 490, a data input / output (I / O) buffer 495, a refresh counter 445 (e.g., a counting circuit), an information register IREG 100, and a prefetch circuit PFC 200.

[0068] The memory cell array 480 may include a plurality of memory bank arrays 480a-480h. The row decoder 460 may include a plurality of memory bank row decoders 460a-460h respectively coupled to the memory bank arrays 480a-480h, the column decoder 470 may include a plurality of memory bank column decoders 470a-470h respectively coupled to the memory bank arrays 480a-480h, and the sense amplifier unit 485 may include a plurality of memory bank sense amplifiers 485a-485h respectively coupled to the memory bank arrays 480a-480h.

[0069] The address register 420 may receive an address ADDR including a bank address BANK_ADDR, a row address ROW_ADDR, and a column address COL_ADDR from the memory controller. The address register 420 may provide the received bank address BANK_ADDR to the bank control logic 430, may provide the received row address ROW_ADDR to the row address multiplexer 440, and may provide the received column address COL_ADDR to the column address latch 450.

[0070] The bank control logic 430 may generate a bank control signal in response to the bank address BANK_ADDR, may activate one of the bank row decoders 460a-460h corresponding to the bank address BANK_ADDR in response to the bank control signal, and may activate one of the bank column decoders 470a-470h corresponding to the bank address BANK_ADDR in response to the bank control signal.

[0071] The row address multiplexer 440 may receive the row address ROW_ADDR from the address register 420 and may receive the refresh row address REF_ADDR from the refresh counter 445. The row address multiplexer 440 may selectively output the row address ROW_ADDR or the refresh row address REF_ADDR as the row address RA. The row address RA output from the row address multiplexer 440 may be applied to the bank row decoders 460a-460h.

[0072] The activated one of the bank row decoders 460a460h may decode the row address RA output from the row address multiplexer 440 and may activate a word line corresponding to the row address RA. For example, the activated bank row decoder may apply a word line driving voltage to the word line corresponding to the row address RA.

[0073] The column address latch 450 may receive the column address COL_ADDR from the address register 420 and may temporarily store the received column address COL_ADDR. In some embodiments, in burst mode, the column address latch 450 may generate a column address incremented from the received column address COL_ADDR. The column address latch 450 may apply the temporarily stored or generated column address to the bank column decoders 470a-470h.

[0074] The activated one of the bank column decoders 470 a ˜ 470 h may decode the column address COL_ADDR output from the column address latch 450 and may control the input / output gating circuit 490 to output data corresponding to the column address COL_ADDR.

[0075] The I / O gating circuit 490 may include a circuit for gating input / output data. The I / O gating circuit 490 may also include a read data latch for storing data output from the memory bank arrays 480a-480h, and a write driver for writing data to the memory bank arrays 480a-480h.

[0076] Data to be read from one of the bank arrays 480a-480h may be sensed by bank sense amplifiers 485a-485h coupled to the one bank array from which data is to be read, and may be stored in read data latches. The data stored in the read data latches may be provided to the memory controller via the data I / O buffer 495. Data DQ to be written to one of the bank arrays 480a-480h may be provided from the memory controller to the data I / O buffer 495. The write driver may write the data DQ to one of the bank arrays 480a-480h.

[0077] The control logic 410 may control the operation of the memory device 400. For example, the control logic 410 may generate a control signal for the memory device 400 to perform a write operation or a read operation. The control logic 410 may include a command decoder 411 that decodes a command CMD received from a memory controller, and a mode register 412 that sets an operation mode of the memory device 400. For example, the command decoder 411 may generate a control signal corresponding to the command CMD by decoding a write enable signal, a row address strobe signal, a column address strobe signal, a chip select signal, etc.

[0078] In an exemplary embodiment, the information register 100 stores indirect memory access information IMAI provided from an external memory controller. In an exemplary embodiment, the indirect memory access information IMAI includes at least a starting address TSADD of a target data array TA and a unit size TSZ of a target element included in the target data array TA. According to an exemplary embodiment, the indirect memory access information IMAI also includes a starting address BSADD of a base array BA, a unit size BSZ of a position element included in the base array BA, a total number NT of position elements, and a read number NR of position elements simultaneously read from the memory cell array 480. A method of using the indirect memory access information IMAI according to an exemplary embodiment will be described below.

[0079] In an exemplary embodiment, the prefetch circuit 200 prefetches a target element corresponding to a position element read from the memory cell array 480 based on the indirect memory access information IMAI. In an exemplary embodiment, the start address TSADD of the target data array TA and the unit size TSZ of the target element in the indirect memory access information IMAI are provided to the prefetch circuit 200 to calculate the target address TADDR, and other information BSADD, BSZ, NT, and NR may be provided to the control logic 410 to perform overall control of the memory device 400. In an exemplary embodiment, the other information BSADD, BSZ, NT, and NR may be stored in the mode register 412.

[0080] Fig. 6A , Figure 6B and Figure 6C is a diagram illustrating an exemplary embodiment of indirect memory access information for in-memory prefetch according to an exemplary embodiment of the inventive concept.

[0081] refer to Fig. 6A , the indirect memory access information IMAI1 includes the start address TSADD of the target data array TA and the unit size TSZ of the target element included in the target data array TA. In an exemplary embodiment, immediately after the write operation of the valid data array AA, the base array BA, and the target data array TA is completed, the indirect memory access information IMAI1 is provided from the memory controller to the memory device. In an exemplary embodiment, the indirect memory access information IMAI1 is provided from the memory controller to the memory device using a command or mode signal indicating indirect memory access. As shown below with reference to Figure 8 As described above, the pre-fetch circuit 200 may calculate the target address TADDR using the start address TSADD of the target data array TA and the unit size TSZ of the target element.

[0082] refer to Figure 6B The indirect memory access information IMAI2 includes the starting address TSADD of the target data array TA and the unit size TSZ of the target elements included in the target data array TA, the starting address BSADD of the base array BA, the unit size BSZ of the position elements included in the base array BA, the total number NT of the position elements, and the read number NR of the position elements read from the memory cell array at the same time.

[0083] refer to Figure 6C The indirect memory access information IMAI3 includes the starting address TSADD of the target data array TA and the unit size TSZ of the target elements included in the target data array TA, the starting address ASADD of the effective data array AA, the unit size ASZ of the effective elements included in the effective data array AA, the starting address BSADD of the base array BA, the unit size BSZ of the position elements included in the base array BA, the total number NT of the position elements and the read number NR of the position elements read simultaneously from the memory cell array.

[0084] As described below, the information TSADD, TSZ, ASADD, ASZ, BSADD, BSZ, NT, and NR may be used when calculating addresses and controlling repeated prefetch operations.

[0085] Figure 7 is a diagram illustrating an example of data parallel arrangement for indirect memory access according to an exemplary embodiment of the inventive concept.

[0086] Typically, a memory device may include multiple memory banks, and Figure 7 A first memory bank MBK1 , a second memory bank MBK2 , and a third memory bank MBK3 are shown as non-limiting examples.

[0087] refer to Figure 7 , the valid data array AA, the base array BA and the target data array TA are stored in different memory banks. For example, the valid data array AA is stored in the first memory bank MBK1, the base array BA is stored in the second memory bank MBK2, and the target data array TA is stored in the third memory bank MBK3.

[0088] The valid data array AA sequentially includes valid elements A(i) (where i is the index of the element) among the elements of the first data. The base array BA sequentially includes position elements B(i) indicating position values ​​corresponding to the valid elements A(i). The target data array TA sequentially includes target elements T(i) of the second data corresponding to the position values ​​of the position elements B(i). Figure 3 , Figure 4A and Figure 4B As described above, the first data may be a sparse matrix, and the second data may be a vector. In this case, the position value of the position element B(i) may be the column position of the valid element A(i).

[0089] Typically, the access delay to different memory banks is less than the access delay to the same memory bank. Data can be read out from different memory banks or different groups of memory banks substantially simultaneously. Thus, by storing the effective data array AA, the base array BA and the target data array TA in different memory banks, the speed of memory access and prefetching can be increased.

[0090] The indirect memory access information IMAI can be used to calculate the address of the data to be read. For example, when the address of the k-th valid element A(k) in the valid data array AA is AADDR{A(k)}, the k-th position element B(k) corresponding to the k-th valid element A(k) in the base array BA can be calculated as Expression 1 and Expression 2.

[0091] Expression 1

[0092] AADDR{A(k)}=ASADD+AOFS=ASADD+(k-1)*ASZ,

[0093] (k-1)=(AADDR{A(k)}-ASADD) / ASZ

[0094] Expression 2

[0095] BADDR{B(k)}=BSADD+BOFS=BSADD+(k-1)*BSZ

[0096] =BSADD+(AADDR{A(k)}-ASADD)*BSZ / ASZ

[0097] Thus, when the address AADDR{A(k)} of the kth valid element A(k) is provided from the memory controller to the memory device, the address BADDR{B(k)} of the corresponding kth position element B(k) can be calculated in the memory device using the stored indirect memory access information IMAI, and the kth position element B(k) can be read out. Therefore, the delay of transmitting the address can be reduced or eliminated, and the speed of pre-fetching and in-memory processing can be further improved. Figure 8 The calculation of the target address using the start address TSADD of the target data array TA and the unit size TSZ of the target element is described.

[0098] Figure 8 is a diagram illustrating an exemplary embodiment of a prefetch circuit included in a memory device according to an exemplary embodiment of the inventive concept.

[0099] refer to Figure 8 , the pre-fetch circuit 200 includes an arithmetic circuit 210 (eg, one or more arithmetic logic units), a target address register TAREG 220, and a target data register TDREG 230. To indicate the data flow, Figure 8 Also shown are an information register IREG 100, a memory bank MBKB of a memory base array BA, and a memory bank MBKT of a memory target data array TA.

[0100] The arithmetic circuit 210 can calculate a target address TADDR{B(i)} corresponding to the position element B(i) read from the memory bank MBKB based on the read position element B(i), the start address TSADD of the target data array TA, and the unit size TSZ of the target element T(i).

[0101] The calculated target address TADDR{B(i)} is stored in the target address register 220 and provided for accessing the memory bank MBKT. The target element T(i) is read out from the target address TADDR{B(i)} and can be stored in the target data register 230.

[0102] In an exemplary embodiment, the arithmetic circuit 210 calculates the target address TADDR{B(i)} using Expression 3.

[0103] Expression 3

[0104] TADDR{T(i)}=TSADD+TSZ*(B(i)-1)

[0105] In Expression 3, T(i) represents the i-th target element, TADDR{T(i)} represents the target address of the i-th target element, TSADD represents the starting address of the target data array, TSZ represents the unit size of the target element, and B(i) represents the i-th position element of the base array.

[0106] The multiplication in Expression 3 can be performed using a shifter (e.g., a shift register). In an embodiment, each target element is multiplied by 2 n bits, where n is an integer, so the shifter can replace the expensive multiplier. In an embodiment, each arithmetic unit (AU) is implemented with a shifter and an adder.

[0107] It is necessary to improve the accuracy of detection and prediction of indirect memory access and optimize prefetch operations to reduce the latency and cost of indirect memory access. Exemplary embodiments are provided to address these two issues, namely, the accuracy of detection and prediction of indirect memory access and the optimization of prefetch operations to ensure additional bandwidth for prefetching. As described above, the minimization of prefetch costs is to prevent delays in other memory accesses caused by prefetching, as well as to ensure the bandwidth of the prefetch itself.

[0108] The memory bandwidth may include external bandwidth and internal bandwidth. Even if the external bandwidth is increased, if the workload is optimized to obtain the increased external bandwidth, the additional bandwidth may not be guaranteed, so the internal bandwidth must be ensured for prefetching. Therefore, prefetching within the memory according to an exemplary embodiment is proposed to ensure the bandwidth for prefetching.

[0109] For in-memory prefetching according to an exemplary embodiment of the inventive concept, the memory controller provides indirect memory access information IMAI to the memory device according to a request of the host device so as to improve the accuracy of detection and prediction of indirect memory access. At least information about the base array BA and the target data array TA may be required for in-memory prefetching.

[0110] In some exemplary embodiments, whenever an indirect memory access is required, the memory controller provides information to the memory device. As described above, with respect to the base array BA, the information may include the starting address BSADD of the base array BA, the unit size BSZ of the position element in the base array BA, the total number NT of the position element, and the read number NR of the position element read from the memory cell array at the same time. With respect to the target array TA, the information may include the starting address TSADD of the target data array TA and the unit size TSZ of the target element in the target data array TA.

[0111] Since the indirect memory access information IMAI can be transmitted before the indirect memory access starts, excessive bandwidth due to error detection is not consumed.

[0112] In order to optimize the prefetch operation, according to an exemplary embodiment of the inventive concept, parallelization of address calculations for indirect memory accesses, i.e., single instruction multiple data (SIMD), may be performed. Figure 8 As shown, the arithmetic circuit 210 may include NR arithmetic units AU configured to provide NR target addresses in parallel based on NR position elements simultaneously read from the memory cell array, where NR is a natural number greater than 1.

[0113] When the total number of position elements is NT, the NR arithmetic units AU can provide NT target addresses by repeatedly performing NT / NR address calculations, where NT is a natural number greater than 1. Although for ease of explanation, Figure 8 Four arithmetic units AU are shown, but exemplary embodiments are not limited thereto.

[0114] The four arithmetic units AU can receive in parallel the four position elements B(k)~B(k+3) of the base array BA read from the memory body MBKB, and provide in parallel four target addresses TADDR{T(k)}~TADDR{(Tk+3)} through the calculation of Expression 3.

[0115] Thereafter, address calculations for the subsequent four position elements B(k+4) to B(k+7) and read operations based on the calculated addresses can be performed. In this way, the operation can be repeated NT / NR times to sequentially pre-fetch the target elements corresponding to all NT position elements. The address calculations and read operations of the target elements can be performed by a pipeline scheme using registers 220 and 230. In other words, the read operation of the previous target element and the subsequent address calculation can be performed simultaneously or can overlap. Through this parallel pipeline operation, the entire pre-fetch time can be further reduced.

[0116] In an exemplary embodiment, the target data register 230 is implemented with a static random access memory (SRAM). Intermediate data such as indirect memory access information IMAI and target address TADDR that are used in calculations and do not need to be stored are not transmitted to external devices, so only target elements used in operations (e.g., SpMV) are stored in the SRAM, and external devices can quickly access data in the SRAM to reduce the overall operation time.

[0117] Fig. 9 is a block diagram illustrating a system including a memory device according to an exemplary embodiment of the inventive concept.

[0118] refer to Fig. 9 , the system 11 includes a host device 50 and a semiconductor memory device 30. The host device may include a processor (not shown), a memory controller MCTR 20, and a cache memory CCM 51. In an exemplary embodiment, the memory controller 20 is implemented as a component different from the host device 50. The semiconductor memory device 30 may include a memory cell array MC 40, an information register IREG 100, a prefetch circuit PFC 200, and a calculation circuit PIMC 300. Fig. 9 include Figure 1 Some elements are omitted hereafter.

[0119] As reference Figure 8 As described above, the target data register 230 in the pre-fetch circuit 200 can be implemented with an SRAM having a fast access speed. In an exemplary embodiment, the memory controller 20 loads the target element pre-fetched in the target data register 230 to the cache memory 51 of the host device 51 according to the request of the host device 50. In an exemplary embodiment, the host device 50 uses the target data register 230 as another cache memory. In this case, the target data register 230 has a lower cache level than the cache memory 51 in the host device 50.

[0120] In an embodiment, the computing circuit 300 performs a process-in-memory (PIM) operation based on the first data and the second data to provide computing result data. As described above, the first data may be a sparse matrix, the second data may be a vector, and the computing circuit 300 may perform SpMV as a PIM operation.

[0121] Fig.10 is a diagram illustrating an example embodiment of an operation unit for performing an in-memory processing operation included in a memory device according to an exemplary embodiment of the inventive concept. Fig. 9 The calculation circuit 300 may include Fig.10 Multiple computing units are shown to perform parallel computing.

[0122] refer to Fig.10, each operation unit 500 includes a multiplication circuit 520 and an accumulation circuit 540. The multiplication circuit 520 includes buffers 521 and 522 and a multiplier 523, and is configured to multiply the effective elements of the effective data array AA corresponding to the first data and the target elements of the target data array TA corresponding to the second data. The accumulation circuit 540 includes an adder 541 (e.g., an adding circuit) and a buffer 542, which are used to accumulate the output of the multiplication circuit 520 to provide corresponding calculation result data DRi. For example, the multiplier 523 can multiply the first effective element of the effective elements by the first target element of the target elements to generate a first result, the adder 541 can add the first result to the initial value (e.g., 0) received from the buffer 542 to generate a first sum to be stored in the buffer 542, the multiplier 523 can multiply the second effective element of the effective elements by the second target element of the target elements to generate a second result, the adder 541 can add the second result to the first sum received from the buffer 542, and so on. The accumulation circuit 540 may be initialized in response to the reset signal RST, and may output corresponding calculation result data DRi in response to the output enable signal OUTEN. For example, in response to the reset signal RST, the value of the buffer 542 may be set to 0. For example, the corresponding calculation result data DRi may be the last output of the ADDER 541 stored in the buffer 542 before the accumulation circuit 540 is reset. Fig.10 The illustrated operation unit 500 can efficiently perform operations such as the above-mentioned SpMV.

[0123] Fig.11 and Fig.12 is an exploded perspective view of a system including a stacked memory device according to an exemplary embodiment of the inventive concept.

[0124] refer to Fig.11 , the system 800 includes a stacked memory device 1000 and a host device 2000 .

[0125] The stacked memory device 1000 may include a base semiconductor die or a logic semiconductor die 1010 , and a plurality of memory semiconductor dies 1070 and 1080 stacked with the logic semiconductor die 1010 . Fig.11 A non-limiting example of one logic semiconductor die and two memory semiconductor dies is shown. Two or more logic semiconductor dies and one, three or more memory semiconductor dies may be included in the stacked structure. In addition, Fig.11 A non-limiting example of memory semiconductor die 1070 and 1080 being stacked vertically with logic semiconductor die 1010 is shown. Fig.13As described, the memory semiconductor dies 1070 and 1080 other than the logic semiconductor die 1010 may be vertically stacked, and the logic semiconductor die 1010 may be electrically connected to the memory semiconductor dies 1070 and 1080 through an interposer and / or a base substrate.

[0126] In an embodiment, logic semiconductor die 1010 includes memory interface MIF 1020 and logic (e.g., logic circuit) to access memory integrated circuits 1071 and 1081 formed in memory semiconductor dies 1070 and 1080. Such logic may include control circuit CTRL 1030, global buffer GBF 1040, and data conversion logic DTL 1050. In an embodiment, logic DTL 1050 is implemented by diode-transistor logic.

[0127] The memory interface 1020 may perform communication with an external device such as a host device 2000 through the interconnect device 12. For example, the memory interface 1020 may be interfaced with a host interface (HIF) 2110 of the host device 2000. The host device 2000 may include a semiconductor die 2100 on which the HIF 2110 is mounted. Additional components CR12120 and CR22130 (e.g., a processor) may be mounted on the semiconductor die 2100 of the host device 2000. The control circuit 1030 may control the overall operation of the stacked memory device 1000. The data conversion logic 1050 may perform a logic operation on data exchanged with the memory semiconductor dies 1070 and 1080 or data exchanged through the memory interface 1020. For example, the data conversion logic may perform a maximum pooling operation, a rectified linear unit (ReLU) operation, channel-by-channel addition, and the like.

[0128] Memory semiconductor dies 1070 and 1080 may include memory integrated circuits 1071 and 1081, respectively. At least one of memory semiconductor dies 1070 and 1080 may be computing semiconductor die 1080 including computing circuit 300. Computing semiconductor die 1080 may include the information register 100 and prefetch circuit 200 described above for performing indirect memory prefetching.

[0129] Fig.12 System 801 and Fig.11 The system 800 is substantially the same as that of FIG. 8 , and repeated descriptions are omitted. Fig.12 , the above-mentioned information register 100 , pre-fetch circuit 200 and calculation circuit 300 may be included in a logic semiconductor die 1010 .

[0130] Memory bandwidth and latency are performance bottlenecks in many processing systems. Memory capacity can be increased by using stacked memory devices, in which multiple semiconductor devices are stacked in the package of a memory chip. The stacked semiconductor die can be electrically connected using through-silicon vias or through-substrate vias (TSVs). This stacking technology can increase memory capacity and also suppress bandwidth and latency costs. Each access of an external device to a stacked memory device involves data communication between stacked semiconductor dies. In this case, for each access, the inter-device bandwidth and inter-device latency costs may occur twice. Therefore, when the task of an external device requires multiple accesses to a stacked memory device, the inter-device bandwidth and inter-device latency may have a significant impact on the processing efficiency and power consumption of the system.

[0131] A stacked memory device and system according to at least one exemplary embodiment may reduce latency and power consumption by efficiently using information registers, pre-fetch circuits, and computation circuits disposed in a logic semiconductor die or a memory semiconductor die, combining memory-intensive or data-intensive data processing and memory access.

[0132] refer to Fig.11 and Fig.12 An exemplary embodiment of the arrangement of the information register 100, the pre-fetch circuit 200, and the calculation circuit 300 is described, but the exemplary embodiments of the inventive concept are not limited thereto. According to an exemplary embodiment of the inventive concept, the information register 100 and the pre-fetch circuit 200 are arranged in a memory semiconductor die, and the calculation circuit 300 is arranged in a logic semiconductor die.

[0133] Fig.13 and Fig.14 is a diagram illustrating a package structure of a stacked memory device according to an exemplary embodiment of the inventive concept.

[0134] refer to Fig.13 , the memory chip 2001 includes an interposer ITP and a stacked memory device stacked on the interposer ITP. The interposer ITP may be implemented by an electrical interface for an interface connection between one socket and another socket or for connection to another socket. The stacked memory device may include a logic semiconductor die LSD and a plurality of memory semiconductor dies MSD1-MSD4.

[0135] refer to Fig.14 , the memory chip 2002 includes a base substrate BSUB and a stacked memory device stacked on the base substrate BSUB. The stacked memory device may include a logic semiconductor die LSD and a plurality of memory semiconductor dies MSD1 ˜ MSD4 .

[0136] Fig.13The structure in which the memory semiconductor dies MSD1 to MSD4 other than the logic semiconductor die LSD are vertically stacked and the logic semiconductor die LSD is electrically connected to the memory semiconductor dies MSD1 to MSD4 through the interposer ITP or the base substrate is shown. Fig.14 The structure in which the logic semiconductor die LSD and the memory semiconductor dies MSD1 - MSD4 are vertically stacked is shown.

[0137] As described above, at least one of the memory semiconductor dies MSD1 ˜ MSD4 may include the above-mentioned information register IREG and prefetch circuit PFC. Although not shown, the above-mentioned calculation circuit may be included in the memory semiconductor dies MSD1 ˜ MSD4 and / or the logic semiconductor die LSD.

[0138] The base substrate BSUB may be the same as or include the interposer ITP. The base substrate BSUB may be a printed circuit board (PCB). External connection elements such as conductive bumps BMP may be formed on the lower surface of the base substrate BSUB, and internal connection elements such as conductive bumps may be formed on the upper surface of the base substrate BSUB. Fig.13 In an exemplary embodiment of the present invention, the logic semiconductor die LSD and the memory semiconductor dies MSD1 ˜ MSD4 may be electrically connected through through-silicon vias. The stacked semiconductor dies LSD and MSD1 ˜ MSD4 may be packaged using a resin RSN.

[0139] Fig.15 is a diagram illustrating an example structure of a stacked memory device according to an exemplary embodiment of the inventive concept.

[0140] refer to Fig.15 , high bandwidth memory (HBM) 1100 can be configured as a stack with multiple DRAM semiconductor dies 1120, 1130, 1140, and 1150. The stacked HBM 1100 can be optimized through multiple independent interfaces called channels. Each DRAM stack can support up to eight channels according to the HBM standard. Fig.15 An example stack including four DRAM semiconductor dies 1120 , 1130 , 1140 , and 1150 is shown, and each DRAM semiconductor die supports two channels CHANNEL0 and CHANNEL1 .

[0141] Each channel provides access to an independent set of DRAM banks. Requests from one channel will not access data attached to other channels. Channels are independently timed and do not require synchronization.

[0142] The HBM 1100 may also include an interface die 1110 or a logic die disposed at the bottom of the stack structure to provide signal routing and other functions. Some functions of the DRAM semiconductor dies 1120 , 1130 , 1140 , and 1150 may be implemented in the interface die 1110 .

[0143] At least one of the DRAM semiconductor dies 1120 , 1130 , 1140 , and 1150 may include the above-described information register 100 and prefetch circuit 200 for performing indirect memory prefetch and calculation circuit 300 for PIM operation, as described above.

[0144] Fig.16 is a block diagram illustrating a computing system according to an exemplary embodiment of the inventive concept.

[0145] refer to Fig.16 , the computing system 4000 includes a processor 4100, an input / output hub IOH 4200, an input / output controller hub ICH 4300, at least one memory module 4400, and a graphic card 4500.

[0146] The processor 4100 may perform various computing functions, such as executing software for performing calculations or tasks. For example, the processor 4100 may be a microprocessor, a central processing unit (CPU), a digital signal processor, etc. The processor 4100 may include a memory controller 4110 for controlling the operation of the memory module 4400. The memory module 4400 may include a plurality of memory devices storing data provided from the memory controller 4110. At least one memory device in the memory module 4400 may include the above-mentioned information register 100 and prefetch circuit 200 for performing indirect memory prefetching. In an embodiment, the processor 4100 is part of a host device.

[0147] The input / output hub 4200 may manage data transfer between the processor 4100 and devices such as a graphics card 4500. In an embodiment, the input / output hub 4200 is implemented by a microchip. The graphics card 4500 may control a display device for displaying an image. The input / output controller hub 4300 may perform data buffering and interface arbitration to efficiently operate various system interfaces. Fig.16 The components in can be coupled to various interfaces, such as an accelerated graphics port (AGP) interface, a peripheral component interface (PCI), a peripheral component interface express (PCIe), a universal serial bus (USB) port, a serial advanced technology attachment (SATA) port, a low pin count (LPC) bus, a serial peripheral interface (SPI), a general purpose input / output (GPIO), etc.

[0148] As described above, the memory device and system according to at least one exemplary embodiment of the present invention can enhance the accuracy and efficiency of in-memory pre-fetching by performing indirect memory access based on indirect memory access information provided from a memory controller. In addition, the memory device and system according to at least one exemplary embodiment of the present invention can reduce the delay of indirect memory access and improve the speed of in-memory pre-fetching by parallelizing the data arrangement for indirect memory access and more effectively utilizing the bandwidth within the memory device. In addition, the memory device and system according to at least one exemplary embodiment can efficiently perform sparse data operations by performing in-memory processing operations using in-memory pre-fetching.

[0149] The present invention can be applied to memory devices and systems including memory devices. For example, the present invention can be applied to systems such as memory cards, solid-state drives (SSDs), embedded multimedia cards (eMMCs), mobile phones, smart phones, personal digital assistants (PDAs), portable multimedia players (PMPs), digital cameras, camcorders, personal computers (PCs), server computers, workstations, notebook computers, digital TVs, set-top boxes, portable game consoles, navigation systems, wearable devices, Internet of Things (IoT) devices, Internet of Everything (IoE) devices, e-books, virtual reality (VR) devices, augmented reality (AR) devices, etc.

[0150] The foregoing is an illustration of example embodiments of the inventive concept and should not be construed as limiting thereof. Although some exemplary embodiments have been described, those skilled in the art will readily appreciate that various modifications may be made in the exemplary embodiments without substantially departing from the inventive concept.

Claims

1. A memory device, comprising: a memory cell array configured to store a valid data array, a base array, and a target data array, the valid data array sequentially including valid elements among elements of the first data, the base array sequentially including position elements indicating position values ​​corresponding to the valid elements, and the target data array sequentially including target elements of the second data corresponding to the position values; an information register configured to store indirect memory access information including a start address of the target data array, a unit size of the target element, and a read number NR of position elements simultaneously read from the memory cell array, wherein NR is a natural number greater than 1; as well as The prefetch circuit is configured to prefetch the target element corresponding to the read number NR of the position elements based on the indirect memory access information.

2. The memory device according to claim 1, wherein: The indirect memory access information is provided to the memory device from an external memory controller.

3. The memory device according to claim 1, wherein: The first data is a sparse matrix, and the second data is a vector.

4. The memory device according to claim 1, wherein: The memory cell array includes a plurality of memory banks, and The valid data array, the base array and the target data array are stored in different storage banks among the multiple storage banks.

5. The memory device according to claim 1, wherein: The pre-fetch circuit comprises: an arithmetic circuit configured to calculate a target address corresponding to the read position element based on the read position element, the start address of the target data array, and the unit size of the target element; a target address register configured to store the target address calculated by the arithmetic circuit; and A target data register is configured to store the target element read from the target address of the memory cell array.

6. The memory device according to claim 5, wherein: The arithmetic circuit calculates an i-th target address among the target addresses by multiplying the unit size by B(i)-1 and adding the start address to the result of the multiplication, where B(i) represents an i-th position element among the position elements in the base array, where i is a natural number greater than or equal to 1.

7. The memory device according to claim 5, wherein: The arithmetic circuit comprises: NR arithmetic units are configured to provide NR target addresses in parallel based on a read number NR of the position elements simultaneously read from the memory cell array.

8. The memory device according to claim 7, wherein: When the total number of the position elements is NT, the NR arithmetic units provide NT target addresses by repeatedly performing NT / NR address calculations, where NT is a natural number greater than 1.

9. The memory device according to claim 5, wherein: The target data register is implemented by a static random access memory SRAM.

10. The memory device of claim 1, wherein: The indirect memory access information also includes a start address of the base array, a unit size of the position element, and a total number of the position elements.

11. The memory device of claim 1 , further comprising: The computing circuit is configured to perform an in-memory processing operation based on the first data and the second data to provide computing result data.

12. The memory device of claim 11, wherein: The first data is a sparse matrix, the second data is a vector, and Wherein, the in-memory processing operation performs sparse matrix-vector multiplication.

13. A memory device comprising: A plurality of memory semiconductor dies stacked in a vertical direction, wherein a memory cell array is formed in the plurality of memory semiconductor dies, wherein the memory cell array stores a valid data array, a base array, and a target data array, wherein the valid data array sequentially includes valid elements among elements of first data, the base array sequentially includes position elements indicating position values ​​corresponding to the valid elements, and the target data array sequentially includes target elements of second data corresponding to the position values; a plurality of through silicon vias electrically connecting the plurality of memory semiconductor dies; an information register configured to store indirect memory access information including a start address of the target data array, a unit size of the target element, and a read number NR of position elements simultaneously read from the memory cell array, wherein NR is a natural number greater than 1; as well as The prefetch circuit is configured to prefetch the target element corresponding to the read number NR of the position elements based on the indirect memory access information.

14. The memory device of claim 13, wherein: The pre-fetch circuit is located in the memory semiconductor die where the memory cell array storing the valid data array, the base array, and the target data array is located.

15. The memory device of claim 13, further comprising: a logic semiconductor die including circuitry for controlling access to said memory cell array, Wherein, the pre-fetch circuit is formed in the logic semiconductor die.

16. The memory device of claim 13, wherein: The indirect memory access information is provided to the memory device from an external memory controller.

17. The memory device of claim 13, wherein: The pre-fetch circuit provides NR target addresses in parallel based on NR of the position elements read simultaneously from the memory cell array.

18. A memory device comprising: a memory cell array configured to store a valid data array, a base array, and a target data array, the valid data array sequentially including valid elements among elements of the first data, the base array sequentially including position elements indicating position values ​​corresponding to the valid elements, and the target data array sequentially including target elements of the second data corresponding to the position values; Random Access Memory RAM; as well as a controller configured to receive indirect memory access information including a start address of the target data array, a unit size of the target element, and a read number NR of position elements to be simultaneously read from the memory cell array from an external device, wherein NR is a natural number greater than 1; The controller performs a prefetch operation in response to receiving the indirect memory access information, wherein the prefetch operation uses the starting address and the unit size to read the target element corresponding to the read number NR of the position elements, and stores the read target element in the RAM.

19. The memory device of claim 18, wherein: The prefetch operation calculates a target address corresponding to the read position element based on the read position element, the start address, and the unit size, and performs reading using the calculated target address.

20. The memory device of claim 19, wherein: The controller calculates an i-th target address among the target addresses by multiplying the unit size by B(i)-1 and adding the start address to the result of the multiplication, where B(i) represents an i-th position element among the position elements in the base array, where i is a natural number greater than or equal to 1.

Citation Information

Patent Citations

  • Intraocular lens inspection

    KR1020180119527A

  • Hardware prefetcher for indirect access patterns

    US20160188476A1

  • Hybrid processor

    US20160224465A1

  • Instruction and Logic for Indirect Accesses

    US20170091103A1