Memory device supporting jump computing mode and operation method thereof

By using processing circuits and index data in memory devices, selective calculations in jump computing mode are realized, and bandwidth and delay problems caused by multiple accesses of stacked memory devices are solved, reducing power consumption and improving processing efficiency.

CN110176260BActive Publication Date: 2025-05-16SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910119977.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-02-21
Filing Date
2019-02-18
Publication Date
2025-05-16
Estimated Expiration
2039-02-18

AI Technical Summary

Technical Problem

In processing systems, memory bandwidth and delay are performance bottlenecks, especially in the case of multiple accesses of stacked memory devices, the inter-device bandwidth and delay penalty has a significant impact on the processing efficiency and power consumption of the system.

Method used

By introducing a processing circuit in the memory device, calculation of broadcast data and internal data is selectively performed in the jump calculation mode based on the index data, and calculation and reading operations regarding invalid data are omitted.

Benefits of technology

This method can reduce power consumption, improve processing efficiency, and reduce bandwidth and delay penalties between devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110176260B_ABST
    Figure CN110176260B_ABST
Patent Text Reader

Abstract

A memory device includes: a memory cell array formed in a semiconductor die, the memory cell array including a plurality of memory cells for storing data; and a computing circuit formed in the semiconductor die. The computing circuit performs computing based on broadcast data and internal data, omits computing about invalid data, and performs computing about valid data based on index data in a skip computing mode, wherein the broadcast data is provided from the outside of the semiconductor die, the internal data is read from the memory cell array, and the index data indicates whether the internal data is valid data or invalid data. Omitting computing and reading operations about invalid data by skip computing mode based on the index data reduces power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the priority of Korean Patent Application No. 10-2018-0020422 filed on February 21, 2018 in the Korean Intellectual Property Office (KIPO), the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] Example embodiments generally relate to semiconductor integrated circuits. For example, at least some example embodiments relate to a memory device supporting a skip computing mode and / or a method of operating a memory device. Background Art

[0004] In many processing systems, memory bandwidth and latency are performance bottlenecks. Memory capacity can be increased by using stacked memory devices, wherein multiple semiconductor devices are stacked in the package of the memory chip. Stacked semiconductor die can be electrically connected by using through silicon vias or through-substrate vias (TSV). This stacking technology can increase memory capacity and can also suppress bandwidth and latency penalties. Each access of an external device to a stacked memory device involves data communication between stacked semiconductor dies. In this case, for each access, inter-device bandwidth and inter-device latency penalties may occur twice. Therefore, when the task of an external device requires multiple accesses to the stacked memory device, inter-device bandwidth and inter-device latency may have a significant impact on the processing efficiency and power consumption of the system. Summary of the invention

[0005] Some example embodiments may provide a memory device capable of efficiently performing process in memory (PIM) and / or a method of operating a memory device.

[0006] According to an example embodiment, a memory device includes: a memory cell array associated with a semiconductor die, the memory cell array including a plurality of memory cells configured to store data; and a processing circuit associated with the semiconductor die, the processing circuit being configured to: selectively perform calculations on broadcast data and internal data in a skip calculation mode based on index data indicating whether the internal data is invalid data or valid data, the broadcast data being provided from outside the semiconductor die and the internal data being read from the memory cell array.

[0007] According to an example embodiment, a memory device includes: a plurality of memory semiconductor dies stacked in a vertical direction; silicon through vias electrically connecting the plurality of memory semiconductor dies; a plurality of memory integrated circuits (ICs) associated with corresponding memory semiconductor dies among the plurality of memory semiconductor dies, the plurality of memory ICs being configured to store data; and a processing circuit associated with one or more computing semiconductor dies among the plurality of memory semiconductor dies, the processing circuit being configured to: selectively perform calculations based on broadcast data and the internal data in a skip calculation mode based on index data indicating that the internal data is invalid data or valid data, the broadcast data being commonly provided to the computing semiconductor die through the silicon through vias, and the internal data being read respectively from the plurality of memory integrated circuits.

[0008] According to example embodiments, a method for operating a memory device is provided, the memory device including a semiconductor die and a processing circuit associated with the semiconductor die, the semiconductor die having a memory cell array. In some example embodiments, the method includes: receiving index data indicating whether internal data is valid data or invalid data in a skip calculation mode; and selectively performing calculations on broadcast data and the internal data by the processing circuit in the skip calculation mode based on the index data indicating whether the internal data is the invalid data or the valid data, the broadcast data being provided from outside the semiconductor die, and the internal data being read from the memory cell array.

[0009] A memory device and a method of operating the memory device according to example embodiments may reduce power consumption by omitting calculation and read operations regarding invalid data through a skip calculation mode based on index data. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Example embodiments of the present disclosure will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.

[0011] Figure 1 is a flowchart illustrating a method of operating a memory device according to example embodiments.

[0012] Figure 2 is a block diagram illustrating a memory system including a memory device according to example embodiments.

[0013] Figure 3 It is shown Figure 2 A diagram of an example embodiment of a memory device included in a system.

[0014] Figure 4 is a block diagram illustrating a memory device according to example embodiments.

[0015] Figure 5 It is shown Figure 4 FIG. 1 is a diagram of an example embodiment of a generation unit included in an index data generator in a memory device of FIG.

[0016] Fig. 6A and Figure 6B It is shown Figure 5 A diagram of an example embodiment of an index storage included in a generation unit.

[0017] Fig. 7A and Figure 7B is a diagram for describing index data according to an example embodiment.

[0018] Figure 8 is a timing diagram illustrating an example write operation of a memory device according to example embodiments.

[0019] Fig. 9 is a block diagram illustrating a memory device according to example embodiments.

[0020] Fig.10 It is shown Fig. 9 FIG. 1 is a diagram of an example embodiment of a computing unit included in a computing circuit in a memory device of FIG.

[0021] Fig.11 It is shown Fig.10 FIG. 1 is a diagram of an example arrangement of calculators included in a computing unit of FIG.

[0022] Fig.12 It is shown Fig.10 FIG. 1 is a diagram of an example embodiment of a calculator included in a computing unit of FIG.

[0023] Fig.13A and Fig. 13B is a timing diagram illustrating example operations of a skip calculation mode in a memory device according to example embodiments.

[0024] Fig.14 is a timing diagram illustrating example operations of a normal computing mode in a memory device according to example embodiments.

[0025] Fig.15 is a diagram illustrating an example embodiment of outputting calculation result data.

[0026] Fig.16 is a diagram illustrating matrix calculation using a calculation circuit according to an example embodiment.

[0027] Fig.17 is an exploded perspective view of a system including a stacked memory device according to example embodiments.

[0028] Fig.18 is a diagram illustrating an example high bandwidth memory (HBM) organization.

[0029] Fig.19 and Fig. 20 is a diagram illustrating a package structure of a stacked memory device according to example embodiments.

[0030] Fig.21 is a block diagram illustrating a mobile system according to an example embodiment. DETAILED DESCRIPTION

[0031] Various example embodiments will be described more fully below with reference to the accompanying drawings, in which some example embodiments are shown. In the accompanying drawings, the same reference numerals always represent the same elements. Repeated descriptions may be omitted.

[0032] Figure 1 is a flowchart illustrating a method of operating a memory device according to example embodiments.

[0033] refer to Figure 1 In operation S100, a computing circuit may perform computing based on broadcast data and internal data. The computing circuit may be formed in a semiconductor die in which a memory cell array is formed. The broadcast data may be provided from outside the semiconductor die, and the internal data may be read from the memory cell array.

[0034] In operation S200, in the skip calculation mode, index data indicating whether internal data is valid data or invalid data in the skip calculation mode is provided.

[0035] In operation S300, in the skip calculation mode, calculation on invalid data is omitted, and calculation on valid data is performed based on index data.

[0036] As such, a method of operating a stacked memory device according to example embodiments may reduce power consumption by omitting calculation and read operations regarding invalid data through a skip calculation mode based on index data.

[0037] Figure 2 is a block diagram illustrating a memory system including a memory device according to example embodiments.

[0038] refer to Figure 2 , the memory system 10 may include a memory controller 20 and a semiconductor memory device 30 .

[0039] The memory controller 20 may control the overall operation of the memory system 10. The memory controller 20 may control the overall data exchange between the external host and the semiconductor memory device 30. For example, the memory controller 20 may write data in the semiconductor memory device 30 or read data from the semiconductor memory device 30 in response to a request from the host. In addition, the memory controller 20 may issue an operation command to the semiconductor memory device 30 for controlling the semiconductor memory device 30.

[0040] In some example embodiments, the semiconductor memory device 30 may be a memory device including a dynamic memory cell, such as a dynamic random access memory (DRAM), a double data rate 4 (DDR4) synchronous DRAM (SDRAM), a low power DDR4 (LPDDR4) SDRAM, or a LPDDR5 SDRAM.

[0041] The memory controller 20 may send a clock signal CLK, a command CMD, and an address (signal) ADDR to the semiconductor memory device 30, and exchange data DQ with the semiconductor memory device 30. In addition, the memory controller 20 may provide a mode signal MD to the semiconductor memory device 30, the mode signal MD indicating a skip calculation mode or a normal calculation mode to be described below. The mode signal MD may be provided as a control signal, or may be provided by a mode register write command for setting a mode register in the semiconductor memory device 30.

[0042] The semiconductor memory device 30 may include a memory cell array MC40 , a calculation circuit CAL100 , and an index data generator IDG200 .

[0043] The memory cell array 40 may include a plurality of memory cells for storing data. The memory cells may be grouped into a plurality of banks, and each bank may include a plurality of data blocks.

[0044] The computing circuit 100 may be formed in a semiconductor die in which the memory cell array 40 is formed. The computing circuit 100 may perform computing based on broadcast data and internal data, wherein the broadcast data is provided from the outside of the semiconductor die and the internal data is read from the memory cell array. In the skip computing mode, the computing circuit 100 may omit computing on invalid data and perform computing on valid data based on index data, wherein the index data indicates whether the internal data is valid data or invalid data.

[0045] In some example embodiments, the computing circuit 100 may generate a skip enable signal SEN to control the selective omission of the computation. Figures 9 to 16 An example embodiment of computing circuit 100 is described.

[0046] The index data generator 200 may generate the index data ID based on the write data stored in the memory cell array 40. Figures 4 to 6B An example embodiment of the index data generator 200 is described.

[0047] In some example embodiments, the index data ID generated by the index data generator 200 may be stored in the memory cell array 40 during a write operation to store write data in the memory cell array 40. The index data ID stored in the memory cell array 40 may be read out from the memory cell array 40 and provided to the calculation circuit 100 in a skip calculation mode.

[0048] Figure 3 It is shown Figure 2 A diagram of an example embodiment of a memory device included in a system.

[0049] Although reference Figure 3 DRAM is described as an example of a memory device formed in a memory semiconductor die, but the memory device can be any of a variety of memory cell architectures including, but not limited to, volatile memory architectures (e.g., DRAM, TRAM, and SRAM) or non-volatile memory architectures (e.g., ROM, flash memory, FRAM, MRAM, etc.).

[0050] refer to Figure 3 , the memory device 400 includes a control logic 410, an address register 420, a storage body control logic 430, a row address multiplexer 440, a column address latch 450, a row decoder 460, a column decoder 470, a memory cell array 480, a computing circuit 100, an input-output (I / O) selection circuit 490, a data input-output (I / O) buffer 495, a refresh counter 445 and an index data generator 200.

[0051] The memory cell array 480 may include a plurality of bank arrays 480a-480h. The row decoder 460 may include a plurality of bank row decoders 460a-460h coupled to the bank arrays 480a-480h, respectively, and the column decoder 470 may include a plurality of bank column decoders 470a-470h coupled to the bank arrays 480a-480h, respectively.

[0052] The computing circuit 100 may include a plurality of computing blocks CB 100a ˜ 100h respectively coupled to memory bank arrays 480a ˜ 480h . Figure 3In the illustrated non-limiting embodiment, the computing circuit 100 is disposed between the memory cell array 480 and the input-output gating circuit 490 , but the input-output gating circuit 490 may also be disposed between the memory cell array 480 and the computing circuit 100 .

[0053] Each of the computing blocks 100a-100h may include a plurality of computing units (not shown) that receive common broadcast data and corresponding internal data from the memory bank arrays 480a-480h.

[0054] The index data generator 200 may generate the index data ID based on the write data WRD stored in the memory cell array 480 during a write operation.

[0055] The address register 420 may receive an address ADDR including a bank address BANK_ADDR, a row address ROW_ADDR, and a column address COL_ADDR from the memory controller. The address register 420 may provide the received bank address BANK_ADDR to the bank control logic 430, may provide the received row address ROW_ADDR to the row address multiplexer 440, and may provide the received column address COL_ADDR to the column address latch 450.

[0056] The bank control logic 430 may generate a bank control signal in response to the bank address BANK_ADDR, may activate one of the bank row decoders 460a-460h corresponding to the bank address BANK_ADDR in response to the bank control signal, and may activate one of the bank column decoders 470a-470h corresponding to the bank address BANK_ADDR in response to the bank control signal.

[0057] The row address multiplexer 440 may receive the row address ROW_ADDR from the address register 420 and may receive the refresh row address REF_ADDR from the refresh counter 445. The row address multiplexer 440 may selectively output the row address ROW_ADDR or the refresh row address REF_ADDR as the row address RA. The row address RA output from the row address multiplexer 440 may be applied to the bank row decoders 460a-460h.

[0058] The activated bank row decoder among the bank row decoders 460a-460h may decode the row address RA output from the row address multiplexer 440 and may activate a word line corresponding to the row address RA. For example, the activated bank row decoder may apply a word line driving voltage to the word line corresponding to the row address RA.

[0059] The column address latch 450 may receive a column address COL_ADDR from the address register 420 and may temporarily store the received column address COL_ADDR. In some example embodiments, in a burst mode, the column address latch 450 may generate a column address incremented from the received column address COL_ADDR. The column address latch 450 may apply the temporarily stored or generated column address to the bank column decoders 470a-470h.

[0060] An activated one of the bank column decoders 470a-470h may decode the column address COL_ADDR output from the column address latch 450, and may control the input / output gating circuit 490 to output data corresponding to the column address COL_ADDR. The I / O gating circuit 490 may include a circuit for gating input data and output data. The I / O gating circuit 490 may also include a read data latch for storing data output from the bank array 480a-480h and a write driver for writing data to the bank array 480a-480h.

[0061] Data to be read from one of the bank arrays 480a-480h may be sensed by a bank sense amplifier coupled to the one bank array from which data is to be read, and may be stored in a read data latch. The data stored in the read data latch may be provided to a memory controller via a data I / O buffer 495. Data DQ to be written to one of the bank arrays 480a-480h may be provided from the memory controller to the data I / O buffer 495. A write driver may write the data DQ to one of the bank arrays 480a-480h.

[0062] The control logic 410 may control the operation of the memory device 400. For example, the control logic 410 may generate a control signal for the memory device 400 to perform a write operation or a read operation. The control logic 410 may include a command decoder 411 that decodes a command CMD received from a memory controller and a mode register 412 that sets an operation mode of the memory device 400. The control logic 410 may control the memory device 400 to selectively operate in a skip calculation mode or in a normal calculation mode in response to the mode signal MD.

[0063] Figure 4 is a block diagram illustrating a memory device according to example embodiments.

[0064] Figure 4 For describing the write operation, only the components used for the write operation are shown, and Figure 4 Other components are omitted. For ease of explanation, Figure 4A configuration corresponding to one memory bank is shown.

[0065] refer to Figure 4 , the memory device 50 may include a plurality of data blocks DB1 ˜DBn, an input / output gating circuit 52 , and an index data generator 200 . Figure 4 The configuration of the first data block DB1 as an example is shown, and the other data blocks DB2_DBn may have the same configuration as the first data block DB1. Each data block may include a plurality of sub-memory cell arrays SARR, and each sub-memory cell array SARR may include a plurality of memory cells. In a write operation, data provided from the outside may be sequentially stored in the memory cells through the global input-output line GIO and the local input-output line LIO. The hierarchical structure of the data blocks may be implemented in various ways.

[0066] The input / output gating circuit 52 can select the local input / output line corresponding to the column address of the write data WRD1-WRDn based on the column selection signal CSL. Figure 3 As described, the column selection signal CSL may be provided from the column decoder 470. The input-output gating circuit 52 may include a plurality of switch circuits MUX1MUXn corresponding to the plurality of data blocks DB1DBn, respectively.

[0067] The index data generator 200 may generate index data ID1 ˜IDn based on the write data WRD1 ˜WRDn stored in the data blocks DB1 ˜DBn, respectively. The index data ID1 ˜IDn may have different values ​​according to the write data WRD1 ˜WRDn.

[0068] The index data generator 200 may include a plurality of generation units GU1~GUn corresponding to a plurality of data blocks DB1~DBn. Each generation unit GUi (i=1~n) may generate corresponding index data IDi based on corresponding write data WRDi stored in the corresponding data block DBi. In other words, the first generation unit GU1 may generate first index data ID1 based on first write data WRD1 stored in the first data block DB1, the second generation unit GU2 may generate second index data ID2 based on second write data WRD2 stored in the second data block DB2, and in this way, the nth generation unit GUn may generate nth index data IDn based on nth write data WRDn stored in the nth data block DBn. ​​During a write operation, the index data ID1~IDn generated by the index data generator 200 may be stored in the data blocks DB1~DBn of the memory cell array together with the write data WRD1~WRDn.

[0069] Figure 5 It is shown Figure 4FIG. 1 is a diagram of an example embodiment of a generation unit included in an index data generator in a memory device of FIG.

[0070] refer to Figure 5 , the generation unit 210 may include a logic gate LG 220 and an index storage IREG 230. The logic gate 220 may perform a logic operation on the corresponding write data WRDi, and the index storage 230 may store the corresponding index data IDi based on an output signal LO of the logic gate 220.

[0071] The corresponding write data IDi may include N data bits B0 to BN-1, and the logic gate 220 may generate an output signal LO by performing a logic operation on the data bits B0 to BN-1 of the corresponding write data WRDi. In some example embodiments, the logic gate 220 may be implemented with an OR logic gate. In this case, when the values ​​of all bits B0 to BN-1 of the corresponding write data WRDi are 0, the output signal LO may have a first value (e.g., "0") indicating that the corresponding write data WRDi is invalid data, and when the value of at least one bit among all bits B0 to BN-1 of the corresponding write data WRDi is 1, the output signal may have a second value (e.g., "1") indicating that the corresponding write data WRDi is valid data.

[0072] The logic gate 220 may sequentially perform a logic operation on the respective write data WRDi corresponding to the plurality of column addresses, and the index storage 230 may sequentially store a plurality of index bits of the respective index data IDi based on an output signal LO of the logic gate 220 .

[0073] Fig. 6A and Figure 6B It is shown Figure 5 A diagram of an example embodiment of an index storage included in a generation unit.

[0074] refer to Figure 5 and Fig. 6A , the index storage 231 can sequentially store the value of the output signal LO of the logic gate 220 based on the pointer signal PT as multiple index bits I0~I7 of the index data IDi. The multiple index bits I0~I7 can correspond to multiple column addresses, and each index bit I0~I7 can indicate whether the internal data read from the corresponding column address is valid data or invalid data. The index storage 231 can output the stored index bits I0~I7 in the form of parallel signals as the index data IDi in response to the output enable signal OEN.

[0075] refer to Figure 5 and Figure 6B, the index storage 232 may be implemented with a shift register, which is configured to perform a shift operation in synchronization with the clock signal CLK to sequentially store the value of the output signal LO of the logic gate 220 as a plurality of index bits I0 to I7 of the index data IDi. In addition, the index storage 232 may perform a shift operation in synchronization with the clock signal CLK to output the stored index bits I0 to I7 in the form of a serial signal as the index data IDi. In some example embodiments, as shown in FIG. Fig. 6A As described above, the index storage 232 may output the stored index bits I0 ˜ I7 as index data IDi in the form of parallel signals in response to the output enable signal OEN.

[0076] Fig. 7A and Figure 7B is a diagram for describing index data according to an example embodiment.

[0077] Fig. 7A shows an example of first write data WRD1 stored in the first data block DB1 and first index data ID1 corresponding to the first write data WRD1, and Figure 7B An example of second write data WRD2 stored in the second data block DB2 and second index data ID2 corresponding to the second write data WRD2 is shown.

[0078] Reference Fig. 7A , the first write data WRD1 may include first to eighth column data D0 to D7 stored at first to eighth column addresses CA0 to CA7, respectively. Each of the first to eighth column data D0 to D7 may include first to eighth bits B0 to B7.

[0079] In the case of the first write data WRD1, the first column data D0 and the fifth column data D4 include at least one bit having a value of "1", and the values ​​of all bits of the other column data D1, D2, D3, D5, D6, and D7 are "0". Figure 5 As described above, logic gates 220 and index storage 230 may be used to generate Fig. 7A The first index data ID1 shown has a first index bit I0 and a fifth index bit I4 of "1" as values, and other bits I1, I2, I3, I5, I6 and I7 of "0" as values.

[0080] The column data D0 ˜ D7 may be stored at the column addresses CA0 ˜ CA7 of the first data block DB1 , and the first index data ID1 may be stored at the base column address CAb of the first data block DB1 .

[0081] Reference Figure 7B, the second write data WRD2 may include first to eighth column data D0 to D7 respectively stored at first to eighth column addresses CA0 to CA7. Each of the first to eighth column data D0 to D7 may include first to eighth bits B0 to B7.

[0082] In the case of the second write data WRD2, the second column data D1, the fourth column data D3 and the seventh column data D6 include at least one bit with a value of "1", and the values ​​of all bits of the other column data D0, D2, D4, D5 and D7 are "0". Therefore, as shown in FIG. Figure 5 As described above, logic gates 220 and index storage 230 may be used to generate Figure 7B The second index data ID2 shown. The values ​​of the second index bit I1, the fourth index bit I3 and the seventh index bit I6 of the second index data ID2 are "1", and the values ​​of the other bits I0, I2, I4, I5 and I7 of the second index data ID2 are "0".

[0083] The column data D0 ˜ D7 may be stored at the column addresses CA0 ˜ CA7 of the second data block DB2 , and the second index data ID2 may be stored at the base column address CAb of the second data block DB2 .

[0084] Although the example illustrates that the corresponding write data includes eight column data corresponding to eight column addresses and the corresponding column data includes eight bits, the number of column data in the corresponding write data and the number of bits in the corresponding column data may be determined in various ways.

[0085] The mapping relationship between column addresses CA0-CA7 and corresponding base column addresses CAb may be determined in various ways. For example, the base column address may have a value "k", and the first to eighth column addresses CA0 to CA7 may have sequentially increasing values ​​"k+1" to "k+8".

[0086] Figure 8 is a timing diagram illustrating an example write operation of a memory device according to example embodiments.

[0087] For ease of explanation, Figure 8 It is shown that the time points t1 ˜ t10 correspond to the rising edge of the clock signal CLK, but example embodiments are not limited thereto.

[0088] refer to Figure 8, the column address signal COL_ADDR may sequentially indicate the first column address CA0 to the eighth column address CA7, so the first column data D0 to the eighth column data D7 included in the corresponding write data WRDi may be stored at the first column address CA0 to the eighth column address CA7 of the corresponding data block DBi during the first time period TP1 to the eighth time period TP8. The column address signal COL_ADDR may indicate the base column address CAb during the ninth time period TP9 after the corresponding write data WRDi is stored, so the corresponding index data IDi may be stored at the base column address CAb of the corresponding data block DBi.

[0089] Fig. 9 is a block diagram illustrating a memory device according to example embodiments.

[0090] Fig. 9 For describing the computing operation, only the components used for the computing operation are shown, and Fig. 9 Other components are omitted. For ease of explanation, Fig. 9 A configuration corresponding to one memory bank is shown.

[0091] refer to Fig. 9 , the memory device 60 may include a plurality of data blocks DB1 -DBn, an input / output gating circuit 62 and a calculation block 300 . Fig. 9 The configuration of the first data block DB1 is shown as an example, and the other data blocks DB2_DBn may have the same configuration as the first data block DB1. Each data block may include a plurality of sub-memory cell arrays SARR, and each sub-memory cell array SARR may include a plurality of memory cells. In a computing operation, the internal data DW1~DWn read out from the data blocks DB1~DBn may be sequentially provided to the computing block 300 through the local input-output line LIO and the global input-output line GIO. The hierarchical structure of the data blocks may be implemented in various ways.

[0092] The input / output strobe circuit 62 can select the local input / output line corresponding to the column address of the internal data DW1 to DWn based on the column selection signal CSL. Figure 3 As described, the column selection signal CSL may be provided from the column decoder 470. The input-output gating circuit 62 may include a plurality of switch circuits MUX1MUXn corresponding to the plurality of data blocks DB1DBn, respectively.

[0093] The computing block 300 may include a plurality of computing units CU1-CUn corresponding to a plurality of data blocks DB1-DBn. As an example, Fig. 9It is shown that each computing unit is assigned to each data block, however, each computing unit may also be arranged in two or more data blocks. Each computing unit CU1-CUn may receive common broadcast data DA and corresponding internal data DW1-DWn read from each data block DB1-DBn. The computing unit CU1-CUn may perform calculations based on the broadcast data DA and the corresponding internal data DW1-DWn to provide calculation result data DR1-DRn, respectively.

[0094] As described below, the computing units CU1-CUn may independently generate a plurality of skip enable signals SEN1-SENn based on the corresponding index data ID1-IDn in the skip computing mode. The skip enable signals SEN1-SENn may be provided to the corresponding switch circuits MUX1-MUXn of the input-output gating circuit 62. When the corresponding skip enable signal is activated, each switch circuit MUX1-MUXn may output valid data from the corresponding data block, and when the corresponding skip enable signal is deactivated, each switch circuit MUX1-MUXn may block invalid data from the corresponding data block.

[0095] Fig.10 It is shown Fig. 9 FIG. 1 is a diagram of an example embodiment of a computing unit included in a computing circuit in a memory device of FIG.

[0096] refer to Fig.10 , the computing unit 310 may include a skip controller SKC 320 and a calculator MAC 330 .

[0097] The skip controller 320 may generate a corresponding skip enable signal SENi based on the corresponding index data IDi corresponding to the corresponding internal data DWi. Fig.13A and Fig. 13B As described above, the skip controller 320 may receive the corresponding index data IDi in response to activation of the index enable signal IEN.

[0098] The calculator 330 may perform calculation based on the broadcast data DA and the corresponding internal data DWi, and omit invalid data in the corresponding internal data DWi based on the corresponding skip enable signal SENi in the skip calculation mode.

[0099] As will be referenced below Fig.13A and Fig. 13BAs described, based on the corresponding index data IDi in the skip calculation mode, when the corresponding internal data DWi is read from the column address corresponding to the valid data, the skip controller 320 may activate the corresponding skip enable signal SENi, and when the corresponding internal data DWi is read from the column address corresponding to the invalid data, the skip controller 320 may deactivate the corresponding skip enable signal SENi. When the corresponding skip enable signal SENi is activated, the calculator 330 may be enabled to perform calculation based on the broadcast data DA and the valid data, and when the corresponding skip enable signal SENi is deactivated, the calculator 330 may be disabled.

[0100] As reference below Fig.14 As described above, the skip controller 320 may activate the corresponding skip enable signal SENi in the normal computing mode regardless of the corresponding index data IDi.

[0101] Fig.11 It is shown Fig.10 FIG. 1 is a diagram of an example arrangement of calculators included in a computing unit of FIG.

[0102] refer to Fig.11 , calculating MAC may include a first input terminal connected to a first node N1 receiving corresponding internal data DWi[N-1:0] and a second input terminal connected to a second node N2 receiving broadcast data DA[N-1:0]. The first node N1 is connected to an output terminal of an input-output sense amplifier IOSA, which amplifies a signal on the global input-output lines GIO and GIOB to output the amplified signal, and the second node N2 is connected to an input terminal of an input-output driver IODRV, which drives the global input-output lines GIO and GIOB.

[0103] During a normal read operation, the calculator MAC is disabled, and the input-output sense amplifier IOSA amplifies the read data provided through the global input-output lines GIO and GIOB to provide the amplified signal to the outside. During a normal write operation, the calculator MAC is disabled, and the input-output driver IODRV drives the global input-output lines GIO and GIOB based on the write data provided from the outside. During a calculation operation, the calculator MAC is enabled to receive the broadcast data DA[N-1:0] and the corresponding internal data DWi[N-1:0]. In this case, the input-output sense amplifier IOSA is enabled to output the corresponding internal data DWi[N-1:0] and the input-output driver IODRV is disabled to prevent the broadcast data DA[N-1:0] from being provided to the internal memory cell.

[0104] In some example embodiments, Fig.11As shown, the output terminal of the calculator MAC providing the corresponding calculation result data DRi can be connected to the first node N1 (i.e., the output terminal of the input-output sense amplifier IOSA). Therefore, the corresponding calculation result data DRi can be provided to the outside through a conventional read path. When the computing unit CU provides the corresponding calculation result data DRi, the input-output sense amplifier IOSA is disabled. In other example embodiments, the output terminal of the calculator MAC may not be connected to the first node N1, and the corresponding calculation result data DRi may be provided through an additional data path different from the conventional read path. In other example embodiments, the output node of the calculator MAC may be connected to the second node N2 to store the corresponding calculation result data DRi in the memory cell through a conventional write path.

[0105] For ease of explanation, Fig.11 Differential global line pairs GIO and GIOB are shown, and each counter MAC can be connected to N global line pairs to receive N-bit broadcast data DA[N-1:0] and N-bit corresponding internal data DWi[N-1:0]. For example, N can be 8, 16 or 21 depending on the operating mode of the stacked memory device.

[0106] Fig.12 It is shown Fig.10 FIG. 1 is a diagram of an example embodiment of a calculator included in a computing unit of FIG.

[0107] refer to Fig.12 , each calculator 500 may include a multiplication circuit 520 and an accumulation circuit 540. The multiplication circuit 520 may include buffers 521 and 522 and a multiplier 523, and the multiplier 523 is configured to multiply the broadcast data DA[N-1:0] with the corresponding internal data DWi[N-1:0]. The accumulation circuit 540 may include an adder 541 and a buffer 542 to accumulate the output of the multiplication circuit 520 to provide corresponding calculation result data DRi. The accumulation circuit 540 may be initialized in response to a reset signal RST, and output the corresponding calculation result data DRi in response to an output enable signal OUTEN. Using the Fig.11 The calculator 500 shown can efficiently perform matrix calculations, such as referring to Fig.16 Described.

[0108] The calculator 500 may be selectively enabled in response to a corresponding skip enable signal SENi. When the corresponding skip enable signal SENi is activated, the calculator 500 may be enabled to perform calculations based on the broadcast data DA and the corresponding internal data DWi[N-1:0] corresponding to the valid data, and when the corresponding skip enable signal SENi is deactivated, the calculator 500 may be disabled.

[0109] Fig.13A and Fig. 13B is a timing diagram illustrating an example operation of a skip calculation mode in a memory device according to an example embodiment, and Fig.14 is a timing diagram illustrating example operations of a normal computing mode in a memory device according to example embodiments.

[0110] For ease of explanation, Fig.13A , Fig. 13B and Fig.14 It is shown that time points t1 ˜ t10 correspond to rising edges of the clock signal CLK, however example embodiments are not limited thereto. For example, a logic high level H of the mode signal MD may represent a skip computing mode, and a logic low level L of the mode signal MD may represent a normal computing mode.

[0111] Fig.13A Shown based on Fig. 7A The jump calculation operation of the first internal data to the first data block DB1 is performed by the first write data WDR1, and Fig.13A Shown based on Figure 7B The jump calculation operation of the second internal data to the second data block DB2 which is the same as the second write data WDR2 is shown.

[0112] refer to Fig.13A , the column address signal COL_ADDR may sequentially indicate the base column address CAb and the first column address CA0 to the eighth column address CA7. The index enable signal IEN may be activated during the first period TP1, and Fig.10 The skip controller 320 in the embodiment may receive the first index data ID1 in response to activation of the index enable signal IEN. The skip controller 320 may be based on the following example: Fig. 7A The first index data ID1 shown generates a first skip enable signal SEN1, so that the first skip enable signal SEN1 can be activated during the second time period TP2 and the sixth time period TP6. Based on the first skip enable signal SEN1, only the first column data D0 and the fifth column data D4 can be output through the global input-output line GIO from the first data block DB1, and other column data D1, D2, D3, D5, D6 and D7 may not be output. In addition, the first calculation result data DR1 may have values ​​VL0, VL11 and VL12 updated at time points t2 and t6, as shown in FIG. Fig.13A This is because Fig.10 The calculator 330 in FIG. 1 performs calculations only for the first column data D0 and the fifth column data D4 .

[0113] refer to Fig. 13B, the column address signal COL_ADDR may sequentially indicate the base column address CAb and the first column address CA0 to the eighth column address CA7. The index enable signal IEN may be activated during the first period TP1, and Fig.10 The skip controller 320 in the embodiment may receive the second index data ID2 in response to activation of the index enable signal IEN. The skip controller 320 may be based on the following example: Figure 7B The second index data ID2 shown generates a second skip enable signal SEN2, so that the second skip enable signal SEN2 can be activated during the third time period TP3, the fifth time period TP5, and the eighth time period TP8. Based on the second skip enable signal SEN2, only the second column data D1, the fourth column data D3, and the seventh column data D6 can be output through the global input-output line GIO from the second data block DB2, and other column data D0, D2, D4, D5, and D7 may not be output. In addition, the second calculation result data DR2 may have the following characteristics: Fig. 13B The values ​​VL0, VL21, VL22 and VL23 are updated at the time points t3, t5 and t8 shown, because Fig.10 The calculator 330 in FIG. 1 performs calculations only on the second column data D1 , the fourth column data D3 , and the seventh column data D6 .

[0114] refer to Fig.14 , the column address signal COL_ADDR may sequentially represent the first column address CA0 to the eighth column address CA7. The index enable signal IEN may be deactivated in a logic low level L, and the corresponding jump enable signal SENi may be always activated in a logic high level H in a normal calculation mode. Based on the activated corresponding jump enable signal SENi, all the first column data D0 to the eighth column data D7 may be sequentially output from the corresponding data block DBi through the global input-output line GIO. The corresponding calculation result data DRi may have the following structure: Fig.14 The values ​​VL0 to VL7 are updated sequentially at the time points t2 to t8 shown in the figure. This is because Fig.10 The calculator 330 in FIG. 1 performs calculations on all the first column data D0 to the eighth column data D7.

[0115] Fig.15 is a diagram illustrating an example embodiment of outputting calculation result data.

[0116] Fig.15It is shown that the calculation result data corresponding to one channel CHANNEL-0 is output. One channel CHANNEL-0 may include a plurality of memory banks BANK0-BANK15, and each of the memory banks BANK0-BANK15 may include a plurality of calculation units CU0-CU15. The memory banks BANK0-BANK15 may be separated by two pseudo channels PSE-0 and PSE-1.

[0117] Each computing semiconductor die forming the computing unit may further include a plurality of memory bank adders 610a-610p. Each memory bank adder 610a-610p may sum the outputs of the computing units CU0-CU15 in each memory bank BANK0-BANK15 to generate each memory bank result signal BR0-BR15. The memory bank result signals BR0-BR15 may be output simultaneously through the data bus DBUS corresponding to each computing semiconductor die. For example, if the data bus corresponding to one computing semiconductor die has a data width of 128 bits and one channel CHANNEL-0 includes sixteen memory banks BANK0-BANK15, the output of each memory bank adder may be output through the data path of the 8-bit or one-byte data bus DBUS. In other words, the storage body result signal BR0 of the first storage body adder 610a can be output through the data path corresponding to the first byte BY0 of the data bus DBUS, the storage body result signal BR1 of the second storage body adder 610b can be output through the data path corresponding to the second byte BY1 of the data bus DBUS, and in this way, the storage body result signal BR15 of the sixteenth storage body adder 610p can be output through the data path corresponding to the sixteenth byte BY15 of the data bus DBUS.

[0118] Fig.16 is a diagram illustrating matrix calculation using a calculation circuit according to an example embodiment.

[0119] Fig.16 Matrix-vector multiplication performed using computing units CU0-0 to CU95-15 in a stacked memory device according to an example embodiment is shown. Fig.16, the computing units Cui-0 to Cui-15 of the i-th row (i=1 to 95) correspond to the i-th memory bank BANKi. For example, the matrix-vector multiplication may be in 32-bit mode, and each memory bank may include sixteen computing units. Assume that each of the four memory semiconductor dies includes two channels, and each channel includes sixteen memory banks. In this case, if one memory semiconductor die is used as the above-mentioned input-output semiconductor die and the other three memory semiconductor dies are used as the above-mentioned computing semiconductor dies, the number of memory banks included in the computing semiconductor die may be 96 (i.e., 6 channels*16 memory banks).

[0120] The first group of broadcast data DA0~DA15 and the second group of broadcast data DA16~DA31 during the first time period T1 are sequentially provided to all computing units in all storage bodies. In this way, the activations can be broadcast sequentially. In addition, the first group of internal data DW0~DW95 during the first time period T1 and the second group of internal data DW96~DW191 as weights are sequentially provided to the computing units. The internal data corresponds to the data read from the corresponding storage body. In this way, the computing unit can perform a dot product operation based on the activations and weights provided sequentially. The computing units in the same storage body provide partial sums of the same output activations. Therefore, after completing the dot product operation, it can be broadcast sequentially. Fig.16 The memory bank adders in the memory bank re-sum the partial sums to provide the final result as the memory bank result signals BR0~BR95.

[0121] like Fig.16 The matrix-vector multiplication shown can correspond to a 1*1 convolution or a fully connected layer. In the case of MLP and RNN, the broadcast data or broadcast activation corresponds to a subarray of the 1D input activation. In the case of CNN, the input activation corresponds to a 1*1 subcolumn of the input activation tensor.

[0122] Fig.17 is an exploded perspective view of a system including a stacked memory device according to example embodiments.

[0123] refer to Fig.17 , the system 800 includes a stacked memory device 1000 and a host device 2000 .

[0124] The stacked memory device 1000 may include a base semiconductor die or a logic semiconductor die 1010 and a plurality of memory semiconductor dies 1070 and 1080 stacked with the logic semiconductor die 1010 . Fig.17A non-limiting example of one logic semiconductor die and two memory semiconductor dies is shown. Two or more logic semiconductor dies and one, three or more memory semiconductor dies may be included in a stacked structure. In addition, Fig.17 A non-limiting example is shown in which memory semiconductor die 1070 and 1080 are stacked vertically with logic semiconductor die 1010. Fig.19 As described, the memory semiconductor dies 1070 and 1080 other than the logic semiconductor die 1010 may be vertically stacked, and the logic semiconductor die 1010 may be electrically connected to the memory semiconductor dies 1070 and 1080 through an interposer and / or a base substrate.

[0125] Logic semiconductor die 1010 may include memory interface MIF 1020 and logic for accessing memory integrated circuits 1071 and 1081 formed in memory semiconductor dies 1070 and 1080. Such logic may include control circuit CTRL 1030, global buffer GBF 1040, and data transformation logic DTL 1050.

[0126] The memory interface 1020 may perform communication with an external device such as the host device 2000 through the interconnect device 12. The control circuit 1030 may control the overall operation of the stacked memory device 1000. The data transformation logic 1050 may perform a logic operation on data exchanged with the memory semiconductor dies 1070 and 1080 or data exchanged through the memory interface 1020. For example, the data transformation logic may perform a maximum pooling, a rectified linear unit (ReLU) operation, a channel-by-channel addition, and the like.

[0127] The memory semiconductor dies 1070 and 1080 may include memory integrated circuits 1071 and 1081, respectively. At least one of the memory semiconductor dies 1070 and 1080 may be a computing semiconductor die 1080 including a computing circuit 100. As described below, the computing circuit 100 may include one or more computing blocks, and each computing block may include one or more computing units. Each computing unit may perform each calculation based on broadcast data and internal data to provide calculation result data. The broadcast data may be provided to the computing semiconductor die in public through silicon vias TSV, and the internal data may be read from the memory integrated circuit of the corresponding computing semiconductor die.

[0128] The host device 2000 may include a host interface HIF 2110 and processor cores CR1 2120 and CR2 2130. The host interface 2110 may perform communication with an external device such as the host device 1000 through the interconnect device 12.

[0129] Fig.18 is a diagram illustrating an example high bandwidth memory (HBM) organization.

[0130] refer to Fig.18 , HBM 1001 may include a stack of multiple DRAM semiconductor dies 1100, 1200, 1300, and 1400. The stacked HBM may be optimized through multiple independent interfaces called channels. Each DRAM stack may support up to 8 channels according to the HBM standard. Figure 3 An example stack including four DRAM semiconductor dies 1100, 1200, 1300, and 1400 is shown, and each DRAM semiconductor die supports two channels CHANNEL0 and CHANNEL1. Figure 4 As shown, the fourth memory semiconductor die 1400 may include two memory integrated circuits 1401 and 1402 corresponding to two channels.

[0131] For example, the fourth memory semiconductor die 1400 may correspond to a computing semiconductor die including a computing unit. Each memory integrated circuit 1401 and 1402 may include a plurality of memory banks MB, and each memory bank MB may include a computing block CB 300. As described above, the computing circuit 100 may include a plurality of computing blocks 300, and each computing block 300 may include a plurality of computing units CU. In this way, the computing units may be distributedly arranged in the memory banks 300 of the computing semiconductor die.

[0132] Each channel provides access to an independent set of DRAM memory banks. Requests from one channel may not be able to access data attached to other channels. Channels are independently timed and do not need to be synchronized. Each memory semiconductor die 1100, 1200, 1300, and 1400 of HBM 1001 can access another memory semiconductor die for transmitting broadcast data and / or calculation result data.

[0133] The HBM 1001 may also include an interface die 1010 or a logic semiconductor die disposed at the bottom of the stack structure to provide signal routing and other functions. Some functions of the DRAM semiconductor dies 1100 , 1200 , 1300 , and 1400 may be implemented in the interface die 1010 .

[0134] Fig.19 and Fig. 20 is a diagram illustrating a package structure of a stacked memory device according to example embodiments.

[0135] refer to Fig.19, the memory chip 2001 may include an interposer ITP and a stacked memory device stacked on the interposer ITP. The stacked memory device may include a logic semiconductor die LSD and a plurality of memory semiconductor dies MSD1 ˜ MSD4 .

[0136] refer to Fig. 20 , the memory chip 2002 may include a base substrate BSUB and a stacked memory device stacked on the base substrate BSUB. The stacked memory device may include a logic semiconductor die LSD and a plurality of memory semiconductor dies MSD1 ˜ MSD4 .

[0137] Fig.19 The following structure is shown, in which the memory semiconductor dies MSD1-MSD4 other than the logic semiconductor die LSD are vertically stacked and the logic semiconductor die LSD is electrically connected to the memory semiconductor dies MSD1-MSD4 through the interposer ITP or the base substrate. In contrast, Fig. 20 The structure in which the logic semiconductor die LSD and the memory semiconductor dies MSD1 - MSD4 are vertically stacked is shown.

[0138] As described above, at least one of the memory semiconductor dies MSD1-MSD4 may be a computing semiconductor die including a computing circuit CAL. The computing circuit CAL may include: a plurality of computing units that perform calculations based on the common broadcast data and corresponding internal data.

[0139] The base substrate BSUB may be the same as or include the interposer ITP. The base substrate BSUB may be a printed circuit board (PCB). External connection elements such as conductive bumps BMP may be formed on the lower surface of the base substrate BSUB, and internal connection elements such as conductive bumps may be formed on the upper surface of the base substrate BSUB. Fig.19 In the example embodiment of FIG. 5 , the logic semiconductor die LSD and the memory semiconductor dies MSD1 ˜ MSD4 may be electrically connected through through silicon vias. The stacked semiconductor dies LSD and MSD1 ˜ MSD4 may be packaged using a resin RSN.

[0140] Fig.21 is a block diagram illustrating a mobile system according to an example embodiment.

[0141] refer to Fig.21 , the mobile system 3000 includes an application processor 3100, a connection unit 3200, a volatile memory device (VM) 3300, a nonvolatile memory device (NVM) 3400, a user interface 3500, and a power supply 3600 connected via a bus.

[0142] The application processor 3100 may execute applications such as a web browser, a game application, a video player, etc. The connection unit 3200 may perform wired or wireless communication with an external device. The volatile memory device 3300 may store data processed by the application processor 3100, or may operate as a working memory. For example, the volatile memory device 3300 may be a DRAM, such as a double data rate synchronous dynamic random access memory (DDR SDRAM), a low power DDR (LPDDR) SDRAM, a graphic DDR (GDDR) SDRAM, a Rambus DRAM (RDRAM), etc. The non-volatile memory device 3400 may store a boot image and other data for booting the mobile system 3000. The user interface 3500 may include at least one input device (e.g., a keypad, a touch screen, etc.) and at least one output device (e.g., a speaker, a display device, etc.). The power supply 3600 may supply a power supply voltage to the mobile system 3000. In some embodiments, the mobile system 3000 may also include a camera image processor (CIP) and / or a storage device, such as a memory card, a solid state drive (SSD), a hard disk drive (HDD), a CD-ROM, and the like.

[0143] The volatile memory device 3300 and / or the nonvolatile memory device 3400 may be configured as described in reference to Figures 17 to 20 The stacked structure can include a plurality of memory semiconductor dies connected by through silicon vias, and the computing unit is formed in at least one memory semiconductor die.

[0144] As described above, a memory device and a method of operating a memory device according to example embodiments can reduce the amount of data exchanged between a stacked memory device, a logic semiconductor die, and an external device to reduce data processing time and power consumption by performing memory-intensive or data-intensive data processing in parallel through multiple computing units included in the memory semiconductor die.

[0145] Furthermore, a memory device and a method of operating the memory device according to example embodiments may reduce power consumption by omitting calculation and read operations regarding invalid data through a skip calculation mode based on index data.

[0146] The exemplary embodiments of the inventive concept can be applied to any device and system including a memory device. For example, the inventive concept can be applied to systems such as mobile phones, smart phones, personal digital assistants (PDAs), portable multimedia players (PMPs), digital cameras, video cameras, personal computers (PCs), server computers, workstations, laptop computers, digital televisions, set-top boxes, portable game consoles, navigation systems, etc.

[0147] According to one or more example embodiments, the above-mentioned units and / or devices include elements of the semiconductor memory device 30, such as the index data generator (IDG) 200 and the calculator circuit (CAL) 100 and sub-elements thereof, such as the generation unit 210 and the calculation unit 310, which can be implemented using hardware, a combination of hardware and software, or a non-transitory storage medium storing software executable to perform its functions, respectively.

[0148] The hardware may be implemented using processing circuitry such as, but not limited to, one or more processors, one or more central processing units (CPUs), one or more controllers, one or more arithmetic logic units (ALUs), one or more digital signal processors (DSPs), one or more microcomputers, one or more field programmable gate arrays (FPGAs), one or more systems on a chip (SoCs), one or more programmable logic units (PLUs), one or more microprocessors, one or more application specific integrated circuits (ASICs), or any other device or devices capable of responding to and executing instructions in a defined manner.

[0149] Software may include computer programs, program codes, instructions, or a combination thereof to independently or collectively instruct or configure hardware devices to operate as desired. Computer programs and / or program codes may include programs or computer-readable instructions, software components, software modules, data files, data structures, etc. that can be implemented by one or more hardware devices (e.g., one or more of the hardware devices mentioned above). Examples of program code include machine code generated by a compiler and higher-level program code executed using an interpreter.

[0150] For example, when the hardware device is a computer processing device (for example, one or more processors, CPUs, controllers, ALUs, DSPs, microcomputers, microprocessors, etc.), the computer processing device can be configured to execute program code by performing arithmetic, logic, and input / output operations according to the program code. Once the program code is loaded into the computer processing device, the computer processing device can be programmed to execute the program code, thereby transforming the computer processing device into a special-purpose computer processing device. In a more specific example, when the program code is loaded into the processor, the processor is programmed to execute the program code and the operation corresponding thereto, thereby transforming the processor into a special-purpose processor. In another example, the hardware device can be an integrated circuit customized in a special-purpose processing circuit (for example, ASIC).

[0151] A hardware device such as a computer processing device can run an operating system (OS) and one or more software applications running on the OS. In addition, the computer processing device can also access, store, manipulate, process, and create data in response to the execution of the software. For simplicity, one or more example embodiments may be illustrated as a computer processing device; however, those skilled in the art will recognize that the hardware device may include multiple processing elements and multiple types of processing elements. For example, the hardware device may include multiple processors or a processor and a controller. In addition, other processing configurations are also possible, such as parallel processors.

[0152] Software and / or data may be implemented permanently or temporarily in any type of storage medium (including but not limited to any machine, component, physical or virtual device, or computer storage medium or device) that can provide instructions or data to a hardware device or can be interpreted by a hardware device. Software may also be distributed on a network-coupled computer system so that the software is stored and executed in a distributed manner. Specifically, for example, software and data may be stored by one or more computer-readable recording media, including tangible or non-transitory computer-readable storage media as discussed herein.

[0153] According to one or more example embodiments, the storage medium may also include one or more storage devices at the unit and / or device. One or more storage devices may be tangible or non-temporary computer-readable storage media, such as random access memory (RAM), read-only memory (ROM), permanent mass storage device (such as disk drive) and / or any other similar data storage mechanism capable of storing and recording data. One or more storage devices may be configured to store computer programs, program codes, instructions or some combinations thereof for one or more operating systems and / or for implementing the example embodiments described herein. A drive mechanism may also be used to load a computer program, program code, instruction or some combination thereof from a separate computer-readable storage medium into one or more storage devices and / or one or more computer processing devices. This separate computer-readable storage medium may include a universal serial bus (USB) flash drive, a memory stick, a Blu-ray / DVD / CD-ROM drive, a memory card and / or other similar computer-readable storage media. A computer program, program code, instruction or some combination thereof may be loaded into one or more storage devices and / or one or more computer processing devices from a remote data storage device via a network interface rather than via a computer-readable storage medium. In addition, a computer program, program code, instruction, or some combination thereof may be loaded into one or more storage devices and / or one or more processors from a remote computing system configured to transmit and / or distribute the computer program, program code, instruction, or some combination thereof via a network. The remote computing system may transmit and / or distribute the computer program, program code, instruction, or some combination thereof via a wired interface, an air interface, and / or any other similar medium.

[0154] One or more hardware devices, storage media, computer programs, program codes, instructions, or some combination thereof may be specially designed and constructed for the purposes of the example embodiments, or they may be known devices that have been altered and / or modified for the purposes of the example embodiments.

[0155] The foregoing is illustrative of exemplary embodiments and should not be construed as limiting thereof. Although a few exemplary embodiments have been described, those skilled in the art will readily appreciate that various modifications may be made in the exemplary embodiments without materially departing from the inventive concept.

Claims

1. A memory device, comprising: a memory cell array associated with the semiconductor die, the memory cell array comprising a plurality of memory cells arranged in a plurality of columns, the plurality of memory cells configured to store internal data; as well as a processing circuit associated with the semiconductor die, the processing circuit being configured to: receiving broadcast data broadcast to the semiconductor die from outside the semiconductor die, reading the internal data from the memory cell array, reading index data including a plurality of index bits, each index bit of the plurality of index bits being associated with a corresponding column of the plurality of columns such that each index bit of the plurality of index bits indicates whether the internal data read from the corresponding column of the plurality of columns is invalid data or valid data, and In the skip calculation mode, calculation is selectively performed using the broadcast data and the internal data as inputs to generate a calculation result based on whether the internal data for a corresponding column among the plurality of columns is invalid data or valid data indicated by the index data.

2. The memory device according to claim 1, wherein: In the skip calculation mode, the memory device is configured to output the valid data from the memory cell array and not output the invalid data from the memory cell array.

3. The memory device according to claim 1, wherein: In a normal computation mode, the processing circuit is configured to perform computations on both the valid data and the invalid data without regard to the index data.

4. The memory device according to claim 1, wherein: When the values ​​of all bits of the internal data read from the corresponding column address among the plurality of column addresses are 0, the corresponding index bit among the plurality of index bits has a first value indicating that the internal data is the invalid data, and When a value of at least one bit of the internal data read from a corresponding column among a plurality of column addresses is 1, a corresponding index bit among the plurality of index bits has a second value indicating that the internal data is the valid data.

5. The memory device according to claim 1, wherein: The memory cell array includes a plurality of data blocks, and The broadcast data is common broadcast data, and the internal data is read from a corresponding data block among the plurality of data blocks.

6. The memory device according to claim 5, wherein: The processing circuit is configured to: generating a corresponding jump enable signal based on the index data corresponding to the internal data, and Calculation is performed based on the broadcast data and the internal data so that the invalid data in the internal data is omitted based on a corresponding skip enable signal in the skip calculation mode.

7. The memory device according to claim 6, wherein: In the skip calculation mode, based on the index data, the processing circuit is configured to: In response to the internal data corresponding to the valid data, activating the corresponding skip enable signal, and In response to the internal data corresponding to the invalid data, the corresponding skip enable signal is deactivated.

8. The memory device according to claim 7, wherein: The processing circuit is configured to: In response to the corresponding skip enable signal being activated, performing a calculation based on the broadcast data and the valid data, and In response to the corresponding skip enable signal being deactivated, the computation is omitted.

9. The memory device according to claim 7, further comprising: An input-output gating circuit is configured to gating data input to the memory cell array and data output from the memory cell array, so that the input-output gating circuit: In response to the corresponding skip enable signal being activated, outputting valid data from a corresponding data block among the plurality of data blocks, and In response to the corresponding skip enable signal being deactivated, invalid data from a corresponding data block of the plurality of data blocks is blocked.

10. The memory device according to claim 6, wherein: The processing circuit is configured to read corresponding bits of the index data from a base column address, the base column address being mapped to a column address of corresponding internal data.

11. The memory device according to claim 6, wherein: The processing circuit is configured to, in a normal computing mode, activate the corresponding skip enable signal regardless of the index data.

12. The memory device according to claim 6, wherein: The processing circuit is configured to independently generate a plurality of skip enable signals based on corresponding bits of the index data in the skip calculation mode.

13. The memory device according to claim 1, wherein: The processing circuit is configured to generate the index data based on write data stored in the memory cell array.

14. The memory device according to claim 13, wherein: The memory cell array includes a plurality of data blocks, and The processing circuit is configured to generate corresponding index data based on corresponding ones of the write data stored in corresponding ones of the plurality of data blocks.

15. The memory device according to claim 14, wherein: The processing circuit is configured to: performing a logic operation on corresponding ones of the write data to generate an output signal; and Based on the output signal, the corresponding index data is stored.

16. The memory device according to claim 15, wherein: The processing circuit is configured to: The logic operation is sequentially performed for respective write data corresponding to a plurality of column addresses, and A plurality of index bits of the corresponding index data are sequentially stored based on the output signal.

17. The memory device according to claim 13, wherein: During a write operation of storing the write data in the memory cell array, the index data is stored in the memory cell array.

18. A memory device comprising: A plurality of memory semiconductor dies are stacked in a vertical direction; Through silicon vias electrically connecting the plurality of memory semiconductor dies; a plurality of memory integrated circuits associated with respective ones of the plurality of memory semiconductor dies, the plurality of memory integrated circuits comprising a plurality of memory cells arranged in a plurality of columns, the plurality of memory cells configured to store internal data; as well as processing circuitry associated with one or more computing semiconductor dies of the plurality of memory semiconductor dies, the processing circuitry being configured to: receiving broadcast data from outside the computing semiconductor die, the broadcast data being commonly broadcast to the computing semiconductor die through the through silicon via, reading the internal data from the plurality of memory cells, reading index data including a plurality of index bits, each index bit of the plurality of index bits being associated with a corresponding column of the plurality of columns such that each index bit of the plurality of index bits indicates whether the internal data read from the corresponding column of the plurality of columns is invalid data or valid data, and In the skip calculation mode, calculation is selectively performed using the broadcast data and the internal data as inputs to generate a calculation result based on whether the internal data for a corresponding column among the plurality of columns is invalid data or valid data indicated by the index data.

19. A method of operating a memory device, the memory device comprising a semiconductor die and processing circuitry associated with the semiconductor die, the semiconductor die having an array of memory cells, the method comprising: receiving broadcast data broadcast to the semiconductor die from outside the semiconductor die; reading internal data from the memory cell array; receiving index data including a plurality of index bits in a skip calculation mode, each index bit of the plurality of index bits being associated with a corresponding column of the plurality of columns such that each index bit of the plurality of index bits indicates whether the internal data read from the corresponding column of the plurality of columns is valid data or invalid data; as well as In the skip calculation mode, calculation is selectively performed by the processing circuit using the broadcast data and the internal data as input to generate a calculation result based on the index data indicating whether the internal data for a corresponding column of the plurality of columns is invalid data or valid data.

Citation Information

Patent Citations

  • A table for photographing radiation images

    KR1020180020422A

  • Method and apparatus for filtering snoop requests using multiple snoop caches

    US20060224839A1

  • Polymorphic Stacked DRAM Memory Architecture

    US20120221785A1