Memory device, operating method thereof, and electronic device

By using the address generator to generate the target address in the memory (PIM) architecture, the data transmission bottleneck problem between the memory and the central processing unit is solved, the performance and operation efficiency of the memory device are improved, and the parallelism and consistency of the in-memory operations are achieved.

CN120234046APending Publication Date: 2025-07-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411870356.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the data transmission bottleneck between the memory and the central processing unit occupies most of the delay in system performance in applications with high memory usage, and the memory address calculated by the host does not match the actual operation sequence, resulting in performance degradation.

Method used

Using an in-memory processing (PIM) architecture, the target address is generated within the memory device through an address generator, and the offset and base address are added sequentially by counters and adders, reducing the need for host computing and realizing operational parallelism in memory.

Benefits of technology

The performance utilization of the memory device is improved, address alignment problems are avoided, memory operation efficiency and consistency are enhanced, and data transmission delay between the host and the memory is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234046A_ABST
    Figure CN120234046A_ABST
Patent Text Reader

Abstract

A memory device, an operating method thereof and an electronic device are provided. The memory device includes a memory array, an address generator, a data register, and a processing unit. The address generator is configured to receive an instruction and a base address of the instruction from a host, and sequentially generate a target address for performing an operation of the instruction by sequentially adding an offset to the base address. The data register is configured to store data values corresponding to one or more of the target addresses. The processing unit is configured to perform one or more of the operations of the instruction based on the data value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application claims priority to Korean Patent Application No. 10-2023-0196766, filed with the Korean Intellectual Property Office on December 29, 2023, the entire disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] The following embodiments relate to a memory device having an address generator and an operation method thereof. Background Art

[0003] A vector matrix multiplication operation (also referred to as a multiply and accumulate (MAC) operation) can be used in various applications. For example, a MAC operation can be performed during machine learning and used to authenticate a neural network including multiple layers. An input signal for an image, a byte stream, or other data set can be used to generate an input vector to be applied to the neural network. The input vector can be multiplied by weights, and an output vector can be obtained based on the result of one or more MAC operations performed by a layer of the neural network on the weighted input vector. The output vector can be provided as an input vector to a subsequent layer of the neural network. Since the MAC operation can be repeatedly used in multiple layers of the neural network, the processing performance of the neural network can be mainly determined by the performance of the MAC operation.

[0004] Processing in memory (PIM) represents a specific type of architecture in which processing elements are placed closer to the memory or integrated within the memory itself. This is intended to reduce the bottleneck caused by data movement between the central processing unit and the memory.

[0005] Since the processing performance of a neural network highly depends on the MAC operation, it can be feasible to improve the performance if the MAC operation can be implemented using PIM. Summary of the Invention

[0006] According to an embodiment, a memory device includes a memory array, an address generator, a data register, and a processing unit. The address generator is configured to receive an instruction and a base address of the instruction from a host, and sequentially generate target addresses for performing an operation of the instruction by sequentially adding an offset to the base address. The data register is configured to store data values corresponding to one or more of the target addresses. The processing unit is configured to perform one or more of the operations of the instruction based on the data values.

[0007] According to an embodiment, an electronic device includes a host, an address generator, a data register, and a processing unit. The address generator is configured to receive an instruction and a base address of the instruction from the host, and sequentially generate target addresses for performing operations of the instruction by sequentially adding an offset to the base address. The data register is configured to store data values corresponding to one or more of the target addresses. The processing unit is configured to perform one or more of the operations of the instruction based on the data values.

[0008] According to an embodiment, a method of operating a memory device includes: receiving an instruction and a base address of the instruction from a host; sequentially generating target addresses for performing operations of the instruction by sequentially adding an offset to the base address; storing data values corresponding to one or more of the target addresses; and performing one or more of the operations of the instruction based on the data values.

[0009] According to an embodiment, a memory device includes a processing unit and an address generator. The address generator includes a first counter, a second counter, an adder, and a selector. The selector is configured to receive an instruction and a base address, increment a first count value of the first counter and provide the first count value as an offset to the adder when the instruction is for storing data in the memory device, and is configured to increment a second count value of the second counter and provide the second count value as the offset to the adder when the instruction is for the processing unit to perform an operation on data in the memory device. The adder is configured to generate a target memory address for accessing the memory device by adding the offset to the base address. The address generator may further include a third counter, wherein when the data is to be stored in a first region of the memory device, the selector increments the first count value, and when the data to be stored in a second region of the memory device different from the first region, the selector increments a third count value of the third counter and provides the third count value as an offset to the adder. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] These and / or other aspects and features of the inventive concept will become apparent and more readily understood from the following description of embodiments in conjunction with the accompanying drawings.

[0011] Figure 1 is a diagram exemplarily showing a configuration of a memory device according to an embodiment.

[0012] Figure 2A is a diagram exemplarily showing a configuration of an address generator of a memory device according to an embodiment.

[0013] Figure 2B is a diagram exemplarily showing a configuration of a counter block of an address generator according to an embodiment.

[0014] Figure 3 is a diagram exemplarily showing corresponding changes in processing times of a host and a memory device caused by address generation of the memory device according to an embodiment.

[0015] Figures 4A to 4C is a diagram exemplarily showing an instruction processing procedure of a memory device according to an embodiment.

[0016] Figure 5 is a diagram exemplarily showing a mapping relationship among weights, input features, and output features according to an embodiment.

[0017] Figure 6A , Figure 6C , Figure 6E and Figure 6G is a diagram exemplarily showing a process of processing a multiply-accumulate (MAC) operation according to an embodiment.

[0018] Figure 6B , Figure 6D , Figure 6F and Figure 6H is a diagram exemplarily showing details of weights, input features, and output features used in the processes of Figure 6A , Figure 6C , Figure 6E and Figure 6G according to an embodiment.

[0019] Figure 7 is a diagram exemplarily showing a configuration of an address generator when a memory bank shares an in-memory processing (PIM) structure according to an embodiment.

[0020] Figure 8 is a flowchart exemplarily showing an operation method of a memory device according to an embodiment.

[0021] Figure 9 is a diagram exemplarily showing a configuration of an electronic device according to an embodiment. DETAILED DESCRIPTION

[0022] Embodiments will now be described more fully hereinafter with reference to the accompanying drawings. Throughout the disclosure, like reference numerals may indicate like components. Here, the embodiments are not to be construed as being limited to the disclosure and should be understood to include all changes, equivalents, and substitutions within the spirit and technical scope of the disclosure.

[0023] It should be noted that if one component is described as being "connected", "coupled", or "joined" to another component, then although the first component may be directly connected, coupled, or joined to the second component, a third component may be "connected", "coupled", or "joined" between the first component and the second component.

[0024] As used herein, unless the context clearly indicates otherwise, the singular forms also include the plural forms.

[0025] As used herein, each of "at least one of A and B", "at least one of A, B, or C", etc. may include any one of the items listed together in the corresponding one of the phrases or all possible combinations of the items.

[0026] Figure 1 is a diagram exemplarily showing the configuration of a memory device according to an embodiment. Referring to Figure 1 , the memory device 100 may include a memory array 110, a row decoder 111 (e.g., a first decoder circuit), a column decoder 112 (e.g., a second decoder circuit), an address generator 130, a data register 140, a processing unit 150 (e.g., a processor), and an instruction buffer 160.

[0027] The memory array 110 may store data. A memory address may be required to access the memory array 110. The row decoder 111 and the column decoder 112 may be used to decode the memory address. The memory address may include row information and column information. The row information may be decoded by the row decoder 111, and the column information may be decoded by the column decoder 112.

[0028] The processing unit 150 may be a processing-in-memory function processing unit (PIM FPU). The processing unit 150 may perform operations. For example, the operations may include multiply-accumulate (MAC) operations. The processing unit 150 may include operation logic for performing operations. The operation logic may temporarily store data for the operations, perform operations using the data for the operations, and generate operation results. The operation logic may correspond to hardware logic (e.g., a logic circuit). The data register 140 may provide a memory space for temporarily storing data for the operations of the processing unit 150. The processing unit 150 may use the data register 140 to perform operations to generate final operation results and store the final operation results in the memory array 110. The host may store host data 103 in the memory array 110 and the data register 140. The host may directly store the host data 103 in the memory array 110 or directly store the host data 103 in the data register 140.

[0029] The memory device 100 may have a PIM architecture including a processing unit 150. The PIM architecture may represent the architecture or operation of a memory with computing capabilities. However, the embodiments are not limited thereto since other architectures (such as, near-memory processing (NMP) and in-memory processing) may be used without being PIM. In a specific system, a bottleneck may occur between the host and the memory. In particular, in memory-intensive applications with high memory usage, data transfer between the host and the memory may account for most of the latency in the overall system performance. The memory device 100 may perform operations internally using the PIM architecture. For example, in the PIM architecture, an in-memory acceleration method based on bank-level parallelism may be provided.

[0030] When the host is computing the memory address for the operation of the processing unit 150, the processing unit 150 may not perform the operation, which may reduce the utilization rate of the PIM architecture and degrade the performance. Additionally, when additional elements (such as, an address alignment mode or a column alignment mode) are used to prevent the memory address calculated by the host from being transmitted to the memory device 100 in a different order than the actual operation order, the additional elements may cause performance degradation. For example, the host may correspond to a central processing unit (CPU) or a graphics processing unit (GPU), and the memory device 100 may correspond to a dynamic random access memory (DRAM), but is not limited thereto.

[0031] According to an embodiment, the target address for performing the operation of the instruction 101 may be generated by the address generator 130 of the memory device 100. When the target address is internally generated in the memory device 100 by the address generator 130 due to receiving the base address 102 from the host, the utilization rate of the PIM architecture may increase, and the performance may increase. Additionally, when the target address is internally generated in the memory device 100, no alignment problem should occur.

[0032] More specifically, the controller 120 may receive the instruction 101 from the host. The address generator 130 may receive the base address 102 of the instruction 101 from the host and sequentially generate the target address for performing the operation of the instruction 101 by sequentially adding the offset to the base address 102. The offset may have a predetermined interval. The address generator 130 may sequentially add the offset to the base address 102 using the offset having the predetermined interval for the base address 102. The predetermined interval may be fixed, and the offset may have the same interval. The memory address of the memory array 110 may be specified using the base address 102 and the offset. The memory address specified using the base address 102 and the offset may be referred to as the target address. The data register 140 may store data values (such as, input elements, weight elements, or output elements) corresponding to one or more of the target addresses. The processing unit 150 may perform one or more of the operations of the instruction 101 based on the data values. The base address 102 may include a base row address and a base column address.

[0033] Instruction 101 can be a direct instruction for directly controlling processing unit 150 or an indirect instruction for indirectly controlling processing unit 150. When a direct instruction is received, controller 120 can control processing unit 150 to execute the operation of instruction 101. Instruction buffer 160 can store auxiliary instructions related to the indirect instruction. When an indirect instruction is received, controller 120 can control processing unit 150 based on the auxiliary instructions related to the indirect instruction located in instruction buffer 160. Controller 120 can load the auxiliary instructions related to the indirect instruction from instruction buffer 160, and control processing unit 150 to execute the operation of the auxiliary instructions.

[0034] For example, instruction 101 can include a store instruction for storing data or auxiliary instructions in memory array 110, data register 140, processing unit 150, or instruction buffer 160, a load instruction for loading data or auxiliary instructions from memory array 110, data register 140, processing unit 150, or instruction buffer 160, an operation instruction for performing an operation using the data (e.g., a MAC operation instruction), etc. For example, auxiliary instructions can be used to configure the operation instruction. However, the type of instruction 101 or the implementation manner of the operation instruction is not limited thereto. The operation logic of processing unit 150 can include a memory space for temporarily storing data used for the operation, and the memory space can be used for loading, storing, and operating through instruction 101.

[0035] Figure 2A is a diagram exemplarily showing the configuration of an address generator of a memory device according to an embodiment. Refer to Figure 2A , address generator 200 can include a counter block 210 (e.g., a logic circuit) for generating an offset, a counter selector 220 (e.g., a selector circuit) for controlling counter block 210, and an address adder 230 for generating a target memory address 203 by adding the offset to the base address 202 of instruction 201. Address generator 200 can be used to implement Figure 1 the address generator 130. The base address 202 can correspond to Figure 1 the base address 102. Instruction 201 can correspond to Figure 1 the instruction 101. In one embodiment, the offset is the count value of the counter provided to address adder 230.

[0036] The counter block 210 may include counters (e.g., control counter 217, source counter 218, and destination counter 219). The control counter 217, source counter 218, and destination counter 219 may be selectively used based on the instruction 201 and / or the base address 202. The counter selector 220 may control the counter block 210 based on the instruction 201 and / or the base address 202. The counter selector 220 may select one of the control counter 217, source counter 218, and destination counter 219 of the counter block 210 based on the type of the instruction 201 and the location indicated by the base address 202 (e.g., memory array or register). For example, the counter selector 220 may control the control counter 217 to increment the count value of the control counter 217 when storing data (e.g., input feature) in the data register, control the source counter 218 to increment the count value of the source counter 218 when the processing unit performs an operation on data (e.g., weight) in the memory array and / or data in the data register (e.g., input feature), and control the destination counter 219 to increment the count value of the destination counter 219 when storing data (e.g., output feature) in the memory array. In one embodiment, the control counter 217 enables the memory device 100 to store a large amount of data in multiple locations within the data register by receiving only a single address from the host without receiving multiple addresses or one or more offsets from the host. In one embodiment, the source counter 218 enables the memory device 100 to perform operations on data in multiple locations within the memory array or data register without receiving multiple addresses or one or more offsets from the host. In one embodiment, the destination counter 219 enables the memory device 100 to store data in multiple locations within the memory array without receiving multiple addresses or one or more offsets from the host. The names and quantities of the counters (such as the control counter 217, source counter 218, and destination counter 219, etc.) are examples and are not limited thereto. For example, if there are two control counters, the first control counter may be used to store a large amount of data in multiple locations within the first region of the data register, and the second control counter may be used to store a large amount of data in multiple locations within the second other region of the data register. For example, if there are two source counters, the first source counter may be used to perform operations on data in multiple locations within the first region of the memory array or data register, and the second source counter may be used to perform operations on data in multiple locations within the second other region of the memory array or data register. For example, if there are two destination counters, the first destination counter may be used to store data in multiple locations within the first region of the memory array, and the second destination counter may be used to store data in multiple locations within the second other region of the memory array.The control counter 217, source counter 218, and destination counter 219 may each include a column counter and a row counter. The count value of the column counter may be referred to as the column count value, and the count value of the row counter may be referred to as the row count value. The column count value may correspond to the column offset of the target memory address 203, and the row count value may correspond to the row offset of the target memory address 203.

[0037] The control counter 217, source counter 218, and destination counter 219 may sequentially increment the count value based on the control of the counter selector 220. A counter selection signal may be used to control the counter selector 220. The control counter 217, source counter 218, and destination counter 219 may each sequentially increment one of the column counter value and the row counter value (e.g., the column counter value) to the maximum value by controlling one of the column counter and the row counter (e.g., the column counter), and when one of the column counter value and the row counter value (e.g., the column counter value) reaches the maximum value, increment the other of the column counter value and the row counter value (e.g., the row counter value). Which of the column counter value and the row counter value to increment first may be determined based on the address configuration of the memory array. If the address configuration of the memory array has a way of incrementing the column address first, the column counter value may be incremented first. If the address configuration of the memory array increments the row address first, the row counter value may be incremented first.

[0038] The address generator 200 may further include a multiplexer block 240. The multiplexer block 240 may include a column multiplexer 241 and a row multiplexer 242. The column multiplexer 241 is used to select the output of one of the column counters among the control counter 217, source counter 218, and destination counter 219 based on the control of the counter selector 220, and the row multiplexer 242 is used to select the output of one of the row counters among the control counter 217, source counter 218, and destination counter 219 based on the control of the counter selector 220. A counter selection signal may be used to control the counter selector 220.

[0039] The column multiplexer 241 and the row multiplexer 242 may output one or more of the target memory address 203, the first target register index 204, and the second target register index 205 to generate the target address 209 based on the control of the counter selector 220 based on the instruction 201, the base address 202, or a combination thereof. The data register of the memory device may include a first register bank and a second register bank. The first target register index 204 may be used to specify a register in the first register bank, and the second target register index 205 may be used to specify a register in the second register bank. However, the configuration of the data register and the configuration of the target register index 204 and the target register index 205 are not limited thereto.

[0040] The access to the memory array and the access to the registers can be synchronized based on the number of registers. The first target register index 204 and the second target register index 205 can be generated by extracting the number of bits that identify the registers in each register group of the data registers from the least significant bits (LSBs) of the output corresponding to the offset from the counter block 210. For example, if the number of registers is "4", then "4" registers can be identified with "2" bits, and thus, "2" bits from the LSBs can be used as the target register index 204 and the target register index 205.

[0041] As described above, the address generator 200 can generate the target address 209 based on the instruction 201 and the base address 202. The target address can be sequentially generated as an offset corresponding to the output of the counter block 210 according to the increase of the count value sequentially added to the base address 202. The offset and the target address can have a predetermined interval corresponding to the interval of adjacent values of the column count value. For example, if the interval of the column count value is "4", then the offset and the target address can have an interval of "4". Since the target address is generated on the memory device side if the instruction 201 and the base address 202 are given, the PIM operation of the processing unit using the memory device can be executed without the host calculating the target address of the base address 202.

[0042] Figure 2B is a diagram exemplarily showing the configuration of the counter block of the address generator according to an embodiment. Refer to Figure 2B FIG. 9, the counter block 210a of the address generator 200a may include a control column counter 211, a control row counter 212, a source column counter 213, a source row counter 214, a destination column counter 215, and a destination row counter 216. The address generator 200a can be used to implement Figure 1 the address generator 130 of FIG. 13. The control column counter 211 and the control row counter 212 may correspond to Figure 2A the column counter and the row counter of the control counter 217 of FIG. 12. The source column counter 213 and the source row counter 214 may correspond to Figure 2A the column counter and the row counter of the source counter 218 of FIG. 14. The destination column counter 215 and the destination row counter 216 may correspond to Figure 2A the column counter and the row counter of the destination counter 219 of FIG. 15.

[0043] The counter selector 220a may control the counter block 210a based on the instruction 201 and / or the base address 202. The counter block 210a may include a column counter group and a row counter group. The column counter group includes a control column counter 211, a source column counter 213, and a destination column counter 215. The row counter group includes a control row counter 212, a source row counter 214, and a destination row counter 216. The counter selector 220a may first control one of the column counter group and the row counter group. Figure 2B This corresponds to an example of first controlling the column counter group, but is not limited thereto. In Figure 2B this example, the counter selector 220a may receive only the base row address of the base address 202. The counter selector 220a may select one of the control column counter 211, the control row counter 212, the source column counter 213, the source row counter 214, the destination column counter 215, and the destination row counter 216 of the counter block 210 based on the type of the instruction 201 and the position indicated by the base address 202 (e.g., a memory array or a register).

[0044] For example, when storing data (e.g., input features) in a data register, the counter selector 220a may control the control column counter 211 to increase the count value of the control column counter 211. When using a processing unit to perform an operation on data (e.g., weights) in a memory array and / or data (e.g., input features) in a data register, the counter selector 220a may control the source column counter 213 to increase the count value of the source column counter 213. And when storing data (e.g., output features) in a memory array, the counter selector 220a may control the destination column counter 215 to increase the count value of the destination column counter 215.

[0045] The control column counter 211, the source column counter 213, and the destination column counter 215 of the column counter group may sequentially increase the column count value based on the control of the counter selector 220a. A counter selection signal may be used to control the counter selector 220a. When the column count value increases to the maximum value, the control row counter 212, the source row counter 214, and the destination row counter 216 of the row counter group may increase the row count value. When the column count value increases to the maximum value, the control column counter 211, the source column counter 213, and the destination column counter 215 may initialize the column count value. For example, the initialization of the column count value may be performed by setting the column count value to zero.

[0046] The column count value and the row count value may be increased at a predetermined interval. For example, the column count value may correspond to the size of a sub-tile of a tile, and the row count value may correspond to a value obtained by multiplying the size of the sub-tile by the number of sub-tiles belonging to a single row of the memory array. For example, the predetermined interval of the column count value may correspond to the size, and the predetermined interval of the row count value may correspond to the value.

[0047] Figure 3 is a diagram exemplarily showing corresponding changes in the processing times of a host and a memory device caused by address generation of a memory device according to an embodiment. Refer to Figure 3 , an existing host 310 performs a process 311 for overall generation of a memory address. Since a memory address is required to perform a data operation of an existing memory device 320, a delay related to the process 311 of the existing host 310 may occur in a process 321 for the data operation of the existing memory device 320. A host 330 according to an embodiment may perform a process 331 for generating a base address among memory addresses. A memory device 340 according to an embodiment generates a memory address by itself based on the base address, and thus may perform a process 341 for a data operation in a state where a delay related to an additional operation of generating a memory address by the host 330 is eliminated.

[0048] Figures 4A to 4C is a diagram exemplarily showing an instruction processing procedure of a memory device according to an embodiment. Refer to Figures 4A to 4C , a bank 400 may include a memory array 410, an address generator 430, a data register 440, and a processing unit 450.

[0049] Refer to Figure 4A , a counter selector 434 of the address generator 430 may receive a first instruction 402a and a first base address 401a of the first instruction 402a from a host 490, and control a first counter group 431 of a counter block to generate an offset of the first base address 401a. The first instruction 402a is for storing a sub-block of an input feature map block of an input feature in a first register group A of the data register 440, and the first base address 401a of the first instruction 402a indicates a starting index of the first register group A. The first instruction 402a may be a store instruction. The first counter group 431 may include a control column counter and a control row counter. The counter selector 434 controls the control column counter.

[0050] When a first target address (e.g., a target register index) of the first register group A is generated by adding an offset of the first base address 401a to the first base address 401a by an address adder of the address generator 430, the first register group A may store a sub-block of the first input feature map block in the first register group A based on the first target address.

[0051] Sub - blocks of the first input feature map block can be stored in the first register bank A through the first instruction 402a without the host 490 performing an operation for generating an offset from the first base address. For example, the host 490 can provide the first instruction 402a to the address generator 430 N times without specifying an offset, and in this process, sub - blocks of the first input feature map block can be stored N times at an interval of M - 1 from the first base address 401a. Here, M can represent the interval of the count value.

[0052] Refer to Figure 4B , the counter selector 434 can receive the second instruction 402b and the second base address 401b of the second instruction 402b, and control the second counter group 432 of the counter block to generate an offset of the second base address 401b. The second instruction 402b is used to store sub - blocks of the first weight map block of the weights in the data space of the processing unit 450. The second base address 401b of the second instruction 402b indicates the starting address of storing weights in the memory array 410. The second instruction 402b can be an operation instruction (e.g., a MAC operation instruction) or a load instruction for loading auxiliary instructions for the operation from the instruction buffer. The second counter group 432 can include a source column counter and a source row counter. The counter selector 434 can control the source column counter.

[0053] When generating the second target address (e.g., the target memory address) of the memory array 410 by adding the offset of the second base address 401b to the second base address 401b by the address adder of the address generator 430, the processing unit 450 can sequentially perform an operation on the sub - block of the first input feature map block loaded into the data space of the processing unit 450 based on the first target address and the sub - block of the first weight map block loaded into the data space of the processing unit 450 based on the second target address to generate an operation result, and generate a sub - block of the first output feature map block by accumulating the operation results. The second register bank B can be used to accumulate the operation results.

[0054] The operation on the sub - block of the first weight map block and the sub - block of the first input feature map block can be executed through the second instruction 402b without the host 490 performing an operation for generating an offset of the second base address 401b. For example, the host 490 can provide the second instruction 402b to the address generator 430 N times without specifying an offset, and in this process, the operation on the sub - block of the first weight map block and the sub - block of the first input feature map block can be executed N times. At this time, the sub - blocks of the first weight map block can be loaded N times at an interval of M - 1 from the second base address 401b. The sub - blocks of the first input feature map block can be loaded N times from the first register bank A using the lower bits of the target memory address of the sub - blocks of the first weight map block.

[0055] Refer toFigure 4C The counter selector 434 can receive a third instruction 402c and a third base address 401c of the third instruction 402c, and control the third counter group 433 of the counter block to generate an offset of the third base address 401c. The third instruction 402c is used to store sub-blocks of the first output feature map block stored in the second register group B of the data register 440 in the memory array 410. The third base address 401c of the third instruction 402c indicates the starting address of the memory array 410 for storing the sub-blocks of the first output feature map block. The third counter group 433 can include a destination column counter and a destination row counter. The counter selector 434 can control the destination column counter.

[0056] When generating a third target address (e.g., a target memory address) of the memory array 410 by adding the offset of the third base address 401c to the third base address 401c by the address adder of the address generator 430, the sub-blocks of the first output feature map block can be stored in the memory array 410 based on the third target address.

[0057] The sub-blocks of the first output feature map block can be stored in the memory array 410 by the third instruction 402c without the host 490 performing an operation for generating the offset of the third base address 401c. For example, the host 490 can provide the third instruction 402c to the address generator 430 N times without specifying an offset, and in this process, the sub-blocks of the first output feature map block can be stored N times at an interval of M - 1 from the third base address 401c. The sub-blocks of the first output feature map block can be loaded N times using the lower bits of the target memory address from the second register group B.

[0058] Figure 5 is a diagram exemplarily showing the mapping relationship between weights, input features, and output features according to an embodiment. Refer to Figure 5 , the output feature 530 can be generated by performing a MAC operation on the weight 510 and the input feature 520. The weight 510 can include weight blocks (such as the weight block 519). Each weight block can include sub-blocks. Although Figure 5 shows an example where the weight block 519 includes four sub-blocks, the embodiment is not limited thereto. Each sub-block in the weight block 519 can include weight elements. The number of weight elements in each sub-block can correspond to the number of sub-operators of the processing unit of the memory device.

[0059] The input feature 520 may include an input tile 521 (e.g., a first input tile 521) and an input tile 522 (e.g., a second input tile 522). The input tile 521 and the input tile 522 may each include a sub-tile. The number of sub-tiles in each of the input tile 521 and the input tile 522 may be equal to the number of sub-tiles in each of the weight tiles. Figure 5 An example is shown in which the input feature 520 includes an input tile 521 and an input tile 522, and the input tile 521 and the input tile 522 each include four sub-tiles, but the embodiment is not limited thereto. Each sub-tile of the input tile 521 and the input tile 522 may include an input element. The number of input elements of each sub-tile may correspond to the number of sub-operators of the processing unit of the memory device.

[0060] The output feature 530 may include an output tile 531 (e.g., a first output tile 531) and an output tile 532 (e.g., a second output tile 532). The output tile 531 and the output tile 532 may each include a sub-tile. The number of sub-tiles of each of the output tile 531 and the output tile 532 may correspond to the product of the number of sub-tiles of each weight tile and each input tile and the number of memory banks. Figure 5 An example in which output feature 530 includes output tile 531 and output tile 532 and an example in which output tile 531 and output tile 532 each include eight sub-tiles are shown. Figure 5 The example corresponds to an example in which the memory device uses two memory banks. However, the foregoing is an example, and embodiments are not limited thereto. Each sub-block of output block 531 and output block 532 may include an output element. The number of output elements of each sub-block may correspond to the number of sub-operators of the processing unit of the memory device.

[0061] The weight 510 may be divided into a first input tile region 511 and a second input tile region 512 based on the input dimension. An operation may be performed on the weight tile in the first input tile region 511 and the first input tile 521, and an operation may be performed on the weight tile in the second input tile region 512 and the second input tile 522. The weight 510 may be divided into a first output tile region 513 and a second output tile region 514 based on the output dimension. A first output tile 531 may be generated based on the operation performed on the weight tile in the first output tile region 513 and the input feature 520, and a second output tile 532 may be generated based on the operation performed on the weight tile in the second output tile region 514 and the input feature 520.

[0062] Figure 6A , Figure 6C , Figure 6E and Figure 6Gis a diagram exemplarily showing a process of processing a MAC operation according to an embodiment, and Figure 6B , Figure 6D , Figure 6F and Figure 6H are diagrams exemplarily showing sub - blocks of weights, input features, and output features used in the processing processes respectively at Figure 6A , Figure 6C , Figure 6E and Figure 6G according to an embodiment.

[0063] Referring to Figure 6A , in a state where a sub - block of a first input tile of an input feature is stored in a first register group A of a data register 612 of a first memory bank 610 based on a first target address, an operation of a processing unit 613 of the first memory bank 610 can be executed. The processing unit 613 can perform an operation on an input element of a first sub - block of a first input tile loaded into a data space of the processing unit 613 and a weight element of a first sub - block of a first weight tile of weights based on a first address in a second target address to generate an operation result, and store the operation result in a second register group B of the data register 612. The operation result can correspond to a first intermediate data of a first sub - block of a first output tile of an output feature. The input element of the first sub - block of the first input tile can be loaded from the first register group A based on a lower bit of the first address in the second target address.

[0064] The first memory bank 610 and the second memory bank 620 can operate in parallel. The first memory bank 610 can generate a part of each output tile (e.g., the first sub - tile to the fourth sub - tile among the first sub - tile to the eighth sub - tile of each output tile), and the second memory bank 620 can generate the remaining part of each output tile (e.g., the fifth sub - tile to the eighth sub - tile among the first sub - tile to the eighth sub - tile of each output tile).

[0065] More specifically, in a state where a sub - block of a first input tile of an input feature is stored in a first register group A of a data register 622 of a second memory bank 620 based on a third target address, an operation of a processing unit 623 of the second memory bank 620 can be executed. The processing unit 623 can perform an operation on an input element of a first sub - block of a first input tile loaded into a data space of the processing unit 623 and a weight element of a first sub - block of a second weight tile of weights based on a first address in a fourth target address to generate an operation result, and store the operation result in a second register group B of the data register 622. The operation result can correspond to a first intermediate data of a fifth sub - tile of a first output tile of an output feature.

[0066] Referring to Figure 6B, the first intermediate data of the first sub-block 671 of the first output block of the output feature 670 can be generated based on the operation performed on the first sub-block 651 of the first weight block of the weight 650 and the first sub-block 661 of the first input block of the input feature 660, and the first intermediate data of the fifth sub-block 672 of the first output block of the output feature 670 can be generated based on the operation performed on the first sub-block 652 of the second weight block of the weight 650 and the first sub-block 661 of the first input block of the input feature 660.

[0067] Refer to Figure 6C , in a state where the sub-block of the first input block of the input feature is stored in the first register group A of the data register 612 of the first memory bank 610 based on the first target address, the operation of the processing unit 613 of the first memory bank 610 can be continuously executed. The processing unit 613 can perform an operation on the input elements of the second sub-block of the first input block loaded into the data space of the processing unit 613 and the weight elements of the second sub-block of the first weight block of the weight based on the second address in the second target address to generate an operation result, and accumulate the operation result in the first intermediate data of the first sub-block of the first output block of the output feature in the second register group B of the data register 612 to generate an accumulated result. The accumulated result can correspond to the second intermediate data of the first sub-block of the first output block of the output feature. The input elements of the second sub-block of the first input block can be loaded from the first register group A based on the lower bits of the second address in the second target address.

[0068] In a state where the sub-block of the first input block of the input feature is stored in the first register group A of the data register 622 of the second memory bank 620 based on the third target address, the operation of the processing unit 623 of the second memory bank 620 can be executed. The processing unit 623 can perform an operation on the input elements of the first sub-block of the first input block loaded into the data space of the processing unit 623 and the weight elements of the second sub-block of the second weight block to generate an operation result, and store the operation result in the second register group B of the data register 622. The operation result can be accumulated in the first intermediate data of the fifth sub-block of the first output block of the output feature in the second register group B of the data register 622 to generate an accumulated result. The accumulated result can correspond to the second intermediate data of the fifth sub-block of the first output block of the output feature.

[0069] Refer to Figure 6D, the second intermediate data of the first sub-block 671 of the first output block of the output feature 670 can be generated based on the operations performed on the second sub-block 653 of the first weight block of the weight 650 and the second sub-block 662 of the first input block of the input feature 660, and the second intermediate data of the fifth sub-block 672 of the first output block of the output feature 670 can be generated based on the operations performed on the second sub-block 654 of the second weight block of the weight 650 and the second sub-block 662 of the first input block of the input feature 660.

[0070] It can be performed Figure 6E and Figure 6G in Figure 6A and Figure 6C the corresponding operations. Referring to Figure 6F , the third intermediate data of the first sub-block 671 of the first output block of the output feature 670 can be generated based on the operations performed on the third sub-block 655 of the first weight block of the weight 650 and the third sub-block 663 of the first input block of the input feature 660, and the third intermediate data of the fifth sub-block 672 of the first output block of the output feature 670 can be generated based on the operations performed on the third sub-block 656 of the second weight block of the weight 650 and the third sub-block 663 of the first input block of the input feature 660. Referring to Figure 6H , the final data of the first sub-block 671 of the first output block of the output feature 670 can be generated based on the operations performed on the fourth sub-block 657 of the first weight block of the weight 650 and the fourth sub-block 664 of the first input block of the input feature 660, and the final data of the fifth sub-block 672 of the first output block of the output feature 670 can be generated based on the operations performed on the fourth sub-block 658 of the second weight block of the weight 650 and the fourth sub-block 664 of the first input block of the input feature 660.

[0071] Figure 7 is a diagram exemplarily showing the configuration of an address generator when a processing-in-memory (PIM) structure is in a bank-shared memory. Referring to Figure 7 , similar to Figure 2A , the address generator 700 similar to the address generator 200 can include a counter block 710, a counter selector 720, an address adder 730, and a multiplexer block 740. The counter block 710, the counter selector 720, the address adder 730, and the multiplexer block 740 can operate similarly to the counter block 210, the counter selector 220, the address adder 230, and the multiplexer block 240 of the address generator 200.

[0072] With Figure 2A the counter selector 220 of Figure 2BUnlike the counter selector 220a, the counter selector 720 also receives the base bank address 708. The base bank address 708 may include the address of one of the banks sharing a PIM structure (e.g., one or more of a controller, an address generator, a data register, and a processing unit). The counter selector 720 may selectively control one of the first source column counter 713a and the second source column counter 713b based on the base bank address 708. The counter selector 720 may select one of the outputs of the first source column counter 713a and the output of the second source column counter 713b by controlling the multiplexer 751 based on the base bank address 708. For example, when the base bank address 708 is a first value, the counter selector 720 may select the first source column counter 713a, and when the base bank address 708 is a second value different from the first value, the counter selector 720 may select the second source column counter 713b. The first value may correspond to the first bank, and the second value may correspond to the second bank.

[0073] When the count value of the first source column counter 713a increases to the maximum value, the first source row counter 714a may operate, and when the count value of the second source column counter 713b increases to the maximum value, the second source row counter 714b may operate. The counter selector 720 may select one of the outputs of the first source column counter 713a and the output of the second source column counter 713b by controlling the multiplexer 752 based on the base bank address 708. The first source column counter 713a and the first source row counter 714a may operate for the first bank, and the second source column counter 713b and the second source row counter 714b may operate for the second bank.

[0074] Figure 8 is a flowchart exemplarily showing an operation method of a memory device according to an embodiment. Refer to Figure 8 , the operation method includes: in operation 810, the memory device receives an instruction and a base address of the instruction from a host. For example, the memory device 100 may receive the instruction 101 and the base address 202. Figure 8 The operation method of also includes: in operation 820, target addresses for performing operations of the instruction are sequentially generated by sequentially adding an offset to the base address. For example, an address generator (e.g., 200, 200a, etc.) may generate the target addresses. Figure 8 The operation method of also includes: in operation 830, data values corresponding to one or more of the target addresses are stored. For example, the data values may be stored in the data register 140. Figure 8 The operation method of also includes: in operation 840, one or more of the operations of the instruction are performed based on the data values. For example, the processing unit 150 may perform one or more operations.

[0075] Figure 9 is a diagram exemplarily showing a configuration of an electronic device according to an embodiment. Referring to Figure 9 , the electronic device 900 may include a host 910 and a memory device 920. The electronic device 900 may also include other devices (such as, a memory, a storage device, an input device, an output device, a network device, etc.). Here, the memory may correspond to a conventional memory having a storage function without a computing function. For example, the electronic device 900 may be implemented as at least a part of a mobile device (such as, a mobile phone, a smart phone, a personal digital assistant (PDA), a netbook, a tablet computer, or a laptop computer), a wearable device (such as, a smart watch, a smart band, or smart glasses), a computing device (such as, a desktop or a server), a household appliance (such as, a TV, a smart TV, or a refrigerator), a security device (such as, a door lock), or a vehicle (such as, an autonomous vehicle or a smart vehicle).

[0076] The host 910 may correspond to a processor (such as, a CPU or a GPU). The host 910 may generate an instruction and a base address, and send the instruction and the base address to the memory device 920. The memory device 920 may include an address generator, a data register, and a processing unit. The address generator is configured to receive the instruction and the base address of the instruction from the host 910 and sequentially generate target addresses for performing operations of the instruction by sequentially adding an offset to the base address. The data register is configured to store data values corresponding to one or more of the target addresses. The processing unit is configured to perform one or more of the operations of the instruction based on the data values.

[0077] The units described herein may be implemented using hardware components, software components, and / or a combination thereof. The processing device may be implemented using one or more general-purpose or special-purpose computers (such as, for example, a processor, a controller, and an arithmetic logic unit (ALU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of responding and executing instructions in a defined manner). The processing device may run an operating system (OS) and one or more software applications running on the OS. The processing device may also access, store, manipulate, process, and create data in response to the execution of the software. For simplicity, the description of the processing device is used as a singular; however, those skilled in the art will understand that the processing device may include multiple processing elements and various types of processing elements. For example, the processing device may include multiple processors, or a single processor and a single controller. In addition, different processing configurations (such as, parallel processors) are feasible.

[0078] Software may include a computer program, a piece of code, instructions, or some combination thereof, to independently or uniformly direct or configure a processing device to operate as required. Software and data may be stored in any type of machine, component, physical or virtual device, or computer storage medium or device capable of providing instructions or data to, or being interpreted by, the processing device. Software may also be distributed over a network-coupled computer system such that the software is stored and executed in a distributed manner. Software and data may be stored by one or more non-transitory computer-readable recording media.

[0079] The method according to the above embodiments may be recorded in a non-transitory computer-readable medium, which includes program instructions for implementing the various operations of the above embodiments. The medium may also include data files, data structures, etc., either alone or in combination with the program instructions. The program instructions recorded on the medium may be program instructions specifically designed and constructed for the purposes of these embodiments. Examples of non-transitory computer-readable media include: magnetic media (such as hard disks, floppy disks, and magnetic tapes); optical media (such as CD-ROM disks, DVDs, and / or Blu-ray disks); magneto-optical media (such as optical disks); and hardware devices specifically configured to store and execute program instructions (such as read-only memory (ROM), random access memory (RAM), flash memory (e.g., USB flash drives, memory cards, memory sticks, etc.)). Examples of program instructions include both machine code generated by a compiler and files containing higher-level code that can be executed by a computer using an interpreter.

[0080] The above hardware devices may be configured to act as one or more software modules for performing the operations of the above examples, and vice versa.

[0081] Although multiple embodiments have been described above, it should be understood that various modifications may be made to these embodiments. For example, suitable results may be achieved if the described techniques are performed in a different order and / or if the components in the described system, architecture, device, or circuit are combined in a different manner and / or replaced or supplemented by other components or their equivalents. Accordingly, other embodiments are within the scope of the appended claims.

Claims

1. A memory device, comprising: Memory array; an address generator configured to receive instructions and base addresses of the instructions from a host, and sequentially generate target addresses for performing operations of the instructions by sequentially adding offsets to the base addresses; a data register configured to store a data value corresponding to one or more of the target addresses; as well as A processing unit is configured to perform one or more of the operations of the instructions based on the data values.

2. The memory device according to claim 1, wherein: The address generator includes: a counter block, configured to generate an offset; a counter selector configured to control the counter block; and An address adder is configured to generate a target memory address in a target address by adding the offset to the base address.

3. The memory device according to claim 2, wherein: The counter block consists of: a column counter group, comprising a plurality of column counters; and A row counter group includes multiple row counters.

4. The memory device of claim 3, wherein The plurality of column counters include: a first column counter configured to sequentially increase a column count value corresponding to a column offset based on control of the counter selector, and The plurality of row counters include: a first row counter configured to increase a row count value corresponding to a row offset when a column count value increases to a maximum value; The first column counter is configured to initialize the column count value when the column count value increases to a maximum value.

5. The memory device according to claim 3, wherein: The address generator also includes: a column multiplexer configured to select an output of one of the plurality of column counters based on control of the counter selector; and The row multiplexer is configured to select an output of one of the plurality of row counters based on control of the counter selector.

6. The memory device according to claim 5, wherein: The column multiplexer and the row multiplexer are configured to output a target memory address or a target register index as the target address based on at least one of the instruction and a base address.

7. The memory device according to claim 6, wherein: The target register index is generated by extracting the number of bits identifying the number of registers in each register bank of data registers from the least significant bits of the output of the counter block.

8. The memory device of claim 2, wherein The counter selector is configured to: receive a first instruction and a first base address of the first instruction, and control the counter block to generate an offset of the first base address, the first instruction being used to store a sub-block of a first input feature block of an input feature in a first register group of a data register, the first base address of the first instruction indicating a starting index of the first register group, The address adder is configured to generate a first target address of the first register group by adding an offset of the first base address to the first base address, and The first register group is configured to store a sub-tile of the first input feature tile in the first register group based on the first target address.

9. The memory device according to claim 8, wherein: The sub-tile of the first input feature tile is stored in the first register group by the first instruction without the host performing an operation for calculating an offset of the first base address.

10. The memory device of claim 8, wherein The counter selector is configured to: receive a second instruction and a second base address of the second instruction, and control the counter block to generate an offset of the second base address, the second instruction is used to store a sub-block of the first weight block of weights in a data space of the processing unit, the second base address of the second instruction indicating a starting address of the memory array for storing the weights, The address adder is configured to generate a second target address of the memory array by adding an offset of the second base address to the second base address, and The processing unit is configured to: sequentially perform operations on sub-blocks of the first weight map block loaded into the data space of the processing unit based on the second target address and sub-blocks of the first input feature map block loaded into the data space of the processing unit based on the first target address to generate operation results, and generate sub-blocks of the first output feature map block by accumulating the operation results.

11. The memory device according to claim 10, wherein: The operations on the sub-block of the first weight map block and the sub-block of the first input feature map block are performed by the second instruction without the need for the host to calculate an offset of the second base address.

12. The memory device of claim 10, wherein The counter selector is configured to: receive a third instruction and a third base address of the third instruction, and control the counter block to generate an offset of the third base address, the third instruction is used to store a sub-block of the first output feature map block stored in the second register group of the data register in the memory array, the third base address of the third instruction indicates a starting address of the memory array for storing the sub-block of the first output feature map block, The address adder is configured to generate a third target address of the memory array by adding an offset of the third base address to the third base address, and The sub-tiles of the first output feature tile are stored in the memory array based on the third target address.

13. The memory device according to claim 12, wherein: The sub-tile of the first output feature tile is stored in the memory array by the third instruction without the host performing an operation for calculating an offset of the third base address.

14. The memory device according to claim 1, wherein: The offsets have predetermined intervals.

15. An electronic device comprising: Host; an address generator configured to receive instructions and base addresses of the instructions from a host, and sequentially generate target addresses for performing operations of the instructions by sequentially adding offsets to the base addresses; a data register configured to store a data value corresponding to one or more of the target addresses; as well as A processing unit is configured to perform one or more of the operations of the instructions based on the data values.

16. The electronic device according to claim 15, wherein: The address generator includes: a counter block, configured to generate an offset; a counter selector configured to control the counter block; and An address adder is configured to generate a target memory address in a target address by adding the offset to the base address.

17. The electronic device according to claim 16, wherein: The counter block consists of: a column counter group, comprising a plurality of column counters; and A row counter group includes multiple row counters.

18. The electronic device according to claim 17, wherein The plurality of column counters include: a first column counter configured to sequentially increase a column count value corresponding to a column offset based on control of the counter selector, and The plurality of row counters include: a first row counter configured to increase a row count value corresponding to a row offset when a column count value increases to a maximum value; The first column counter is configured to initialize the column count value when the column count value increases to a maximum value.

19. The electronic device according to claim 17, wherein: The address generator also includes: a column multiplexer configured to select an output of one of the plurality of column counters based on control of the counter selector; and The row multiplexer is configured to select an output of one of the plurality of row counters based on control of the counter selector.

20. A method for operating a memory device, the method comprising: receiving an instruction and a base address of the instruction from a host; sequentially generating target addresses for performing operations of the instructions by sequentially adding offsets to base addresses; storing data values ​​corresponding to one or more of the target addresses; as well as One or more of the operations of the instruction are performed based on the data value.

Citation Information

Cited By

  • Register management method and device, electronic equipment and storage medium

    CN121116394A