Memory circuit and operation method thereof

By introducing a CIM array and decoder into the memory circuit, the compression weight set is converted into a differential signal set, which solves the problems of increased processing volume and decreased power efficiency caused by changes in wire resistance, and achieves higher power efficiency and memory capacity.

CN120766735APending Publication Date: 2025-10-10TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510729768.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-11
Filing Date
2025-06-03
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

As semiconductor integrated circuits shrink and increase in complexity, variations in wire resistance within digital devices affect operating voltages and overall IC performance, leading to increased processing throughput and decreased power efficiency.

Method used

A computation-in-memory (CIM) array and decoder are used to compress the weight set into a differential signal set, reduce processing volume and improve power efficiency. The encoder and decoder are used for signal compression and decompression, and the memory unit array is combined for calculation.

Benefits of technology

By reducing the amount of processing and memory resources used, power efficiency and memory capacity are improved, the number of accesses to the external buffer is reduced, and the overall performance of the memory circuit is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766735A_ABST
    Figure CN120766735A_ABST
Patent Text Reader

Abstract

The invention discloses a memory circuit and an operation method thereof. The memory circuit comprises a memory internal computing (CIM) array. The CIM array includes an array of memory cells for storing a first set of data. The first data set comprises a first weight set or a second data set. The first data set is an exponential portion of a corresponding floating-point number. The second set of data is a compressed version of the first set of weights. The first set of weights has a first data length, and the second set of data has a second data length less than the first data length. The CIM array further includes a decoder coupled to the array of memory cells and configured to generate a first set of output signals in response to the first set of input signals, the first set of data, and the flag signal. According to the memory circuit, the processing capacity executed by the memory circuit can be reduced, and the power efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a memory circuit and an operating method thereof. BACKGROUND

[0002] The semiconductor integrated circuit (IC) industry has produced a wide variety of digital devices to address many different applications. Some of these digital devices, such as memory arrays, are configured to store data. As ICs become smaller and more complex, the resistance of the wires within these digital devices also changes, affecting the operating voltage of these digital devices and the overall IC performance. SUMMARY

[0003] One embodiment of the disclosure includes a memory circuit, comprising: a compute-in-memory (CIM) array, comprising: a memory cell array to store a first set of data, the first set of data comprising a first set of weights or a second set of data, the first set of data corresponding to an exponent portion of a floating point number, the second set of data being a compressed version of the first set of weights, the first set of weights having a first data length, the second set of data having a second data length that is less than the first data length; and a decoder coupled to the memory cell array and configured to generate a first set of output signals in response to a first set of input signals, the first set of data, and a flag signal.

[0004] Another embodiment of the disclosure includes a memory circuit, comprising: a compute-in-memory (CIM) array, comprising: a memory cell array to store a first set of exponent data and a first set of mantissa data, the first set of exponent data comprising a first set of weights or a second set of exponent data, the first set of exponent data corresponding to an exponent portion of a floating point number, and the second set of exponent data being a compressed version of the first set of weights, the first set of mantissa data being a second set of weights, and the first set of mantissa data corresponding to a mantissa portion of the floating point number; a first adder circuit coupled to the memory cell array and to generate a first set of output signals in response to a first set of input signals, the first set of exponent data, and a flag signal; and a set of multipliers coupled to the memory cell array and to generate a second set of output signals in response to the first set of input signals and the first set of mantissa data.

[0005] Another embodiment of the present disclosure includes a method for operating a memory circuit, the method comprising the following steps: receiving a first weight set by an encoder, the first weight set being in a floating point format; compressing the first weight set into a first differential signal set by the encoder, the first weight set including a first data length, the first differential signal set including a second data length that is smaller than the first data length; performing a read operation on a memory cell array in the in-memory computing array by an in-memory computing array, thereby outputting a first differential signal set, the in-memory computing array being coupled to the encoder; and generating a first output signal set by a decoder in response to the first input signal set and the first differential signal set. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The aspects of this disclosure are in accordance with the accompanying Figure 1 The following detailed description is best understood when read together. It should be noted that, in accordance with standard practice in the industry, various features are not drawn to scale. In fact, the dimensions of the various features may be arbitrarily increased or reduced for clarity of discussion.

[0007] Figure 1 is a block diagram of a memory circuit according to some embodiments;

[0008] Figure 2 is a block diagram of a memory circuit according to some embodiments;

[0009] Figure 3 is a block diagram of a memory circuit according to some embodiments;

[0010] Figure 4 is a schematic diagram of a digit according to some embodiments;

[0011] Figure 5A is a flow chart of a method of operating a memory circuit according to some embodiments;

[0012] Figure 5B According to some embodiments of the present invention Figure 5A a schematic diagram that graphically illustrates one or more operations of a method;

[0013] Figure 6 is a flow chart of a method of operating a memory circuit according to some embodiments;

[0014] 7A to 7B is a corresponding block diagram of a corresponding schematic diagram according to some embodiments;

[0015] Figure 8A is a circuit diagram of a decoder circuit according to some embodiments;

[0016] Figure 8B is a block diagram of a schematic diagram according to some embodiments;

[0017] Figure 8C According to some embodiments Figure 6 a schematic diagram graphically illustrating at least a portion of a method;

[0018] Figure 9A is a schematic diagram of a memory device according to some embodiments;

[0019] Figure 9B is a schematic diagram of a neural network according to some embodiments;

[0020] Figure 9C is a schematic diagram of an integrated circuit (IC) device according to some embodiments.

[0021]

Explanation of symbols

[0022] 100:Memory circuit

[0023] 102: Encoder

[0024] 104:CIM Array

[0025] 106:Decoder

[0026] 110:CIM Macro

[0027] 200:Memory circuit

[0028] 202: memory cell array

[0029] 202A~202D: Memory partition

[0030] 210AR:Memory Cell Array

[0031] 210AT: Adder Tree

[0032] 210L: Memory Group

[0033] 210M: FP multiplication circuit

[0034] 210U: Memory Group

[0035] 212:Memory device

[0036] 300:Memory circuit

[0037] 302: Memory Macro

[0038] 304:Memory Array

[0039] 306:Memory Array

[0040] 308: Adder Set

[0041] 310: multiplier set

[0042] 400: digit

[0043] 402a: sign

[0044] 404a: exponent

[0045] 406a: mantissa

[0046] 500A: method

[0047] 500B: diagram

[0048] 502-514: operations

[0049] 520: weight set

[0050] 530: first base value

[0051] 540: delta set

[0052] 550: zone

[0053] 600: method

[0054] 602-616: operations

[0055] 700A: diagram

[0056] 700B: diagram

[0057] 710: memory zone

[0058] 712: flag field

[0059] 714: base value field

[0060] 716: data field

[0061] 720: memory zone

[0062] 730: data field

[0063] 750a: first delta set

[0064] 750b: second delta set

[0065] 800A: decoder / decoder circuit

[0066] 800B: diagram

[0067] 800C: diagram / decoder

[0068] 802: adder set

[0069] 804: adder set

[0070] 804(0)~804(15): Adder

[0071] 808: Input data field

[0072] 810: Multiplexer Set

[0073] 810(0)~810(15):Multiplexer

[0074] 812: register set

[0075] 812(0)~812(15): register

[0076] 814: Multiplexer Set

[0077] 814(0)~814(15):Multiplexer

[0078] 816: Data field

[0079] 900A:Memory device

[0080] 900B: Neural Networks

[0081] 900C:IC device

[0082] 902-908: Memory Macro

[0083] 911: Input Data

[0084] 912~918: Matrix

[0085] 919: Output data

[0086] 920:Memory controller

[0087] 932:Hardware Processor

[0088] 934:Memory device

[0089] 936: Bus DETAILED DESCRIPTION

[0090] The following disclosure provides many different embodiments, or examples, for implementing the different features of the provided subject matter. Specific examples of components, materials, values, steps, configurations, or the like are described below to simplify the disclosure. Other components, materials, values, steps, configurations, or the like are contemplated. Of course, these are merely examples and are not intended to be restrictive. For example, in the following description, the formation of a first feature above or on a second feature may include an embodiment in which the first feature and the second feature are formed in direct contact, and may also include an embodiment in which an additional feature may be formed between the first feature and the second feature so that the first feature and the second feature may not be in direct contact. In addition, the disclosure may repeat reference numbers and / or letters in various examples. This repetition is for simplicity and clarity purposes and does not, in itself, indicate the relationship between the various embodiments and / or configurations discussed.

[0091] Furthermore, for ease of description, spatially relative terms, such as "below," "beneath," "lower," "above," "upper," and the like, may be used herein to describe the relationship of one element or feature to another element or feature illustrated in the figures. Spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations), and the spatially relative descriptors used herein should be similarly interpreted accordingly.

[0092] According to some embodiments, the memory circuit includes a compute-in-memory (CIM) array.

[0093] In some embodiments, the CIM array includes a memory cell array. In some embodiments, the memory cell array is used to store a first data set.

[0094] In some embodiments, the first data set includes a first weight set or a second data set. In some embodiments, the first data set is an exponent portion of a corresponding floating-point number. In some embodiments, the first weight set is compressed into a second data set, wherein the first weight set has a first data length and the second data set has a second data length that is less than the first data length.

[0095] In some embodiments, the CIM array further includes a decoder coupled to the memory cell array. In some embodiments, the decoder is configured to generate a first set of output signals in response to a first set of input signals, a first set of data, and a flag signal. In some embodiments, the decoder is configured to generate a first set of output signals in response to the first set of input signals, the flag signal, and at least the first set of weights or the second set of data.

[0096] In some embodiments, by compressing the first set of weights into the second set of data, the memory circuit can reduce the amount of processing performed by the memory circuit compared to other approaches. In some embodiments, reducing the amount of processing performed by the memory circuit can result in improved power efficiency compared to other approaches having vector multiplier accumulator (MAC) units.

[0097] Figure 1 is a block diagram of a memory circuit 100 according to some embodiments.

[0098] The memory circuit 100 includes an encoder 102 and a compute in-memory (CIM) macro 110 .

[0099] The CIM macro 110 includes a CIM array 104 and a decoder 106 .

[0100] Encoder 102 is coupled to CIM array 104. An input of encoder 102 is configured to receive a weight set W. An output of encoder 102 is configured to output data set FP1. In some embodiments, each received signal in weight set W is in floating point format. In some embodiments, each signal in data set FP1 is in floating point format. In some embodiments, weight set W includes 16 words, each of which is 8 bits long. Other numbers of words or word lengths within weight set W are also within the scope of the present disclosure.

[0101] Encoder 102 is configured to generate data set FP1 in response to weight set W. In some embodiments, data set FP1 includes weight set W or delta set D. In some embodiments, each signal in delta set D has a floating point format. In some embodiments, delta set D includes at least one of delta set D1 or delta set D2. In some embodiments, delta set D is a compressed version of weight set W. In some embodiments, at least one of delta set D1 or delta set D2 is a compressed version of weight set W.

[0102] In some embodiments, the weight set W includes at least one of the weight set W1 or the weight set W2. In some embodiments, the delta set D1 is a compressed version of the weight set W1 and the delta set D2 is a compressed version of the weight set W2.

[0103] In some embodiments, the encoder 102 is used to compress the weight set W when generating the difference set D. In some embodiments, the encoder 102 is used to compress the weight set W into at least one of the difference set D1 or the difference set D2. In some embodiments, compressing the first signal into the second signal includes changing the first size of the first signal to the second size of the second signal. In some embodiments, compressing the data includes reducing the size of the data. In some embodiments, the size of the data includes the length of the data. For example, in some embodiments, the weight set W includes a data length L2, and the difference set D1 or D2 includes a data length L1. In these embodiments, the data length L1 is less than the data length L2. In other words, since the data length L1 is less than the data length L2, the difference set D1 or D2 is compressed relative to the weight set W. In some embodiments, the encoder 102 is also referred to as a compressor.

[0104] In some embodiments, the difference set D1 or D2 is the exponent portion of the corresponding floating point number.

[0105] In some embodiments, the delta set D includes 32 words, and each word is 4 bits long. Other numbers of words or word lengths in the delta set D are also within the scope of the present disclosure.

[0106] In some embodiments, delta set D1 includes 16 words, each of which is 4 bits long, and delta set D2 includes 16 words, each of which is 4 bits long. Other numbers of words or word lengths within delta set D1 or D2 are also within the scope of the present disclosure.

[0107] In some embodiments, if the data set FP1 is equal to the weight set W, the encoder 102 is configured to pass (eg, without compression) the weight set W as the data set FP1.

[0108] Other configurations of encoder 102 are also within the scope of this disclosure.

[0109] The CIM array 104 is coupled to the output of the encoder 102 and the input of the decoder 106. The input of the CIM array 104 is coupled to the output of the encoder 102. The output of the CIM array 104 is coupled to the input of the decoder 106. In some embodiments, the CIM array 104 includes an array of memory cells (e.g., Figure 3 shown).

[0110] The memory cell array in CIM array 104 is configured to store signal set FP1. CIM array 104 is configured to generate signal set FP2 in response to signal set FP1. In some embodiments, signal set FP2 is identical to signal set FP1. In some embodiments, data set FP2 includes a weight set W or a difference set D. In some embodiments, each signal in data set FP2 is in floating point format.

[0111] In some embodiments, the signal set FP1 includes an index signal set FPE (e.g. Figure 3 As shown) and tail signal collection FME ( Figure 3 In some embodiments, the signal set FP2 includes an exponential signal set FE (e.g. Figure 3 As shown) and the mantissa signal set FM (as Figure 3 In some embodiments, the exponent signal set FE is equal to the exponent signal set FPE. In some embodiments, the mantissa signal set FM is equal to the mantissa signal set FME.

[0112] Other configurations or formats of at least the signal set FP2 are within the scope of the present disclosure.

[0113] In some embodiments, the memory cell array in CIM array 104 is a volatile memory cell array including volatile memory cells. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a static random-access memory (SRAM) cell. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a dynamic random-access memory (DRAM) cell.

[0114] In some embodiments, memory cell array 102 is a non-volatile memory cell array including non-volatile memory cells. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a magnetoresistive random-access memory (MRAM) cell. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a phase-change RAM (PRAM) cell. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a phase-change RAM (PRAM) cell. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a ferroelectric RAM (FeRAM) cell. In some embodiments, each memory cell in the memory cell array of CIM array 104 corresponds to a ferroelectric field effect transistor (FeFET) cell.

[0115] Other configurations or other types of memory cells in the memory cell array of the CIM array 104 are also within the scope of the present disclosure.

[0116] CIM array 104 and decoder 106 are part of memory macro 110. In some embodiments, memory macro 110 is used to perform vector multiplication of data set FP1 and input signal set XIN. In some embodiments, memory macro 110 performs one or more multiply-accumulate (MAC) operations.

[0117] In some embodiments, the memory circuit 100 is part of a neural network, the input signal set XIN corresponds to an input vector, the signal set FP2 corresponds to a weight vector, and the memory macro 110 is used to multiply the input vector by the weight vector to generate the output signal set D_OUT.

[0118] In some embodiments, the input vector corresponds to data values ​​in one or more neural networks based on the type of application. In some embodiments, the weight vector corresponds to the values ​​of one or more trained filter coefficients within a particular layer of the one or more neural networks.

[0119] Other configurations of the CIM array 104 are also within the scope of this disclosure.

[0120] The decoder 106 is coupled to the CIM array 104. A first input of the decoder 106 is to receive the signal set FP2. A second input of the decoder 106 is to receive the input signal set XIN. A third input of the decoder 106 is to receive the flag signal F. An output of the decoder 106 is to output the output signal set D_OUT.

[0121] The decoder 106 is to generate the output signal set D_OUT in response to at least one of the signal set FP2, the input signal set XIN, or the flag signal F. In some embodiments, the output signal set D_OUT has a floating-point number format.

[0122] In some embodiments, the decoder 106 is to decompress the signal set FP2 and perform at least one of an addition or a multiplication of the decompressed signal set FP2 and the input set XIN.

[0123] In some embodiments, the flag signal F can be used by the decoder 106 to determine whether to decompress the difference signal set D in generating the output signal set D_OUT or whether to use the weight set W in generating the output signal set D_OUT.

[0124] In some embodiments, decompressing a signal is opposite to compressing a signal performed by the encoder 102. In some embodiments, decompressing a first signal into a second signal comprises changing a length of the first signal to a length of the second signal. For example, in some embodiments, the decoder 106 is to change a length of the first difference set D1 or the second difference set D2 to a length of the weight set W. In some embodiments, the decoder 106 is also referred to as a decompressor.

[0125] Other configurations of the decoder 106 are also within the scope of the present disclosure.

[0126] In some embodiments, at least two or more of the encoder 102, the CIM array 104, or the decoder 106 are combined into a single circuit.

[0127] In some embodiments, by compressing the weight set W into the difference set D, the memory macro 110 is able to reduce an amount of processing performed by the memory macro 110 compared to other methods. In some embodiments, reducing the amount of processing performed by the memory macro 110 results in power efficiency improvement compared to other methods having vector multiplier accumulator (MAC) units.

[0128] In some embodiments, by compressing the weight set W into the difference set D, the CIM array 104 is able to utilize less memory resources compared to other methods, thereby increasing a memory capacity of the CIM array 104.

[0129] In some embodiments, by compressing the weight set W into the delta set D, the CIM array 104 can reduce the number of memory accesses to one or more external buffers compared to other approaches.

[0130] In some embodiments, by decompressing the set of differences D that are exponents of floating point numbers, the decoder 106 can perform decompression of the data by utilizing fewer logical resources compared to other methods, thereby reducing the energy used to perform the decompression.

[0131] Other configurations or numbers of components in the memory circuit 100 are also within the scope of the present disclosure.

[0132] Figure 2 is a block diagram of a memory circuit 200 according to some embodiments.

[0133] For illustration purposes, the Figure 2 In some embodiments, the memory circuit 200 also includes Figure 2 Various elements other than those depicted may be configured in other ways to perform the operations described below.

[0134] The memory circuit 200 is an embodiment of the memory macro 110 , and thus similar detailed description is omitted.

[0135] The memory circuit 200 is an integrated circuit (IC) including memory partitions 202A- 202D and an adder tree 210AT.

[0136] Each memory partition 202A-202D includes memory groups 210U and 210L. The memory groups 210U and 210L are adjacent to the adder tree 210AT.

[0137] Each memory bank 210U and 210L includes a memory cell array 210AR and a floating point (FP) multiplication circuit 210M.

[0138] In some embodiments, memory groups 210U and 210L and adder tree 210AT are embodiments of memory macro 110, and similar detailed descriptions are omitted. In some embodiments, adder tree 210AT is an embodiment of decoder 106, and similar detailed descriptions are omitted. In some embodiments, memory cell array 210AR is an embodiment of CIM array 104, and similar detailed descriptions are omitted.

[0139] A memory partition, such as memory partitions 202A-202D, is a portion of the memory circuit 200 that includes a subset of memory devices ( Figure 2and adjacent circuitry to selectively access subsets of the memory devices in program and read operations. In Figure 2 In embodiments, the memory circuit 200 includes a total of four partitions. In some embodiments, the memory circuit 200 includes a total number of partitions that is greater than or less than four.

[0140] Each memory bank 210U and 210L includes a corresponding memory cell array 210AR that includes memory cells or memory devices 212 to be accessed by adjacent local input output (LIO) circuitry (not shown) in program and read operations.

[0141] Each memory cell array 210AR includes an array of memory devices 212 having N columns and M rows, where M and N are positive integers. Columns of cells in the memory cell array 202 are configured along a first direction X. Rows of cells in the memory cell array 202 are configured along a second direction Y. The second direction Y is different than the first direction X. In some embodiments, the second direction Y is perpendicular to the first direction X. In some embodiments, each memory cell array 210AR is divided into an upper region and a lower region (not shown). In some embodiments, each row of memory devices 212 in the memory cell array 210AR is coupled to a corresponding FP multiplication circuit 210M and a corresponding adder tree 210AT.

[0142] Memory devices 212 are shown in the memory banks 210U and 210L of the memory partition 202A. For ease of illustration, memory devices 212 are not shown in the memory banks 210U and 210L of the memory partitions 202B, 202C, and 202D.

[0143] A memory device 212 is an electrical, electromechanical, electromagnetic, or other device to store bit data represented by a logic state. At least one logic state of a memory device 212 can be programmed in a write operation and detected in a read operation. In some embodiments, a logic state corresponds to a voltage level of a charge stored in a given memory device 212. In some embodiments, a logic state corresponds to a physical property of a component of a given memory device 212, such as a voltage, a current, a resistance, or a magnetic orientation.

[0144] In some embodiments, the memory device 212 includes one or more single port (SP) static random-access memory (SRAM) cells. In some embodiments, the memory device 212 includes one or more dual port (DP) SRAM cells. In some embodiments, the memory device 212 includes one or more multi-port SRAM cells. Different types of memory cells in the memory device 212 are also within the contemplated scope of the present disclosure. In some embodiments, the memory device 212 includes one or more dynamic random-access memory (DRAM) cells. In some embodiments, the memory device 212 includes one or more one-time programmable (OTP) memory devices, such as an electronic fuse (eFuse) or anti-fuse device, a flash memory device, a random-access memory (RAM) device, a resistive RAM device, a ferroelectric RAM device, a magnetoresistive RAM device, an erasable programmable read only memory (EPROM) device, an electrically erasable programmable read only memory (EEPROM) device, or the like. In some embodiments, the memory device 212 is an OTP memory device including one or more OTP memory cells.

[0145] In some embodiments, each FP multiplication circuit 210M is to perform multiplication between a set of input signals XIN( Figure 1 ) and a set of weights W( Figure 1 ). In some embodiments, each FP multiplication circuit 210M includes one or more multipliers (e.g., a set of multipliers 310 in Figure 3 ).

[0146] In some embodiments, the adder tree 210AT is to perform addition between a set of input signals XIN( Figure 1 ) and a set of weights W( Figure 1 ). In some embodiments, the adder tree 210AT includes one or more adders (e.g., a set of adders 308 in Figure 3 ).

[0147] Region 202 is part of memory circuit 200. In some embodiments, region 202 includes adder tree 210AT, FP multiplication circuit 210M, and a portion of memory cell array 210AR.

[0148] Other configurations of memory circuit 200 are within the scope of this disclosure.

[0149] Figure 3 is a block diagram of memory circuit 300 according to some embodiments.

[0150] For purposes of illustration, details Figure 3 In some embodiments, memory circuit 300 also includes various elements in addition to, or otherwise configured to perform the operations described below, the elements depicted in FIG. 3. Figure 3

[0151] Memory circuit 300 is an embodiment of region 202 of Figure 2 , and similar detailed descriptions are omitted.

[0152] Memory circuit 300 includes memory macro 302.

[0153] Memory macro 302 includes memory array 304, memory array 306, adder set 308, and multiplier set 310.

[0154] In some embodiments, memory array 304 is an embodiment of the first portion of memory cell array 210AR of Figure 2 , and memory array 306 is an embodiment of the second portion of memory cell array 210AR of Figure 2 , and similar detailed descriptions are omitted.

[0155] In some embodiments, adder set 308 is an embodiment of adder tree 210AT of Figure 2 , and multiplier set 310 is an embodiment of FP multiplication circuit 210M of Figure 2 , and similar detailed descriptions are omitted.

[0156] In some embodiments, at least one of memory array 304 or 306 is an embodiment of CIM array 104 of Figure 1 , and similar detailed descriptions are omitted. In some embodiments, at least one of adder set 308 or multiplier set 310 is an embodiment of decoder 106 of Figure 1 , and similar detailed descriptions are omitted.

[0157] ​The memory array 304 is coupled to the adder set 308. The memory array 304 is configured to receive or store the exponent signal set FPE. In some embodiments, the exponent signal set FPE is the exponent portion of the signal set FP1.

[0158] The memory array 304 includes memory cell columns ranging from N columns to 2*N columns, where N is an integer corresponding to the number of columns in the memory array 306. In some embodiments, if the signal set FP1 is compressed by the encoder 102, a single column of memory cells in the memory array 304 is used to store the first set of deltas D1 and the second set of deltas D2 in corresponding columns "column 1" and "column N+1." In other words, in some embodiments, when the signal set FP1 is compressed by the encoder 102, the memory cell column "column 1" in the memory array 304 is used to store the first set of deltas D1, and the memory cell column "column N+1" in the memory array 304 is used to store the second set of deltas D2.

[0159] During a read operation of the memory array 304, the memory array 304 is configured to output an index signal set FE. In some embodiments, the index signal set FE corresponds to the index signal set FPE. In some embodiments, the index signal set FE is equal to the index signal set FPE.

[0160] In some embodiments, the exponential signal set FE includes one or more of the exponential signals FE(0), FE(1), ..., FE(X), where X is an integer corresponding to the number of signals in the exponential signal set FM.

[0161] In some embodiments, the index signal set FE includes a first index signal set FE1 (not labeled) and a second index signal set FE2 (not labeled). In some embodiments, the first index signal set FE1 is a difference set D1. In some embodiments, the second index signal set FE2 is a difference set D2.

[0162] In some embodiments, the index signal set FE corresponds to a portion of a single column. For example, in some embodiments, the difference set D1 is stored in column 1 of the memory array 304, and the difference set D2 is stored in column N+1 of the memory array 304.

[0163] Other configurations of the memory array 304 are also within the scope of this disclosure.

[0164] The memory array 306 is coupled to the multiplier set 310. The memory array 306 is configured to receive or store the mantissa signal set FPM. In some embodiments, the mantissa signal set FPM is the mantissa portion of the signal set FP1.

[0165] Memory array 306 includes N columns of memory cells, where N is an integer corresponding to a number of columns in memory array 306. In some embodiments, a single column of memory cells in memory array 306 is used to store a set of mantissa signals FPM.

[0166] During a read operation of memory array 306, memory array 306 is used to output a set of mantissa signals FM. In some embodiments, the set of mantissa signals FM corresponds to the set of mantissa signals FPM. In some embodiments, the set of mantissa signals FM is equal to the set of mantissa signals FPM.

[0167] In some embodiments, the set of mantissa signals FM includes one or more of mantissa signals FM(0), FM(1),..., FM(Y), where Y is an integer corresponding to a number of signals in the set of mantissa signals FM. In some embodiments, integer Y is different from integer X. In some embodiments, integer Y is the same as integer X.

[0168] In some embodiments, the set of mantissa signals FM includes a first set of mantissa signals FM1 (not labeled) and a second set of mantissa signals FM2 (not labeled). In some embodiments, the first set of mantissa signals FM1 corresponds to the first set of exponent signals FE1. In some embodiments, the second set of mantissa signals FM2 corresponds to the second set of exponent signals FE2.

[0169] In some embodiments, the set of mantissa signals FM corresponds to a single column of memory array 306. In some embodiments, the set of mantissa signals FM corresponds to more than one column of memory array 306.

[0170] Other configurations in memory array 306 are within the scope of the present disclosure.

[0171] Adder set 308 is coupled to memory array 304 and multiplier set 310.

[0172] Adder set 308 is used to generate a set of exponent output signals DE in response to a set of input signals XIN and a set of exponent signals FE. In some embodiments, the set of exponent output signals DE is a sum of the set of input signals XIN and the set of exponent signals FE.

[0173] A first input of adder set 308 is used to receive the set of input signals XIN.

[0174] The second set of inputs of the adder set 308 is configured to receive the exponent signal set FE. The second set of inputs of the adder set 308 is coupled to the memory array 304. Each input of the second set of inputs of the adder set 308 is configured to receive a corresponding exponent signal FE(0), FE(1), ..., FE(X) of the exponent signal set FE. In some embodiments, each input of the second set of inputs of the adder set 308 is configured to receive a corresponding exponent signal FE(0), FE(1), ..., FE(X) of the exponent signal set FE from a corresponding memory cell in the memory array 304.

[0175] The output set of the adder set 308 is used to output or generate the output exponent signal set DE. Each output of the output set of the adder set 308 is used to output or generate a corresponding output exponent signal DE(0), DE(1), . . . , DE(X) of the output exponent signal set DE.

[0176] In some embodiments, the output exponential signals DE(0), DE(1), ..., DE(X) of the output exponential signal set DE are equal to the sum of the corresponding input signals of the input signal set XIN and the corresponding exponential signals FE(0), FE(1), ..., FE(X) of the exponential signal set FE.

[0177] The multiplier set 310 is coupled to the memory array 306 and the adder set 308 .

[0178] The multiplier set 310 is configured to generate a set of mantissa output signals DM in response to the set of input signals XIN and the set of mantissa signals FM. In some embodiments, the set of mantissa output signals DM is the product of the set of input signals XIN and the set of mantissa signals FM.

[0179] A first input of the multiplier set 310 is configured to receive the input signal set XIN.

[0180] The second set of inputs to the set of multipliers 310 is configured to receive the set of mantissa signals FM. The second set of inputs to the set of multipliers 310 is coupled to the memory array 306. Each input in the second set of inputs to the set of multipliers 310 is configured to receive a corresponding mantissa signal FM(0), FM(1), ..., FM(Y) from the set of mantissa signals FM. In some embodiments, each input in the second set of inputs to the set of multipliers 310 is configured to receive a corresponding mantissa signal FM(0), FM(1), ..., FM(Y) from the set of mantissa signals FM of a corresponding memory cell in the memory array 306.

[0181] The output set of the multiplier set 310 is used to output or generate the output mantissa signal set DM. Each output of the output set of the multiplier set 310 is used to output or generate a corresponding output mantissa signal DM(0), DM(1), . . . , DM(Y) of the output mantissa signal set DM.

[0182] In some embodiments, the output mantissa signals DM(0), DM(1), ..., DM(Y) of the output mantissa signal set DM are equal to the product of the corresponding input signal of the input signal set XIN and the corresponding mantissa signals FM(0), FM(1), ..., FM(Y) of the mantissa signal set FM.

[0183] In some embodiments, the output signal set D_OUT includes an output exponent signal set DE and an output mantissa signal set DM.

[0184] In some embodiments, one of the outputs of the set of adders 308 or the outputs of the set of multipliers 310 is coupled to an accumulator (not shown).

[0185] Other configurations of the memory circuit 300 are also within the scope of this disclosure.

[0186] Figure 4 is a schematic diagram of digit 400 according to some embodiments.

[0187] The digit 400 is Figure 1 The embodiment of at least one received signal in the received signal set FP1 or FP2 is omitted, so similar detailed description is omitted.

[0188] To Figures 1 to 9C The same or similar components in one or more of the embodiments are given the same reference numerals, and thus detailed description thereof is omitted.

[0189] Digit 400 is a floating-point number in base 2. Digit 400 includes a sign 402a, an exponent 404a, and a mantissa 406a. Sign 402a corresponds to the sign of the floating-point number (e.g., digit 400). Exponent 404a corresponds to the exponent of the floating-point number (e.g., digit 400). Mantissa 406a corresponds to the mantissa of the floating-point number (e.g., digit 400).

[0190] In some embodiments, digit 400 corresponds to one or more floating point numbers of the present application. In some embodiments, symbol 402a corresponds to one or more symbols of the present application. In some embodiments, exponent 404a corresponds to one or more exponents of the present application. In some embodiments, mantissa 406a corresponds to one or more mantissas of the present application.

[0191] In some embodiments, the floating point format of digit 400 includes half precision (e.g., "FP16 format"). In some embodiments, FP16 includes 16 bits. Other floating point formats of digit 400 are also within the scope of the present disclosure. For example, in some embodiments, the floating point format of digit 400 includes one or more of a 32-bit, 64-bit, 128-bit, or 256-bit floating point format. In some embodiments, the floating point format of digit 400 includes one or more of the Institute of Electrical and Electronics Engineers (IEEE)-754 floating point formats. In some embodiments, the floating point format of digit 400 includes one or more of FP8 (E4M3), FP8 (E5M2), FP16 (E5M10), or BF16 (E8M7).

[0192] Other numbers of bits in the floating point format of digit 400 are also within the scope of this disclosure.

[0193] Other types of floating point formats for the number 400 are also within the scope of this disclosure.

[0194] Other configurations of the digital device 400 are also within the scope of this disclosure.

[0195] Figure 5A is a flow chart of a method 500A of operating a memory circuit according to some embodiments.

[0196] In some embodiments, Figure 5A Yes Operation Figure 1 In some embodiments, Figure 5A Yes Operation Figure 1 The memory circuit 100 or Figure 9C Flowchart of the method of IC device 900C. It should be understood that Figure 5A Additional operations are performed before, during, and / or after the method 500A shown, and some other operations may only be briefly described herein. In some embodiments, other operation orders of the method 500A are within the scope of this disclosure. In some embodiments, one or more operations of the method 500A are not performed.

[0197] Method 500A includes illustrative operations, but these operations are not necessarily performed in the order shown. Operations may be added, replaced, changed in order, and / or eliminated as appropriate, in accordance with the spirit and scope of the disclosed embodiments. It is understood that method 500A utilizes Figure 1 Encoder 102, Figure 1 The memory circuit 100, or Figure 9C Features of one or more of the IC devices 900.

[0198] It can be understood that method 500A utilizes Figure 2 Memory circuit 200, Figure 4 The digit 400, Figure 5B Diagram 500B, Figure 7A Diagram 700A, Figure 7B Diagram 700B, or Figure 8B Features of one or more of the diagrams 800B.

[0199] In some embodiments, method 500A is repeated for each weight set W. For example, if the weight set W includes a weight set W1 and a weight set W2, method 500A is performed on the weight set W1 to obtain a first difference set D1 (e.g., Figures 8A to 8C ), execute method 500A on the weight set W2 to obtain a second difference set D2 (as shown in FIG. Figures 8A to 8C shown).

[0200] In operation 502 of method 500A, a first set of weights is received.

[0201] In some embodiments, the first set of weights is received by a controller. In some embodiments, the controller of method 500A is processor 932. In some embodiments, the first set of weights is received by encoder 102.

[0202] In some embodiments, the first set of weights of method 500A is as follows: Figure 5B The weight set 520 in is shown.

[0203] In operation 504 of method 500A, a first base value for the first set of weights is determined.

[0204] In some embodiments, at least one of operations 504, 506, 508, 510, 512, or 514 is performed by processor 932. In some embodiments, at least one of operations 504, 506, 508, 510, 512, or 514 is performed by hardware (not shown).

[0205] In some embodiments, the first base value BV is the minimum value of the first set of weights.

[0206] In some embodiments, the first base value BV of method 500A includes Figure 5B The base value is 530.

[0207] In some embodiments, the first base value BV of method 500A is shown as Figure 5B The first base value in is 530.

[0208] In operation 506 of method 500A, a first set of differentials (D1 or D2) is determined from the first base value of the first set of weights and the first set of weights.

[0209] In some embodiments, the first set of differentials of method 500A includes the set of differentials D1 or the set of differentials D2 in Figure 5B

[0210] In some embodiments, the first set of differentials of method 500A includes the set of differentials D1 or the set of differentials D2 in Figures 8A to 8C

[0211] In some embodiments, the first set of differentials is equal to the difference between the first base value and the first set of weights.

[0212] In some embodiments, the first set of differentials of method 500A is shown as the set of differentials 540 in Figure 5B

[0213] In operation 508 of method 500A, a maximum differential value MaxD is determined from the first set of differentials.

[0214] In some embodiments, the maximum differential value of method 500A is the maximum differential value in the first set of differentials.

[0215] In some embodiments, the maximum differential value of method 500A is shown as the maximum differential value Max(D1) in region 550 of Figure 5B

[0216] In operation 510 of method 500A, it is determined whether the maximum differential value is less than a first threshold value FT.

[0217] In some embodiments, the first threshold value FT of method 500A is a threshold value used to determine whether to compress the set of weights W to a set of differentials D in Figure 1

[0218] In some embodiments, the first threshold value FT corresponds to the number of bins in the set of weights W. For example, in some embodiments, the set of weights W includes 16 bins, and each bin is 8 bits long. In these embodiments, the first threshold value FT is equal to 2 4 or 16. In some embodiments, the first threshold value FT is less than 2 (the number of bits per bin / 2).

[0219] Other numbers of bins or bin lengths within the set of weights W are within the scope of the present disclosure. Other values of the first threshold value of the set of weights W are within the scope of the present disclosure.

[0220] ​​​​​In some embodiments, the first threshold value is selected by a user of IC device 900. In some embodiments, the first threshold value is pre-programmed into memory device 934. In some embodiments, the first threshold value is dynamically adjusted by processor 932 or a user of IC device 900.

[0221] In some embodiments, if the maximum difference value is less than the first threshold value FT, the result of operation 510 is “true” and the method 500A proceeds to operation 512 .

[0222] In some embodiments, if the maximum difference value is not less than the first threshold value FT, the result of operation 510 is “false” and the method 500A proceeds to operation 514 .

[0223] In operation 512 of method 500A, in response to a maximum delta value MaxD in the first delta set being greater than a first threshold value FT, the first delta set is written into the memory cell array.

[0224] In some embodiments, performance of operation 512 results in compression of the first set of weights into a first set of deltas.

[0225] In operation 514 of method 500A, in response to a maximum difference value MaxD in the first difference set being less than a first threshold value, the first weight set is written into the memory cell array.

[0226] In some embodiments, performance of operation 512 results in the first set of weights being uncompressed into the first set of deltas.

[0227] By performing at least method 500A, the memory circuit is operated to achieve one or more advantages of the present application.

[0228] Figure 5B According to some embodiments of the present invention Figure 5A 500B is a graphical illustration of one or more operations of method 500A.

[0229] Diagram 500B includes a set of weights 520 , a first base value 530 , a set of deltas 540 , and a region 550 .

[0230] According to some embodiments, the set of weights 520 corresponds to the set of weights following operation 502 of method 500A.

[0231] According to some embodiments, first base value 530 corresponds to the first base value after operation 504 of method 500A.

[0232] According to some embodiments, the set of deltas 540 corresponds to the first set of deltas 540 following operation 506 of method 500A.

[0233] Area 550 includes a maximum difference value MaxD, a first threshold value FT, and a compression result field “Compressed (True)”.

[0234] According to some embodiments, the maximum difference value MaxD corresponds to the maximum difference value MaxD after operation 508 of method 500A.

[0235] According to some embodiments, the first threshold value FT corresponds to the first threshold value FT of operation 510 of method 500A.

[0236] According to some embodiments, the compression result field "Compressed (True)" corresponds to the result of method 500A after operation 512 of method 500A.

[0237] In some embodiments, weight set 520 is an embodiment of weight set W, and thus similar detailed description is omitted.

[0238] In some embodiments, the first difference set 540 is an embodiment of the difference set D, and thus similar detailed description is omitted.

[0239] Other values ​​in the weight set 520 or the format for the weight set 520 are also within the scope of the present disclosure.

[0240] Other values ​​for the first base value 530 or the format for the first base value 530 are also within the scope of the present disclosure.

[0241] Other values ​​in the delta set 540 or the format for the delta set 540 are also within the scope of the present disclosure.

[0242] Other values ​​for field 550 or formats for field 550 are also within the scope of the present disclosure.

[0243] Other configurations of diagram 500B are also within the scope of the present disclosure.

[0244] Figure 6 is a flow chart of a method 600 of operating a memory circuit according to some embodiments.

[0245] In some embodiments, Figure 6 Yes Operation Figure 1 Decoder 106, Figure 8A Decoder 800A, or Figure 8C In some embodiments, Figure 6 Yes Operation Figure 1 The memory circuit 100 or Figure 9C Flowchart of the method of IC device 900C. It can be understood that Figure 6Additional operations are performed before, during, and / or after the method 600 shown, and some other operations are only briefly described here. In some embodiments, other operation orders of the method 600 are within the scope of this disclosure. In some embodiments, one or more operations of the method 600 are not performed.

[0246] Method 600 includes illustrative operations, but these operations are not necessarily performed in the order shown. Operations may be added, replaced, changed in order, and / or eliminated as appropriate, in accordance with the spirit and scope of the disclosed embodiments. It is understood that method 600 utilizes Figure 1 106 or Figure 8A Decoder 800A or Figure 8C Features of one or more of decoders 800C.

[0247] It can be understood that method 600 utilizes Figure 1 The memory circuit 100, Figure 2 Memory circuit 200, Figure 3 Memory circuit 300, Figure 4 The digit 400, Figure 5B Diagram 500B, Figure 7A Diagram 700A, Figure 7B Diagram 700B, or Figure 8B Features of one or more of the diagrams 800B.

[0248] In some embodiments, method 600 is repeated for each weight set W. For example, if the weight set W includes a first weight set W1 and a second weight set W2, then for each weight set W having a first difference set D1 and a second difference set D2 (e.g., Figures 8A to 8C ) executes method 600 for a first weight set W1 having a first difference set D1 and a second difference set D2 (as shown in FIG. Figures 8A to 8C Method 600 is executed with a second weight set W2 (shown).

[0249] In operation 602 of method 600 , a first data set is read from a memory array.

[0250] In some embodiments, the first data set is read by at least one of the decoder 106 , the adder 308 , the multiplier 310 , or the decoder 800A, the decoder 800C, or the memory controller 920 .

[0251] In some embodiments, the first data set of method 600 includes a signal set FP2.

[0252] In some embodiments, the first data set of method 600 includes a weight set W or a difference set D.

[0253] In some embodiments, the first data set of method 600 includes at least one of a first set of weights W1 or a second set of weights W2.

[0254] In some embodiments, the first data set of method 600 includes at least one of a first delta set D1 or a second delta set D2.

[0255] In some embodiments, the memory array of method 600 includes CIM array 104, memory cell array 210AR, memory array 304, memory array 306, memory region 710, memory region 720, or diagram 800B.

[0256] In some embodiments, the first data set of method 600 is as follows: Figure 5B The weight set 520 is shown.

[0257] In some embodiments, the first data set of method 600 is as follows: Figure 5B The difference set 540 is shown.

[0258] In some embodiments, the first data set of method 600 is Figures 8A to 8B The differences D1(0), ..., D1(15) of the first difference set D1 are shown in FIG.

[0259] In some embodiments, the first data set of method 600 is Figures 8A to 8B The differences D2(0), ..., D2(15) of the first difference set D2 are shown in FIG.

[0260] In some embodiments, the first data set of method 600 is shown as Figures 8A to 8B The weights W(0), ..., W(15) of the weight set W1 in .

[0261] In operation 604 of method 600 , a first sum value FSV is determined in response to the first input signal set XIN and the first base value BV of the first weight set W1 . In some embodiments, the first sum value FSV is equal to the sum of the first input signal set XIN and the first base value BV of the first weight set W1 .

[0262] In some embodiments, operation 604 of method 600 is performed by a first set of adders. In some embodiments, the first set of adders of method 600 includes adder set 802.

[0263] In some embodiments, the first set of input signals of method 600 includes a set of input signals XIN.

[0264] In some embodiments, the first set of weights of method 600 includes a first set of weights W1.

[0265] In operation 606 of method 600, in response to address signal ADDR, first difference set Dl or second difference set D2 is selected as first signal set C.

[0266] In some embodiments, operation 606 is performed by a first multiplexer set. In some embodiments, the first multiplexer set of operation 606 comprises multiplexer set 810. In some embodiments, first signal set C is output by the first multiplexer set.

[0267] In some embodiments, address signal ADDR can be used by the first multiplexer set as a selection signal to select first difference set Dl or second difference set D2 as an output signal (e.g., first signal set C).

[0268] In some embodiments, second difference set D2 is a compressed version of second weight set W2.

[0269] In operation 608 of method 600, it is determined whether flag F is equal to first value FV. In some embodiments, flag F is a single bit. In some embodiments, flag F is more than a single bit.

[0270] In some embodiments, first value FV is a single bit. In some embodiments, first value FV is more than a single bit.

[0271] In some embodiments, first value FV is equal to a logical high (e.g., logical 1). In some embodiments, first value FV is equal to a logical low (e.g., logical 0).

[0272] In some embodiments, flag can be used by a second multiplexer set (e.g., multiplexer set 814) as a selection signal.

[0273] In some embodiments, if flag F is equal to first value FV, the result of operation 608 is "true" and method 600 proceeds to operation 610.

[0274] In some embodiments, if flag F is not equal to first value FV, the result of operation 608 is "false" and method 600 proceeds to operation 614.

[0275] In operation 610 of method 600, in response to flag F, first signal set C is selected as second signal set B.

[0276] In some embodiments, at least one of operations 610 or 614 is performed by a second multiplexer set. In some embodiments, the second multiplexer set of method 600 comprises multiplexer set 814. In some embodiments, second signal set B is output by the second multiplexer set.

[0277] In some embodiments, the second signal set B includes the first signal set C or the first weight set W1.

[0278] In some embodiments, flag F can be used by the second set of multiplexers as a selection signal to select the second signal set B or the first weight set W1 as the output signal (eg, the second signal set B).

[0279] In some embodiments, operation 610 is performed due to the first set of weights W1 being previously compressed by a compressor (eg, encoder 102 ), and thus the first set of deltas D1 or the second set of deltas D2 is decompressed.

[0280] In operation 612 of method 600, the first sum value FSV is added to each difference value in the second signal set B to form a first output signal set. In other words, in operation 612, the first output signal set is determined by adding the first sum value FSV to each difference value in the second signal set B. In some embodiments, operation 612 includes determining the first output signal set by adding the first sum value FSV to each difference value in the first difference set D1 or the second difference set D2.

[0281] In some embodiments, operation 612 of method 600 is performed by a second set of adders. In some embodiments, the second set of adders of method 600 includes adder set 804.

[0282] In some embodiments, the first set of output signals of method 600 includes a set of exponential output signals DE.

[0283] In operation 614 of method 600 , in response to flag F, the first weight set W1 is selected as the second signal set B.

[0284] In some embodiments, operation 614 is performed due to the first set of weights W1 not being compressed by a compressor (eg, encoder 102 ), and therefore the first set of weights W1 is not decompressed.

[0285] In operation 616 of method 600, the first sum value FSV is added to each weight value in the second signal set B to form a first output signal set. In other words, in operation 616, the first output signal set is determined by adding the first sum value FSV to each weight value in the second signal set B. In some embodiments, operation 616 includes determining the first output signal set by adding the first sum value FSV to each weight value in the first weight set W1.

[0286] In some embodiments, operation 616 of method 600 is performed by a second set of adders.

[0287] By performing at least method 600 , the memory circuit is operated to achieve one or more advantages of the present application.

[0288] 7A to 7B are corresponding block diagrams of the corresponding diagrams 700A-700B according to some embodiments.

[0289] For illustration purposes, the 7A to 7B In some embodiments, diagram 700A or 700B includes 7A to 7B Various elements other than those depicted in the drawings may be included or otherwise configured to perform the operations described below.

[0290] Diagram 700A is Figure 3 The embodiments of one or more columns of the memory array 304 are omitted, and thus similar detailed description is omitted.

[0291] Diagram 700A includes memory array 304 .

[0292] Diagram 700A further includes region 710 .

[0293] In some embodiments, the region 710 is an embodiment of one or more columns of the memory array 304 , and thus similar detailed description is omitted.

[0294] In some embodiments, region 710 corresponds to executing Figure 5A The one or more rows of the memory array 304 are processed after operation 512 of the method 500 .

[0295] In some embodiments, region 710 corresponds to executing Figure 6 The method 600 may include one or more rows of the memory array 304 prior to operation 602 .

[0296] Region 710 includes a flag field 712 , a base value field 714 , and a data field 716 .

[0297] In some embodiments, the flag field 712 is Figure 6 、 Figure 8A 、 Figure 8B ,and Figure 8C The flag F, base value field 714 is Figure 5A 、 Figure 5B 、 Figure 6 、 Figure 8A 、 Figure 8B ,and Figure 8C The base value BV of the data field 716 is the difference set D, so similar detailed description is omitted.

[0298] In some embodiments, the length of the flag field 712 is 1 bit. In some embodiments, the length of the flag field 712 is greater than 1 bit.

[0299] In some embodiments, the base value field 714 or base value BV is 8 bits long. In some embodiments, the base value field 714 or base value BV is greater than 8 bits long. In some embodiments, the base value field 714 or base value BV is less than 8 bits long.

[0300] In some embodiments, the length of the data field 716 is 128 bits. In some embodiments, the length of the data field 716 is greater than 128 bits. In some embodiments, the length of the data field 716 is less than 128 bits.

[0301] In some embodiments, the data field 716 includes a first set of deltas D1 750a and a second set of deltas D2 750b.

[0302] In some embodiments, the first difference set D1 750a includes difference values ​​D1(0), D1(1), ..., D1(15). In some embodiments, the first difference set D1 includes 16 difference values. Other numbers of values ​​in the first difference set D1 are also within the scope of the present disclosure.

[0303] In some embodiments, the second difference set D2 750b includes difference values ​​D2(0), D2(1), ..., D2(15). In some embodiments, the second difference set D2 includes 16 difference values. Other numbers of values ​​in the second difference set D2 are also within the scope of the present disclosure.

[0304] In some embodiments, each difference value D1(0), D1(1), ..., D1(15) in the first difference set D1 has a length of 4 bits. Other numbers of bits for each difference value D1(0), D1(1), ..., D1(15) in the first difference set D1 are also within the scope of the present disclosure.

[0305] In some embodiments, each difference value D2(0), D2(1), ..., D2(15) in the second difference set D2 has a length of 4 bits. Other numbers of bits for each difference value D2(0), D2(1), ..., D2(15) in the second difference set D2 are also within the scope of the present disclosure.

[0306] Other configurations of diagram 700A are also within the scope of the present disclosure.

[0307] Figure 7B is a block diagram of diagram 700B according to some embodiments.

[0308] Diagram 700B is Figure 3 The embodiments of one or more columns of the memory array 304 are omitted, and thus similar detailed description is omitted.

[0309] Diagram 700B includes memory array 304 .

[0310] Diagram 700B further includes region 720 .

[0311] In some embodiments, the region 720 is an embodiment of one or more columns of the memory array 304 , and thus similar detailed description is omitted.

[0312] In some embodiments, region 720 corresponds to executing Figure 5A The one or more rows of the memory array 304 are stored after operation 512 of the method 500 .

[0313] In some embodiments, region 720 corresponds to executing Figure 6 The method 600 may include one or more rows of the memory array 304 prior to operation 602 .

[0314] Area 720 includes a flag field 712 , a base value field 714 , and a data field 730 .

[0315] In some embodiments, the data field 730 is a weight set W, and thus similar detailed description is omitted.

[0316] In some embodiments, the length of the data field 730 is 128 bits. In some embodiments, the length of the data field 730 is greater than 128 bits. In some embodiments, the length of the data field 730 is less than 128 bits.

[0317] In some embodiments, the data field 730 includes a set of weights W. In some embodiments, the data field 730 includes a first set of weights W1 or a second set of weights W2.

[0318] In some embodiments, the weight set W includes weight values ​​W(0), W(1), ..., W(15). In some embodiments, the weight set W includes 16 weight values. Other numbers of values ​​of the weight set W are also within the scope of the present disclosure.

[0319] In some embodiments, the first weight set W1 includes weight values ​​W1(0), W1(1), ..., W1(15). In some embodiments, the first weight set W1 includes 16 weight values. Other numbers of values ​​in the first weight set W1 are also within the scope of the present disclosure.

[0320] In some embodiments, the second weight set W2 includes weight values ​​W2(0), W2(1), ..., W2(15). In some embodiments, the second weight set W2 includes 16 weight values. Other numbers of values ​​in the second weight set W2 are also within the scope of the present disclosure.

[0321] In some embodiments, each weight value W(0), W(1), ..., W(15) in the weight set W has a length of 8 bits. Other numbers of bits for each weight value W(0), W(1), ..., W(15) in the weight set W are also within the scope of the present disclosure.

[0322] Other configurations of diagram 700B are also within the scope of the present disclosure.

[0323] Figure 8A is a circuit diagram of a decoder circuit 800A according to some embodiments.

[0324] Decoder circuit 800A is Figure 1 The decoder circuit 800A is at least Figure 3 In some embodiments, the decoder 800A is Figure 2 The embodiments of the adder tree 210AT are described in detail, and thus similar detailed description is omitted.

[0325] For illustration purposes, the Figure 8A In some embodiments, the decoder circuit 800A also includes Figure 8A Various elements other than those depicted may be configured in other ways to perform the operations described below.

[0326] The decoder circuit 800A includes a set of adders 802 , a set of multiplexers 810 , a set of registers 812 , a set of multiplexers 814 , and a set of adders 804 .

[0327] The adder set 802 is coupled to the adder set 804. The multiplexer set 810 is coupled to the register set 812 and the multiplexer set 814. The register set 812 and the multiplexer set 814 are coupled to the adder set 804.

[0328] The adder set 802 is configured to receive an input signal set XIN and a first base value BV of the first weight set W1 or the second weight set W2. The adder set 802 is configured to output a first sum value FSV to the adder set 804. The adder set 802 is configured to generate a first sum value FSV in response to the input signal set XIN and the first base value BV of the first weight set W1 or the second weight set W2. In some embodiments, the adder set 802 is configured to determine the first sum value FSV in response to the input signal set XIN and the first base value BV of the first weight set W1 or the second weight set W2. In some embodiments, the first base value of the first weight set W1 or the second weight set W2 is the minimum value in the first weight set W1 or the second weight set W2. In some embodiments, the adder set 802 determines the first sum value FSV for each first base value BV in the first weight set W1 or the second weight set W2.

[0329] In some embodiments, the adder set 802 is used to perform at least operation 604 of the method 600 , and thus similar detailed description is omitted.

[0330] In some embodiments, a first input terminal of the adder set 802 is coupled to a source of the input signal XIN, and a second input terminal of the adder set 802 is coupled to a source of a first base value (eg, Figure 1 CIM memory array 104, Figure 2 Memory array 210AR, or Figure 3 In some embodiments, the output of the adder set 802 is coupled to the first set of input terminals of the adder set 804. In some embodiments, the output of the adder set 802 is used to output the first sum value FSV to the first set of input terminals of the adder set 804.

[0331] In some embodiments, the adder set 804 includes at least adder 804a. Other numbers of adders in the adder set 804 are also within the scope of the present disclosure.

[0332] Other configurations of adder set 804 are also within the scope of this disclosure.

[0333] The multiplexer set 810 is coupled to Figure 1 CIM memory array 104, Figure 2 Memory array 210AR, or Figure 3 The memory array 304 is described in detail, and thus similar detailed description is omitted.

[0334] In some embodiments, the multiplexer set 810 is used to perform at least operations 606 and 608 of the method 600 , and thus similar detailed descriptions are omitted.

[0335] Multiplexer set 810 is configured to receive first difference set D1, second difference set D2, and address signal ADDR. Multiplexer set 810 is configured to output signal set C in response to address signal ADDR. In some embodiments, address signal ADDR can be used by multiplexer set 810 to select either first difference set D1 or second difference set D2 as signal set C.

[0336] In some embodiments, based on the value of the address signal ADDR, the signal set C is the first difference set D1 or the second difference set D2 .

[0337] In some embodiments, when the address signal ADDR is equal to logic low (eg, logic 0), the signal set C is equal to the first difference set D1. In some embodiments, when the address signal ADDR is equal to logic high (eg, logic 1), the signal set C is equal to the second difference set D2.

[0338] In some embodiments, the multiplexer set 810 includes at least one of multiplexers 810(0), 810(1), ..., 810(15). Other numbers of multiplexers in the multiplexer set 810 are also within the scope of the present disclosure.

[0339] In some embodiments, signal set C includes at least one of signals C[0], C[1], ..., C

[15] . Other numbers of signals in signal set C are also within the scope of the present disclosure.

[0340] In some embodiments, each multiplexer 810(0), 810(1), ..., 810(15) in the multiplexer set 810 is used to receive the corresponding differential values ​​D1[0], D1[1], ..., D1

[15] of the first differential set D1, the corresponding differential values ​​D2[0], D2[1], ..., D2

[15] of the second differential set D2, and the address signal ADDR.

[0341] In some embodiments, each multiplexer 810(0), 810(1), ..., 810(15) in the multiplexer set 810 is configured to output a corresponding signal C[0], C[1], ..., C

[15] of the signal set C in response to the address signal ADDR.

[0342] In some embodiments, based on the address signal ADDR, each signal C[0], C[1], ..., C

[15] in the signal set C is equal to the corresponding difference value D1[0], D1[1], ..., D1

[15] of the first difference set D1 or the corresponding difference value D2[0], D2[1], ..., D2

[15] of the second difference set D2.

[0343] Other configurations of the address signal ADDR are also within the scope of the present disclosure. For example, in some embodiments, when the address signal ADDR is equal to a logic high (e.g., a logic 1), the signal set C is equal to the first difference set D1. For example, in some embodiments, when the address signal ADDR is equal to a logic low (e.g., a logic 0), the signal set C is equal to the second difference set D2.

[0344] In some embodiments, the first difference set D1 is equal to the difference between the first base value BV of the first weight set W1 and the first weight set W1.

[0345] In some embodiments, the second difference set D2 is equal to the difference between the second base value BV of the second weight set W2 and the second weight set W2. In some embodiments, the second base value of the second weight set W2 is the minimum value in the second weight set W2.

[0346] Other configurations of the multiplexer set 810 are also within the scope of this disclosure.

[0347] The first input terminal of register set 812 is coupled to a source of a tie-low signal TIEL. In some embodiments, tie-low signal TIEL is a 4-bit signal comprising a logic low (e.g., logic 0) signal. Other numbers of bits in tie-low signal TIEL are also within the scope of the present disclosure. In some embodiments, tie-low signal TIEL is a signal comprising one or more bits of a logic high (e.g., logic 1) signal.

[0348] The second input terminal set of the register set 812 is coupled to the output terminal of the multiplexer set 810. The output terminal of the register set 812 is used to output the signal set PD. The register set 812 is used to generate the signal set PD. In some embodiments, the signal set PD is a combination of the tied low signal TIEL and the signal set C. In some embodiments, the signal set PD is a zero-filled version of the signal set C and has the same length as the first weight set W1 or the second weight set W2. For example, according to some embodiments, if the first weight set W1 or the second weight set W2 has a length equal to 8 bits, then the signal set PD has a length equal to 8 bits. In this example, according to some embodiments, if the signal set C has a length equal to 4 bits, then the tied low signal TIEL has a length equal to 4 bits (because 8-4 is equal to 4 bits). In this example, according to some embodiments, if the signal set C has a length equal to 5 bits, then the tied low signal TIEL has a length equal to 3 bits (because 8-5 is equal to 3 bits).

[0349] In some embodiments, the tie-low signal TIEL may be used by the register set to add a series or a plurality of zeros to the front end of the signal set C when generating the signal set PD.

[0350] In some embodiments, register set 812 includes at least one of registers 812(0), 812(1), ..., 812(15). Other numbers of registers in register set 812 are also within the scope of the present disclosure.

[0351] In some embodiments, at least one of the registers 812(0), 812(1), ..., 812(15) of register set 812 is a shift register. In some embodiments, at least one of the registers 812(0), 812(1), ..., 812(15) of register set 812 is a memory element for storing at least one bit of data. In some embodiments, at least one of the registers 812(0), 812(1), ..., 812(15) of register set 812 is a memory element for storing at least one bit of data.

[0352] In some embodiments, the signal set PD includes at least one of the signals PD[0], PD[1], ..., PD

[15] . Other numbers of signals in the signal set PD are also within the scope of the present disclosure.

[0353] In some embodiments, each register 812(0), 812(1), ..., 812(15) in the register set 812 is used to receive a corresponding signal of the tie-low signal TIEL or a corresponding signal C[0], C[1], ..., C

[15] of the signal set C.

[0354] In some embodiments, each register 812 ( 0 ), 812 ( 1 ), . . . , 812 ( 15 ) in the register set 812 is configured to output a corresponding signal PD[ 0 ], PD[ 1 ], . . . , PD[ 15 ] of the signal set PD.

[0355] In some embodiments, signals PD[0], PD[1], ..., PD

[15] of signal set PD are equal to corresponding signals of tie-low signal TIEL or corresponding signals C[0], C[1], ..., C

[15] of signal set C.

[0356] Other configurations of register set 812 are also within the scope of this disclosure.

[0357] The multiplexer set 814 is coupled to Figure 1 The register set 814 and the CIM memory array 104, Figure 2 Memory array 210AR, or Figure 3 The memory array 304 is described in detail, and thus similar detailed description is omitted.

[0358] In some embodiments, the multiplexer set 814 is configured to perform at least operation 610 or 614 of the method 600 , and thus similar detailed description is omitted.

[0359] The multiplexer set 814 is configured to receive the signal set PD, the first weight set W1, and the flag signal F. The multiplexer set 814 is configured to output the signal set B in response to the flag signal F. In some embodiments, the flag signal F can be used by the multiplexer set 814 to select the signal set PD or the first weight set W1.

[0360] In some embodiments, based on the value of the flag signal F, the signal set B is the signal set PD or the first weight set W1 as the signal set B.

[0361] In some embodiments, when the flag signal F is equal to logic high (eg, logic 1), the signal set B is equal to the signal set PD. In some embodiments, when the flag signal F is equal to logic low (eg, logic 0), the signal set B is equal to the first weight set W1.

[0362] In some embodiments, the multiplexer set 814 includes at least one of multiplexers 814(0), 814(1), ..., 814(15). Other numbers of multiplexers in the multiplexer set 814 are also within the scope of the present disclosure.

[0363] In some embodiments, signal set B includes at least one of signals B[0], B[1], ..., B

[15] . Other numbers of signals in signal set B are also within the scope of the present disclosure.

[0364] In some embodiments, each multiplexer 814(0), 814(1), ..., 814(15) in the multiplexer set 814 is used to receive the corresponding signals PD[0], PD[1], ..., PD

[15] of the signal set PD, the corresponding weight values ​​W[0], W[1], ..., W

[15] of the first weight set W1, and the flag signal F.

[0365] In some embodiments, each multiplexer 814(0), 814(1), ..., 814(15) in the multiplexer set 814 is configured to output a corresponding signal B[0], B[1], ..., B

[15] of the signal set B in response to the flag signal F.

[0366] In some embodiments, based on the flag signal F, each signal B[0], B[1], ..., B

[15] in the signal set B is equal to the corresponding signal PD[0], PD[1], ..., PD

[15] of the signal set PD or the corresponding weight value W[0], W[1], ..., W

[15] of the first weight set W1.

[0367] Other configurations of flag signal F are also within the scope of the present disclosure. For example, in some embodiments, when flag signal F is equal to a logic low (e.g., a logic 0), signal set B is equal to signal set PD. For example, in some embodiments, when flag signal F is equal to a logic high (e.g., a logic 1), signal set B is equal to the first weight set W1.

[0368] Other configurations of the multiplexer set 814 are also within the scope of this disclosure.

[0369] The adder set 804 is coupled to the adder set 802 and the multiplexer set 814 .

[0370] In some embodiments, the adder set 804 is further coupled to Figure 1 CIM memory array 104, Figure 2 Memory array 210AR, or Figure 3 The memory array 304 is described in detail, and thus similar detailed description is omitted.

[0371] The adder set 804 is configured to receive the first sum value FSV and a signal set B. The adder set 804 is configured to output an output exponent signal set DE. The adder set 804 is configured to generate the output exponent signal set DE in response to the signal set B and the first sum value FSV. In some embodiments, the adder set 804 is configured to determine the output exponent signal set DE in response to the signal set B and the first sum value FSV. In some embodiments, the output exponent signal set DE is the sum of the signal set B and the first sum value FSV.

[0372] In some embodiments, the adder set 804 is used to perform at least operation 612 or 616 of the method 600 , and thus similar detailed description is omitted.

[0373] In some embodiments, the adder set 804 includes at least one of adders 804(0), 804(1), ..., 804(15). Other numbers of adders in the adder set 804 are also within the scope of the present disclosure.

[0374] In some embodiments, the output index signal set DE includes at least one of the output index signals DE(0), DE(1), ..., DE(X) in the output index signal set DE. Other numbers of output index signals in the output index signal set DE are also within the scope of the present disclosure.

[0375] In some embodiments, the input ends of the corresponding adders 804(0), 804(1), ..., 804(15) of the adder set 814 are coupled to the corresponding output ends of the corresponding multiplexers 814(0), 814(1), ..., 814(15) of the multiplexer set 814, and the output end of the adder set 804.

[0376] In some embodiments, each adder 804(0), 804(1), ..., 804(15) in the adder set 814 is used to receive the corresponding signal B[0], B[1], ..., B

[15] of the signal set B and the first sum value FSV.

[0377] In some embodiments, each adder 804(0), 804(1), ..., 804(15) in the adder set 814 is used to respond to the corresponding signals B[0], B[1], ..., B

[15] of the signal set B and the first sum value FSV, and output the corresponding output exponent signals DE(0), DE(1), ..., DE(X) of the output exponent signal set DE.

[0378] In some embodiments, each output index signal DE(0), DE(1), ..., DE(X) in the output index signal set DE is equal to the corresponding sum of the corresponding signal B[0], B[1], ..., B

[15] of the signal set B and the first sum value FSV.

[0379] Other configurations of adder set 804 are also within the scope of this disclosure.

[0380] Other configurations of the decoder circuit 800A are also within the scope of the present disclosure.

[0381] Figure 8B is a block diagram of diagram 800B according to some embodiments.

[0382] In some embodiments, a portion 802 of diagram 800B is Figure 3 The embodiments of one or more columns of the memory array 304 are omitted, and thus similar detailed description is omitted.

[0383] Diagram 800B includes an input data field 808 , an address field 810 , a flag field 812 , a base value field 814 , and a data field 816 .

[0384] In some embodiments, the input data field 808 is Figure 1 、 Figure 3 、 Figure 6 、 Figure 8A 、 Figure 8B ,and Figure 8C The input signal XIN, address field 810 is Figure 6 、 Figure 8A 、 Figure 8B ,and Figure 8C The address signal ADDR, flag field 812 is Figure 6 、 Figure 8A 、 Figure 8B ,and Figure 8CThe flag F, base value field 814 is Figure 5A 、 Figure 5B 、 Figure 6 、 Figure 8A 、 Figure 8B ,and Figure 8C The base value BV of the data field 816 is the difference set D, so similar detailed description is omitted.

[0385] In some embodiments, the length of the input data field 808 is 8 bits. In some embodiments, the length of the input data field 808 is different than 8 bits.

[0386] In some embodiments, the address field 810 is 1 bit in length. In some embodiments, the address field 810 is greater than 1 bit in length.

[0387] In some embodiments, the length of the flag field 812 is 1 bit. In some embodiments, the length of the flag field 812 is greater than 1 bit.

[0388] In some embodiments, the base value field 814 or base value BV is 8 bits long. In some embodiments, the base value field 814 or base value BV is greater than 8 bits long. In some embodiments, the base value field 814 or base value BV is less than 8 bits long.

[0389] In some embodiments, the length of the data field 816 is 128 bits. In some embodiments, the length of the data field 816 is greater than 128 bits. In some embodiments, the length of the data field 816 is less than 128 bits.

[0390] In some embodiments, the data field 816 includes a first set of deltas D1 .

[0391] In some embodiments, the first difference set D1 includes difference values ​​D1(0), D1(1), ..., D1(15). In some embodiments, the first difference set D1 includes 16 difference values. Other numbers of values ​​in the first difference set D1 are also within the scope of the present disclosure.

[0392] In some embodiments, each difference value D1(0), D1(1), ..., D1(15) in the first difference set D1 has a length of 4 bits. Other numbers of bits for each difference value D1(0), D1(1), ..., D1(15) in the first difference set D1 are also within the scope of the present disclosure.

[0393] Other configurations of diagram 800B are also within the scope of the present disclosure.

[0394] Figure 8C is based on Figure 6Diagram 800C is a graphical illustration of at least a portion of method 600 of some embodiments.

[0395] In some embodiments, diagram 800C corresponds to determining the Figure 8A FIG. 8 is a graphical illustration of the exponential output signal set DE of the decoder circuit 800A, and thus similar detailed description is omitted.

[0396] For illustration purposes, the Figure 8C .

[0397] Illustrated 800C includes Figure 8A The decoder circuit 800A and Figure 8B The values ​​of Figure 800B (when applied to Figure 8A decoding circuit 800A).

[0398] For example, in some embodiments, when the input signal XIN is 15 and the first base value BV is 13, the first sum value FSV (eg, the output signal of the adder set 802 ) is equal to 28.

[0399] For example, in these embodiments, when the address signal ADDR is 0 or logic low, the multiplexer set 810 is configured to output the first difference set D1 as the signal set C. For example, in these embodiments, when the address signal ADDR is 0 or logic low, the multiplexers 810(0), ..., 810(15) of the multiplexer set 810 are configured to output corresponding signals D1[0], ..., D

[15] as corresponding signals C[0], ..., C

[15] of the signal set C. For example, in these embodiments, the signal C[0] is equal to 2, and the signal C

[15] is equal to 4.

[0400] In these embodiments, when the length of the weight signal set W1 is equal to 8 bits and the length of the first difference set D1 is equal to 4 bits, the length of the tied low signal is 4 bits and has a value of 0000. In these embodiments, when the signal C[0] is equal to 2 and the signal C

[15] is equal to 4, the corresponding registers 812(0), ..., 812(15) of the register set 812 are configured to output the corresponding signal PD[0] equal to 2 and the signal PD

[15] equal to 4.

[0401] For example, in these embodiments, when flag signal F is 1 or logic high, multiplexer set 814 is configured to output corresponding signals PD[0], ..., PD

[15] of signal set PD as corresponding signals B[0], ..., B

[15] of signal set B. For example, in these embodiments, when flag signal F is 1 or logic high, multiplexers 812(0), ..., 812(15) of multiplexer set 812 are configured to output corresponding signals PD[0], ..., PD

[15] as corresponding signals B[0], ..., B

[15] of signal set B. For example, in these embodiments, signal B[0] is equal to 2, and signal B

[15] is equal to 4.

[0402] For example, in these embodiments, when the first sum value FSV (e.g., the output signal of the adder set 802) is 28, and when the signal B[0] is equal to 2 and the signal B

[15] is equal to 4, the output exponent signals DE[0], ..., DE

[15] of the output exponent signal set DE are equal to 30, ..., 32.

[0403] Other values ​​or configurations of diagram 800C are also within the scope of the present disclosure.

[0404] Figure 9A is a schematic diagram of a memory device 900A according to some embodiments.

[0405] Memory device 900A includes memory macros 902, 904, 906, 908 and a memory controller 920. In some embodiments, one or more of memory macros 902, 904, 906, 908 correspond to memory macro 110, and / or memory controller 920 corresponds to encoder 102. In some embodiments, one or more of memory macros 902, 904, 906, 908 correspond to CIM memory cell array 104 and / or decoder 106, and / or memory controller 920 corresponds to encoder 102.

[0406] exist Figure 9A In the example configuration, memory controller 920 is a common memory controller for memory macros 902, 904, 906, and 908. In at least one embodiment, at least one of memory macros 902, 904, 906, and 908 has its own memory controller. The number of four memory macros in memory device 900A is an example. Other configurations are within the scope of various embodiments.

[0407] The memory macros 902, 904, 906, and 908 are coupled to each other in sequence, wherein the output data of the previous memory macro is the input data of the next memory macro. For example, the input data DIN is input to the memory macro 902. The memory macro 902 is based on the input data DIN and the weight data or difference set 716 (such as the weight set W stored in the memory macro 902). Figure 7A ) performs one or more CIM operations and generates output data DOUT2 as a result of the CIM operation. The output data DOUT2 is provided as input data DIN4 of the memory macro 904. The memory macro 904 ... Figure 7A ) performs one or more CIM operations and generates output data DOUT4 as a result of the CIM operation. The output data DOUT4 is provided as input data DIN6 of the memory macro 906. The memory macro 906 is based on the input data DIN6 and the weight data or difference set 716 (as shown in FIG. Figure 7A ) performs one or more CIM operations and generates output data DOUT6 as a result of the CIM operation. The output data DOUT6 is provided as input data DIN8 of the memory macro 908. The memory macro 908 generates a weight data or a difference set 716 (as shown in FIG. 1 ) based on the input data DIN8 and the weight set W stored in the memory macro 908. Figure 7A One of the CIM operations (shown) performs one or more CIM operations and generates output data DOUT as a result of the CIM operation.

[0408] One or more of the input data DIN, DIN4, DIN6, DIN8 corresponds to Figure 1 The data set FP1, and / or one or more of the output data DOUT2, DOUT4, DOUT6, DOUT corresponds to Figure 1 The output data set D_OUT is described, and therefore similar detailed description is omitted. In at least one embodiment, the configuration of memory macros 902, 904, 906, and 908 implements a neural network. In at least one embodiment, one or more advantages described herein can be achieved by memory device 900A.

[0409] Other configurations or quantities of components in the memory device 900A are also within the scope of the present disclosure.

[0410] Figure 9B is a schematic diagram of a neural network 900B according to some embodiments.

[0411] Neural network 900B includes a plurality of layers A through E, each containing a plurality of nodes (or neurons). Nodes in successive layers of neural network 900B are connected to each other via a matrix or array of connections. For example, nodes in layers A and B are connected via connections in matrix 912, nodes in layers B and C are connected via connections in matrix 914, nodes in layers C and D are connected via connections in matrix 916, and nodes in layers D and E are connected via connections in matrix 918. Layer A is the input layer that receives input data 911. Input data 911 propagates from one layer to the next through neural network 900B via the corresponding connection matrices between the layers. As data propagates through neural network 900B, it undergoes one or more computations and is output as output data 919 from layer E, which is the output layer of neural network 900B. Layers B, C, and D between input layer A and output layer E are sometimes referred to as hidden layers or intermediate layers. Figure 9B The number of layers, number of connection matrices, and number of nodes in each layer are merely examples. Other configurations are within the scope of various embodiments. For example, in at least one embodiment, neural network 900B does not include any hidden layers and has an input layer connected to an output layer via a connection matrix. In one or more embodiments, neural network 900B has one, two, or more than three hidden layers.

[0412] In some embodiments, matrices 912, 914, 916, and 918 are implemented by memory macros 902, 904, 906, and 908, respectively. Input data 911 corresponds to input data DIN, and output data 919 corresponds to output data DOUT, so similar detailed descriptions are omitted. Specifically, in matrix 912, the connection between a node in layer A and another node in layer B has a corresponding weight. For example, the connection between node A1 and node B1 has a weight W(A1, B1), which corresponds to a weight value or difference set 716 (such as W) stored in the memory array of memory macro 902. Figure 7A ). Memory macros 904, 906, and 908 are configured in a similar manner. When performing machine learning using neural network 900B, for example, weight data or difference set 716 (e.g., weight set W) in one or more of memory macros 902, 904, 906, and 908 is updated by processor and memory controller 920. Figure 7A According to some embodiments, one or more advantages described herein may be achieved in a neural network 900B that is implemented in whole or in part by one or more memory macros and / or memory devices.

[0413] Other configurations or numbers of elements in neural network 900B are also within the scope of the present disclosure.

[0414] Figure 9C is a schematic diagram of an integrated circuit (IC) device 900C in accordance with some embodiments.

[0415] IC device 900C is an embodiment of memory device 100 or Figure 1 memory device 900A of FIG. 1, and similar detailed descriptions are omitted. Figure 9A

[0416] IC device 900C includes one or more hardware processors 932, one or more memory devices 934 coupled to processors 932 by one or more buses 936. In some embodiments, one or more hardware processors 932 can be used as one or more components in encoder 102 of Figure 1 memory controller 920 of FIG. 1, and similar detailed descriptions are omitted. In some embodiments, one or more memory devices 934 can be used as memory circuit 102 of Figure 9A Figure 1 memory macro 110 of FIG. 1, or Figure 1 memory macro 902, 904, 906, or 908 of FIG. 1, and similar detailed descriptions are omitted. Figure 9A

[0417] ​​​In some embodiments, IC device 900C includes one or more further circuits, including, but not limited to, a cellular transceiver, a global positioning system (GPS) receiver, network interface circuitry for one or more of Wi-Fi, USB, Bluetooth, or the like. Examples of processor 932 include, but are not limited to, a central processing unit (CPU), a multi-core CPU, a neural processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), other programmable logic devices, a multimedia processor, an image signal processor (ISP), or the like. Examples of memory device 934 include one or more memory devices and / or memory macros described herein. In at least one embodiment, each of processors 932 is coupled to a corresponding memory device in memory device 934.

[0418] Because one or more memory devices 934 are CIM memory devices, various computations can be performed in the memory devices, which reduces the computational workload of the corresponding processor, reduces memory access time, and improves performance. In at least one embodiment, IC device 900C is a system-on-a-chip (SOC). In at least one embodiment, one or more advantages described herein can be achieved by IC device 900C.

[0419] Other configurations or quantities of components in IC device 900C are also within the scope of the present disclosure.

[0420] In some embodiments, at least a portion of method 500A is implemented as a standalone software application executed by a processor. In some embodiments, at least a portion of method 500A is implemented as a software application that is part of an additional software application. In some embodiments, at least a portion of method 500A is implemented as a plug-in to a software application. In some embodiments, at least a portion of method 500A is implemented as a software application that is part of a neural network tool. In some embodiments, at least a portion of method 500A is implemented as a software application used by a neural network tool.

[0421] In some embodiments, one or more of the operations of method 500A or 600 are not performed. Figures 1 to 9C The various logic circuits shown in FIG are for illustrative purposes. The embodiments of the present disclosure are not limited to specific logic circuits, and Figures 1 to 9C One or more of the logic circuits shown in the figure may be replaced with one or more corresponding logic circuits having different or equivalent functions. Similarly, the low or high logic values ​​of the various signals used in the above description are for illustrative purposes. The embodiments of the present disclosure are not limited to specific logic values ​​when the signals are activated and / or deactivated. It is within the scope of the various embodiments to select different logic values. Figures 1 to 9C It is also within the scope of various embodiments to select a different number of logic circuits.

[0422] It will be readily apparent to one skilled in the art that one or more of the disclosed embodiments satisfy one or more of the advantages set forth above. After reading the foregoing description, one skilled in the art will be able to effect various changes, equivalent substitutions, and various other embodiments broadly disclosed herein. Accordingly, the protection granted herein is subject only to the definitions contained in the appended claims and their equivalents.

[0423] One aspect of the present disclosure relates to a memory circuit. The memory circuit includes a memory cell array configured to store a first data set, the first data set including a first weight set or a second data set, the first data set corresponding to an exponent portion of a floating-point number, the second data set being a compressed version of the first weight set, the first weight set having a first data length, and the second data set having a second data length less than the first data length. In some embodiments, the CIM array includes a decoder coupled to the memory cell array and configured to generate a first output signal set in response to a first input signal set, the first data set, and a flag signal.

[0424] In some embodiments, the decoder includes: a first adder coupled to the memory cell array, configured to receive a first input signal set and a first base value of a first weight set, and to determine a first sum value in response to the first input signal set and the first base value of the first weight set, wherein the first base value of the first weight set is a minimum value in the first weight set. In some embodiments, the decoder further includes: a first multiplexer set coupled to the memory cell array, the first multiplexer set configured to receive a first difference signal set, a second difference signal set, and an address signal, and to output a first signal set in response to the address signal, wherein the first difference signal set is a second data set; the first difference signal set is equal to the difference between the first base value of the first weight set and the first weight set; the second difference signal set is equal to the difference between the second base value of the second weight set and the second weight set, wherein the second base value of the second weight set is a minimum value in the second weight set; and the address signal can be used by the first multiplexer set to select the first difference signal set or the second difference signal set as the first signal set. In some embodiments, the decoder further comprises: a first register set coupled to the first multiplexer set, the first register set configured to receive the first signal set and the first input signal set, and configured to output a second signal set, wherein the second signal set is a combination of the first input signal set and the first signal set and is a zero-padded version of the first signal set having the same length as the first weight set. In some embodiments, the first input signal set is a series of logic 0s. In some embodiments, the decoder further comprises: a second multiplexer set coupled to the memory cell array and the first register set, the second multiplexer set configured to receive the second signal set, the first weight set, and a flag signal, and configured to output a third signal set in response to the flag signal, wherein the flag signal can be used by the second multiplexer set to select either the second signal set or the first weight set as the third signal set. In some embodiments, the decoder further includes: a first adder set coupled to the memory cell array, the first adder, and the second multiplexer set, the first adder set being configured to receive the first sum and the third signal set and to generate a first output signal set, the first output signal set being the sum of the first sum and the third signal set, wherein each output signal in the first output signal set is equal to the sum of the first sum and a corresponding signal of the third signal set. In some embodiments, the memory circuit further includes: an encoder coupled to the in-memory computation array and being configured to receive the first weight set and to generate the first data set. In some embodiments, the in-memory computation array further includes: a multiplier set coupled to the memory cell array and being configured to multiply the mantissa portion of the first data set by the first input signal set.

[0425] Another aspect of the disclosure is directed to a memory circuit. The memory circuit includes a compute in-memory (CIM) array. In some embodiments, the CIM array includes a memory cell array to store a first set of exponent data and a first set of mantissa data, the first set of exponent data including a first set of weights or a second set of exponent data, the first set of exponent data corresponding to an exponent portion of a floating-point number, the second set of exponent data being a compressed version of the first set of weights, the first set of mantissa data being a second set of weights, and the first set of mantissa data corresponding to a mantissa portion of the floating-point number. In some embodiments, the CIM array further includes a first adder circuit coupled to the memory cell array and to generate a first set of output signals in response to a first set of input signals, the first set of exponent data, and a flag signal. In some embodiments, the CIM array further includes a set of multipliers coupled to the memory cell array and to generate a second set of output signals in response to the first set of input signals and the first set of mantissa data.

[0426] In some embodiments, the first output signal set is equal to the sum of the first input signal set and the first exponent data set; and the second output signal set is equal to the product of the first input signal set and the first mantissa data set. In some embodiments, the first adder circuit includes: a first adder coupled to the memory cell array, configured to receive the first input signal set and a first base value of the first weight set, and to determine a first sum value in response to the first input signal set and the first base value of the first weight set, wherein the first base value of the first weight set is a minimum value in the first weight set. In some embodiments, the first adder circuit further includes: a first multiplexer set coupled to the memory cell array, the first multiplexer set being used to receive a first differential signal set, a second differential signal set, and an address signal, and being used to output a first signal set in response to the address signal, wherein the first differential signal set is a second index data set; the first differential signal set is equal to the difference between the first base value of the first weight set and the first weight set; the second differential signal set is a third index data set; the second differential signal set is equal to the difference between the second base value of the second weight set and the second weight set, wherein the second base value of the second weight set is the minimum value in the second weight set; and the address signal can be used by the first multiplexer set to select the first differential signal set or the second differential signal set as the first signal set. In some embodiments, the first adder circuit further includes: a first register set coupled to the first multiplexer set, the first register set being configured to receive the first signal set and the first input signal set, and being configured to output a second signal set, wherein the second signal set is a combination of the first input signal set and the first signal set and is a zero-padded version of the first signal set having the same length as the first weight set. In some embodiments, the first input signal set is a series of logic 0s. In some embodiments, the first adder circuit further includes: a second multiplexer set coupled to the memory cell array and the first register set, the second multiplexer set being configured to receive the second signal set, the first weight set, and a flag signal, and being configured to output a third signal set in response to the flag signal, wherein the flag signal can be used by the second multiplexer set to select either the second signal set or the first weight set as the third signal set. In some embodiments, the first adder circuit further includes: a first adder set coupled to the memory cell array, the first adder, and the second multiplexer set, the first adder set being used to receive the first sum value and the third signal set, and being used to generate a first output signal set, the first output signal set being the sum of the first sum value and the third signal set, wherein each output signal in the first output signal set is equal to the corresponding sum of the first sum value and the corresponding signal of the third signal set.

[0427] Still another aspect of the present specification relates to a method for operating a memory circuit. In some embodiments, the method includes receiving a first set of weights by an encoder, the first set of weights being in a floating point format. In some embodiments, the method further includes compressing the first set of weights into a first set of differential signals by the encoder, the first set of weights including a first data length, the first set of differential signals including a second data length that is less than the first data length. In some embodiments, the method further includes performing a read operation of a memory cell array in a compute in-memory (CIM) array through a CIM array, thereby outputting the first set of differential signals, the CIM array being coupled to the encoder. In some embodiments, the method further includes generating a first set of output signals by a decoder in response to the first set of input signals and the first set of differential signals.

[0428] In some embodiments, the step of compressing the first weight set into a first differential signal set includes the following steps: receiving the first weight set by a controller; determining a first base value of the first weight set, the first base value being the minimum value of the first weight set; determining a first differential set from the first base value of the first weight set and the first weight set, the first differential set being equal to the difference between the first base value and the first weight set; determining a maximum differential value in the first differential set; and at least: in response to the maximum differential value in the first differential set being greater than a first threshold value, writing the first differential set to a memory cell array; or in response to the maximum differential value in the first differential set being less than a first threshold value, writing the first weight set to the memory cell array. In some embodiments, the step of generating a first output signal set in response to a first input signal set and a first differential signal set includes the following steps: determining a first sum value by a first adder set in response to a first base value of the first input signal set and a first weight set; selecting a first differential value set or a second differential value set as a first signal set by a first multiplexer set in response to an address signal, the second differential value set being a compressed version of the second weight set; and selecting the first signal set as a second signal set by a second multiplexer set in response to the flag in response to a determination that a flag is equal to a first value; and adding the first sum value to each differential value in the second signal set by a second adder set as a first output signal set; or selecting the first weight signal set as a second signal set by a second multiplexer set in response to the flag in response to a determination that the flag is not equal to the first value; and adding the first sum value to each weight value in the second signal set by a second adder set as a first output signal set.

[0429] The foregoing summarizes the features of several embodiments so that those skilled in the art may better understand the aspects of the present disclosure. Those skilled in the art will appreciate that they may readily use this disclosure as a basis for designing or modifying other processes and structures for implementing the same purposes and / or achieving the same advantages of the embodiments introduced herein. Those skilled in the art will also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and replacements may be made herein for such equivalent constructions without departing from the spirit and scope of the present disclosure.

Claims

1. A memory circuit, characterized in that: Include: A compute-in-memory (CIM) array comprising: A memory cell array for storing a first data set, the first data set including a first weight set or a second data set, the first data set being a plurality of exponent parts corresponding to a plurality of floating-point numbers, the second data set being a compressed version of the first weight set, the first weight set having a first data length, and the second data set having a second data length smaller than the first data length; and A decoder is coupled to the memory cell array and configured to generate a first output signal set in response to a first input signal set, the first data set, and a flag signal.

2. The memory circuit according to claim 1, wherein: The decoder contains: A first adder is coupled to the memory cell array, for receiving the first input signal set and a first base value of the first weight set, and for determining a first sum value in response to the first input signal set and the first base value of the first weight set, wherein the first base value of the first weight set is a minimum value in the first weight set.

3. The memory circuit according to claim 2, wherein: The decoder further comprises: a first multiplexer set coupled to the memory cell array, the first multiplexer set being configured to receive a first differential signal set, a second differential signal set, and an address signal, and to output a first signal set in response to the address signal; wherein the first difference signal set is the second data set; The first difference signal set is equal to a difference between the first base value of the first weight set and the first weight set; The second difference signal set is equal to a difference between a second base value of a second weight set and the second weight set, wherein the second base value of the second weight set is a minimum value in the second weight set; and The address signal is used by the first multiplexer set to select the first differential signal set or the second differential signal set as the first signal set.

4. The memory circuit according to claim 3, wherein: The decoder further comprises: a first register set coupled to the first multiplexer set, the first register set being configured to receive the first signal set and a first input signal set, and to output a second signal set; The second signal set is a combination of the first input signal set and the first signal set, and is a zero-padded version of the first signal set having the same length as the first weight set.

5. The memory circuit according to claim 4, wherein: The decoder further comprises: a second multiplexer set coupled to the memory cell array and the first register set, the second multiplexer set being configured to receive the second signal set, the first weight set, and the flag signal, and to output a third signal set in response to the flag signal; The flag signal is used by the second multiplexer set to select the second signal set or the first weight set as the third signal set.

6. A memory circuit, characterized in that: Include: A compute-in-memory (CIM) array comprising: a memory cell array for storing a first exponent data set and a first mantissa data set, the first exponent data set comprising a first weight set or a second exponent data set, the first exponent data set being a plurality of exponent parts of a plurality of corresponding floating-point numbers, the second exponent data set being a compressed version of the first weight set, the first mantissa data set being a second weight set, and the first mantissa data set being a plurality of mantissa parts of the plurality of corresponding floating-point numbers; a first adder circuit coupled to the memory cell array and configured to generate a first output signal set in response to a first input signal set, the first index data set, and a flag signal; and A multiplier set is coupled to the memory cell array and is configured to generate a second output signal set in response to the first input signal set and the first set of mantissa data.

7. The memory circuit according to claim 6, wherein: The first adder circuit comprises: A first adder is coupled to the memory cell array, for receiving the first input signal set and a first base value of the first weight set, and for determining a first sum value in response to the first input signal set and the first base value of the first weight set, wherein the first base value of the first weight set is a minimum value in the first weight set.

8. The memory circuit according to claim 7, wherein: The first adder circuit further comprises: a first multiplexer set coupled to the memory cell array, the first multiplexer set being configured to receive a first differential signal set, a second differential signal set, and an address signal, and to output a first signal set in response to the address signal; wherein the first difference signal set is the second index data set; The first difference signal set is equal to a difference between the first base value of the first weight set and the first weight set; The second difference signal set is a third index data set; The second difference signal set is equal to a difference between a second base value of a second weight set and the second weight set, wherein the second base value of the second weight set is a minimum value in the second weight set; and The address signal can be used by the first multiplexer set to select the first set of differential signals or the second set of differential signals as the first set of signals.

9. A method of operating a memory circuit, characterized in that: The method comprises the following steps: A first weight set is received by an encoder, wherein the first weight set is in a floating point format; compressing the first weight set into a first difference signal set by the encoder, wherein the first weight set includes a first data length, and the first difference signal set includes a second data length that is smaller than the first data length; performing a read operation on a memory cell array in an in-memory calculation array by an in-memory calculation array, thereby outputting the first difference signal set, wherein the in-memory calculation array is coupled to the encoder; and A decoder generates a first output signal set in response to a first input signal set and the first difference signal set.

10. The method according to claim 9, wherein The step of compressing the first weight set into the first difference signal set comprises the following steps: Receiving the first weight set by a controller; determining a first base value of the first weight set, where the first base value is a minimum value of the first weight set; determining a first difference set from the first base value of the first weight set and the first weight set, the first difference set being equal to a difference between the first base value and the first weight set; determining a maximum difference value in the first difference set; and At least: In response to the maximum difference value in the first difference set being greater than a first threshold value, writing the first difference set into the memory cell array; or In response to the maximum difference value in the first difference set being less than the first threshold value, writing the first weight set into the memory cell array.