Memory system, memory device and operating method thereof

By encoding the weights into a double sign bit format, reducing the number of low-resistance state memory cells, the high power consumption problem during ReRAM weight storage and access is solved, and a significant energy consumption reduction is achieved.

CN120279962APending Publication Date: 2025-07-08TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285834.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-02
Filing Date
2025-03-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The high power consumption problem during ReRAM weight storage and access is due to the low-resistance state unit consuming a large amount of current during access, resulting in high energy consumption.

Method used

The encoder circuit is used to encode the weights into a double-sign bit format, reducing the number of low-resistance state memory cells, and perform in-memory calculations and result decoding to reduce power consumption through the combination of encoder, memory array, accumulator circuit and analog-to-digital converter.

Benefits of technology

By reducing the number of low-resistance state memory cells, the current consumption of memory access is reduced, with an average energy consumption improving by about 1.33 times and a peak energy consumption improving by about 1.55 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279962A_ABST
    Figure CN120279962A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a memory system, a memory device, and a method of operating the memory device. A memory device includes an encoder circuit, a memory array, and an accumulator circuit. An encoder circuit converts the first bit of each weight to a sign bit according to flag data to generate an encoded weight. The memory array includes memory cells. The memory cells arranged in the same column store bits of encoding weights and flag data having the same index number. The memory array performs an in-memory computation (CIM) operation on the encoding weight and the input to generate a plurality of CIM results. Each of the plurality of CIM results corresponds to a column of the memory array. The accumulator circuit decodes the CIM result according to the flag data to generate a decoded CIM result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of electronic circuits, and more particularly, to storage systems, storage devices, and methods of operating the same. Background Art

[0002] Resistive random access memory (ReRAM) cells are commonly used as memory cells for in-memory computing devices of machine learning models. A ReRAM cell can be programmed to a low resistance state or a high resistance state to store the weight bits of a machine learning model. However, some methods of ReRAM weight storage and access face high power consumption, which is caused by a high incidence of low resistance state cells that consume a large amount of current when accessed. Summary of the Invention

[0003] An embodiment of the present invention provides a storage device, including: an encoder circuit configured to convert a first bit of each of a plurality of weights into a sign bit according to flag data to generate a plurality of encoded weights; a memory array including a plurality of memory cells, wherein the memory cells arranged in the same column are configured to store bits having the same index number of the encoded weights and the flag data, wherein the memory array is configured to perform an in-memory computing (CIM) operation on the encoded weights and a plurality of inputs to generate a plurality of CIM results, wherein each of the plurality of CIM results corresponds to a column of the memory array; and an accumulator circuit configured to decode the plurality of CIM results according to the flag data to generate a plurality of decoded CIM results.

[0004] Another embodiment of the present invention provides a storage system, including: an encoder circuit configured to convert a plurality of weights in two's complement form into a plurality of encoded weights in double sign bit form according to flag data, wherein each of the encoded weights has two sign bits; a memory array including a plurality of memory cells, wherein a first row of the memory cells is configured to store the flag data, and a plurality of second rows of the memory cells are configured to store the encoded weights, wherein the memory array is configured to perform a dot product operation on the encoded weights and a plurality of inputs on a word line of the memory array to generate a plurality of dot product results on a bit line of the memory array; an analog-to-digital converter circuit configured to convert the dot product results into a plurality of digital dot product results; and an accumulator circuit configured to decode the digital dot product results according to the flag data to generate a plurality of decoded dot product results.

[0005] Another embodiment of the present invention provides a method for operating a memory device, including: converting the magnitude bits of each of a plurality of weights into sign bits according to flag data to generate a plurality of encoded weights, wherein the flag data indicates the index numbers of the magnitude bits; performing a plurality of in-memory computing (CIM) operations on the encoded weights and a plurality of inputs transmitted to a memory array storing the weights to generate a plurality of CIM results; decoding the CIM results to generate a plurality of decoded CIM results; and accumulating the decoded CIM results to generate a final result. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Aspects of the present invention are best understood from the following detailed description when read with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, various components are not drawn to scale. In fact, for the sake of clarity of discussion, the dimensions of various components may be increased or decreased arbitrarily.

[0007] Figure 1 is a schematic diagram of a system according to some embodiments of the present disclosure.

[0008] Figure 2A is a schematic diagram of a two's complement format according to some embodiments of the present disclosure.

[0009] Figure 2B is a schematic diagram of a double sign bit format according to some embodiments of the present disclosure.

[0010] Figure 3 is according to some embodiments of the present disclosure showing corresponding to Figures 2A - 2B is a table of examples of the two's complement format and the double sign bit format.

[0011] Figure 4 is according to some embodiments of the present invention showing corresponding to Figure 2A and Figure 2B is an equation showing the relationship between the two's complement format and the double sign bit format.

[0012] Figure 5A and Figure 5B show examples of encoding a signed number in two's complement format into a double sign bit format corresponding to Figures 2A - 2B according to some embodiments of the present disclosure.

[0013] Figure 6A shows an example of generating Figure 1 the bit line current of a memory device in according to some embodiments of the present disclosure using the two's complement format.

[0014] Figure 6B shows an example of generating Figure 1An example of the bit line current of the memory device in

[0015] Figure 7 is a schematic diagram of an accumulator circuit of a memory device corresponding to Figure 1 , Figure 6A and Figure 6B according to some embodiments of the present disclosure.

[0016] Figure 8 is a schematic diagram of an accumulator circuit of a memory device corresponding to Figure 1 , Figures 6A - 6B and Figure 7 according to some embodiments of the present disclosure.

[0017] Figure 9 is a flowchart of a method for operating a memory device according to some embodiments of the present disclosure.

[0018] Figure 10 is a flowchart of a method for operating a memory device according to some embodiments of the present disclosure.

[0019] Figure 11 is a flowchart of a method for operating a memory device according to some embodiments of the present disclosure.

[0020] Figure 12 is a flowchart of a method for operating a memory device according to some embodiments of the present disclosure. Detailed Description

[0021] The following disclosure provides many different embodiments or examples to implement different features of the present invention. Specific examples of components, materials, values, steps, arrangements, etc. are described below to simplify the present disclosure. Of course, these are only examples and are not intended to be limiting. Other components, materials, values, steps, arrangements, etc. can be expected. For example, in the following description, forming the first component above or on the second component can include embodiments where the first component and the second component are in direct contact, and can also include embodiments where additional components are formed between the first component and the second component such that the first component and the second component are not in direct contact. Moreover, the present invention may repeat reference numerals and / or letters in various examples. This repetition is for the sake of brevity and clarity, but it does not itself indicate the relationship between the various embodiments and / or configurations discussed.

[0022] In addition, for ease of description, spatial relative terms such as "under", "below", "lower", "above", "upper", etc. may be used herein to describe the relationship of one element or component to another (or other) element or component as shown in the figures. In addition to the orientations shown in the figures, the spatial relationship terms are intended to include different orientations during the use or operation of the device. The device may be positioned in other ways (rotated 90 degrees or in other orientations), and the spatial relationship descriptors used herein may be interpreted accordingly. The terms mask, photolithographic mask, optical mask, and reticle are used to refer to the same item.

[0023] The terms throughout the following description and claims generally have the ordinary meanings that are clearly defined in the art or in the specific context in which each term is used. Those of ordinary skill in the art will understand that a component or process may be referred to by different names. The numerous different embodiments detailed in this specification are merely illustrative and do not limit the scope and spirit of the present disclosure or any of the example terms in any way.

[0024] It is noted that the words used herein to describe various elements or processes, such as "first" and "second", are intended to distinguish one element or process from another. However, the elements, processes, and their order should not be limited by these words. For example, the first element may be called the second element, and the second element may similarly be called the first element, without departing from the scope of the present disclosure.

[0025] In the following discussion and claims, the words "comprising", "including", "containing", "having", "involving", etc. should be understood as open-ended, i.e., should be interpreted as including but not limited to. As used herein, the phrase "and / or" is not mutually exclusive, but includes any relevant listed items and all combinations of one or more of the relevant listed items.

[0026] In the field of computing, signed numbers are typically encoded in two's complement format. However, for machine learning applications, weights encoded in two's complement format may contain a large number of "1" bits, which may increase the power consumption of the memory storing the weights. For example, the weights of a machine learning model typically contain a large number of small-magnitude negative numbers, and the two's complement format of small-magnitude negative numbers contains a large number of "1" bits. In this case, a memory such as a resistive random access memory will consume a large amount of power because the current for accessing a "1" bit is greater than the current for accessing a "0" bit. An encoding format with fewer occurrences of the "1" bit helps reduce the power consumption of the memory.

[0027] Now referring to Figure 1 。 Figure 1 is a schematic diagram of system 10 according to some embodiments of the present disclosure. As Figure 1As shown, system 10 includes a storage device 100 and a control circuit 200 coupled to the storage device 100. In some embodiments, the control circuit 200 is a processor, such as a central processing unit (CPU) and / or a microcontroller unit (MCU). In some embodiments, the storage device 100 is configured as a in-memory computing (CIM) device for multiplication and accumulation (MAC) operations. For example, the storage device 100 includes an encoder 110, a memory writer 120, a memory array 130, a multiplexer (MUX) circuit 140, an analog-to-digital converter (ADC) circuit 150, an accumulator circuit 160, and a register circuit 170. In some embodiments, the encoder 10 encodes the weights W according to the flag data FG indicating the encoding scheme to generate encoded weights WE. The memory writer 120 writes the encoded weights WE and the flag data FG into the memory array 130. The memory array 130 multiplies the weights WE and the input IN and generates an output. The multiplexer circuit 140 transmits the output to the analog-to-digital converter circuit 150. The analog-to-digital converter circuit 150 converts the output into a digital signal as a partial MAC result pMACV. The register circuit 170 retrieves the flag data FG from the memory array 130. The accumulator circuit 160 decodes the partial MAC result pMACV according to the flag data FG in the register circuit 170 to generate a decoded partial MAC result pMACV D . In some embodiments, the accumulator circuit 160 processes multiple decoded partial MAC results pMACV D and accumulates them to generate a MAC result MACV. In some embodiments, the MAC result MACV corresponds to the output of a computing node of a machine learning model (such as a neural network).

[0028] For ease of illustration, the encoder 110 is coupled to the memory writer 120. The memory writer 120 is also coupled to the memory array 130. The memory array 130 is also coupled to the multiplexer circuit 140 and the register circuit 170. The multiplexer circuit 140 is also coupled to the analog-to-digital converter circuit 150. The analog-to-digital converter circuit 150 and the register circuit 170 are also coupled to the accumulator circuit 160.

[0029] The memory array 130 includes memory cells MC, word lines WL, and bit lines BL. Each memory cell MC is located at the intersection of a row and a column in the memory array 130. As Figure 1 shown, the memory cells MC arranged in the same row are coupled to the same word line WL. The memory cells MC arranged in the same column are coupled to the same bit line BL.

[0030] In some embodiments, the memory array 130 can be a non-volatile memory array. For example, the memory cell MC is operable to store one bit of data therein, i.e., "1" or "0". In some embodiments, the memory cell MC is a resistive random access memory (ReRAM) cell having a low resistance state (LRS) or a high resistance state (HRS). In some embodiments, the memory cell MC storing the bit "1" has a low resistance state, while the memory cell MC storing the bit "0" has a high resistance state. In some embodiments, the memory cell MC having a low resistance state consumes more power than the memory cell MC having a high resistance state.

[0031] In some embodiments, to reduce power consumption, the encoder 110 encodes the weight W to reduce the memory cells MC having a low resistance state. Specifically, the encoder 110 encodes the weight into an encoded weight WE, where the number of occurrences of the bit "1" in the encoded weight WE is lower than the number of occurrences of the bit "1" in the weight W. Therefore, the number of memory cells MC having a low resistance state in the memory cells MC storing the encoded weight WE is less than the number of memory cells MC having a low resistance state in the memory cells MC storing the weight W. In other words, the power consumption of the memory cells MC storing the encoded weight WE is less than the power consumption of the memory cells MC storing the weight W. In some embodiments, the weight W is in two's complement format. In some embodiments, the encoded weight WE is in double sign bit (DSB) format. In some embodiments, when the flag data FG indicates the format of the weight W (i.e., the weight W and the encoded weight WE have the same format), the encoder 110 does not encode the weight W. For example, the encoder 110 directly outputs the weight W in two's complement format as the encoded weight WE without encoding the weight W according to the flag data FG indicating the two's complement format.

[0032] Details regarding the DSB format and the two's complement format for expressing the encoded weight WE will be described below with reference to Figures 2A - 2B 、 Figures 3 - 4 and Figures 5A - 5B 。

[0033] Now refer to Figure 2A and Figure 2B 。 According to some embodiments of the present disclosure, Figure 2A is a schematic diagram of the two's complement format, Figure 2B is a schematic diagram of the double sign bit format.

[0034] The two's complement format uses binary digits to represent signed numbers. In the two's complement format, the most significant bit (MSB) is the sign bit corresponding to the sign, indicating whether the signed number is positive or negative. For example, when the most significant bit is 1, the signed number is marked as negative; when the most significant bit is 0, the signed number is marked as positive.

[0035] In the two's complement format, the bits from the second most significant bit to the least significant bit (LSB) are the magnitude bits representing the magnitude of the signed number. As Figure 2A shown, in the "(N + 1)-bit" two's complement format ("N" being an integer), the bits with bit values from 2 N-1 to 2 0 are used to represent the magnitude of the signed number, and the bit with bit value -2 N is used to represent the sign of the signed number. Figure 2A The decimal value DV of the corresponding signed number can be obtained through the following function, where each bit from X0 to X N is either "1" or "0".

[0036]

[0037] Similarly, Figure 2B the double sign bit format in Figure 2B uses binary digits to represent signed numbers. The difference between the double sign bit format and the two's complement format is that the double sign bit format uses two bits to represent the sign. Compared with the two's complement format, the double sign bit format also uses a second sign bit different from the first sign bit (MSB) to represent the sign. For example, as N and M shown, in the "(N + 1)-bit" double sign bit format, two bits with bit values of -2 Figure 2B and -2 N are used to represent the sign, where "M" is an integer less than "N".

[0038]

[0039] In some embodiments, the weight W includes numbers that cannot be represented by the double sign bit format. For example, values greater than or equal to "2 N - 2 M " cannot be represented by the double sign bit format. In some embodiments, weights including these non - representable values will not be encoded.

[0040] Now refer to Figure 3 . Figure 3 is shown according to some embodiments of the present disclosure corresponding to Figures 2A - 2BTable 300 of examples of two's complement format and double sign bit format.

[0041] For example, Table 300 shows multiple numerical values "-8 to 7" and their corresponding binary numbers in 4-bit two's complement format, 4-bit double sign bit format (where "M" is 1), and 4-bit double sign bit format (where "M" is zero). Table 300 also shows the occurrence of zero corresponding to each format, and the more the occurrence of zero, the better the power consumption reduction performance.

[0042] The "Number of Zeros" column shows the number of bit "0" corresponding to the binary number of each numerical value in two's complement format or double sign bit format.

[0043] The symbol "×" indicates that the corresponding numerical value cannot be represented for the double sign bit format.

[0044] The "Distribution" column shows the distribution of numerical values. According to some embodiments, the distribution of numerical values corresponds to the computing nodes of a machine learning model.

[0045] The second column from the right shows the product of the distribution corresponding to the two's complement format and the number of bit "0". The rightmost column shows the product of the distribution corresponding to the double sign bit format and the number of bit "0". The "SUM" row shows the sum of the values in these two columns. These sums represent the occurrence of zero corresponding to the two's complement format and the double sign bit format.

[0046] The "Improvement (%)" row shows the percentage improvement in the occurrence of zero of the double sign bit format relative to the two's complement format. As shown in Table 300, in terms of the occurrence of zero, the double sign bit format is superior to the two's complement format. In other words, in the case of Table 300, the double sign bit format consumes less power than the two's complement format.

[0047] Now refer to Figure 4 . Figure 4 Shows an equation corresponding to the relationship between the two's complement format and the double sign bit format according to some embodiments of the present invention. Figure 2A and Figure 2B of the two's complement format and the double sign bit format.

[0048] For illustration, the two's complement format of a signed number corresponding to Figure 2A is equal to the double sign bit format plus a conversion term. According to some embodiments, the signed number in two's complement format can be encoded into the double sign bit format according to the conversion term.

[0049] The following shows the Figure 4 proof of the equivalence shown.

[0050]

[0051] As shown in the above formula, it is proved that the two's complement format is equal to the double sign-bit format plus a conversion term. The following paragraphs will refer to Figures 5A - 5B to discuss the details of encoding a signed number into the double sign-bit format according to the conversion term.

[0052] Now refer to Figure 5A and Figure 5B . Figure 5A and Figure 5B are schematic diagrams of examples of encoding a signed number in two's complement format into a double sign-bit format corresponding to Figures 2A - 2B .

[0053] To encode a signed number X (2's) in two's complement format into a double sign-bit format with the "(M - 1)"-th bit from the LSB side as the sign bit, first convert the bit value 2 (2's) of the signed number X M to -2 M to generate a signed number X (DSB) in double sign-bit format. Then, add the conversion term "(2 M+1 )(X M )" to the signed number X (DSB) to generate a signed number Y (DSB) in double sign-bit format. The signed number Y (DSB) is the result of encoding the signed number X (2's) from two's complement format into double sign-bit format. The signed number X (2's) and the signed number Y (DSB) correspond to the same numerical value.

[0054] In the example shown in Figure 5A , the signed number X (2's) has 8 bits "11111111", corresponding to the numerical value "-1". To generate a signed number Y (DSB) with the third bit from the LSB side as the sign bit (i.e., "M" equals 4), convert the bit value 2 2 of the third bit to -2 2 to generate a signed number X (DSB) . Then, add the bit "1000" corresponding to the conversion term "(2 M+1 )(X M )" with "M" equal to 4 to the signed number X (DSB) to generate a signed number Y (DSB) with bits "00000111". According to the function corresponding to the double sign-bit format, the bits "00000111" correspond to the numerical value "-2 2 +2 1 +2 0 =-1".

[0055] Similarly, in Figure 5B the example shown, the signed number X (2's) has 8 bits of "10011001", corresponding to the numerical value "-103". To generate a signed number Y (DSB) with the third bit from the LSB side as the sign bit (i.e., "M" equals 4), the bit value 2 2 of the third bit is 2 converted to -2 (DSB) to generate the signed number X M+1 . Then, the conversion term corresponding to "M" equals 4 of the conversion term "(2 M )(X (DSB) )", which is the bit "0000", is added to the signed number X (DSB) to generate the signed number Y with the bit "10011001" 7 . According to the function corresponding to the double sign bit format, the bit "10011001" corresponds to the numerical value "-2 4 +2 3 +2 0 =-103".

[0056] In some embodiments, after encoding the weight W into the encoded weight WE, the encoded weight WE is stored in Figure 1 the memory array 130 to generate the bit line current IBL.

[0057] Now refer to Figure 1 , Figure 6A and Figure 6B . Figure 6A shows an example of generating the bit line current IBL of the storage device 100 in Figure 1 in two's complement format according to some embodiments of the present disclosure. Figure 6B shows an example of generating the bit line current of the storage device 100 in Figure 1 using the double sign bit format according to some embodiments of the present disclosure.

[0058] With respect to Figure 1 , Figures 2A - 2B , Figures 3 - 4 and Figures 5A - 5B of the embodiments, for ease of understanding, Figures 5A - 5B the similar elements in

[0059] are designated with the same reference numerals. Figure 6A and Figure 6BAs shown, the memory array 130 stores flag data FG and multiple encoded weights WE (encoded weights WE1 to WEK). In some embodiments, the flag data FG is stored in the first row of memory cells MC. In some embodiments, the encoded weights WE1 to WEK are stored in the second to the "K"th row of memory cells MC. In some embodiments, each bit of the flag data FG and the encoded weights WE1 to WEK is stored in the column of the memory cell MC corresponding to the index number of that bit. For example, each of the flag data FG and the encoded weights WE1 to WEK has "N + 1" bits. The bits from the LSB to the MSB have index numbers "0" to "N" respectively. The bits with index numbers "0" to "N" are stored in the "N + 1" columns from the rightmost column to the leftmost column of the memory cell MC respectively.

[0060] In some embodiments, in the flag data FG, bit 1 represents the sign bit. Specifically, the index of bit 1 in the flag data FG represents the index of the sign bit in the encoded weights WE. For example, in Figure 6A the flag data FG, the bit 1 with an index number of 8 (i.e., the 8th bit starting from the LSB side) represents that the corresponding bit with an index number of 8 in each encoded weight WE is the sign bit.

[0061] In some embodiments, the MSB of the flag data FG is bit 1 and the other bits are bit 0, indicating that the encoded weights WE are in two's complement format. In some embodiments, the MSB and the "M"th bit starting from the LSB side of the flag data FG are bit 1 and the other bits are bit 0, indicating that the encoded weights WE are in double sign bit format, where the MSB is the first sign bit and the "M"th bit starting from the LSB side is the second sign bit.

[0062] For example, Figure 6A the flag data FG with the bit "10000000" in Figure 6A represents the two's complement format. The flag data FG with the bit "10000100" in

[0063] represents the two's complement format, where the third bit starting from the LSB side is the second sign bit.

[0064] In some embodiments, multiple inputs IN (inputs IN1 to INK) are respectively input to the rows of the memory cells MC storing the encoded weights WE1 to WEK through word lines WL coupled to the rows. The memory array 130 performs a CIM operation (e.g., MAC operation) on the inputs IN and the encoded weights WE1 to WEK. In some embodiments, the memory array 130 multiplies the inputs IN1 to INK and the encoded weights WE1 to WEK respectively.

[0064] In some embodiments, the results corresponding to the same bit values of the multiplications of IN1 to INK and the encoded weights WE1 to WEK are accumulated and output through the corresponding bit lines BL.

[0065] According to various embodiments, the zero occurrence rate of the two-sign bit format can be greater than the zero occurrence rate of the two's complement format. Accordingly, the magnitude of the current IBL corresponding to the accumulation result on the bit line BL of the two-sign bit format can be less than that of the two's complement format. For example, although Figure 6A the numerical values of the encoded weights WE1 to WEK stored in the memory array 130 in Figure 6B are equal to the numerical values of the encoded weights WE1 to WEK stored in the memory array 130 in Figure 6A the current IBL corresponding to the memory array 130 of the two's complement format can be greater than Figure 6B the current IBL corresponding to the memory array 130 of the two-sign bit format.

[0066] After the current IBL is generated, Figure 1 the multiplexer circuit 140 of

[0067] transmits the current IBL to the analog-to-digital converter circuit 150. The analog-to-digital converter circuit 150 performs analog-to-digital conversion on the current IBL to generate a partial MAC result pMACV.

[0068] The partial MAC result pMACV is generated based on the current IBL. Accordingly, although the numerical values of the encoded weights WE1 to WEK stored in the memory array 130 corresponding to the two's complement format are equal to the numerical values of the encoded weights WE1 to WEK stored in the memory array 130 corresponding to the two-sign bit format, the partial MAC result pMACV corresponding to the memory array 130 of the two's complement format may be greater than the partial MAC result pMACV corresponding to the memory array 130 of the two-sign bit format.

[0068] In some embodiments, the accumulator circuit 160 decodes the partial MAC result pMACV in the two-sign bit format into a decoded partial MAC result pMACV D . The decoded partial MAC result pMACV D corresponds to the sign of the partial MAC result pMACV multiplied by the bit value corresponding to the partial MAC result pMACV.

[0069] In some embodiments, the register circuit 170 includes a plurality of flip-flops. In some embodiments, the flip-flops are D-type flip-flops. In some embodiments, the bits of the flag data FG are respectively stored in the flip-flops.

[0070] Now refer to Figure 1 , Figure 6A , Figure 6B and Figure 7 . Figure 7 is corresponding to Figure 1 , Figure 6A andFigure 6B Schematic diagram of the accumulator circuit 160 of the storage device 100. Relative to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B and Figures 6A - 6B For the embodiments of, for ease of understanding, Figure 7 Similar elements in are designated with the same reference numerals.

[0071] In some embodiments, the accumulator circuit 160 includes one or more recovery circuits 710 and a shifter adder circuit 720. The recovery circuit 710 receives partial MAC results pMACV corresponding to a column of memory cells MC. For example, the analog-to-digital converter circuit 150 outputs a plurality of partial MAC results pMACV, including partial MAC results pMACV[0] to pMACV[N]. The partial MAC results pMACV[0] to pMACV[N] correspond to "N" bits of the flag data FG and the encoded weight WE from the LSB to the MSB. For example, the partial MAC result pMACV[i] corresponds to the "i"th bit of the flag data FG and the encoded weight WE, where "i" is one of the integers from "1" to "N". The partial MAC result pMACV[i] is transmitted through the bit line BL of the column corresponding to the "i"th bit of the memory cell MC (i.e., the "i"th column from the right).

[0072] In some embodiments, the recovery circuit 710 decodes the partial MAC results pMACV according to the flag data FG in the register circuit 170. For example, to decode the partial MAC result pMACV[i], when the "i"th bit FG[i] of the flag data FG in the register circuit 170 is bit 0, the recovery circuit 710 directly outputs the partial MAC result pMACV[i] as the decoded partial MAC result pMACV D [i]. When the "i"th bit FG[i] of the flag data FG in the register circuit 170 is bit 1, the recovery circuit 710 inverts the partial MAC result pMACV[i], and outputs the inverted result of the partial MAC result pMACV[i] as the decoded partial MAC result pMACV D [i].

[0073] In some embodiments, the shifter adder circuit 720 shifts and combines a plurality of decoded partial MAC results pMACV D (e.g., the decoded partial MAC results pMACV D [0] to pMACV D [N]) according to the corresponding bit values in two's complement format to generate the MAC result MACV. For example, the shifter adder circuit 720 shifts each decoded partial MAC result pMACV D[i]Shift by “i” bits and combine all shifted and decoded partial MAC results pMACV D [i]to generate the MAC result MACV.

[0074] Now refer to Figure 1 、 Figure 6A 、 Figure 6B 、 Figure 7 and Figure 8 。 Figure 8 is a schematic diagram of the accumulator circuit 160 of the storage device 100 corresponding to Figure 1 、 Figure 6A - also 6B and Figure 7 according to some embodiments of the present disclosure. For ease of understanding, Figure 1 、 Figures 2A - 2B 、 Figures 3 - 4 、 Figures 5A - 5B 、 Figures 6A - 6B and Figure 7 in the embodiments, for ease of understanding, Figure 8 similar elements in

[0075] are designated with the same reference numerals. Figure 8 As shown in

[0076] In some embodiments, the accumulator circuit 160 includes switches 811, 812, 821, and 822 and an inverter 830. In some embodiments, switches 811 and 821 are n-type transistors. Switches 812 and 822 are p-type transistors. In some embodiments, switches 811 and 812 are used as transmission gates. Switches 821 and 822 are used as transmission gates. For example, the source terminals of switches 811 and 812 and the input terminal of the inverter 830 are coupled to the analog-to-digital converter circuit 150 to receive the partial MAC result pMACV[i]. The control terminals of switches 812 and 821 are coupled to the register circuit 170 to receive the bit FG[i]. The control terminals of switches 811 and 822 are coupled to the register circuit 170 to receive the inverted result of the bit FG[i] D [i]output to the shifter adder circuit 720.

[0077] The inverter 830 inverts the partial MAC result pMACV[i]. When the bit FG[i] is bit 0, the switches 811 and 812 are turned on. When the bit FG[i] is bit 1, the switches 811 and 812 are turned off. Conversely, when the bit FG[i] is bit 1, the switches 821 and 822 are turned on. When the bit FG[i] is bit 0, the switches 821 and 822 are turned off. Thus, when the bit FG[i] is bit 0, the recovery circuit 710 directly outputs the partial MAC result pMACV[i] as the decoded partial MAC result pMACV D [i]. When the bit FG[i] is bit 1, the recovery circuit 710 outputs the inverted result of the partial MAC result pMACV[i] as the decoded partial MAC result pMACV D [i].

[0078] Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 8 are given for illustrative purposes only. Various embodiments are within the scope of the present disclosure. For example, in some embodiments, the register circuit 170 further includes an inverter for generating the inverted result .

[0079] Now refer to Figure 9 . Figure 9 is a flowchart of a method 900 for operating a memory device corresponding to, for example, a memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 8 in accordance with some embodiments of the present disclosure. It should be understood that additional operations may be provided before, during, and after the processes shown in Figure 9 , and for additional embodiments of the method 900, some of the operations described below may be replaced or deleted. The order of the operations / processes may be interchanged. Throughout the various views and exemplary embodiments, the same reference numerals are used to designate the same elements. The method 900 includes operations 901 to 903, which will be described below with reference to a memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 8 .

[0080] In some embodiments, the encoder 110 performs an encoding operation according to the method 900.

[0081] In operation 901, the control circuit 200 determines the value of "M" (the index number of the second sign bit) in the double sign bit format to avoid the inability to represent the weight W. Specifically, the control circuit 200 scans all the weights W (e.g., the weights of the kernels of a convolutional neural network) to calculate the MAC result MACV to find the maximum value of "M" that satisfies the limitation that the numerical value of each weight W is less than "2 N -2 M ".

[0082] In operation 902, the encoder 110 encodes the weight W into the double sign bit format by changing the bit with the index number of "M" from the numerical bit to the sign bit.

[0083] In operation 903, the encoder 110 marks the weight W as the encoded weight and records the index number of the second sign bit in the flag data FG. In some embodiments, when there is no "M" that satisfies the limitation that the numerical value of each weight W is less than "2 N -2 M ", the encoder 110 does not perform encoding and clears the valid state of the flag data FG.

[0084] Now refer to Figure 10 . Figure 10 is a flowchart of a method 1000 for operating a memory device corresponding to, for example, a memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 9 in accordance with some embodiments of the present disclosure. It should be understood that additional operations may be provided before, during, and after the process shown in Figure 10 , and for additional embodiments of the method 1000, some of the operations described below may be replaced or deleted. The order of the operations / processes may be interchanged. Throughout the various views and exemplary embodiments, the same reference numerals are used to designate the same elements. The method 1000 includes operations 1001 to 1007, which will be described below with reference to the memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 9 .

[0085] In some embodiments, encoder 110 performs encoding operations according to method 1000. In some embodiments, control circuit 200 performs operations 1001-1007 to determine the optimal index number of the second sign bit of encoder 110. In some embodiments, control circuit 200 performs operations 1001 to 1007 and provides input IN, weight W, and flag data FG to storage device 100 to perform MAC operations.

[0086] In operation 1001, the maximum value W among the numerical values of all weights W used to calculate the MAC result MACV is determined. max .

[0087] In operation 1002, the value of "M" (the index number of the second sign bit) for testing is optimized for the double sign bit format. In some embodiments, the value of "M" for testing is set to the maximum value among all possible values. In some embodiments, the value of "M" for testing is initially set to the index number of the MSB minus two. After operation 1002, tests are performed using operations 1003 to 1007 to determine whether the tested value of "M" is the optimal value of the index number of the second sign bit.

[0088] In operation 1003, it is determined whether the maximum W max is greater than or equal to "2 N -2 M " to avoid overflow. In other words, it is determined whether the maximum W max is greater than or equal to "2 N -2 M " to ensure that the maximum W max can be represented by the double sign bit format. In some embodiments, a comparison is performed between the maximum W max and "2 N -2 M -1" to determine whether the maximum W max is greater than "2 N -2 M -1" to avoid overflow.

[0089] In some embodiments, after determining that the maximum value W max is not greater than or equal to "2 N -2 M " (or not greater than "2 N -2 M -1"), operation 1004 is performed. In operation 1004, the weight W is converted into an encoded weight WE in the double sign bit format, where the index number of the second sign bit is "M".

[0090] In operation 1005, the sparsity of the encoding weight WE is determined, and a comparison between the sparsity and the maximum sparsity is performed. In some embodiments, the sparsity of the encoding weight WE is a number proportional to the occurrence of zeros, as described above with reference to Figure 3 as described. In some embodiments, the first determined sparsity of the encoding weight WE corresponding to the first test value "M" is determined as the maximum sparsity in the first iteration of the test. In some embodiments, it is determined whether the sparsity of the encoding weight WE is greater than the maximum sparsity to ensure an increase in the occurrence rate of zeros.

[0091] In some embodiments, operation 1006 is performed after it is determined in operation 1005 that the sparsity of the encoding weight WE is greater than the maximum sparsity. In operation 1006, the maximum sparsity is updated with the sparsity of the encoding weight WE. The temporary best index number of the second symbol bit is updated with the tested "M" value.

[0092] In operation 1007, it is determined whether all possible values of "M" have been tested. In some embodiments, in operation 1003, it is determined that the maximum value W max is greater than or equal to "2 N -2 M "(or greater than "2 N -2 M -1") and then operation 1007 is performed. In some embodiments, operation 1007 is performed after it is determined in operation 1005 that the sparsity of the encoding weight WE is not greater than the maximum sparsity.

[0093] When the determination in operation 1007 indicates that all possible values of "M" have been tested, the temporary best index number of the second symbol bit is determined as the best index number of the second symbol bit.

[0094] When the determination in operation 1007 indicates that all possible values of "M" have not been tested, operation 1002 is repeated. In the repeated operation 1002, a new "M" is determined for testing among all the possible values of "M" that have not been tested. In some embodiments, the value of the tested "M" is set to the value of the previously tested "M" minus one.

[0095] Now refer to Figure 11 . Figure 11 is a flowchart of a method 1100 for operating a memory device 100 corresponding to, for example, corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 10 in accordance with some embodiments of the present disclosure. It should be understood that it can be in Figure 11Provide additional operations before, during, and after the process shown, and for additional embodiments of the method 1100, some of the operations described below may be replaced or deleted. The order of operations / processes may be swapped. Throughout the various views and exemplary embodiments, the same reference numerals are used to designate the same elements. The method 1100 includes operations 1101 to 1107, which will be described below with reference to the memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 10 .

[0096] In some embodiments, the memory device 100 performs a MAC operation on the encoded weight WE and the input IN according to the method 1100 and generates a MAC result.

[0097] In operation 1101, the encoder 110 receives flag data FG. In some embodiments, the flag data FG includes bits FG[0] to FG[N], where the bit corresponding to the sign bit has a value of "1" and the bit corresponding to the magnitude bit has a value of "0". The memory writer 120 writes the bits FG[i] of the flag data FG to the corresponding columns of the corresponding bit lines BL, as shown in Figure 6A and 6B . The register circuit 170 reads the bits FG[i] corresponding to each bit line BL and stores the bits FG[i].

[0098] In operation 1102, the memory array 130 performs a bitwise dot product on the input IN and the bits of the encoded weight W. The memory array 130 performs a dot product on the input IN and the bits of the encoded weight W stored in the same column, and generates a current IBL on the corresponding bit line BL as the dot product result.

[0099] The multiplexer circuit 140 and the analog-to-digital converter circuit 150 generate a partial MAC result pMACV[i] for each bit line BL. Each partial MAC result pMACV[i] is generated according to the current IBL on the corresponding bit line BL.

[0100] In operation 1103, when the bit FG[i] corresponding to the bit line BL is bit 1, the restoration circuit 710 inverts the sign of the partial MAC result pMACV[i] of the bit line BL. Conversely, when the bit FG[i] corresponding to the bit line BL is bit 0, the restoration circuit 710 does not invert the sign of the partial MAC result pMACV[i] of the bit line BL.

[0101] In operation 1104, the recovery circuit 710 sends the partial MAC result pMACV[i] to the shifter adder circuit 720. The shifter adder circuit 720 accumulates a plurality of partial MAC results pMACV[i] to generate the MAC result MACV.

[0102] Now refer to Figure 12 . Figure 12 is a flowchart of a method 1200 for operating a memory device corresponding to, for example, a memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 11 in accordance with some embodiments of the present disclosure. It should be understood that additional operations may be provided before, during, and after the process shown in Figure 12 , and for additional embodiments of the method 1200, some of the operations described below may be replaced or deleted. The order of operations / processes may be exchanged. Throughout the various views and exemplary embodiments, the same reference numerals are used to designate the same elements. The method 1200 includes operations 1201 to 1204, which will be described below with reference to a memory device 100 corresponding to Figure 1 , Figures 2A - 2B , Figures 3 - 4 , Figures 5A - 5B , Figures 6A - 6B and Figures 7 - 11 .

[0103] In operation 1201, the encoder 110 encodes the weight W in the form of double sign bits. The encoder 110 converts the magnitude bits of each weight W into sign bits according to the flag data FG to generate the encoded weight WE. The flag data indicates the index number of the magnitude bits to be converted.

[0104] In operation 1202, the memory array 130, the multiplexer circuit 140, and the analog-to-digital circuit perform a CIM operation on the encoded weight WE and the input IN transmitted to the memory array 130 to generate a CIM result (e.g., a partial MAC result pMACV).

[0105] In operation 1203, the accumulator circuit 160 decodes the CIM result according to the flag data FG to generate a decoded result (e.g., the decoded partial MAC results pMACVD[0] to pMACVD[N]).

[0106] In some embodiments, the accumulator circuit 160 inverts a CIM result (e.g., partial MAC result pMACV[i]) in the CIM results according to a corresponding bit (e.g., bit FG[i]) of the flag data FG being one, to generate one of the decoded CIM results (e.g., decoded partial MAC result pMACVD[i]).

[0107] In operation 1204, the accumulator circuit 160 accumulates the decoded results to generate a final result (e.g., MAC result MACV).

[0108] In some embodiments, the shifter adder circuit 720 performs a shift and add operation on each decoded CIM result according to the position of the column to generate a final result. For example, in the shift and add operation, the shifter adder circuit 720 shifts the decoded partial MAC result pMACV D [i] according to the position of the column corresponding to pMACV D [i] (e.g., the "(i + 1)"-th column starting from the right) to generate a shifted decoded partial MAC result pMACV D [i]. Then, the shifter adder circuit 720 adds the shifted decoded partial MAC result pMACV D [i] to a temporary result. The temporary result is the result of adding the shifted decoded partial MAC results. After performing the shift and add operations on all the decoded partial MAC results, the temporary result is determined as the MAC result MACV.

[0109] In some embodiments, the encoder 110 converts the magnitude bits of each weight W having a test index number into sign bits to generate test encoded weights. According to the sparsity of the test encoded weights, the test index number is determined as the index number of the bits to be converted into sign bits in the double sign bit format.

[0110] In some embodiments, the control circuit 200 determines the index number of the bits to be converted into sign bits according to the maximum value of the weight W. For example, the index number is configured to satisfy the limitation that the value "2 N -2 M " is less than the maximum value of the weight W. "M" is the index number of the bits to be converted into sign bits, and "N" is the index number of the MSB of the weight W.

[0111] In some embodiments, the control circuit 200 generates flag data FG, where the bits with a value of "1" represent the indices of the sign bits, and the bits with a value of "0" represent the indices of the magnitude bits.

[0112] In summary, the storage device and its operation method help to reduce power consumption in memory access. Weights encoded in a two-sign-bit format have a higher zero occurrence rate. Therefore, the memory array will have more cells in a high-resistance state, consuming less current for access. Compared with some methods, the average power consumption of 8-bit integer and 8-bit binary floating-point weights is improved by about 1.33 times and 1.29 times respectively. The peak power consumption of 8-bit integer and 8-bit binary floating-point weights is improved by about 1.55 times and 1.39 times respectively.

[0113] A storage device is also disclosed. The storage device includes: an encoder circuit configured to convert the first bit of each weight into a sign bit according to flag data to generate an encoded weight; a memory array including memory cells, wherein the memory cells arranged in the same column are configured to store bits of the encoded weight and flag data having the same index number, wherein the memory array is configured to perform an in-memory computing (CIM) operation on the encoded weight and an input to generate a CIM result, wherein each CIM result corresponds to a column of the memory array; and an accumulator circuit configured to decode the CIM result according to the flag data to generate a decoded CIM result.

[0114] In some embodiments, the flag data indicates a first index number. The encoder circuit converts the first bit having the first index number into a sign bit and adds the value of the first bit to the second bit of each weight. The second bit has a second index number that is one greater than the first index number.

[0115] In some embodiments, the weight is in two's complement format and the encoded weight is in a two-sign-bit format with two sign bits.

[0116] In some embodiments, the index number of the bit with a value of one in the flag data corresponds to the index number of the sign bit of the encoded weight.

[0117] In some embodiments, the memory array is further configured to generate a current corresponding to a column of the memory array according to the CIM operation, and the storage device further includes: an analog-to-digital converter configured to perform analog-to-digital conversion on the current to generate a CIM result.

[0118] In some embodiments, the storage device further includes a register circuit configured to retrieve the flag data from the memory array, wherein the accumulator circuit includes: a recovery circuit configured to decode the CIM result according to the flag data from the register circuit to generate a decoded CIM result; and a shifter adder circuit configured to accumulate the decoded CIM results to generate a final result.

[0119] In some embodiments, each recovery circuit is configured to decode a corresponding CIM result in the CIM results according to a corresponding flag bit of the flag data, wherein each recovery circuit includes: a first transmission gate configured to turn on in response to the value of the corresponding flag bit being zero to output the corresponding CIM result as the corresponding result in the decoded CIM results; an inverter configured to invert the corresponding CIM result to generate an inverted result; and a second transmission gate configured to turn on in response to the value of the corresponding flag bit being one to output the inverted result as the corresponding result in the decoded CIM results.

[0120] In some embodiments, the index number of the sign bit is set to be equal to the number "N", where "2 M -2 N " is less than the maximum value of the weights, where the number "M" corresponds to the maximum index number of the bits of the weights.

[0121] Also disclosed is a storage device. The storage device includes: an encoder circuit configured to convert weights in two's complement form into encoded weights in double sign bit form according to flag data, wherein each encoded weight has two sign bits; a memory array including memory cells, wherein the first row of memory cells is configured to store flag data, the second row of memory cells is configured to store encoded weights, and the memory array is configured to perform a dot product operation on the encoded weights and inputs on the word lines of the memory array to generate a dot product result on the bit lines of the memory array; an analog-to-digital converter circuit configured to convert the dot product result into a digital dot product result; and an accumulator circuit configured to decode the digital dot product result according to flag data to generate a decoded dot product result.

[0122] In some embodiments, the encoder circuit is further configured to convert the first bit of each weight into a sign bit and add the value of the first bit to the second bit of each weight, wherein the index number of the second bit is one greater than the index number of the first bit.

[0123] In some embodiments, the first sign bit of the two sign bits is the most significant bit of each encoded weight, and the storage device further includes: a processor configured to determine the index number of the second sign bit of the two sign bits, where "2 M -2 N " is less than the maximum value of the weights, where the number "M" corresponds to the maximum index number of the bits of the weights, and the number "N" corresponds to the index number.

[0124] In some embodiments, the processor is further configured to encode the weights into test encoded weights with a test index number for the second sign bit, and the processor is further configured to determine the test index number as the index number according to the sparsity of the test encoded weights.

[0125] In some embodiments, the accumulator circuit includes: a recovery circuit, where each recovery circuit includes: an inverter configured to receive a first digital dot product result of a digital dot product result and generate an inverted result of the first digital dot product result; a first switch configured to turn on in response to the first bit of the flag data having a value of 1 to transmit the inverted result as the output of the recovery circuit; and a second switch configured to turn on in response to the first bit of the flag data having a value of 0 to transmit the first digital dot product result as the output of the recovery circuit.

[0126] In some embodiments, the accumulator circuit further includes: a shifter adder circuit configured to shift and add the output of the recovery circuit to generate a multiply-accumulate result of the weight and the input.

[0127] A method of operating a memory device is also disclosed. The method includes: converting the magnitude bits of each weight into sign bits according to flag data to generate encoded weights, where the flag data indicates the index number of the magnitude bits; performing an in-memory computing (CIM) operation on the encoded weights and an input transmitted to a memory array storing the weights to generate a CIM result; decoding the CIM result to generate a decoded CIM result; and accumulating the decoded CIM results to generate a final result.

[0128] In some embodiments, the method further includes: converting the magnitude bits having a test index number of each weight into sign bits to generate test encoded weights; and determining the test index number as the index number according to the sparsity of the test encoded weights.

[0129] In some embodiments, the method further includes: determining the index number according to the maximum value of the weights.

[0130] In some embodiments, the method further includes: generating flag data with the first bit being one, where the first bit corresponds to the index number.

[0131] In some embodiments, decoding includes: inverting a first CIM result in the CIM result to generate a first decoded CIM result in the decoded CIM result according to the first bit of the flag data being one.

[0132] In some embodiments, each decoded CIM result corresponds to a column of the memory array, where accumulating includes: performing a shift and add operation on each decoded CIM result according to the position of the column to generate a final result.

[0133] The components of several embodiments have been described above so that those skilled in the art can better understand the various embodiments of the present invention. Those skilled in the art should understand that it is easy to use the present invention as a basis to design or change other processes and structures for achieving the same purposes and / or realizing the same advantages as the embodiments introduced in the present invention. Those skilled in the art should also realize that these equivalent structures do not depart from the spirit and scope of the present invention, and various changes, substitutions, and alterations can be made without departing from the spirit and scope of the present invention.

Claims

1. A storage device, comprising: An encoder circuit configured to convert the first bit of each of a plurality of weights into a sign bit according to flag data to generate a plurality of encoded weights; A memory array including a plurality of memory cells, wherein the memory cells arranged in the same column are configured to store bits with the same index number of the encoded weights and the flag data, Wherein the memory array is configured to perform in-memory computing (CIM) operations on the encoded weights and a plurality of inputs to generate a plurality of CIM results, Wherein each of the plurality of CIM results corresponds to a column of the memory array; and An accumulator circuit configured to decode the plurality of CIM results according to the flag data to generate a plurality of decoded CIM results.

2. The memory device according to claim 1, wherein, The flag data indicates a first index number, Wherein the encoder circuit is further configured to convert the first bit with the first index number into the sign bit and add the value of the first bit to the second bit of each of the plurality of weights, and Wherein the second bit has a second index number that is one greater than the first index number.

3. The memory device according to claim 1, wherein, The plurality of weights are in two's complement format, and the plurality of encoded weights are in a double sign bit format with two sign bits.

4. The memory device according to claim 1, wherein The index number of the bit having a value in the flag data corresponds to the index number of the sign bit of the encoded weights.

5. The storage device according to claim 1, wherein, The memory array is further configured to generate a plurality of currents corresponding to the columns of the memory array according to the CIM operations, Wherein the storage device further comprises: An analog-to-digital converter configured to perform analog-to-digital conversion on the plurality of currents to generate the plurality of CIM results.

6. A storage system, comprising: An encoder circuit configured to convert a plurality of weights in two's complement form into a plurality of encoded weights in double sign bit form according to flag data, wherein each of the encoded weights has two sign bits; A memory array including a plurality of memory cells, wherein the first row of the memory cells is configured to store the flag data, and a plurality of second rows of the memory cells are configured to store the encoded weights, Wherein the memory array is configured to perform a dot product operation on the encoded weights and a plurality of inputs on the word lines of the memory array to generate a plurality of dot product results on the bit lines of the memory array; An analog-to-digital converter circuit configured to convert the dot product results into a plurality of digital dot product results; and An accumulator circuit configured to decode the digital dot product results according to the flag data to generate a plurality of decoded dot product results.

7. The system according to claim 6, wherein the encoder circuit is further configured to convert a first bit of each of the plurality of weights to a sign bit and add a value of the first bit to a second bit of each of the plurality of weights, wherein, The index number of the second bit is one greater than the index number of the first bit.

8. A method of operating a storage device, comprising: Converting the magnitude bit of each of a plurality of weights into a sign bit according to flag data to generate a plurality of encoded weights, wherein the flag data indicates the index number of the magnitude bit; Performing a plurality of in-memory computing (CIM) operations on the encoded weights and a plurality of inputs transmitted to a memory array storing the weights to generate a plurality of CIM results; Decoding the CIM results to generate a plurality of decoded CIM results; and Accumulate the decoded CIM results to generate a final result.

9. The method according to claim 8, further comprising: Convert the magnitude bits having the test index numbers of each of the plurality of weights into the sign bits to generate a plurality of test encoded weights; and Determine the test index number as the index number according to the sparsity of the plurality of test encoded weights.

10. The method according to claim 8, further comprising: Determine the index number according to the maximum value of the plurality of weights.