Compute-in-memory system and method of operating the same

TWI931783BActive Publication Date: 2026-07-11TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW113126319
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2024-07-12
Publication Date
2026-07-11
Estimated Expiration
2044-07-11

Smart Images

  • Figure IMG-2_DRAW_113126319-A0305-14-0001-1
    Figure IMG-2_DRAW_113126319-A0305-14-0001-1
  • Figure IMG-2_DRAW_113126319-A0305-14-0002-2
    Figure IMG-2_DRAW_113126319-A0305-14-0002-2
  • Figure IMG-2_DRAW_113126319-A0305-14-0003-3
    Figure IMG-2_DRAW_113126319-A0305-14-0003-3
Patent Text Reader

Abstract

The memory-based arithmetic system includes: a first component of memory cells correspondingly used to store units, and an array of multipliers and a first error bit detector. A plurality of first memory cells are arranged in a corresponding first array and used to store the first unit. A plurality of second memory cells are arranged in a corresponding second array and used to store corresponding bits of the first unit. For each of the first groups, each includes one of the corresponding first arrays, one of the second arrays, one of the multipliers, and one of the first error bit detectors. The multiplier is used to multiply the input bit and a corresponding unit of the first unit, and the first error bit detector is used to detect error bits in the corresponding first unit based on the corresponding corresponding bits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a bit error detection method within a memory-based arithmetic system. Prior Technology

[0002] The semiconductor integrated circuit industry produces a wide variety of analog and digital devices to address a range of different issues. Advances in semiconductor process technology have progressively reduced component size and space requirements, while increasing transistor density. Integrated circuit dimensions have become smaller. Summary of the Invention

[0003] This disclosure includes an in-memory computing system comprising, in a first region of a semiconductor die, a first assembly comprising memory cells corresponding to store units, and an array comprising a multiplier and a first error bit detector; the first memory cells are arranged in a corresponding first array and are used to store a first bit; the second memory cells are arranged in a corresponding second array and are used to store corresponding bits of the first bit; and each of the first group comprises a first array of the first array, a second array of the second array, a multiplier of the multiplier, and a first error bit detector of the first error bit detector, wherein the multiplier is used to multiply the input bits and corresponding multiples of the first bit, and the first error bit detector is used to detect error bits corresponding to the first bit based on the corresponding bits of the first bit.

[0004] This disclosure includes a method for operating an in-memory computing system, the method comprising, for each of a first group, a multiplier comprising a multiplier and a corresponding first error bit detector, the in-memory computing system comprising a first component in a first region of a semiconductor die, the first component comprising memory cells correspondingly used for storing units and arranged in an array, an array of multipliers and a first error bit detector, a first memory cell disposed in the first array for storing a first bit, and a second memory cell disposed in a second array for storing a corresponding bit in the same position as the first bit: performing a multiplication of an input bit and the first corresponding bit; and performing a first error detection of an error bit in the corresponding first bit based on the corresponding bit in the same position, the first error detection being performed simultaneously with the multiplication.

[0005] This disclosure includes an in-memory computing system, comprising a first region of a semiconductor die, a first component including memory cells for storing characters, a multiplier, and a trajectory inference data generator; the first set of memory cells are arranged in a first array and used to store first characters; each of the first sets includes a first array of a corresponding first array, a multiplier of a corresponding multiplier, and a trajectory inference data generator, and the operation of each of the first sets is related to a corresponding number of first characters, the multiplier generating the first characters of the first array by performing (A) multiplication of an input character and an associated first checksum character and (B) a corresponding weight character and an associated second checksum character, and the trajectory inference data generator generating an error bit trajectory inference signal based on a selected number of first characters. Simple Explanation of the Diagram

[0006] The nature of this disclosure is best understood when read in conjunction with the accompanying drawings in the following detailed description. It should be noted that, in accordance with industry standard practice, the various features are not drawn to scale. In fact, for clarity of discussion, the dimensions of the various features may be arbitrarily increased or decreased. Figures 1A and 1C through 1E are functional block diagrams of a system illustrated according to some embodiments. Figure 1B is a functional block diagram illustrating a memory configuration according to some embodiments. Figures 2A to 2D are schematic diagrams illustrated according to some embodiments. Figures 3A to 3M are schematic diagrams illustrated according to some embodiments. Figures 4A to 4G are block diagrams illustrated according to some embodiments. Figures 5A to 5E are flowcharts illustrating various methods according to some embodiments. Figure 6 is a functional block diagram of an electronic design automation (EDA) system according to some embodiments. Figure 7 is a functional block diagram illustrating an integrated circuit (IC) manufacturing system and its associated manufacturing process according to some embodiments. Implementation

[0007] The following disclosure provides numerous different embodiments or examples for implementing various features of the provided subject matter. Specific examples of components and configurations are described below to simplify this disclosure. Of course, these components and configurations are merely examples and are not intended to be limiting. For example, in the following description, forming a first feature on or on a second feature may include embodiments in which the first and second features are formed in direct contact, and may also include embodiments in which an additional feature is formed between the first and second features such that the first and second features are not in direct contact. Furthermore, reference numerals and / or letters may be repeated in various instances of this disclosure. This repetition is for simplicity and clarity and does not in itself determine the relationship between the various embodiments and / or configurations discussed.

[0008] Additionally, for ease of description, spatial relative terms such as "below," "under," "lower," "above," "upper," "top," "bottom," and similar terms may be used herein to describe the relationship between one element or feature illustrated in the figures and one or more other elements or features. Besides the orientation depicted in the figures, spatial relative terms are intended to cover different orientations of elements in use or operation. Devices may be oriented in other ways (rotated 90 degrees or otherwise), and the spatial relative descriptive terms used herein will be interpreted accordingly.

[0009] In some embodiments, a first compute-in-memory (CIM) system includes, in a first region of a first semiconductor die, a plurality of first components comprising a plurality of memory cell groups for storing a plurality of units, and a plurality of arrays comprising a plurality of multipliers and a plurality of first error bit detectors. The plurality of first groups of memory cells are configured with corresponding plurality of first arrays and are used to store a plurality of first bits. The plurality of second groups of memory cells are configured with corresponding plurality of second arrays and are used to store a plurality of corresponding bits for the plurality of first bits. For each plurality of first groups, each comprising one of the respective plurality of first arrays, one of the plurality of second arrays, one of the plurality of multipliers, and one of the plurality of first error bit detectors: the multipliers are used to multiply the plurality of input bits and corresponding bits of the plurality of first bits, and the first error bit detectors are used to detect an error bit corresponding to one of the plurality of first bits based on the corresponding plurality of corresponding bits. In some embodiments, an instance of one of the error bit detectors is used to detect the presence of an error bit and generate a similar indicative signal. For example, a flag signal FLG1 is generated as shown in Figure 2A, or other similar signals, wherein the flag FLG1 can announce the presence of an error bit. In some embodiments, the in-memory computing system is included as part of an artificial intelligence (AI) system. According to another approach, a reference memory system corresponding to the in-memory computing system performs error bit detection in the reference memory of the first array before bit multiplication (pre-multiplication detection). Pre-multiplication error bit detection according to the other approach uses two operation cycles. In contrast, according to at least some embodiments of the in-memory computing system, a first error bit detector 210 (Q-1) is included to perform error bit detection synchronously with multiplication in the multiplier, using one operation cycle, which is one operation cycle faster than the other approach.

[0010] In some embodiments, the second memory-in-cell processing system includes, in a second region of the second semiconductor die, a plurality of first components comprising a plurality of memory cell groups for storing a plurality of characters, a plurality of multipliers, and a trajectory inference data generator. The plurality of first groups within the memory cell groups are configured in a plurality of corresponding first arrays and are used to store a plurality of first characters. For each of the plurality of first arrays, a plurality of multipliers, and a trajectory inference data generator, and for each of the plurality of first groups, each operating in conjunction with a corresponding plurality of characters among the plurality of first characters: the multiplier generates a plurality of first characters of the first array by performing (A) a plurality of input characters and associated plurality of first checksum characters and (B) a plurality of weight characters and associated plurality of second checksum characters; and the trajectory inference data generator generates one or more error bit trajectory inference signals based on selected characters among the plurality of first characters, wherein the error bit inference signals are used to infer the location of an error bit, for example, a memory cell located at the intersection of one of the identified columns and identified bars in the array. Multiple examples of signals inferred from multiple error bit trajectories of the computing system within the second memory include column sum vectors (e.g., Row_Sum 304J shown in Figure 3J), column sum vectors (e.g., Col_Sum 304L shown in Figure 3L), or other similar signals. In some embodiments, the computing system within the second memory is incorporated as part of an artificial intelligence system. According to another approach, in a corresponding reference memory system to the computing system within the second memory, (1) a weight array is stored in a first region of the first die, and (2) error bit detection, localization, and correction (DLC) is performed entirely by a processor and associated random access memory (RAM) on at least one corresponding second die. According to the above-described alternative approach, in order to perform error bit DLC, a large amount of data is transferred from the weight array on the first die to the processor on the second die (out-of-die transfer), resulting in significant transmission delays and thus significantly reducing the speed of error bit DLC. In contrast, in at least some embodiments, the computing system within the second memory substantially reduces the amount of error bit detection, location, and correction (DLC) transmitted off-chip compared to the other approach, and thus achieves substantially faster error bit DLC compared to the other approach.

[0011] According to some embodiments, Figure 1A is a functional block diagram illustrating the computing system 100A within the digital memory.

[0012] In Figure 1A, the in-memory computing system 100A includes a semiconductor die 102A, which includes at least one in-memory computing system (CIM) region 103A and a second region 114A. Multiple components in the CIM region 103A (refer to Figures 2A-2C, 4A-4B, or other similar illustrations) include: a weighted and parallel (W&P) array 104A, multiple multipliers 107A, multiple adder trees 108A, multiple error bit detectors 110, a trajectory inference data generator 112, and multiple parallel encoders 158. Region 114A includes an error bit corrector 116 (refer to Figure 2D). In some embodiments, region 114 is a processor, and the error bit corrector 116 is a function executed by the processor. In some embodiments, instances of the error bit detectors 110 are used to detect the presence of error bits and generate a similar indicative signal. For example, a flag signal FLG1 as shown in Figure 2A, or other similar signals, is generated, wherein the flag FLG1 can announce the presence of an error bit. The trajectory inference data generator 112 is used to generate one or more error bit trajectory inference signals to infer the location of the error bit, for example, a memory cell located at one intersection of an array of identified columns and rows. In some embodiments, examples of multiple error bit trajectory inference signals include the pointer signal EPT and the flag signal FLG2 in Figure 2C, or other similar signals. In some embodiments, the in-memory computing system 100A is included as part of an artificial intelligence system.

[0013] Because multiple error bit detectors 110 and trajectory inference data generators 112 are located in the same region—that is, multiple other components in CIM region 103A are located on the same die as CIM region 103A—these components in CIM region 103A are physically closer to each other than if error bit detectors 110 and trajectory inference data generators 112 were located on a different die than in CIM region 103A. This relatively closer proximity of the components in CIM region 103A results in several benefits, including increased operating speed, such as faster error bit detection by detector 110 and faster generation of trajectory inference data by generator 112.

[0014] Figure 1B is a functional block diagram of the W&P array 104B according to some embodiments.

[0015] In some embodiments, the W&P array 104B is an example of one of the W&P arrays 104A in Figure 1A, or other similar arrays. The W&P array 104B is arranged in N columns (refer to Figure 2A) and Q slices 118(0) to 118(Q-1), where N and Q are corresponding positive integers. Each column in the W&P array 104B has a corresponding segment in slices 118(0) to 118(Q-1). Each slice 118(0) to 118(Q-1) contains a corresponding plurality of first and second arrays. With respect to slice 118(0), the first subarray is a two-dimensional weight array 120(0) and a one-dimensional isotope array 122(0). With respect to slice 118(Q-1), the first subarray is a two-dimensional weight array 120(Q-1) (refer to Figure 2A) and a one-dimensional isotope array 122(Q-1) (refer to Figure 2A). In some embodiments, each of the co-position arrays 122(0) to 122(Q-1) is a two-dimensional array.

[0016] Figure 1C is a functional block diagram of a digital CIM system 100C according to some embodiments.

[0017] CIM system 100C is similar to CIM system 100A in Figure 1A. For the sake of simplicity, the following discussion will focus on the differences between CIM system 100C and CIM system 100A, rather than their similarities.

[0018] In Figure 1C, the CIM system 100C includes two semiconductor dies, namely a first semiconductor die 102C(1) and a second semiconductor die 102C(2), while the CIM system 100A includes a semiconductor die 102A. The first semiconductor die 102C(1) includes at least one CIM region 103C and a region 155(1). The second semiconductor die 102C(2) includes region 155(2) and region 144C. In some embodiments, the in-memory computing system 100C is included as part of an artificial intelligence system.

[0019] The CIM region 103C and the first semiconductor die 102C(1) in Figure 1C are corresponding counterparts to the CIM region 103A and the semiconductor die 102A in Figure 1A. However, the co-position encoder 158 and the trajectory inference data generator 112 are contained in region 155(1) of the CIM system 100C, rather than in CIM region 103A, which is one difference from the CIM system 100A. In some embodiments, as indicated by the imaginary lines (dashed lines), one or more co-position encoders 158 or trajectory inference data generators 112 are contained in region 155(2) of the second semiconductor die 102C(2), rather than in CIM region 103C of the second semiconductor die 102C(1), which is another difference from the CIM system 100A.

[0020] Region 114C in Figure 1C is a counterpart to region 114A in Figure 1A. However, region 114C is a region within the second semiconductor die 102C(2), and not within the first semiconductor die 102C(1). Therefore, the error bit corrector 116 is included as part of the second semiconductor die 102C(2), and not as part of the first semiconductor die 102C(1), which is another difference from the CIM system 100A. In some embodiments, region 114C is a processor and the error bit corrector 116 is a function executed by the processor.

[0021] Because the error bit detector 110 and the trajectory inference data generator 112 are located on the same die as the CIM region 103C, the components of the error bit detector 110 and the trajectory inference data generator 112 in the CIM region 103C are physically closer to each other than if the error bit detector 110 and the trajectory inference data generator 112 were located on a die different from the CIM region 103C. This relative proximity of the error bit detector 110 and the trajectory inference data generator 112 in the CIM region 103C can facilitate several benefits, including increased operating speed, such as faster error bit detection by the detector 110 and faster generation of trajectory inference data by the generator 112.

[0022] Figure 1D is a functional block diagram of a digital CIM system 100D according to some embodiments.

[0023] CIM system 100D is similar to CIM system 100A in Figure 1A. For simplicity, the following discussion will focus on the differences between CIM system 100D and CIM system 100A, rather than their similarities.

[0024] CIM system 100D includes semiconductor die 102D, which includes CIM regions 124D and 114D. CIM region 124D and semiconductor die 102D in Figure 1D correspond to CIM region 103A and semiconductor die 102A in Figure 1A. In Figure 1D (and also refer to Figures 3A to 3L), W&P array 104B, multiplier 107B, and adder tree 108B correspond to W&P array 104A, multiplier 107A, and adder tree 108A in Figure 1A. Multiple trajectory inference data generators 126 (refer to Figures 3G to 3L) correspond to trajectory inference data generator 112. However, the multiple trajectory inference data generators 126 further include multiple column sum generators 127 (refer to Figure 3I) and multiple column sum generators (refer to Figure 3K). The trajectory inference data generator 126 generates an error bit trajectory inference signal to infer the location of an error bit, for example, a memory cell located at the intersection of an array of identified columns and rows. Examples of the error bit trajectory inference signal include a column sum vector (e.g., Row_Sum 304J in Figure 3J), a column sum vector (e.g., Col_Sum 304L in Figure 3L), or other similar signals. In some embodiments, the in-memory computing system 100D is included as part of an artificial intelligence system.

[0025] CIM area 124D further includes input data and checksum array 105 (refer to Figures 3A to 3C), inner product array 106 (refer to Figures 3G to 3H), and checksum generator 178 (refer to Figures 3H to 3L). Checksum generator 178 further includes column checksum generator 127 (refer to Figures 3A to 3C) and column checksum generator 128 (refer to Figures 3D to 3F).

[0026] CIM region 124D does not include a counterpart to the co-encoder 158 and the error bit detector 110. However, region 114D in the semiconductor die 102D includes an error bit generator, locator, and corrector 130 corresponding to the error bit detector 110 (refer to Figure 3M). In some embodiments, region 114D represents the processor and the error bit detector, locator, and corrector 130 as corresponding functions executed by the processor.

[0027] Because the checksum generator 178 and the trajectory inference data generator 126 are located on the same die as the other components of the CIM region 124D, the components of the CIM region 124D are physically closer to each other than if the checksum generator 178 and the trajectory inference data generator 126 were located on dies different from those of the other components of the CIM region 124D. This relatively closer proximity of the multiple components within the CIM region 124D results in several benefits, including increased operating speed, such as faster checksum generation by generator 178 and faster trajectory inference data generation by generator 126.

[0028] Figure 1E is a functional block diagram of a digital CIM system 100E according to some embodiments.

[0029] CIM system 100E is similar to CIM system 100D in Figure 1D. For simplicity, the following discussion will focus on the differences between CIM system 100E and CIM system 100D, rather than their similarities.

[0030] In Figure 1E, the CIM system 100E includes two semiconductor dies, namely a first semiconductor die 102E (1) and a second semiconductor die 102E (2), while the CIM system 100D includes a single semiconductor die 102D. The first semiconductor die 102E (1) includes at least one CIM region 124E and a region 155 (3). The second semiconductor die 102E (2) includes region 114E, and in some embodiments includes region 155 (4). In some embodiments, the in-memory computing system 100E is included as part of an artificial intelligence system.

[0031] CIM region 124E and first semiconductor die 102E(1) in Figure 1E are corresponding counterparts to CIM region 124D and semiconductor die 102D in Figure 1D. However, checksum generator 178, which includes column checksum generator 179 and column checksum generator 180, is included in region 114D of CIM system 100E, rather than in CIM region 124D, which is one difference from CIM system 100D. In some embodiments, as indicated by the imaginary lines (dashed lines), one or more checksum generators 178, which include column checksum generator 179 and column checksum generator 180, are included in region 155(4) of second semiconductor die 102E(2), rather than in CIM region 124E of first semiconductor die 102E(1), which is another difference from CIM system 100D.

[0032] Region 114E in Figure 1E is the counterpart to region 114D in Figure 1D. However, region 114E is a region within the second semiconductor die 102E(2), and not within the first semiconductor die 102E(1). Therefore, the error bit generator, locator, and corrector 130 is included as a part of the second semiconductor die 102E(2), and not as a part of the first semiconductor die 102E(1), which is another difference from the CIM system 100D. In some embodiments, region 114E is a processor and the error bit generator, locator, and corrector 130 is a function executed by the processor.

[0033] Since the checksum generator 178 is located in region 155(3) and therefore on the same die as CIM region 124E, the components of checksum generator 178 and CIM region 124E are physically closer to each other than if checksum generator 178 were located on a die different from CIM region 124E. The relatively closer proximity of checksum generator 178 and multiple components in CIM region 124E can facilitate several benefits including increased operating speed, such as faster checksum generation by generator 178.

[0034] Figure 2A is a schematic diagram illustrating CIM region 203(1) of one of the digital CIM systems according to some embodiments.

[0035] CIM region 203(1) is an example of a portion of region 103A in Figure 1A, an example of a portion of region 103C in Figure 1C, or other similar regions. Multiple instances of CIM region 203(1) include CIM region 103A, CIM region 103C, or other similar regions.

[0036] CIM region 203(1) includes: a weighted bit and same bit (W&P) array 204A containing slices 218(0) to 218(Q-1); an EM array 236 containing enhanced multiplier (EM) blocks 238(0) to 238(Q-1), each containing a multiplier (e.g., 251(Q-1)) and an error bit detector (e.g., 210(Q-1)); and an addition tree 208, wherein Q is a positive integer. In some embodiments, Q is a quadratic value. In some embodiments, Q is the value 64. In some embodiments, Q is the quadratic value of a positive integer other than the value 64.

[0037] On a column-by-column basis, EM array 236 is used to receive a given column data bit from one of the W&P arrays 204A. Accordingly, EM blocks 238(0) to 238(Q-1) are used to receive the corresponding segments of a given column. The EM array in Figure 2A is further used to receive a given column input bit from one of the input arrays XIN1, which corresponds to a given column data bit from one of the W&P arrays 204A, and the corresponding two are multiplied together to obtain a Q-product PRD1(0) to PRD(Q-1). Addition tree 208 adds the Q-products PRD1(0) to PRD(Q-1) to generate an output signal Out_1. The output signal Out_1 is operated according to, for example, an error bit corrector 216 (refer to Figure 2D).

[0038] Taking EM block 238(Q-1) as an example of one of EM blocks 238(0) to 238(Q-1), EM block 238(Q-1) includes multiplier 251(Q-1), error bit detector (e.g. 210(Q-1)), and data multiplexer (MUX) 254(Q-1). One of W&P array 204A and EM array 236, component 242, includes a slice 218(Q-1) of W&P array 204A and EM block 238(Q-1). A component 244(1) of component 242 is shown in Figure 2A in a magnified view for more detail.

[0039] In component 244(1), as shown in an enlarged view, slice 218(Q-1) includes a two-dimensional weighted array 220(Q-1) of a one-bit memory cell 245 and a one-dimensional co-occurrence bit 221(Q-1) of a one-bit memory cell 246. In Figure 2A, memory cells 245 and 246 are assumed to be static random access memory (SRAM) cells. In some embodiments, memory cells 245 and 246 are memory cells other than static random access memory cells.

[0040] Memory cells 245 and 246 of slice 218 (Q-1) are arranged in columns and rows. For simplicity, some, but not all, of the signal lines shown in slice 218 (Q-1) are mentioned in the text. 218 (Q-1) is used to ensure that data bits are read one column at a time at any given time. The selection of a given column of slice 218 (Q-1) is controlled by the corresponding read word lines RWL[0] to RWL[N-1], where N is a positive integer. This given column is selected not only in slice 218 (Q-1) but also in each of the other slices 218 (0) to 218 (Q-2). The columns of slice 218 (Q-1) have corresponding read bit lines RBL[0] to RBL

[12] .

[0041] The weight array 220(Q-1) is configured relative to the read bit lines RBL[0]~RBL

[11] . Therefore, the weight array 220(Q-1) is a An array, where N is a positive integer. The parity array 221(Q-1) is set relative to the read bit line RBL

[12] . Therefore, the parity array 221(Q-1) is a... Array. Accordingly, slice 218(Q-1) is a Array.

[0042] For simplicity, Figure 2A assumes that each column of the weight array 220(Q-1) is stored in a 12-bit character. Generally, the weight array 220(Q-1) stores K-bit characters, where K is a positive integer and is assumed to be 12 in Figure 2A. Therefore, the weight array 220(Q-1) is... Array. In some embodiments, K is a square of a positive integer, for example... , or other similar quadratic values. In some embodiments, K is a positive integer other than 8, 12, 16, or 32. It is worth noting that the input array XIN1 is An array, where L is a positive integer.

[0043] Repeatingly, on a column-by-column basis, EM block 238 (Q-1) generates an output signal PRD1 (Q-1) to represent an inner product of the multiplication of a column of the N×K weight array 220 (Q-1) with the corresponding column of the K×L array XIN1. Specifically, repeatingly, multiplier 251 (Q-1) receives a column of data bits from the weight array 220 (Q-1) as the multiplicand and a column of input bits from the input array XIN1 as the multiplicand, and multiplies the multiplicand and the multiplicand to obtain the inner product PRD1 (Q-1). Accordingly, multiplier 251 (Q-1) receives a column of K data bits from read bit lines RBL[0]~RBL

[11] as the multiplicand and a column of input bits from the input array XIN1 as the multiplicand, and multiplies the multiplicand and the multiplicand to obtain the inner product PRD1 (Q-1). PRD1(Q-1) is a single character with K+K=2K bits. In the example in Figure 2A, PRD1(Q-1) has K+K=12+12=24 bits. The operation of multiplier 251(Q-1) is also discussed in the description of Figure 4B.

[0044] Using slice 218(Q-1) as a representative of slices 218(0)~218(Q-1) of W&P array 204A, a single error bit occurs when the value is stored in one of the selected columns of slice 218(Q-1) among multiple memory cells 245, that is, when reading the value of one of the bit lines RBL[0]~RBL

[11] , representing a one-bit error. A double error bit occurs when the value is stored in two of the selected columns of slice 218(Q-1) among multiple memory cells 245, that is, when reading the corresponding two values ​​of bit lines RBL[0]~RBL

[11] , representing two bit errors. The probability of a single error bit occurring in slice 218(Q-1) is relatively low. The probability of a double-error bit occurrence in slice 218(Q-1) is substantially lower than the probability of a single-error bit occurrence in slice 218(Q-1). In practice, Figure 2A assumes that a double-error bit occurrence will not occur in slice 218(Q-1). Accordingly, Figure 2A is used to detect and correct single-error bit occurrences in slice 218(Q-1), but not to detect and correct double-error bit occurrences in slice 218(Q-1). At least in some other embodiments disclosed herein, it is used to detect and correct single-error bit occurrences, rather than double-error bit occurrences, in other slices compared to slice 218(Q-1).

[0045] On a column-by-column basis, the error bit detector 210 (Q-1) receives K data on read bit lines RBL[0]~RBL

[11] . Based on the values ​​on read bit lines RBL[0]~RBL

[11] , the error bit detector 210 (Q-1) determines whether one of these read bit lines RBL[0]~RBL

[11] has an error bit, and generates an output signal representing a flag signal FLG1 accordingly. The flag signal FLG1 represents one of the output signals of EM block 238 (Q-1) and is internally provided to data multiplexer 254 (Q-1). It is worth noting that read bit line RBL

[12] is also provided to trajectory inference data generator 212 (refer to Figure 2C).

[0046] The flag signal FLG1 can declare the presence of an error bit, that is, the error bit detector 210 (Q-1) detects an error bit in the corresponding column of data bits from the weighted array 220 (Q-1). Figure 2A assumes the following declaration of the flag signal FLG1: when the flag signal FLG1 is not declared, that is, when FLG1=0, no error occurs in the read bit lines RBL[0]~RBL

[11] , and when the flag signal FLG1 is declared, that is, when FLG1=1, an error occurs in one of the read bit lines RBL[0]~RBL

[11] .

[0047] Error bit detector 210 (Q-1) includes a mutex OR (Q-1) gate 253 (Q-1) to receive K+1 inputs. Generally, for any multi-input mutex OR gate, its output is true (or logic one) when an odd number of inputs are true. In some embodiments, the declaration of the flag signal FLG1 described above is assumed. The operation of the mutex OR gate 253 (Q-1) is also discussed in the description of Figure 4B.

[0048] According to another approach, in a reference memory system corresponding to the computational system within the memory in which the CIM region 203(1) constitutes a component, error bit detection for one of the error bits of the weight array 220(Q-1) is performed before the multiplication operation (pre-multiplication detection). According to another approach, two operation cycles are used for error bit detection before the aforementioned multiplication operation. Conversely, in at least some embodiments, the error bit detector 210(Q-1) encompassing the CIM region 210(Q-1) uses one operation cycle to perform the aforementioned multiplication operation synchronously with the multiplier 251(Q-1), which is one operation cycle faster than the other approach.

[0049] Another approach uses a Q-bit weight array as opposed to weight array 220 (Q-1). Furthermore, where Q is 12 for each column in this weight array, this approach uses 5 check bits to perform pre-multiplication detection, which will cause significant disadvantages to the bare die on an area basis, namely, loss (increased footprint), power loss, the flexibility of signal segments (refer to block 547 in Figure 5C) and / or PG segments (refer to block 547 in Figure 5C), or other similar disadvantages. Conversely, including error bit detector 210 (Q-1) and a similar approach of storing a single corresponding bit in corresponding array 221 (Q-1) reduces the number of error bit detection bits by four compared to the other approach, thus reducing the values ​​of parameters including footprint, power loss, flexibility, or other similar factors compared to the other approach.

[0050] Data multiplexer 254 (Q-1) is used to output a selection column by column. As a selection input, data multiplexer 254 (Q-1) receives the inner product PRD1 (Q-1) from multiplier 251 (Q-1) and a predetermined 2K-bit word representing a reference value REF. In Figure 2A, the reference value REF is assumed to be zero. In some embodiments, the reference value REF has a non-zero value. As a control input, data multiplexer 254 (Q-1) receives a flag signal FLG1 from mutex gate 253 (Q-1). Based on the flag signal FLG1, when no error bit occurs, data multiplexer 254 (Q-1) selects the inner product PRD1 (Q-1), and when an error bit occurs, data multiplexer 254 (Q-1) selects the reference value word.

[0051] Based on column-by-column, the EM array 236 in Figure 2A receives a column of bits from the input array XIN1 and multiplies it with the corresponding column of data bits in the W&P array 204A to obtain Q inner products PRD1(0)~PRD(Q-1). The adder tree 208 adds the Q inner products PRD1(0)~PRD(Q-1) to generate an output signal Out_1. The output signal Out_1 operates according to, for example, an error bit detector 216 (refer to Figure 2D).

[0052] Adder tree 208 is used to receive Q inner products PRD1(0) to PRD(Q-1) from EM array 236 and add similar ones. Adder tree 208 has J adders 240 with processes (crs(0), ..., crs(J-1)), where J is a positive integer less than Q. In some embodiments, the value J of the process of adder 208 is related to the value Q of the slices 218(0) to 218(Q-1) as follows: Q equals 2 to the power of J. In such embodiments, EM array 236 produces Q inner product characters, where W&P array 204A has Q slices to supply Q characters to EM array 236. Correspondingly, adder tree 208 includes adders 240 with J processes (crs(0), ..., crs(J-1)) and produces a single character as output signal Out_1. Each adder 240 is used to receive two single character inputs. For example, the enhanced multiplier array 236 has J=6 processes, where Q=64.

[0053] In some embodiments, each character in the Q character is represented by 2L bits, such that a single character output Qut_1 can be represented by Q*2K = 2J*2k bits. In some such embodiments, 2K = 24, such that a single character output Qut_1 can be represented by 1536 = 26*24 bits.

[0054] Figure 2B is a schematic diagram of component 244(2) of component 242 of CIM region 203(1) of digital CIM system according to some embodiments.

[0055] Component 244(2) includes a cutout 218 (Q-1) of the W&P array and a co-position encoder 259 (Q-1). Figure 2B enlarges the component 244(2) of Figure 2A, where K=12. In some embodiments, the co-position encoder 259 (Q-1) is one of the plurality of co-position encoders 259 (Q-1) in Figure 1A or other similar examples.

[0056] Based on a single column, the co-occurrence encoder 259 (Q-1) generates a co-occurrence bit (co-occurrence value) corresponding to the bit value of that column and writes this value into a corresponding cell in memory cell 246. Each memory cell 246 is further used to selectively write the co-occurrence value from the co-occurrence encoder 259 (Q-1) together with the corresponding value.

[0057] A given column of a slice 218 (Q-1) contains K=12 instances of memory cells 245 in a weighted array 220 (Q-1), these instances corresponding to read bit lines RBL[0]~RBL

[11] and one instance of memory cell 246 in a parity array 221 (Q-1). Accordingly, for a given column of slice 218 (Q-1), a parity encoder 259 (Q-1) is further used to receive K data bits on read bit lines RBL[0]~RBL

[11] from the weighted array 220 (Q-1) of slice 218 (Q-1). Based on the bit values ​​on read bit lines RBL[0]~RBL

[11] , the parity encoder 259 (Q-1) generates a parity value corresponding to the given column. The parity value of the given column is then written into the instance of memory cell 246 contained in the given column.

[0058] The co-position encoder 259 (Q-1) includes a mutex OR gate 260 (Q-1) for receiving K inputs corresponding to the bit values ​​on the read bit lines RBL[0] to RBL

[11] and generating an output signal representing the corresponding co-position bit. The operation of the mutex OR gate 260 (Q-1) of the co-position encoder 259 (Q-1) is also discussed in the description of Figure 4A.

[0059] Figure 2C is a schematic diagram illustrating the CIM area 203(2) of a digital CIM system according to some embodiments.

[0060] CIM region 203(2) is a portion of CIM region 103A in Figure 1A, a portion of CIM region 103C in Figure 1C, or other similar examples. Multiple instances of CIM region 203(2) include CIM region 103A, CIM region 103, or other similar examples. Trajectory inference data generator 212 is an example of trajectory inference data generator 112 in Figure 1A, Figure 1C, or other similar examples. Figure 2C enlarges the example in Figure 2A, where K=12.

[0061] The trajectory inference data generator 212 receives Q instances of the flag signal FLG1 (Q flag FLG1) from the EM array 236 and Q corresponding bits of the W&P array 204A, and generates an error bit trajectory inference signal accordingly. The error bit trajectory inference signal generated by the trajectory inference data generator 212 includes an index signal (index) EPT and a flag signal (flag) FLG2. The index EPT and flag FLG2 operate according to, for example, an error bit corrector 216 (refer to Figure 2D). It should be recalled that the Q number of the flag FLG1 of the W&P array 204A corresponds to the Q number of the slices 218(0) to 218(Q-1).

[0062] In Figure 2C, the trajectory inference data generator 212 includes a Q:P encoder 262 and an error slicing detector 213.

[0063] Encoder 262 is used to receive Q instances of flag FLG1 from EM array 236 and generate an index EPT, where the index EPT is a P-bit character, P is a positive integer, P is less than Q, and Q = 2P. In some embodiments, encoder 262 receives Q instances of flag FLG1 as a concatenation of Q instances of flag FLG1. The operation of encoder 262 is also discussed in the description of Figure 4B.

[0064] In some embodiments, the Q flag FLG1 is provided to the encoder 262 as a character with Q bits, such that Q = {f(0), f(1), ..., f(i-1), f(Q-1)}. Extending to the example of the declaration of the flag FLG1 discussed in Figure 2A, where no error bits appear in the Q slices of the W&P array 204A, each bit f(0) to f(Q-1) is set to a logic zero value. However, if one of the given slices 218(i) has an error bit, then the bit f(i) of the character with Q bits is set / declared to a logic one value, that is, f(i) = 1, and the remaining bits f(0), f(1), ..., f(i-1), f(i+1), ..., f(Q-2), f(Q-1) are not declared, that is, the remaining bits are set to a logic zero value.

[0065] The faulty shard detector 213 in Figure 2C receives Q instances of flag FLG1 from the EM array 236 and generates a flag signal (flag) FLG2. Flag FLG2 can declare the presence of a faulty shard, that is, the detector 213 detects a faulty shard. OR logic gate 263 is included within the faulty shard detector 213. OR logic gate 263 receives Q flags FLG1 from the EM array 236 and generates flag FLG2.

[0066] Regarding the flag FLG2, the example of the declaration of the flag FLG1 discussed in Figure 2A can be extended as follows: When the flag FLG2 is not declared, that is, when FLG1=0, there are no erroneous clips in clips 218(0)~218(Q-1). When the flag FLG2 is declared, that is, when FLG2=1, there are erroneous clips in clips 218(0)~218(Q-1).

[0067] Figure 2D is a schematic diagram illustrating a digital CIM system 200D according to some embodiments.

[0068] CIM system 200D is an example of CIM system 100A in Figure 1A, CIM system 100C in Figure 1C, or other similar systems. CIM system 200D includes CIM region 203(3) and error bit detector 216.

[0069] CIM region 203(3) is an example of CIM region 103A in Figure 1A, and a combination of CIM region 103C and region 155(1) or other similar regions in Figure 1C. Error bit corrector 216 is an example of error bit corrector 116 in Figure 1A and Figure 1C or other similar regions.

[0070] Error bit corrector 216 receives the output signal Out_1, index EPT, and flag FLG2 from CIM region 203(3) and operates according to flowchart 290. In Figure 2D, flowchart 290 shows an enlarged view 244D of error bit corrector 216. The operation of error bit corrector 216 is also discussed in the description of Figure 4B.

[0071] Flowchart 290 contains blocks 265(1) to 265(5). In block 265(1), a decision is made based on whether the flag FLG2 is declared, that is, FLG2=0, then the process proceeds to block 265(2).

[0072] At block 265(2), the process stops because there is no erroneous truncation. Therefore, the output signal Qut_1 does not need to be corrected for erroneous bits. If the result of block 265(1) is yes, that is, FLG=1, then an erroneous truncation requires the erroneous bit corrector and the process proceeds to block 265(3).

[0073] In block 265(3), an erroneous slice is located. It should be recalled that the output signal Out_1, the index EPT, and the flag FLG2 are repeatedly generated iteratively on a column-by-column basis, and the output signal Out_1, the index EPT, and the flag FLG2 in each iteration are based on the current multiplicand, that is, based on the corresponding one of the multiple input columns of the input array XIN1 and the current multiplicand, that is, the corresponding one of the multiple data columns in the W&P array 204A. Once the flag FLG1 is declared, the corrupted column in slice 218(Q-1) is identified as the current multiplicand. Then the erroneous slice is reserved for block 265(3) for identification. Therefore, in block 265(3), the index EPT is examined to determine which of the bits f(0)~f(Q-1) is declared, that is, the value is set to 1. In the bits f(0)~f(Q-1), the value set to 1 indicates that there is an erroneous slice in the corresponding slice of slice 218(0)~218(Q-1). The process proceeds from block 265(3) to block 265(4) and block 265(5).

[0074] A single stub failure occurs when one of the stubs 218(0) to 218(Q-1) experiences a single faulty bit. A double stub failure occurs when both of the stubs 218(0) to 218(Q-1) experience a single faulty bit. The probability of a single stub failure occurring in stubs 218(0) to 218(Q-1) is low. The probability of a double stub failure occurring in stubs 218(0) to 218(Q-1) is significantly lower than the probability of a single stub failure occurring in stubs 218(0) to 218(Q-1). Accordingly, typically, the cascaded bit sequence of the Q flags FLG1 has only one unit bit, whose value is set to logic one. In block 265(3), regardless of the total number of bits whose value is set to logic 1 in the sequence of Q flags FLG1, each bit whose value is set to logic 1 can also be identified as having experienced a bit error in the corresponding slice.

[0075] In block 265(4), each slice 218(0) to 218(Q-1) in the W&P array 204A identified by block 265(3) will be updated. In some embodiments, error bit detector 216 is used to update each corrupted slice by writing uncorrupted data bit values ​​to memory cells 245 of the column identified in block 265(3), for example, by copying the corresponding uncorrupted data bit values ​​from a source copy or permanent copy in the W&P array 204A. For example, the source copy or permanent copy is stored outside the CIM area containing the W&P array 204A.

[0076] In block 265(5), the output value of output signal Out_1 is corrected. In some embodiments, output signal Out_1 is stored in a first register. In some embodiments, error bit corrector 216 is used to multiply the current multiplicand, that is, one of the input columns of input array XIN1, with the current multiplicand, that is, multiply with one of the data columns in W&P array 204A, to form a corrected inner product, and write the correct inner product into the first register.

[0077] In some embodiments, block 265(4) and block 265(5) are executed substantially simultaneously. In some embodiments, block 265(4) is executed before block 265(5). In some embodiments, block 265(5) is executed before block 265(4).

[0078] Figure 3A is a schematic diagram illustrating the CIM region 324 of a digital CIM system according to some embodiments.

[0079] CIM region 324 is a part of CIM region 124D in Figure 1D, a part of CIM region 124E in Figure 1E, or another similar example. Multiple instances of CIM region 324 include CIM region 124D, CIM region 124E, or others similar to it. CIM region 324 is similar to CIM region 203(1) in Figure 2A. For simplicity, the following discussion will focus on the differences between CIM region 324 and CIM region 203(1), rather than their similarities.

[0080] CIM region 324 includes: a weighted and checksum (W&C) array 304A containing slices 318(0) to 318(Q-1), a multiplier array 337 containing multipliers 352(0) to 352(Q-1), an adder tree 308, a C1 inner product array 306A, and a trajectory inference data (LID) generator 326A, where Q is a positive integer. In some embodiments, Q is a power of two. In some embodiments, Q = 64. In some embodiments, Q is the square of a positive integer other than 64.

[0081] A component 342 comprising the W&C array 304A and the multiplier array 337 includes a slice 318 (Q-1) of the W&C array 304A and a multiplier 352 (Q-1) of the multiplier array 337. Components in component 342 include slice 318 (Q-1), multiplier 352 (Q-1), and multiplier 364 (Q-1). Although multiplier 364 (Q-1) is indicated in Figure 3A, it is noteworthy that multipliers 364(0) to 364(Q-2) are not indicated in Figure 3A for the sake of simplicity.

[0082] The slice 318 (Q-1) includes a two-dimensional weight array 320 (Q-1) and a two-dimensional checksum (CHK) array 323 (Q-1). The weight array 320 (Q-1) includes a CHK array 323 (Q-1) of one-bit memory cells 349 and one-bit memory cells 350. In Figure 3A, memory cells 349 and 350 are assumed to be SRAM cells. In some embodiments, memory cells 349 and 350 are memory cell types other than SRAM.

[0083] On a column-by-column basis, multiplier array 337 is used to receive a segment of a given column of weight bits from W&C array 304A. Accordingly, multipliers 352(0)~352(Q-1) of each multiplier array 337 are used to receive the corresponding segment of weight bits of that given column. Multiplier array 337 is further used to receive a given column of input bits from input XIN2 (refer to Figure 3B) corresponding to a given column of bits from W&C array 304A, and multiply them to obtain Q first inner products PRD2(0)~PR2(Q-1). The first inner products PRD2(0)~PR2(Q-1) are also provided to addition tree 308A.

[0084] Adder tree 308A adds Q first inner products PRD2(0)~PR2(Q-1) to generate output signal Out_2. Output signal Out_2 operates according to, for example, error bit detector, locator, and corrector unit 316 (refer to Figure 3M). Adder tree 308A is similar to adder tree 208 in Figure 2A. The operation of multiplier 352(Q-1) is also discussed in the descriptions of Figures 3E to 3G and Figure 4C.

[0085] In CIM region 324 of Figure 3A, multipliers 352(0)~352(Q-1) correspond to multipliers 251(0)~251(Q-1) of EM blocks 238(0)~238(Q-1) in Figure 2A. However, CIM region 324 does not include error bit detectors 210(0)~210(Q-1) and the multiplier array 254(0)~254(Q-1) of EM blocks 238(0)~238(Q-1) in Figure 2A. Instead, CIM region 324 includes multipliers 364(0)~364(Q-1), a C1 inner product array 306A, and an LID generator 326A, which are not present in Figure 2A.

[0086] On a column-by-column basis, multiplier 364 (Q-1) receives a column of weighted and checksum bits from slice 318 (Q-1). Multiplier 364 (Q-1) further receives a given column of input bits from input array XIN3 (refer to Figure 3D) corresponding to a given column of bits from slice 318 (Q-1), and multiplies them to obtain a second inner product D[i][j], where the second inner product D[i][j] represents a word of C1 array 306A, and i and j are non-negative integers.

[0087] On a column-by-column basis, the second inner product D[i][j] is cumulatively stored by multiplier 364(Q-1) and multipliers 364(0)~364(Q-2) in the other C1 array 306A (refer to the C1 array 306H in Figure 3H). The operation of multiplier 364 is also discussed in the description of Figure 3H and Figure 4C.

[0088] LID generator 326A operates on C1 array 306A and generates an Error Bit Trajectory Inference (BELI) signal containing the column sum signal Row_Sum and the column sum signal Col_Sum. The operation of LID generator 326A is also discussed in the descriptions of Figures 3I to 3M and Figures 4D to 4E. In Figure 3A, LID generator 326A is used to receive column (i) of C1 in C1 array 306.

[0089] For simplicity, Figure 3A assumes the following: there are N columns in slice 318 (Q-1), and also in weight array 320 (Q-1) and CHK array 323 (Q-1), with each column in weight array 320 (Q-1) storing K=12 bits. This assumption in Figure 3A is similar to that in Figure 2A.

[0090] Figure 3A also assumes that the CHK array 323(Q-1) has N columns, and each column of the CHK array 323(Q-1) stores an 11-bit word corresponding to the read bit word lines RBL

[12] ~RBL

[22] , such that the CHK array 323(Q-1) is Array. For a given column containing multiple segments in each of 318(0) to 318(Q-1), there is also a weight array 320(Q-1) and a CHK array 323(Q-1) containing multiple segments in each of 318(0) to 318(Q-1), where the 11-bit word / segment in the CHK array 323(Q-1) represents a checksum of one of the 12-bit words / segments in the weight array 320(Q-1).

[0091] The CHK array 323(Q-1) stores an S-bit word, where S is a positive integer and is assumed to be S=11 in Figure 3A. For further understanding, the CHK array 323(Q-1) is... Array. Accordingly, slice 318(Q-1) is... Array. It is worth noting that the input array XIN3 (refer to Figure 3D, 305D) is... The array, where S and Z are corresponding positive integers. In some embodiments, S is a positive integer other than 11. To concisely indicate the positions of the input array XIN3 in Figure 3D and the input XIN2 in Figures 3B to 3C, K+S is represented by V (e.g., the input array 305V in Figure 3A), that is, V=K+S, where V is a positive integer.

[0092] Regarding multiplier 352 (Q-1), iteratively (column-by-column), multiplier 352 (Q-1) is used to generate the output signal PRD2 (Q-1) to represent The inner product is the product of a column of weight array 320 (Q-1) and the corresponding column of input XIN2. More specifically, iteratively, multiplier 352 (Q-1) receives a column of weight bits from weight array 320 (Q-1) as the multiplicand and a column of input bits from input XIN2 as the multiplicand, and multiplies the multiplicand and multiplicand to obtain an inner product PRD2 (Q-1). Multiplier 352 (Q-1) receives a column of K data bits on read bit word lines RBL[0]~RBL

[11] and a column of K input bits on input XIN2, and multiplies them accordingly to obtain an inner product PRD2 (Q-1). The inner product PRD2 (Q-1) is a single character with K+K=2K bits. In the example of Figure 3A, the inner product PRD2 (Q-1) has 2K=12+12=24 bits.

[0093] Regarding multiplier 364(Q-1), iteratively (column-by-column), multiplier 364(Q-1) is used to generate the output signal PRD2(Q-1) to represent... One of the columns of slice 318 (Q-1) and The second inner product D[i][j] is obtained by multiplying the corresponding column of the input array XIN3. More specifically, iteratively, multiplier 364(Q-1) is used to receive a column of bits from slice 318(Q-1) as the multiplicand and a column of input bits from the input array XIN3 as the multiplicand, and multiplies the multiplicand and the multiplicand to obtain a second inner product D[i][j], wherein the latter is indicated to be provided to the C2 array 306A. Multiplier 364(Q-1) is used to receive a column of K+S data bits on the read bit word lines RBL[0]~RBL

[22] and a column of K+S input bits on the input array XIN3, and multiplies them accordingly to obtain a second inner product D[i][j].

[0094] Similar to Figure 2A, in practice, Figure 3A assumes that the double-error bit situation will not occur in slice 318 (Q-1). Accordingly, Figure 3A is used to detect and correct the single-error bit situation in slice 318 (Q-1), but not to detect and correct the double-error bit situation in slice 318 (Q-1). At least in some other embodiments disclosed herein, it is used to detect and correct the single-error bit situation, rather than the double-error bit situation, in other slices compared to slice 318 (Q-1).

[0095] However, the discussion in Figure 3A focuses on the number of bits in segments of multiple bit lines and multiple columns and bars, while Figures 3B to 3M focus on the discussion of characters.

[0096] Figure 3B is a schematic diagram illustrating the input XIN2 array 305B according to some embodiments.

[0097] The input XIN2 array 305B is operated according to a column checksum generator 379 to produce a column vector 386D, which is appended to the input XIN2 array 305B to form the input XIN3 array 305D in the 3D graph. The column checksum generator 379 is an example of one of the column checksum generators 179 in the 1D graph.

[0098] Input XIN2 array 305B is The array is given by the input XIN2 array 305B, where F and G are the corresponding positive integers. Each position (i,j) in the input XIN2 array 305B represents a corresponding character A[i][j], where i and j are the corresponding non-negative integers. Thus, the input XIN2 array 305B contains positions A[0][0], ..., A[F-1][G-1].

[0099] Figure 3C is a schematic diagram illustrating a column checksum generator 379 according to some embodiments.

[0100] Column checksum generator 379 is a column checksum generator used to generate a column checksum R_ChkSum_1. Column checksum generator 379 is column checksum generator 179 in Figure 1D or another similar example. Column checksum generator 379 is an array of G recursive adders 377 corresponding to the G-column input XIN2 array 305B. Each recursive adder 377 generates the sum of the characters in the corresponding column of the input XIN2 array 305B by adding the characters A[i][j] in the corresponding column on a column-by-column basis.

[0101] Each recursive adder 377 includes an adder 240 and a register 382. For each recursive adder 377, adder 240 receives characters from the corresponding column in register 382 and characters in register 382. Each instance of register 382 is initialized to store the value zero. At time t=0, adder 240 adds the corresponding character in column 0 of the input XIN2 array 305B to the character in register 382 (which was previously initialized to zero), and stores / overwrites the sum in register 382 at t=0. At time t=1, adder 240 adds the corresponding character in column 1 of the input XIN2 array 305B to the character in register 382 at time t=0, and stores / overwrites the sum in register 382 at t=1. At time t=F-1, adder 240 adds the corresponding character in column V-2 of the input XIN2 array 305B to the character corresponding to register 382 at time t=F-2, and stores / overwrites the sum in register 382 at t=F-1. The character at t=F-1 in the G instances of register 382 represents a vector R_ChkSum_1, which is appended to the input XIN2 array 305B in graph 3B to form the input XIN3 array 305D in graph 3D.

[0102] Figure 3D is a schematic diagram illustrating the input XIN3 array 305D according to some embodiments.

[0103] The input XIN3 array 305D is the result of appending the R_ChkSum_1 vector 386D to the input XIN2 array 305B. Therefore, the input XIN3 array 305D is... Array.

[0104] Figure 3E is a schematic diagram illustrating the weighted W1 array 304E according to some embodiments.

[0105] Weight array W1 304E is a weight array included as part of weight array 320 (Q-1) in Figure 3A, or another similar example. Weight array W1 304E operates according to column checksum generator 380 to generate weight arrays W2 304G and 305G in Figure 3G. Column checksum generator 380 is column checksum generator 180 in Figure 1D, or another similar example.

[0106] The weighted W1 array 304E is The array is defined by E, where E is a positive integer. Each position (i,j) of the weight array W1304E represents a corresponding character B[i][j]. Accordingly, the weight array W1304E contains positions B[0][0], ..., B[E-1][F-1].

[0107] Figure 3F is a schematic diagram illustrating a column checksum generator 380 according to some embodiments.

[0108] Column checksum generator 380 is a column checksum generator used to generate a column vector 387G to represent the checksum C_ChkSum_1, which is appended to the weight W1 array 304E to form the weight W2 array 304G in the 3G graph. Column checksum generator 380 is an example of column checksum generator 180 in the 1D graph or another similar example. Column checksum generator 380 contains an addition tree 308F. On a column-by-column basis, column checksum generator 380 is used to generate the sum of characters corresponding to a given column of the weight W1 array 304E and stores the sum in the corresponding column of the checksum C_ChkSum_1 column vector 387G in the 3G graph.

[0109] The output of the column checksum generator 380 is Where x is a non-negative integer variable, at time t=0, the column checksum generator 380 adds the corresponding characters in the 0th column of the weight W1 array 304E and stores the sum as the characters in the 0th column of the checksum column vector 387G. At time t=1, the column checksum generator 380 adds the corresponding characters in the first column of the weight W1 array 304E and stores the sum as the characters in the 0th column of the checksum column vector 387G. At time t=E-1, the column checksum generator 380 adds the corresponding characters in the (E-1)th column of the weight W1 array 304E, and stores the sum as the characters in the (E-1)th column of the checksum column vector 387G. In summary, ,…, The checksum column vector 387G is attached to the weight array W1 304E in the 3E graph to form the W2 array 304G in the 3G graph.

[0110] Figure 3G is a schematic diagram illustrating the weighted W2 array 304G according to some embodiments.

[0111] The weight array W2 306G is the result of adding a column of vectors (e.g., C_ChkSum_1 vector 387G) to the weight array W1 304E. Therefore, the weight array W2 306E is... Array.

[0112] Figure 3H is a schematic diagram illustrating the C1 inner product array 306H according to some embodiments.

[0113] C1 inner product array 306H is an example of one of the C1 inner product arrays 306A in Figure 3A. C1 inner product array 306H is... The array contains multiple memory cells 374. Each position (i,j) in the C1 inner product array 306H represents a corresponding character C[i][j]. Accordingly, the C1 inner product array 306H contains positions C[0][0],…,C[E-1][G-1].

[0114] The C1 inner product array 306H is the inner product of the weight W2 array 306E and the input XIN3 array 305D, i.e., C1 = W2 * XIN3. The C1 inner product array 306H is an example of an array whose columns are generated column-by-column by multiplier 352 in Figure 3A or similar arrays. The lower column 388 of the C1 inner product array 306H (refer to Figure 3J) represents the corresponding R_ChkSum_2 column vector to the R_ChkSum_1 column vector 386D in Figure 3D. The rightmost column 389 of the C1 inner product array 306H (refer to Figure 3J) represents the corresponding C_ChkSum_2 column vector to the C_ChkSum_1 column vector 387G in Figure 3G.

[0115] The C1 inner product array 306H operates according to at least the following description: a column vector generator (refer to Figures 3I to 3J) generates a column vector Row_Sum for use by the error bit detector, locator, and corrector (refer to Figure 3M), and a column vector generator (refer to Figures 3K to 3L) generates a column vector Col_Sum for use by the error bit detector, locator, and corrector (refer to Figure 3M).

[0116] Figure 3I is a schematic diagram illustrating a column sum generator 326I according to some embodiments.

[0117] Column sum generator 326I is a column checksum generator used to generate the column sum 304J in Figure 3J. Column sum generator 326I is the column sum generator 127 in Figure 1D or another similar example. Column sum generator 326I is an array of G recursive adders 377 corresponding to the G columns of C1 array 306H. Each recursive adder 377 generates the sum of the characters in the corresponding column of C1 array 306H by adding the characters C[i][j] in the corresponding column on a column-by-column basis. At time t=0, adder 240 adds the corresponding character in column 0 of C1 array 306H with the character in the corresponding register 382 (previously initialized to zero) and stores / overwrites the corresponding register 382 at t=0 sum. At time t=1, adder 240 adds the corresponding character in the first column of C1 array 306H and the character corresponding to register 382 at time t=0, and stores / overwrites the corresponding register 382 at t=1 for the sum. At time t=N-1, adder 240 adds the corresponding character in the (N-1)th column of C1 array 306H and the character corresponding to register 382 at time t=N-2, and stores / overwrites the corresponding register 382 at t=N-1 for the sum. The character at t=N-1 in the G instances of register 382 represents a column vector Row_Sum, which is labeled 304J in Figure 3J.

[0118] Figure 3J is a schematic diagram illustrating the Row_Sum vector 304J according to some embodiments.

[0119] In Figure 3J, the Row_Sum vector 304J is adjacent to and aligned below the C1 array 306H. In some embodiments, the Row_Sum vector 304J is a one-dimensional array of one-bit memory cells 356.

[0120] Figure 3K is a schematic diagram illustrating the column sum generator 328K according to some embodiments.

[0121] Column sum generator 328K is a column checksum generator used to generate the column vector sum Col_Sum in the 3L diagram. Column sum generator 328K is an example of column sum generator 128 in the 1D diagram or another similar example. Column sum generator 328K contains an addition tree 308K(Q-1). On a column-by-column basis, column sum generator 328K is used to generate the sum of characters in a given column corresponding to C1 array 306H, and stores the sum in the corresponding column of column vector Col_Sum 304L in the 3L diagram. At time t=0, column sum generator 328K adds the corresponding characters in column 0 of C1 array 306H and stores the resulting sum in characters. The sum is stored in column 0 of column vector Col_Sum 304L. At time t=1, column sum generator 328K adds the corresponding characters in column 1 of C1 array 306H and outputs the sum as a character. The sum is stored in the first column of column vector Col_Sum 304L. At time t=E-1, column sum generator 328K adds the corresponding characters in the (E-1)th column of C1 array 306H and outputs the sum as a character. It is stored in the (E-1)th column of the column vector Col_Sum 304L. In summary, ,…, The column vector Col_Sum 304L is indicated as Col_Sum 304L in the 3L diagram.

[0122] Figure 3L is a schematic diagram illustrating Row_Sum vector 304J and Col_Sum vector 304L according to some embodiments.

[0123] In Figure 3J, the Col_Sum vector 304L is adjacent to and aligned to the right of the C1 array 306H. Similarly, in Figure 3J, the Row_Sum vector 304J is adjacent to and aligned below the C1 array 306H. In some embodiments, the Col_Sum vector 304L is a one-dimensional array of one-bit memory cells 357.

[0124] Figure 3M is a schematic diagram illustrating a digital CIM system 300D according to some embodiments.

[0125] CIM system 300D is CIM system 100D in Figure 1D, CIM system 100E in Figure 1E, or other similar examples. CIM system 300D includes CIM region 324 and error bit detector, locator, and corrector (DLC) unit 316.

[0126] CIM region 324 is CIM region 124D in Figure 1D, a combination of CIM region 124E and region 114 in Figure 1E, or another similar example. DLC unit 316 is an example of bit error detector, locator and corrector 130 in Figure 1D, Figure 1E, or other similar examples.

[0127] DLC unit 316 is used to receive output signals Out_2, R_ChkSum, C_ChkSum, Row_Sum, and Col_Sum from CIM region 304M and operates these signals according to flowchart 383. In Figure 3M, flowchart 383 is shown as an enlarged view of DLC unit 316. The operation of DLC unit 316 is also discussed in the description of Figures 4C to 4G.

[0128] Flowchart 383 includes blocks 384(1) to 384(5). In some embodiments, the flow proceeds from block 560 of the 5D diagram to block 384(1), as indicated by page break reference 392.

[0129] At block 384(1), a decision will be made regardless of whether the following conditions are met: (1) R_ChkSum = Row_Sum and (2) C_ChkSum = Col_Sum. If the decision at block 384(1) is yes, the process proceeds to block 384(2).

[0130] At block 384(2), the process stops because there are no bit errors, so the output signal Out_2 does not need to be corrected for bit errors. However, if block 384(1) determines the result to be no, it means that there are bit errors, making it necessary to correct for bit errors, and the process proceeds to block 384(3).

[0131] In block 384(3), the faulty bit is located. Figures 4C through 4F provide an example illustrating how this location is performed. From block 384(3), the process proceeds to block 384(4) and block 384(5).

[0132] In block 384(4), the erroneous bits identified in block 384(3) in the W&C array 304A are updated. In some embodiments, the DLC unit 316 updates the characters identified as having erroneous bits in block 384(3) by writing undamaged characters to the memory unit 349, for example, by copying undamaged data bits from a source copy or permanent copy in the W&C array 304A. For example, the source copy or permanent copy is stored outside the CIM area containing the W&C array 304A.

[0133] In block 384(5), the signal value of output signal Out_2 is corrected. In some embodiments, output signal Out_2 is stored in the first temporary register. In some embodiments, DLC unit 316 is used to multiply the current multiplicand, that is, a corresponding input column of input XIN1, and the current multiplicand, that is, a corresponding data column in W&C array 304A, to form a corrected inner product, and write the correct inner product into the first temporary register.

[0134] In some embodiments, block 384(4) and block 384(5) are executed substantially simultaneously. In some embodiments, block 384(4) is executed before block 384(5). In some embodiments, block 384(5) is executed before block 384(4).

[0135] According to another approach, in contrast to CIM region 324 or other similar CIM systems, and with DLC unit 316 comprising a component, (1) a weight array is stored in a first region of a first die, and (2) correspondingly on at least one second die, error bit detection, location, and correction (DLC) is performed entirely by the processor and associated RAM. According to this other approach, for error bit DLC to be performed, a large amount of data is transferred from the weight array on the first die to the processor on the second die (out-of-die transfer), resulting in significant transmission delays and thus significantly reducing the speed of error bit DLC. In contrast, in at least some embodiments, the amount of out-of-die transfer associated with error bit detection, location, and correction (DLC) is substantially reduced compared to the other approach, thus achieving substantially faster error bit DLC compared to the other approach.

[0136] In other words, according to at least some embodiments, the column checksum generator 379, the input XIN3 array 305D, the column checksum generator 380, the weight W2 array 304G, the C1 array 306H, the column sum generator 326I, the column vector Row_Sum 304J, the column sum generator 328K, and the column vector Col_Sum 304L are contained within a CIM region (e.g., CIM region 324) as a W&C array 304A, a multiplier array 337, and an adder tree 308A, which increases the proximity of arithmetic operators to storage (AOS) compared to another approach. According to at least some embodiments, the proximity of (1) storage locations to (2) arithmetic units accessing / manipulating the storage locations within the aforementioned similar CIM regions effectively improves AOS proximity compared to another approach, thereby substantially reducing the amount of off-chip transfers included as part of the error bit DLC according to at least some embodiments, and thus achieving a substantially faster DLC compared to another approach.

[0137] More specifically, according to at least some embodiments, the proximity of (1) storage locations (represented by input XIN3 array, weight W2 array 304G, C1 array 306H, column vector Row_Sum 304J, and column vector Col_Sum 304L, or others similar) to (2) arithmetic units (represented by column checksum generator 379, column checksum generator 380, column sum generator 326I, and column sum generator 328K) in the aforementioned similar CIM regions effectively improves AOS proximity compared to another approach, thereby substantially reducing the amount of off-chip transmissions included as part of the error bit DLC according to at least some embodiments, and thus achieving substantially faster DLC compared to another approach.

[0138] Figure 4A is a block diagram illustrating a simple example of one of the same bit generators 466 in some embodiments.

[0139] The parity bit generator 466 is a simplified example of how the parity encoder 259 (Q-1) or other similar mutex or logic gate 260 (Q-1) generates parities stored in the parity array 221 (Q-1) in Figure 2B as one of the sources.

[0140] In Figure 4A, the parity array 422 stores the parity bits corresponding to the data bits in the weight array 420. The weight values ​​stored in the weight array 420 are represented in decimal, with their equivalent binary representation enclosed in parentheses. For example, the weight value at the intersection of the second column and the second bar in the weight array 420 (position (2,2)) is represented as 3 in decimal, and 11 in binary. That is, position (1,1) in the parity array 422 corresponds to position (1,1) in the weight array 420. Extending the above example, the value stored at position (1,1) in the parity array 422 is 0. This means that a logical mutual exclusion OR operation is applied to the value 3 stored at position (1,1) in the weight array 420, which is represented as (1,1) in binary, i.e., 1^1 = 0.

[0141] The corresponding value stored at each position in the corresponding array 422 represents the result of a logical OR operation on the binary representation of the corresponding position in the weight array 420. For example, position (1,1) in the corresponding array 422 corresponds to position (1,1) in the weight array 420. Extending this example, the value stored at position (1,1) in the corresponding array 422 is 0, which means that a logical OR operation is applied to the value 3 stored at position (1,1) in the weight array 420, which is represented by binary (1,1), i.e., 1^1 = 0.

[0142] Figure 4B is a block diagram illustrating a simple example of one of the bit error detection 468 embodiments.

[0143] Error detection 468 refers to how error bit detector 210 (Q-1) or something similar detects bit errors in a corresponding column of weight array 220 (Q-1).

[0144] Figure 4B contains a weight array 472 corresponding to weight array 420 in Figure 4A, the difference being that weight array 472 is assumed to have erroneous bits. More specifically, the weight value at position (1,1) in weight array 472 is corrupted and incorrectly displays 2 (10) instead of 3 (11).

[0145] Figure 4B further includes: the same-position array 422, the input array XIN1 470, the inner product array 474, and the error bit detection array 476 from Figure 4A. The inner product array 474 represents the result of multiplying the input array XIN1 470 and the weight array 472. The error bit detection array 476 represents the result of performing a logical mutually exclusive OR operation on the binary representation shown at the corresponding position in the weight array 472 and the same-position bit value at the corresponding position in the same-position array 422.

[0146] A value of 0 at a given position in one of the error bit detection arrays 476 indicates that there is no error bit at the corresponding position in the weight array 472. Conversely, a value of 1 at a given position in one of the error bit detection arrays 476 indicates that there is an error bit at the corresponding position in the weight array 472.

[0147] Regarding the error bit detection array 476, for example, position (1,1) of the error bit detection array 476 stores the result of applying a logical mutual exclusion OR operation to the value in position (1,1) of the weight array 472 and the value in position (1,1) of the parity array 422. Specifically, the value stored in position (1,1) of the error bit detection array 476 is 1, which represents the result of applying a logical mutual exclusion OR operation to the damaged value 2 stored in position (1,1) of the weight array 420 in binary representation (10) and the parity value 1 in position (1,1) of the parity array 422, that is, 1^0^0=1, where the insertion (tone symbol) symbol (^) is used to indicate that a logical mutual exclusion OR operation is applied to the bits of the bit string 100. Conversely, if the position (1,1) of the weight array 420 is not damaged, that is, if the position (1,1) stores 3 (11) instead of 2 (10), then the position (1,1) of the parity array 422 displays the parity value 0, that is, 1^1^0=0, which means there is no erroneous bit.

[0148] Figure 4C is a block diagram illustrating a simple example of checksum generator and array multiplication according to some embodiments.

[0149] In Figure 4C, array C2 is the inner product of input array XIN3 and weight array W2, where C2 = XIN3 * W2. As shown in C2, there are erroneous bits as discussed below.

[0150] The input array XIN3 shows the result of adding column vector R_ChkSum_1 to the input array XIN2, such that... , where the symbol Used to represent append operators and text string formats This is used to represent B being appended to A. For example, in the input array XIN3, XIN3 position (3,1)=4, which means XIN3 position (1,1)=1 plus XIN3 position (2,1)=3. The third / bottom column of the input array XIN3 represents the column checksum R_ChkSum_1.

[0151] The weight array W2 shows the result of applying the additional column vector C_ChkSum_1 to the weight array W1, which makes... For example, in the weighted array W2, position (3,1)=3, which means position (1,1)=1 plus position (1,2)=4. The third / rightmost column of the weighted array W2 represents the column checksum C_ChkSum_1.

[0152] Regarding the multiplication C2=XIN3*W2, for example, consider the C2 array with C2 position (2,1)=11. The character at C2 position (2,1)=11 is the sum of (1) the inner product of (XIN3 position (2,1)=3 and W2 position (1,1)=1 plus (2) the inner product of (XIN3 position (2,2)=4 and W2 position (1,2)=2.

[0153] The third / bottom column of array C2 represents the column checksum R_ChkSum_1 of the bottom column of input array XIN3, compared to column checksum R_ChkSum_2. The third / rightmost column of array C1 306H(Q-1) (refer to Figure 3J) represents the column checksum C_ChkSum_1 vector 387G of Figure 3G, compared to column checksum C_ChkSum_2 vector.

[0154] For the purpose of identifying, locating, and correcting faulty bits (refer to Figures 4D to 4G), the C2 position (1,1) shown in Figure 4C is a potential faulty bit, that is, it appears after the initial C2 array is stored in the memory cell. Without a faulty bit, the C2 position (1,1) should be C2 position (1,1) = 5. However, the example above assumes that a potential faulty bit makes the C2 position (1,1) C2 position (1,1) = 4.

[0155] Error columns at positions C2 (1,1) or (2,1) can be detected based on the checksum character at position (3,1), as shown in Figure 4D. The identification of an error column implicitly suggests that an error bit occurred at one of the multiple non-checksum characters stored in the corresponding column. Similarly, error columns at positions C2 (1,2) or (2,2) can be detected based on the checksum character at position (3,2).

[0156] Error columns at positions C2 (1,1) or (2,1) can be detected based on the checksum character at position (3,1), see Figure 4E. Similarly, error columns at positions C2 (1,2) or (2,2) can be detected based on the checksum character at position (3,2). The identification of an error column implicitly suggests that the error bit occurred at one of the multiple non-checksum characters stored in the corresponding column.

[0157] Error bits in the C2 array can be located when they are at the intersection of a column with an identified error field and a column with an identified error column, as shown in Figure 4F. Once located, the error bit can be corrected, as shown in Figure 4G.

[0158] Figure 4D is a block diagram illustrating a simple example of error column location based on some embodiments.

[0159] Figure 4D (an example extending from Figure 4C) shows a determination of block 384(3)D corresponding to block 384(3)D in Figure 3M or another similar block 484(3)D. In block 484(3)D, on a character-by-character basis, it is determined whether the sum (3,j) matches position (3,j) in the column checksum R_ChkSum in the C2 array.

[0160] To generate the sum (3,1), the current column-by-column sum is generated using the non-checksum characters in column 1 of the C2 array. That is, the incorrect C2 positions (1,1)=4 and (2,1)=11 are used to generate the sum (3,1)=15. Here, the phrase "currently" means that the sum is based on the current version of the C2 array. Similarly, to generate the sum (3,2), the column-by-column sum is generated using the non-checksum characters in column 2 of the C2 array. And to generate the sum (3,3), the column-by-column sum is generated using the non-checksum characters in columns 1 and 2 of column 3 of the C2 array.

[0161] The sum (3,2) is used to match (whether equal) the checksum character at position (3,2) in C2 to indicate that there is no error column in the second column of the C2 array. The sum (3,3) is used to match (whether equal) the checksum character at position (3,3) in C2 to indicate that there is no error column in the third column of the C2 array. However, the sum (3,1) is used to mismatch (whether unequal) the checksum character at position (3,1) in C2 to indicate that there is an error column in the first column of the C2 array.

[0162] In Figure 4E, the sum (3,1) = 15 represents the first mismatch, where the checksum character at position C2 (3,1) = 16 does not match the sum (3,1) = 15. The reason for this first mismatch is that the character at position C2 (1,1) represents a potentially erroneous bit. This first mismatch is used to locate and correct the erroneous bit at position C2 (1,1) (see Figures 4D, 4F, and 4G).

[0163] Figure 4E is a block diagram illustrating a simple example of error bar location according to some embodiments.

[0164] Figure 4E (an example extending from Figure 4C) shows a determination of block 384(3)E corresponding to block 384(3) in Figure 3M or another similar block 484(3)E. In block 484(3)E, on a character-by-character basis, it is determined whether the sum (i,3) matches the position (i,3) in the column checksum C_ChkSum in the C2 array.

[0165] To generate the sum (1,3), the current column-by-column sum is generated using the non-checksum characters in the first column of the C2 array. That is, the incorrect C2 positions (1,1)=4 and (1,2)=10 are used to generate the sum (1,3)=14. Similarly, to generate the sum (2,3), the column-by-column sum is generated using the non-checksum characters in the second column of the C2 array. And to generate the sum (3,3), the column-by-column sum is generated using the non-checksum characters in the first and second columns of the third column of the C2 array.

[0166] In block 484(3), the sum (2,3) is used to match the column checksum character at position (2,3) of C2 to indicate that there is no error column in the second column of the C2 array. The sum (3,3) is used to match the column checksum character at position (3,3) of C2 to indicate that there is no error column in the third column of the C2 array. However, the sum (1,3) is used to mismatch the column checksum character at position (1,3) of C2 to indicate that there is an error column in the first column of the C2 array.

[0167] In Figure 4E, the sum (1,3) represents a second mismatch where the checksum character at position C2 (1,3) = 15 does not match (1,3) = 14. The reason for this second mismatch is that the character at position C2 (1,1) represents a potential erroneous bit. This second mismatch is used to locate and correct the erroneous bit at position C2 (1,1) (refer to Figures 4E through 4G).

[0168] Figure 4F is a block diagram illustrating a simple example of error bit localization according to some embodiments.

[0169] Figure 4F (an example extending from Figure 4D to Figure 4E) shows a portion corresponding to block 384(4) in Figure 3M or another similar block 484(4). In block 484(4), an intersection point can be determined by the following descriptions: the column in Figure 4D where an error column is identified, namely column 1 of array C2, and the column in Figure 4E where an error bar is identified, namely bar 1 of array C2. The intersection point of column 1 and bar 1 of array C2 is position C2 (1,1). Therefore, the intersection point of column 1 and bar 1 of array C2 locates the error element at position C2 (1,1).

[0170] Figure 4G is a block diagram illustrating a simple example of error bit correction according to some embodiments.

[0171] Figure 4G (an example extending from Figure 4D to Figure 4F) shows a portion corresponding to block 384(5) in Figure 3M or another similar block 484(5). In block 484(5), for the non-checksum position (i,j) in the C2 array determined in block 484(4) of Figure 4D, the difference is... The decision rests between the sum (3,1) and the position of C2 (3,1), such that... =Sum(3,1) - C2 position(3,1). Or, difference. The decision rests between the sum (1,3) and the position of C2 (1,3), such that... =Sum(1,3)-C2 position(1,3). Next, the correction value C'[1][1] used to store / overwrite the value at C2 position is obtained by using the difference. Add the error value C[1][1] at position (1,1) of C2, such that C'[1][1] = C[1][1] + .

[0172] Figure 5A is a flowchart 500 illustrating a method for manufacturing a memory device according to some embodiments.

[0173] The method of flowchart 500 can be implemented, for example, using an electronic design automation (EDA) system 600 (refer to Figure 6 in the discussion below) and an IC manufacturing system 700 (refer to Figure 7 in the discussion below). Examples of devices that can be produced according to the method of flowchart 500 include devices similar to those described herein, based on this disclosure and illustrations.

[0174] In Figure 5A, the method of flowchart 500 includes blocks 502 through 504. In block 502, a layout diagram is generated, specifically including a CIM system and / or areas corresponding to the bare die disclosed herein, or one or more similar layout diagrams. Block 502 may be implemented, for example, using EDA system 600 (refer to Figure 6 discussed below). The process proceeds from block 502 to block 504.

[0175] In block 504, according to the layout diagram, at least one of the following is performed: (A) generating one or more lithography exposures, or (B) fabricating one or more lithography masks, or (C) fabricating one or more components in a layer of a device (e.g., a semiconductor device). Refer to the discussion of IC manufacturing system 700 in Figure 7 below.

[0176] Figure 5B is a flowchart 508 illustrating a method of operating a CIM system according to some embodiments.

[0177] Some examples of CIM systems can be operated according to the method in flowchart 508. The CIM system includes CIM system 100A in Figure 1A, CIM system 100C in Figure 1C, or other similar CIM systems. Flowchart 508 includes blocks 510 to 540.

[0178] In block 520, for weight bits in the weight array, the corresponding parity bits are encoded by a parity encoder and stored in the parity array. An example of a weight array is weight array 220 (Q-1) of Figure 2B, or other similar arrays. An example of a parity array is parity array 221 (Q-1) of Figure 2B, or other similar arrays. An example of a parity encoder is parity encoder 259 (Q-1) of Figure 2B, or other similar encoders. Block 510 contains block 512. Within block 510, the process proceeds to block 512.

[0179] In block 512, the peer encoder encodes the peer bits by performing a logical OR operation on the weighted bits using a mutex OR gate. An example of a mutex OR gate is mutex OR gate 260 (Q-1), or other similar arrays. The process leaves block 512 and proceeds to block 514.

[0180] In some embodiments, the execution of block 514 and the execution of block 510 occur nearly simultaneously. In some embodiments, the execution of block 514 and the execution of block 510 occur at different times.

[0181] In block 514, a column of weight segments from the weight and parity array and a column of input from the input array are received by the multiplier, the column of segments and the column of input representing the multiplicand and the multiplier, respectively. An example of a weight and parity array is slice 218(Q-1) of the W&P array 204A in Figure 2A, or other similar arrays. An example of a column of weight and parity array is composed of data bits from any N columns of slice 218(Q-1) selected from one of the read word lines RWL[0] to RWL[N-1], or other similar composition. An example of an input array is the input array XIN1 in Figure 2A, or other similar arrays. Examples of multipliers include multipliers 251(0) to 251(Q-1) located in the corresponding EM blocks 238(0) to 238(Q-1) of the EM array 236 in Figure 2A, or other similar multipliers. The process proceeds from block 514 to block 516 and then to block 518.

[0182] In block 516, the multiplicand and multiplier are multiplied by a multiplier to form an inner product. Examples of some inner products include the Q inner products PRD1(0) to PRD1(Q-1) generated by the EM array 236, or other similar inner products. The process proceeds from block 516 to block 522 (which will be explained after the discussion of blocks 518 to 520).

[0183] In block 518, error bits in the multiplicand are detected by an error bit detector, and the result of the error bit detection is displayed by generating a first flag signal. Examples of some error bit detectors include error bit detectors 210(0) to 210(Q-1) corresponding to EM blocks 238(0) to 238(Q-1) in Figure 2A, or other similar detectors, wherein each error bit detector 210(0) to 210(Q-1) generates one instance of a corresponding first error flag signal. Examples of some first error flag signals include Q instances of flag FLG1 generated accordingly by error bit detectors 212(0) to 212(Q-1), or other similar instances. Within block 518, the process proceeds to block 520.

[0184] In block 520, a logical OR operation is performed on the multiplicand using a mutex or gate to generate a first error flag. Examples of mutexes or gates include mutexes or gates 253(0) to 253(Q-1) in the corresponding EM blocks 238(0) to 238(Q-1) in Figure 2A, or other similar mutexes or gates. An example of the result of a logical OR operation is the state of flag FLG1 generated by mutex or gate 253(Q-1), i.e., whether flag FLG1 is declared (FLG1=1) to indicate the presence of an error bit or whether flag FLG1 is not declared (FLG1=0) to indicate the absence of an error bit, or other similar states. The process leaves block 518 and proceeds to block 522.

[0185] Block 518 is performed concurrently with block 516. According to another method corresponding to block 518, error bit detection is performed before the multiplication operation, comparing the error bits located in the memory cells with respect to the weight array 220 (Q-1). According to another method, error bit detection before the multiplication operation requires two operation cycles. Conversely, performing blocks 516 and 518 concurrently, according to at least some embodiments, requires only one operation cycle, which is one operation cycle faster than the other method.

[0186] In block 522, a decision will be made regardless of whether an error bit has been detected. An example of determining whether an error bit has been detected is whether the flag FLG1 has been declared, that is, whether the flag FLG1 is set to 1 by the mutex or logic gate 253 (Q-1), or some other similar state. Based on the decision made in block 522, the process proceeds to block 524 or block 528.

[0187] If the decision in block 522 is negative, meaning no error bit was detected, the process proceeds to block 524. In block 524, the inner product PRD1 is selected by a selector, not a reference value. An example of a reference value is the reference value REF in Figure 2A, or other similar reference values. An example of a selector is the multiplier array 251 (Q-1) in Figure 2A, or other similar arrays, where the multiplier array 251 (Q-1) receives the inner product PRD1 (Q-1) and the reference value REF as inputs, and the flag FLG1 as a control signal, and further selects the inner product PRD1 (Q-1) when FLG=0. The process proceeds from block 524 to block 526, and ends at block 526.

[0188] If the decision in block 522 is yes, meaning an error bit has been detected, the process proceeds to block 528. In block 528, the reference value is selected by the selector, not the inner product PRD1. Extending the example of block 524, when FLG=1, the multiplier array 251(Q-1) is further used to select the reference value REF. The process then proceeds from block 528 to block 530.

[0189] In block 530, multiple trajectory inference signals are generated by a trajectory inference data generator. An example of a trajectory inference data generator is trajectory inference data generator 212 in Figure 2A, or other similar generators, wherein trajectory inference data generator 212 is used to generate error indicators and second error flags. An example of an error indicator is indicator PRD1(Q-1) in Figure 2C, or other similar indicators. An example of a second error flag is flag FLG2, or other similar flags. Within block 530, the process proceeds to blocks 532 and 534.

[0190] In block 532, the Q:P encoding of the Q instances of the first flag is performed by an encoder to generate an error index. Examples of some of the Q instances of the first flag are the Q instances of flag FLG1 generated by the corresponding error bit detectors 212(0) to 212(Q-1) in EM blocks 238(0) to 238(Q-1). An example of an encoder is the Q:P encoder 262 in Figure 2C, or other similar encoders, where encoder 262 generates an example of an error index, namely the index EPT. The process proceeds from block 532 to block 538 (to be described after the discussion of blocks 534 to 536).

[0191] In block 534, an erroneous fragment is detected by an erroneous fragment detector, which generates a second error flag based on Q instances of the first flag. An example of an erroneous fragment detector is erroneous fragment detector 213 in Figure 2C, or other similar detectors, which generate one example of the second error flag (i.e., flag FLG2). Within block 534, the process proceeds to block 536.

[0192] In block 536, an OR gate performs a logical OR operation on Q instances of the first flag to generate a second error flag. An example of an OR gate is OR gate 263 of the trajectory inference data generator 212, or other similar gates, which generate one example of the second flag (i.e., flag FLG2). An example of the result of a logical OR operation is the state of flag FLG2 generated by OR gate 263, that is, whether flag FLG2 is declared (FLG2=1) to indicate the presence of a truncating error or whether flag FLG2 is not declared (FLG2=0) to indicate the absence of a truncating error, or other similar states. The process leaves block 534 from block 536 and proceeds to block 538.

[0193] In block 538, in response to the second flag indicating the presence of an erroneous fragment, flag FLG2 is declared (FLG2=1) to indicate the presence of an erroneous fragment, which will be located by an error bit corrector. An example of an error bit corrector is error bit corrector 216 in the second-dimensional diagram, or other similar correctors. An example of error bit location is block 265(3) in flowchart 290 of the second-dimensional diagram, or other similar blocks. The process proceeds from block 538 to block 540.

[0194] In block 540, the error bits are corrected by the error bit corrector. An example of an error bit correction is one or more blocks 265(4)~265(5) in flowchart 290 of Figure 2D, or other similar blocks. The process proceeds from block 538 to block 526, and ends at block 526.

[0195] Figure 5C is a flowchart 543 illustrating a method for manufacturing a CIM system according to some embodiments.

[0196] Flowchart 543 is an example of block 504 in Figure 5A. The method of flowchart 543 can be implemented, for example, using IC manufacturing system 700 (refer to Figure 7 discussed below) according to some embodiments. Examples of digital CIM systems that can be manufactured according to flowchart 543 include CIM systems based on the content of this disclosure, or other similar systems.

[0197] Flowchart 543 includes blocks 545 to 547. In block 545, a first structure including a first component is formed in a first region of a first semiconductor die. The first component includes a plurality of memory cells corresponding to storage units, a multiplier, and a first error bit detector. Furthermore, some of the plurality of memory cells are configured as corresponding plurality of first arrays for storing first data bits. Some of the plurality of memory cells are configured as corresponding plurality of second arrays for storing corresponding bits of the first data bits. The first component is organized into a plurality of first groups, wherein each first group includes a corresponding first array, a second array, a multiplier, and a first error bit detector.

[0198] Regarding block 545, some examples of the first structures include structures of semiconductor devices, such as transistors, structures that facilitate transistor coupling, or other similar structures. In some embodiments, the aforementioned structures including transistors and structures that facilitate transistor coupling are both formed on one or more first layers, and are collectively referred to as transistor layers. Examples of transistors include field-effect transistors (FETs) such as P-type metal-oxide-semiconductor (PMOS) field-effect transistors (PFETs), N-type metal-oxide-semiconductor (NMOS) field-effect transistors, or other similar transistors (NFETs).

[0199] Some structures that include transistors include: multiple active regions of a semiconductor layer, well regions around selected active regions, source / drain (S / D) regions in the active regions, channel regions in the active regions between corresponding S / D region pairs, gate structures on some of the active regions, and (selectively) buried gate (BG) structures under some of the active regions, or other similar structures.

[0200] Examples of structures that facilitate transistor coupling include: a metal-to-source (MD) junction coupled above and to the S / D region and (optionally) a contrast buried MD (BMD) coupled below and to the S / D region; a metal-to-gate (MG) junction coupled to a gate structure and (optionally) a contrast buried MG (BMG) coupled to a BG structure; a via-to-MD (VD) junction coupled to an MD junction and a contrast buried VD (BVD) coupled to a BMD junction; a via-to-MG (VG) junction coupled to an MG junction and a contrast buried VG (BVG) coupled to a BMG junction; a region interconnect (LI) structure coupled to, for example, an MD junction and / or a gate structure; and a buried LI (BLI) structure coupled to, for example, a BMD junction and / or a BG gate structure, or other similar structures.

[0201] Regarding block 545, an example of a first set of multiple memory cells includes memory cell 245 of diagram 2A, or other similar cells. An example of a second set of multiple memory cells includes memory cell 246 of diagram 2A, or other similar cells. An example of a first array is weight array 220 (Q-1) of diagram 2A, or other similar arrays. An example of a second array is weight array 221 (Q-1) of diagram 2A, or other similar arrays. An example of a multiplier is multiplier 251 (Q-1) of diagram 2A, or other similar multipliers. An example of a first error bit detector is error bit detector 212 (Q-1) of diagram 2A, or other similar detectors. An example of a first group is a group containing slice 218 (Q-1) and EM block 238 (Q-1) of diagram 2A, or other similar combinations. The process proceeds from block 545 to block 547.

[0202] In block 547, multiple interconnections consist of multiple first components, thus producing at least: for each first group, a multiplier is used to multiply the input data bits with corresponding first data bits, and for each first group, a first error bit detector is used to detect error bits in the corresponding first data bits based on the correlation of the corresponding co-bit bits.

[0203] Regarding block 547, examples of interconnected formation include signal segments and / or PG segments formed in metallization layers correspondingly above and (selectively) below the transistor. In some embodiments, the signal segments may be conductive and used to carry multiple signals including input / output (I / O) signals, control signals, or other similar signals. In such embodiments, the signal segments are correspondingly coupled to a VD junction, an MG junction, (selectively) a BVD junction, (selectively) a BVG junction, or other similar junctions. In some embodiments, the PG segments may be conductive and used to be energized by a corresponding reference voltage of the power grid (PG). In such embodiments, the PG segments are correspondingly coupled to a VD junction, an MG junction, (selectively) a BVD junction, (selectively) a BVG junction, or other similar junctions. For example, some of the aforementioned PG segments are energized with a first reference voltage, such as VDD, and some of the PG segments are energized with a second reference voltage, such as VSS.

[0204] In some embodiments, with respect to block 547, the formation of mutual coupling in the first component further produces at least for each first group (e.g. 244(1)), a first error bit detector (e.g. 210(Q-1)) for detecting error bits in the corresponding first data bits based on the corresponding first data bits and the corresponding bit.

[0205] In some embodiments, with respect to block 547, the formation of mutual coupling in the first component (e.g., block 547) further generates at least for each first group (e.g., 244(1)), a first error bit detector (e.g., 210(Q-1)) is further used to detect errors while multiplying with a multiplier (e.g., 251(N-1)).

[0206] In some embodiments, with respect to block 547, the formation of mutual coupling in the first components (e.g., block 547) further generates at least for each first group (e.g., 244(1)), the CIM system is used to locate the error bit after detection.

[0207] In some embodiments, with respect to block 547, the formation of mutual coupling in the first component (e.g., block 547) further produces, at least for each first group (e.g., 244(1)), the CIM system to correct erroneous bits after positioning.

[0208] In some embodiments, flowchart 543 further includes a first block located in a first region (e.g., 103) of a first semiconductor die (e.g., 102A and 102C(1)), a second structure including a second component, and a second component including a multiplier array (e.g., 254A). In such an embodiment, each first group (e.g., 244(1)) further includes a corresponding one of the multiplier arrays (e.g., 254A). In such an embodiment, flowchart 543 further includes a second block in which a plurality of components are mutually coupled to at least a second component or a first component and generate, for at least for each first group (e.g., 244(1)), a multiplier array (e.g., 254A) for selecting, for example (i) an inner product generated by a multiplier (e.g., 251(N-1)), or (ii) a preset value based on an output signal generated by a first error bit detector (e.g., 210(Q-1)).

[0209] In some embodiments, flowchart 543 further includes a first block encompassing a first region (e.g., 103) or a second region (e.g., 155(1)) of a first semiconductor die (e.g., 102A and 102C(1)) or a first region (e.g., 155(2)) of a second semiconductor die (e.g., 102C(2)), a second structure including a second component, and a second structure including a co-encoder (e.g., 158). In such an embodiment, each first group (e.g., 244(1)) further includes a corresponding co-encoder (e.g., 158). In such an embodiment, flowchart 543 further includes a second block in which mutual coupling formed by at least a second component or a first component generates at least one co-encoder (e.g., 158) in each first group (e.g., 244(1)) for encoding a corresponding co-encoder bit based on a corresponding first data bit.

[0210] In some embodiments, flowchart 543 further includes a first region (e.g., 103) encompassing a first semiconductor die (e.g., 102A and 102C(1)), a third structure including a third component, a third component including a mutually exclusive OR (e.g., XOR) logic gate (e.g., 260(x)), and a first block correspondingly including a mutually exclusive OR logic gate (e.g., 260(x)) as part of a co-position encoder (e.g., 158). In such embodiments, each first group (e.g., 244(1)) further includes a plurality of mutually exclusive OR logic gates (e.g., 260(x)) of one of them. In such an embodiment, flowchart 543 further includes a second block comprising at least a third component and a first or second component that are mutually coupled, and generates, for at least each first group (e.g., 244(1)), a mutex or logic gate (e.g., 253(x)) on a column-by-column basis for performing operations including receiving a column of first data bits as input, generating corresponding parity bits, and storing the parity bits in a corresponding column of a second array (e.g., 221(N-1)).

[0211] In some embodiments, flowchart 543 further includes a first block in which a second structure is formed in a first region of a first semiconductor die (e.g., 102A and 102C(1)), the second structure including a second component, the second component including a mutex or logic gate (e.g., 253(x)), the mutex or logic gate (e.g., 253(x)) being included as part of a corresponding first error bit detector (e.g., 210(Q-1)). In such an embodiment, each first group (e.g., 244(1)) further includes one of a plurality of mutex or logic gates (e.g., 253(x)). In such an embodiment, flowchart 543 further includes a second block in which at least the second component or the first component is formed to be mutually coupled and to generate, at least for each first group (e.g., 244(1)), its mutex or logic gate (e.g., 253(x)) is used to receive a first data bit and a corresponding bit as input and thereby generate an output signal representing a first flag signal (e.g., FLG1), this flag signal can declare the presence of an error bit.

[0212] In some embodiments, flowchart 543 further includes a first block in which a second structure comprising a second component is formed in a first region (e.g., 103) or a second region (e.g., 155(1)) of a first semiconductor die (e.g., 102A and 102C(1)) or a first region (e.g., 155(2)) of a second semiconductor die (e.g., 102C(2)). The second structure includes a trajectory inference data generator (e.g., 112). In such an embodiment, flowchart 543 further includes a second block in which at least the second component or the first component forms an interconnected trajectory inference signal generator (e.g., 112) for generating one or more error bit trajectory inference signals (e.g., EPT and FLG2 of diagram 2C).

[0213] In some embodiments, flowchart 543 further includes a first block in which a third structure comprising a third component is formed in a second region (e.g., 155(1)) of a first region (e.g., 103) or a first semiconductor die (e.g., 102A and 102C(1)). The third component includes a Q:P encoder (e.g., 262) and a second error bit detector (e.g., 213), which are included in a trajectory inference signal generator (e.g., 112). In such an embodiment, flowchart 543 further includes a second block in which at least the third component, the second component, or the first component are mutually coupled and generate at least: for each first group (e.g., 244(1)), a first error bit detector (e.g., 210(Q-1)) is further used to generate an output signal representing a first flag signal (FLG1), which can declare the presence of an error bit. For each first group (e.g., 244(1)), the first array (e.g., 220(N-1)) is configured with multiple columns and Q columns, where Q is a positive integer. For each first group (e.g., 244(1)), there are Q first groups (e.g., 244(1)) and Q instances of the first flag signal (e.g., FLG1). For each first group (e.g., 244(1)), a trajectory inference signal generator (e.g., 112) is used to receive the Q instances of the first flag signal (e.g., FLG1). For each first group (e.g., 244(1)), a Q:P encoder (e.g., 262) is used to encode P-bit signals into the Q instances of the first flag signal (e.g., FLG1) to represent an error index (e.g., EPT), where the error index (e.g., EPT) is the first of one or more error-bit trajectory inference signals, and P is a positive integer. For each first group (e.g. 244(1)), a second error bit detector (e.g. 213) generates a second flag signal (FLG2) based on Q instances of a first flag signal (e.g., FLG1), the second flag signal (FLG2) being able to declare an error index (e.g., EPT) pointing to an error bit, and the second flag signal (FLG2) being the second of one or more error bit trajectory inference signals.

[0214] In some embodiments, flowchart 543 further includes a first block in which a fourth structure including a fourth component is formed in a first region (e.g., 103) of a first semiconductor die (e.g., 102A and 102C(1)). The fourth component includes an OR logic gate (e.g., 263), or the logic gate (e.g., 263) is correspondingly included in a second error bit detector (e.g., 213).

[0215] In such an embodiment, flowchart 543 further includes a second block in which at least a fourth component, a third component, a second component, or a first component are mutually coupled and generate at least Q instances of a first flag signal (e.g., FLG1) for each first group (e.g., 244(1)) or logic gate (e.g., 263) to receive the first flag signal (e.g., FLG1) and thereby generate an output signal representing the second flag signal (FLG2).

[0216] Figure 5D is a flowchart 550 illustrating a method for manufacturing a CIM system according to some embodiments.

[0217] Examples of digital CIM systems that can be manufactured according to flowchart 550 include CIM system 100D of diagram 1D, CIM system 100E of diagram 1E, or other similar systems. Flowchart 550 includes blocks 552 to 564.

[0218] In block 552, for each input character of the input array, the corresponding checksum is generated by a column checksum generator, which can be appended to the input array to form an input and checksum (I&C) array. An example of a column checksum generator is column checksum generator 379 in Figure 3C, or other similar generators. An example of an input array is represented by input XIN2 array 305B in Figure 3B, or other similar arrays. An example of an I&C array is input XIN3 array 305D in Figure 3D, or other similar arrays. The process proceeds from block 552 to block 554.

[0219] In block 554, for the weight characters of the weight array, the corresponding checksum is generated by a column checksum generator, which can be appended to the input array to form an input and checksum (W&C) array. An example of a column checksum generator is shown as column checksum generator 380 in Figure 3F, or other similar generators. An example of a weight array is shown as weight arrays 320(0)~320(Q-1) in Figure 3A, or other similar arrays. An example of a checksum is shown as checksum arrays 323(0)~323(Q-1) in Figure 3A, the checksum... This includes the column checksum C_ChkSum_1 387G of graph 3G, or other similar arrays. An example of a W&C array is the weight W2 array 304G of graph 3G, or other similar arrays. The process proceeds from block 554 to block 556.

[0220] In block 556, the column segments of the weight array and their associated checksums (as multiplicands) are received from the W&C array by the multiplier, and one column of the input and its associated checksum (as multiplicands) is received from the I&C array by the multiplier. An example of a multiplier is represented by multiplier combinations 364(0) to 364(Q-1) in Figure 3A, or other similar multipliers. The process proceeds from block 556 to block 558.

[0221] In block 558, iteratively, the multiplicands are multiplied by a multiplier to form an inner product column of the inner product array. Some examples of inner product arrays and inner product columns are, for instance, the weight W1 array 304E and an i-th column in graph 3E, or other similar arrays. The process proceeds from block 558 to block 560.

[0222] In block 560, the trajectory inference signal is generated by the trajectory inference signal generator.

[0223] Examples of trajectory inference signals include the column sum vector Row_Sum 304J of Figure 3J, the column sum vector Col_Sum 304L of Figure 3L, or other similar vectors. Examples of trajectory inference signal generators include the column sum generator 326I of Figure 3I, the column sum generator 328K of Figure 3K, or other similar generators. Block 560 includes blocks 562 and 564.

[0224] In block 562, one of the multiple trajectory inference signal generators performs column-by-column summation to form a column sum. An example of one of the trajectory inference signal generators is column sum generator 326I in Figure 3I, or other similar generators. The process proceeds from block 562 to block 564.

[0225] In block 564, one of the multiple trajectory inference signal generators performs column-by-column summation to generate the corresponding character for the column sum. An example of a column sum is the column sum vector Col_Sum 304L of the 3L graph, or other similar vectors. An example of one of the trajectory inference signal generators is the column sum generator 328K of the 3K graph, or other similar generators. The process leaves block 560 and proceeds to block 384(1) of the 3M graph, as indicated in the aforementioned page break reference 392.

[0226] In some embodiments, block 564 is executed before block 562. In some embodiments, block 564 and the block before block 562 are executed substantially simultaneously.

[0227] Figure 5E is a flowchart 573 illustrating a method for manufacturing a CIM system according to some embodiments.

[0228] Flowchart 573 is an example of block 504 in Figure 5A. The method in flowchart 573 can be implemented, for example, using IC manufacturing system 700 (refer to Figure 7 discussed below) according to some embodiments. Examples of digital CIM systems that can be produced according to the method in flowchart 573 include 100D in Figure 1D, 100E in Figure 1E, or other similar systems.

[0229] Flowchart 573 contains blocks 575 to 577.

[0230] In block 575, in a first region of the first semiconductor die, a first structure including a first component is formed. The first component includes a plurality of memory cells correspondingly used to store units, a multiplier, and a trajectory inference data (LID) generator. Furthermore, some of the plurality of memory cells are configured as corresponding first arrays and used to store first characters. The first component is organized into multiple first groups, each including one of the corresponding first arrays, a multiplier, and an LID generator. Regarding block 575, some examples of the first structure include the example of the first structure discussed in block 545 of Figure 5C of the specification, or other similar structures discussed.

[0231] Examples of the first of a plurality of memory cells include memory cell 349 of Figure 3A, or other similar memory cells. An example of a first array is inner product array 306A of Figure 3A, or other similar arrays. An example of a multiplier is multiplier 251(Q-1) of Figure 2A, or other similar multipliers. An example of an LID generator is LID generator 326A, or other similar generators. An example of a first group is a group including slice 318(Q-1) of Figure 3A, multiplier 364, and LID generator 326A, or other similar groups. The flow proceeds from block 575 to block 577.

[0232] In block 577, interconnected components are formed and generate at least: for each first group, a multiplier performs a multiplication of one or more (i) input characters and associated first checksum characters and (ii) corresponding weight characters and associated second checksum characters; and for each first group, an LID generator generates one or more LID signals based on selected first characters. Regarding block 577, some examples of interconnectedness are included in the examples of interconnectedness discussed in block 547 of Figure 5C of the specification, or other similar discussions.

[0233] In some embodiments, flowchart 573 further includes a first block in which a second structure comprising a second component is formed in a first region (e.g., 124) or a second region (e.g., 155(3)) of a first semiconductor die (e.g., 102E(1)) or a first region (e.g., 155(4)) of a second semiconductor die (e.g., 102E(2)). The second component includes a column checksum generator (e.g., 379). In such an embodiment, flowchart 573 further includes a second block in which at least the second component or one or more first components are formed mutually coupled and generate a column checksum generator (e.g., 379) based at least on an input character (e.g., input XIN2 array 305B) to generate a first checksum character (e.g., a character in R_ChkSum_1).

[0234] In some embodiments, flowchart 573 further includes a first block in which a third structure is formed within the same region as the column checksum generator (e.g., 379), comprising a third component including a recursive adder (e.g., 379) contained within the column checksum generator (e.g., 379). In such an embodiment, a second portion (e.g., 347) of a plurality of memory cells is configured as a corresponding second array (e.g., input XIN2 array 305B) for storing input characters (e.g., input XIN2 array 305B). In such an embodiment, flowchart 573 further includes a second block in which at least the first or second component is formed to be mutually coupled, and further generates at least a second array (e.g., input XIN2 array 305B) configured with a first column and one or more first columns, with the input character (e.g., input XIN2 array 305B) correspondingly positioned at the intersection of the first column and one or more first columns (e.g., A[x][y]). In such an embodiment: the second array (e.g., 305B) is a first portion of the third array (e.g., input XIN3 array 305D), a third portion (e.g., 348) of a plurality of memory cells is configured in the second portion (e.g., 386D) of the third array (e.g., input XIN3 array 305D) and is accordingly used to store a second character representing a first checksum (first checksum character) (e.g., a character in R_ChkSum_1), and the second portion (e.g., 386D) of the third array (e.g., input XIN3 array 305D) is configured in a second column and a second column, with the first checksum character (e.g., R_ChkSum_1) correspondingly positioned at the intersection of the second column and the second column. In such an embodiment, flowchart 573 further includes a second block in which at least a third component, a first component, or a second component are mutually coupled and produce at least one of the first checksum characters (e.g., characters in R_ChkSum_1) by recursively adding the corresponding input characters (e.g., A[x][i]) in the first column column.

[0235] In such an embodiment, flowchart 573 further includes a first block in which a second structure comprising second components is formed in a first region (e.g., 124) or a second region (e.g., 155(3)) of a first semiconductor die (e.g., 102E(1)) or a first region (e.g., 155(4)) of a second semiconductor die (e.g., 102E(2)). The second structure includes column checksum generators (e.g., 180 and 380). In such an embodiment, each first group further includes one corresponding column checksum generator (e.g., 380). In such an embodiment, flowchart 573 further includes a second block in which at least the first components or the second components are formed to be mutually coupled, and at least one or each first group (e.g., circuit 344 and generator 326A (e.g., Q-1)) is generated by a column checksum generator (e.g., 380) based on a weighted character (e.g., weighted W1 character 304E) to generate a second checksum character (e.g., C_ChkSum_1).

[0236] In some embodiments, flowchart 573 further includes a first block, wherein located in the same area as the column checksum generators (e.g., 180 and 380), a third structure is formed including a third component, the third component including an addition tree (e.g., 308F), the addition tree (e.g., 308F) being included as part of the corresponding column checksum generators (e.g., 180 and 380). In such an embodiment: a second portion (e.g., 349) of a plurality of memory cells is configured as a corresponding second array (e.g., weighted W1 array 304E) and used to store weight characters (e.g., weighted W1 array 304E); the second array (e.g., weighted W1 array 304E) is configured with a first column and a first column, and the weight characters (e.g., weighted W1 array 304E) are respectively positioned at the intersection of the first column and the first column, respectively representing a first character (e.g., B[x][y]); the second array (e.g., weighted W1 array 304E) is the first part of a third array (e.g., 304F); a third portion (e.g., 350) of a plurality of memory cells is configured as the second part (e.g., 387) of the third array (e.g., 304F) and is respectively used to store a second checksum character (e.g., C_ChkSum_1); and the second part (e.g., 387) of the third array (e.g., 304F) is configured with a second column and a second column, and the second checksum character (e.g., C_ChkSum_1) is respectively... The corresponding position is at the intersection of the second column and the second column.

[0237] In such an embodiment, flowchart 573 further includes a second block in which at least a third component, a first component, or a second component are mutually coupled and generate at least for each first group (e.g., circuit 344 and generator 326A (e.g., Q-1)), the addition tree (e.g., 308F) generates a second checksum character by adding the weight characters (e.g., B[i][y]) in the corresponding first column of the second array (e.g., weight W1 array 304E).

[0238] In some embodiments, flowchart 573 further includes a first block, located in the same region as the trajectory inference data generator (e.g., 326I and 328K), forming a second structure including a second component, the second component including a column sum generator (e.g., 326I), which is included as a corresponding part of the trajectory inference data generator (e.g., 326I and 328K). In such an embodiment: a first array (e.g., the 3H diagram C1 array 306A) is configured with a first column and a first bar, a first character is correspondingly located at the intersection of the first column and the first bar, and the first character (e.g., the 3H diagram C1 array) represents an inner product character. In such an embodiment, flowchart 573 further includes a second block in which at least a second component or a first component is formed to be mutually coupled and generates, for at least each first group (e.g., circuit 344 and generator 326A (e.g., Q-1)), a column sum generator (e.g., 326I) based on inner product characters to generate a column sum (e.g., Row_Sum in Figure 3J).

[0239] In some embodiments, the column sum (e.g., Row_Sum) is a column vector including the second character. In such an embodiment, flowchart 573 further includes a first block in which a third structure is formed including a third component, the third component including a recursive adder (e.g., 379), the recursive adder (e.g., 379) being included as a corresponding part of a column sum generator (e.g., 326I). In such an embodiment, flowchart 573 further includes a second block in which at least the third component, the second component, or the first component are mutually coupled and generate at least for each first group (e.g., circuit 344 and generator 326A (e.g., Q-1)), the recursive adder (e.g., 379) generates a corresponding second character in the column sum (e.g., Row_Sum) by recursively adding the second characters in the corresponding first column (e.g., the C1 array of Figure 3H) column by column.

[0240] In some embodiments, a first character in a selected first column (e.g., 388) represents a third checksum (e.g., Figure 3J R_ChkSum_2). In such an embodiment, flowchart 573 further includes a first block in which a third structure comprising a third component is formed in a second region (e.g., 114D) of a first semiconductor die (e.g., 102D) or a first region (e.g., 114E) of a second semiconductor die (e.g., 102E(2)), the third component comprising a processor (e.g., 114D, 114E, and 316 of Figure 3M). In such an embodiment, flowchart 573 further includes a second block in which at least the third component, the second component, or the first component are mutually coupled and generate at least a processor (e.g., 316) to compare the column sum (e.g., Figure 3J Row_Sum) with one of the corresponding third checksums (e.g., Figure 3J R_ChkSum_2) (e.g., 384(1)) to indicate that the first column has an error bit.

[0241] In some embodiments, flowchart 573 further includes a first block, located in the same region as the trajectory inference data generator (e.g., 326I and 328K), forming a second structure including a second component, the second component including a column sum generator (e.g., 328K), which is included as a corresponding part of the trajectory inference data generator. In such an embodiment, a first array (e.g., C1 array 306A) is configured with a first column and a first column, a first character correspondingly located at the intersection of the first column and the first column, and the first character representing an inner product character. In such an embodiment, flowchart 573 further includes a second block, wherein at least the second component or the first component is formed to be mutually coupled and generates a column sum (e.g., Col_Sum in the 3L figure) based on the inner product character for at least each first group (e.g., circuit 344 and generator 326A (e.g., Q-1)).

[0242] In some embodiments, the column sum (e.g., Col_Sum) is a column vector containing the second character. In such an embodiment, flowchart 573 further includes a first block, wherein a third structure is formed within the same region as the column sum generator (e.g., 328K(Q-1)), comprising a third component including an addition tree (e.g., 308K), which is included as a corresponding part of the column sum generator (e.g., 328K). In such an embodiment, flowchart 573 further includes a second block, wherein at least the third component, the second component, or the first component are mutually coupled, and generate, for at least each first group (e.g., circuit 344 and generator 326A(e.g., Q-1)), the addition tree (e.g., 308K) generates the second character in the column sum (e.g., Col_Sum) by adding the first characters in the corresponding column of the first column column column by column.

[0243] In some embodiments, the first character in a selected first column (e.g., 389 in Figure 3L) represents the fourth checksum (e.g., C_ChkSum_2 in Figure 3L).

[0244] In such an embodiment, process 573 further includes a first block in which a third structure comprising a third component is formed in a second region (e.g., 114D) of a first semiconductor die (e.g., 102D) or a first region (e.g., 114E) of a second semiconductor die (e.g., 102E(2)). The third component includes a processor (e.g., 114D, 114E and 316 of the 3M figure).

[0245] In such an embodiment, flowchart 573 further includes a second block in which at least a third component, a second component, or a first component are mutually coupled and generate at least a processor (e.g., 316) to compare the column sum (e.g., the third L graph Col_Sum) with one of the corresponding fourth checksums (e.g., the third L graph C_ChkSum_2) to indicate that there is an error bit in the first column.

[0246] In some embodiments, the mutual coupling formed in the first component further generates at least for each first group (e.g., circuit 344 and generator 326A (e.g., Q-1)), the trajectory inference data generator (e.g., generator 326A) is further used to generate one or more error bit trajectory inference signals (e.g., Row_Sum for Figures 3I to 3J and Col_Sum for Figures 3L to 3M) following one or more multiplications of the multiplier (e.g., 251(N-1)).

[0247] Figure 6 is a functional block diagram illustrating an electronic design automation (EDA) system 600 according to some embodiments.

[0248] In some embodiments, EDA system 600 includes an Automatic Placement and Rewind (APR) system. In some embodiments, the EDA system is a general-purpose computing device including a hardware processor 602 and a non-transitory computer-readable storage medium 604. The storage medium 604 is encoded (i.e., stored) in particular by, for example, computer program code 600 (i.e., a set of executable instructions). The instructions (computer program code 600) executed by the hardware processor 602 represent EDA tools that implement some or all of the methods according to one or more embodiments (hereinafter referred to as processes and / or methods), such as the method of generating the illustrated layout in this disclosure, such as the layout diagram in this disclosure or the layout diagram corresponding to the apparatus in this disclosure, or other similar methods.

[0249] Storage medium 604 stores, in particular, layout diagram 611, such as the layout diagram in this disclosure, or other similar layout diagrams.

[0250] Processor 602 is electrically coupled to computer-readable storage medium 604 via bus 608. Processor 602 is further electrically coupled to input / output interface 610 via bus 608. Network interface 612 connects to network 614, enabling processor 602 and computer-readable storage medium 604 to be connected to external components via network 614. Processor 602 is used to execute computer program code 606 encoded in computer-readable storage medium 604, so that EDA system 600 can be used to perform some or all of the processes and / or methods. In one or more embodiments, processor 602 is a central processing unit (CPU), a multi-core processor, a distributed computing system, an application-specific integrated circuit (ASIC), and / or a suitable processor.

[0251] In one or more embodiments, the computer-readable storage medium 604 is an electronic, magnetic field, photoelectric, electromagnetic, infrared, and / or semiconductor system (or instrument or device). For example, the computer-readable storage medium 604 includes semiconductor or solid-state memory, magnetic tape, removable floppy disk, random access memory (RAM), read-only memory (ROM), hard disk, and / or optical disc. In one or more embodiments using optical discs, the computer-readable storage medium 604 includes a read-only optical disc (CD-ROM), a rewritable optical disc (CD-R / W), and / or a digital versatile disc (DVD).

[0252] In one or more embodiments, computer-readable storage medium 604 stores computer program code 606 for enabling an EDA system (execution of which represents at least a portion of the EDA tools) to perform some or all of the processes and / or methods. In one or more embodiments, computer-readable storage medium 604 further stores information that facilitates the performance of some or all of the processes and / or methods. In one or more embodiments, computer-readable storage medium 604 stores a library 607 of standard units, which are included in the standard units disclosed herein. In some embodiments, computer-readable storage medium 604 stores one or more layout diagrams 611.

[0253] EDA system 600 includes an input / output interface 610. The input / output interface 610 is coupled to external circuitry. In one or more embodiments, the input / output interface 610 includes a keyboard, numeric keypad, mouse, trackball, trackpad, touchscreen, and / or cursor arrow keys to communicate information and instructions to processor 602.

[0254] EDA system 600 further includes a network interface 612 coupled to processor 602. Network interface 612 allows EDA system 600 to communicate with a network 614 connected to one or more computer systems. Network interface 612 includes a wireless network interface, such as Bluetooth, Wi-Fi, WiMAX, GPRS, or WCDMA, or a wired network interface, such as Ethernet, USB, or IEEE-1364. In one or more embodiments, some or all of the processes and / or methods are implemented in two or more EDA systems 600.

[0255] EDA system 600 receives information through input / output interface 610. The information received through input / output interface 610 includes one or more instructions, data, design rules, standard cell libraries, and / or other parameters processed by processor 602. This information is transmitted to processor 602 via bus 608. EDA system 600 also receives user interface (UI) related information through input / output interface 610. This information is stored as UI 642 on computer-readable storage media 604.

[0256] In some embodiments, some or all processes and / or methods are implemented as separate application software and executed by a processor. In some embodiments, some or all processes and / or methods are implemented as application software as part of additional application software. In some embodiments, some or all processes and / or methods are implemented plugged into application software. In some embodiments, at least one process and / or method is implemented as application software as part of an EDA tool. In some embodiments, some or all processes and / or methods are implemented as application software and used by an EDA system 600. In some embodiments, a layout comprising standard cells is generated using a tool, such as VIRTUOSO® from CADENVE DESIGN SYSTEMS Inc., or other suitable layout generation tools.

[0257] In some embodiments, a processor may be understood as a function of a program stored on a nontransitory computer-readable storage medium. Examples of nontransitory computer-readable storage media include, but are not limited to, external / removable and / or internal / embedded storage or memory units, such as one or more optical discs (e.g., DVDs), magnetic disks (e.g., hard disks), semiconductor memory (e.g., ROM, RAM, memory cards, or similar components).

[0258] Figure 7 is a functional block diagram illustrating an integrated circuit (IC) manufacturing system 700 and the associated IC manufacturing process according to some embodiments.

[0259] In some embodiments, based on the layout diagram generated from block 502 of Figure 5A, the IC manufacturing system 700 implements block 504 of Figure 5A, wherein at least (A) one or more semiconductor photomasks, or (B) at least one component in one layer of an unfinished semiconductor integrated circuit, is manufactured using the IC manufacturing system 700. In some embodiments, the IC manufacturing system 700 implements one or more flowcharts of this disclosure.

[0260] In Figure 7, the IC manufacturing system 700 includes departments such as a design plant 720, a photomask plant 730, and an IC fabrication / fab plant 750. These departments exchange services related to the design, development, and production cycle and / or the production of IC devices 760. The departments in the IC manufacturing system 700 are connected via a communication network. In some embodiments, the communication network is a single network. In some embodiments, the communication network is multiple different networks, such as an intranet and an extranet. The communication network includes wired and / or wireless communication channels. Each department communicates with one or more other departments and provides and / or receives services from one or more other departments. In some embodiments, two or more design plants 720, photomask plants 730, and IC fabrication / fab plants 750 are owned by a single large company. In some embodiments, two or more design plants 720, photomask plants 730, and IC fabrication / fab plants 750 coexist in shared facilities and use shared resources.

[0261] Design plant (or design team) 720 generates IC design layout 722. IC design layout 722 contains various geometric patterns designed for IC device 760. Geometric patterns corresponding to the metal, oxide, or semiconductor layers constituting various components of IC device 760 are manufactured. Multiple layers are combined to form various IC features. For example, a portion of IC design layout 722 includes different IC features, such as active regions, gate terminals, source and drain terminals, vias for metal lines or inner layer interconnects, and openings for adapters, formed on a semiconductor substrate (e.g., a silicon wafer) and in various metal layers deployed on the semiconductor substrate. Source and drain regions may be individual sources or drains, as discussed in the specification. Design plant 720 implements an appropriate design flow to form IC design layout 720. The design process includes one or more logic designs, physical designs, or placement and routing. IC design layout 722 is presented in one or more data files containing geometric pattern information. For example, IC design layout 722 is represented in GDSII or DFII file format.

[0262] The photomask department 730 includes photomask data preparation 732 and photomask fabrication 734. The photomask department 730 uses an IC design layout 722 and produces one or more photomasks 735 for various multilayers within an IC device 760, based on the IC design layout 722. The photomask department 730 performs photomask data preparation 732, in which the IC design layout 722 is translated into a Representative Data File (RDF). The photomask data preparation 732 provides the RDF to the photomask fabrication 734. The photomask fabrication 734 includes a photomask writer. The photomask writer converts the RDF into a pattern on a substrate, such as a photomask (crosshairs) or a semiconductor wafer. The IC design layout is manipulated by the photomask data preparation 732 to conform to the specific characteristics required by the photomask writer and / or the IC manufacturing plant 750. In Figure 7, the photomask data preparation 732, photomask fabrication 734, and photomask 735 are shown as separate components. In some embodiments, photomask data preparation 732 and photomask manufacturing 734 collectively refer to photomask data preparation.

[0263] In some embodiments, mask preparation 732 includes optical proximity correction (OPC), which utilizes lithographic enhancement techniques to compensate for pattern errors, such as those caused by diffraction, interference, other process effects, and similar factors. OPC adjusts the IC design layout 722. In some embodiments, mask preparation 732 further includes resolution enhancement techniques (RET), such as morphing illumination techniques, sub-resolution adjustment features, phase-transfer masks, other suitable techniques, and combinations thereof. In some embodiments, reverse lithography (ILT) is further used, treating OPC as a reverse imaging problem.

[0264] In some embodiments, mask preparation 732 includes mask criterion checking (MRC), which uses a set of mask creation criteria to check the IC design layout that has entered the OPC process. The mask creation criteria include specific geometric and / or connectivity constraints to ensure sufficient margins to handle the variability of semiconductor manufacturing processes, or for other similar reasons. In some embodiments, MRC adjusts the IC design layout to compensate for constraints during mask fabrication 734. The MRC adjustment may revert some modifications made during OPC to comply with the mask creation criteria.

[0265] In some embodiments, mask preparation 732 includes lithography process inspection (LPC), which simulates the process of manufacturing IC device 760 by IC fabrication / manufacturing plant 750. LPC simulates this process based on IC design layout 722, creating a simulated production device, such as IC device 760. The parameters of the LPC simulation process may include parameters related to various processes in the IC manufacturing cycle, parameters related to the tools used in the manufacturing process of ICs and / or other aspects. LPC takes into account various factors, such as spatial image contrast, depth of focus (DOF), mask error enhancement factor (MEEF), other suitable factors, and combinations thereof. In some embodiments, after the device simulated by LPC is manufactured, if the simulated device's shape does not meet design criteria, OPC and / or MRC will be repeated to further refine the IC design layout 722.

[0266] The above description of photomask preparation 732 is for clarity and simplification. In some embodiments, photomask preparation 732 includes additional features, such as logic operation procedures (LOPs), to adjust the IC design layout according to manufacturing guidelines. Furthermore, the above-described processes applied to IC design layout 722 during photomask preparation 732 can be performed in various different sequences.

[0267] After photomask data preparation 732 and during photomask fabrication 734, one or more photomasks 735 are fabricated based on a modified IC design layout. In some embodiments, an electron beam or electron beam mechanism based on the modified IC design layout is used to form the pattern on the photomask (crosshairs). In some embodiments, the photomask is formed using a variety of different technologies. In some embodiments, the photomask is formed using a binary technology. In some embodiments, the photomask pattern includes opaque areas and transparent areas. A radiation beam, such as an ultraviolet (UV) beam, used to expose a patterned photosensitive material (e.g., photoresist) coated on the wafer is shielded by the opaque areas and transmitted to the transparent areas. In one example, a binary photomask includes a transparent substrate (e.g., mixed quartz) and an opaque material (e.g., chromium) plated on the opaque areas of the photomask. In another example, the photomask is formed using a phase transfer technique. In a phase transfer photomask (PSM), various features in the pattern formed on the photomask are used to have appropriate phase differences to improve resolution and imaging quality. In various examples, the phase-transfer mask is an attenuation type PSM or an alternating type PSM. The mask produced by mask fabrication 734 is used in a variety of processes. For example, the above-mentioned mask is used in ion implantation processes to form various doped regions in semiconductor wafers, etching processes to form various etched regions in semiconductor wafers, and / or other suitable processes.

[0268] IC manufacturing plant 750 is an IC manufacturer that includes one or more production facilities to manufacture a variety of different IC products. In some embodiments, IC manufacturing plant 750 is a semiconductor foundry. For example, there may be one production facility that performs front-end manufacturing (FEOL) of multiple IC products, another production facility that supplies back-end manufacturing (BEOL) of IC product interconnects and packaging, and a third production facility that supplies other services to the foundry.

[0269] IC manufacturing plant 750 uses photomasks (multiple photomasks) 735 manufactured by photomask factory 730 to manufacture IC device 760 using manufacturing tool 752. In some embodiments, semiconductor wafer 753 manufactured by IC manufacturing plant 750 is used to form IC device 760 using photomasks (multiple photomasks) 735. Semiconductor wafer 753 includes a silicon substrate or other suitable substrate having material layers formed thereon. Semiconductor wafer further includes one or more multiple doped regions, dielectric features, multilayer interconnects, and other similar features (formed in subsequent manufacturing steps).

[0270] In some embodiments, the in-memory computing (CIM) system includes: a plurality of first components comprising a plurality of memory cells correspondingly used to store a plurality of units in a first region of a semiconductor die, and a plurality of arrays comprising a plurality of multipliers and a first error bit detector. The plurality of first components of the plurality of memory cells are configured in a plurality of corresponding first arrays and are used to store a plurality of first bits. The plurality of second components of the plurality of memory cells are configured in a plurality of corresponding second arrays and are used to store a plurality of co-occurrence bits corresponding to one of the plurality of first arrays. Each of the first group includes a first array of a plurality of first arrays, a second array of a plurality of second arrays, a multiplier of a plurality of multipliers, and a first error bit detector of a plurality of first error bit detectors. The multiplier is used to multiply a plurality of input bits and a plurality of corresponding first bits, and the first error bit detector is used to detect error bits in the respective plurality of first bits based on the respective co-occurrence bits.

[0271] In some embodiments, for each first group, the first error bit detector is further used to detect multiplication synchronously with the multiplier.

[0272] In some embodiments, the arrays of the plurality of multipliers and the plurality of first error bit detectors further include: a plurality of data multiplexers, and each first group further includes one of the plurality of data multiplexers. For each first group, the data multiplexer is used to select (i) the inner product generated by the multipliers or (ii) a preset value generated by the first error bit detectors based on the output signal.

[0273] In some embodiments, the CIM system further includes a plurality of co-encoders, wherein: each first group further includes one of the plurality of co-encoders corresponding to the co-encoders, and the co-encoders of each first group are used to encode to a plurality of co-encoders corresponding to the plurality of co-encoders based on the respective plurality of first bits.

[0274] In some embodiments, for each first group, the first error bit detector includes a mutex or logic gate for receiving a plurality of first bits and a plurality of corresponding bits as inputs and thereby generating an output signal representing a first flag signal that can be declared to indicate the presence of an error bit.

[0275] In some embodiments, the CIM system further includes: a trajectory inference data generator for generating one or more erroneous bit trajectory inference signals based on multiple co-occurrence bits.

[0276] In some embodiments, for each first group: a first error bit detector is further configured to generate an output signal representing a first flag signal that definitively indicates the presence of an error bit. The first array is configured with multiple columns and Q columns, where Q is a positive integer, and has Q groups and Q instances corresponding to the first flag signal. A trajectory inference data generator is configured to receive the Q instances of the first flag signal and includes a Q:P encoder for encoding the Q instances of the first flag signal into a P-bit signal representing an error index, the error index being the first of one or more error bit trajectory inference signals and P being a positive integer. A second error bit detector generates a second flag signal based on the Q instances of the first flag signal, the second flag signal being a definitively indicated error index pointing to an error bit, and the second flag signal being the second of one or more error bit trajectory inference signals.

[0277] In some embodiments, the second error bit detector includes an OR logic gate for receiving Q instances of the first flag signal and generating an output signal representing the second flag signal thereon.

[0278] In some embodiments, the CIM system further includes an error bit corrector for determining that one of the corresponding memory cells in one of the plurality of first arrays is a data-corrupted cell to represent the location of an error bit based on an error index and a second flag signal.

[0279] In some embodiments, a method (a method of operating a CIM system) includes: a CIM system comprising a plurality of first components in a first region comprising a first semiconductor die, a plurality of first components comprising a plurality of memory cells correspondingly configured to store a plurality of units and arranged in an array, an array of a plurality of multiplier arrays and an array of a plurality of first error bit detectors, a plurality of first units of the plurality of memory cells being arranged in a first array and used to store a plurality of first bits, a plurality of second units of the plurality of memory cells being arranged in a second array and used to store a plurality of co-occurrence bits corresponding to the plurality of first bits; performing a multiplication of a plurality of input bits and a plurality of first bits respectively, and performing a first error detection based on error bits in the respective plurality of co-occurrence bits in the respective plurality of first bits, the first error detection being performed simultaneously with the multiplication.

[0280] In some embodiments, the above method further includes, for each first group, performing a first error detection comprising: performing a logical mutually exclusive OR operation on a plurality of first bits and a plurality of co-occurring bits as input to generate a first flag signal that can declare the presence of an error bit.

[0281] In some embodiments, the above method further includes, for each first group, generating one or more error bit trajectory inference signals based on multiple co-occurring bits.

[0282] In some embodiments, for each first group, performing a first error detection includes: generating a first flag signal that definitively indicates the presence of an error bit; a first array configured with multiple columns and Q columns, where Q is a positive integer; and having Q first groups and Q instances corresponding to the first flag signal. For each first group, generating one or more error bit trajectory inference signals includes: representing an error index by encoding the Q instances of the first flag signal into a P-bit signal using Q:P encoding, where the error index is the first of the one or more error bit trajectory inference signals, where P is a positive integer. For each first group, performing a second error detection involves generating a second flag signal based on the Q instances of the first flag signal, where the second flag signal definitively indicates that the error index points to an error bit, and the second flag signal is the second of the one or more error bit trajectory inference signals.

[0283] In some embodiments, for each first group, performing a second error detection includes: performing a logical OR operation on Q instances of the first flag signal to generate an output signal representing the second flag signal.

[0284] In some embodiments, the CIM system includes: in a first region of a semiconductor die, a plurality of first components including a plurality of memory cells correspondingly used to store a plurality of characters, a plurality of multipliers, and a trajectory inference data generator. The plurality of first components of the plurality of memory cells are configured in a plurality of first arrays and used to store the plurality of first characters. Each of the plurality of first groups includes one of the corresponding plurality of first arrays, one of the plurality of multipliers, and a trajectory inference data generator, and the operation of each first group relates to a corresponding plurality of the plurality of first characters. The multipliers generate the plurality of first characters of the first array by performing one or more (A) multiple input characters and associated multiple first checksum characters and (B) corresponding multiple weight characters and associated multiple second checksum characters. The trajectory inference data generator generates one or more error bit trajectory inference signals based on a selected plurality of the plurality of first characters.

[0285] In some embodiments, for each first group: the first array is configured with a plurality of first columns and a plurality of first bars, the plurality of first characters are respectively positioned at the intersection of the plurality of first columns and the plurality of first bars, the plurality of first characters represent a plurality of inner product characters, and for each first group, the trajectory inference data generator includes: a column sum generator for generating a column sum based on the plurality of inner product characters.

[0286] In some embodiments, for each first group, the column sum is a column vector comprising a plurality of second characters, and the column sum generator comprises: a plurality of recursive adders corresponding to a plurality of first columns in the first array, each recursive adder generating a plurality of second character corresponding to a plurality of first columns by recursively adding a plurality of first characters in one of the plurality of first columns column by column to generate a plurality of second character corresponding to a column sum.

[0287] In some embodiments, for each first group: the first array is configured with a plurality of first columns and a plurality of first bars, the plurality of first characters are respectively positioned at the intersection of the plurality of first columns and the plurality of first bars, the plurality of first characters represent a plurality of inner product characters, and for each of the plurality of first groups, the trajectory inference data generator includes: a bar sum generator for generating a bar sum based on the plurality of inner product characters.

[0288] In some embodiments, for each first group: the column sum is a column vector comprising a plurality of second characters, and the column sum generator comprises: an addition tree that generates a plurality of second characters for the column sum on a column-by-column basis by adding a plurality of first characters in one of the plurality of first columns.

[0289] In some embodiments, for each first group: the trajectory inference data generator is then multiplied by one or more multipliers to further generate one or more erroneous bit trajectory inference signals.

[0290] The foregoing outlines the features of several embodiments, enabling those skilled in the art to better understand the nature of this disclosure. Those skilled in the art will understand that they can readily use this disclosure as a basis for designing or modifying other processes and structures for implementing the same purpose and / or achieving the advantages of the embodiments described herein. Those skilled in the art will also recognize that such equivalent constructions do not depart from the spirit and scope of this disclosure, and that various changes, substitutions, and modifications can be made herein without departing from the spirit and scope of this disclosure.

[0291] 100A, 100C, 100D, 100E: In-Memory Computing (CIM) System 102A, 102C(1), 102C(2), 102D, 102E(1), 102E(2): Semiconductor bare die 103A, 103C: In-Memory Computing System (CIM) Area 104A, 104B: Weighted and Same-Place (W&P) Arrays 105: Checksum Array 106: Inner Product Array 107A, 107B: Multipliers 108A, 108B: Addition Tree 110: Error Bit Detector 112: Trajectory Inference Data Generator 114A, 114C, 114D, 114E: Areas 116: Error Bit Corrector 118(0)~118(Q-1): Cutout 120(0)~120(Q-1): Weight array 122(0)~122(Q-1): Corresponding array 124D, 124E: Computational Information Modeling (CIM) area within memory 126: Trajectory Inference Data Generator 127: Column Sum Generator 128: Column Checksum Generator 130: Error bit detectors, locators, and correctors 155(1), 155(2), 155(3), 155(4): Region 158: Co-position encoder 178: Checksum Generator 179: Column Checksum Generator 180: Column Checksum Generator 200D: In-Memory Computing (CIM) System 203(1), 203(2), 203(3): Computational Information System (CIM) region within memory 204A: Weighted and Peer-to-Peer (W&P) Array 208: Addition Tree 210(Q-1): Error Bit Detector 212: Trajectory Inference Data Generator 213: Screenshot Error Detector 216: Error Bit Corrector 218(0)~218(Q-1): Cutout 220(Q-1): Weighted array 221(Q-1): Corresponding array 236: Enhanced Multiplier Array 238(0)~238(Q-1): Strong multiplier block 240: Adder 242: Composition 244(1), 244(2): Components 245, 246: Memory units 251(Q-1): Multiplier 253(Q-1): XOR gate 254(Q-1): Data Multiplexer (MUX) 259(Q-1): Co-position encoder 260(Q-1): XOR gate 262: Encoder 263: OR logic gate 265(1)~265(5): Blocks 290: Flowchart 300D: In-Memory Computing (CIM) System 304E, 304G: Weighted arrays 304I, 304J: Column summation 304L: Total of columns 305V, 305B, 305D: Input Array 306A, 306H: Inner Product Array 308A, 308F, 308K: Addition Tree 316: Error bit detector, locator and corrector unit 318(0)~318(Q-1): Cutout 320(Q-1): Weighted array 323(Q-1): Checksum Array 324: Computational Information Modeling (CIM) area within memory 326A, 326I: Trajectory Inference Data Generator 328K: Column Sum Generator 337: Multiplier Array 342: Composition 344: Circuit 347, 349, 350: Memory units 352(0)~352(Q-1): Multiplier 356, 357: Memory units 364(Q-1): Multiplier 374: Memory Unit 379: Column Checksum Generator 380: Column Checksum Generator 382: Temporary Register 383: Flowchart 384(1)~384(5): Blocks 386D: Column vector 387G: Column Vector 388: column 389: column 392: Page Break Reference 420: Weighted Array 422: Isotopic array 466: Same-position bit generator 468: Error Detection 470: Input Array 472: Weighted Array 474: Inner Product Array 476: Error Bit Detection Array 484(3)D, 484(3)E, 484(4), 484(5): Blocks 500: Flowchart 502, 504: Blocks 508: Flowchart Blocks 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, and 540: 543: Flowchart Blocks 545 and 547 550: Flowchart Blocks 552, 554, 556, 558, 560, 562, and 564 573: Flowchart Blocks 575 and 577 600: Electronic Design Automation (EDA) System 602: Processor 604: Storage Media 606: Computer code 607: Function Library 608: Busbar 610: Input / Output Interface 611: Layout Diagram 612: Network Interface 614: Internet 642: User Interface (UI) 700: IC Manufacturing System 720: Design Factory 722: IC Design Layout 730: Photomask Factory 732: Preparation of Photomask Data 734: Photomask Manufacturing 735: Photomask 750: IC manufacturing plant 752: Manufacturing Tools 753: Semiconductor wafer 760: IC device

[0292] Domestic storage information (please note in order of storage institution, date, and number) none Overseas storage information (please note in the order of storage country, institution, date, and number) none

Claims

1. A memory-based computing system, comprising: In a first region of a semiconductor die, there are a plurality of first components corresponding to a plurality of memory cells for storing a plurality of units, and a plurality of arrays including a plurality of multipliers and a plurality of first error bit detectors; a plurality of first components in the memory cells are arranged in a plurality of corresponding first arrays and are used to store a plurality of first bits; a plurality of second components in the memory cells are arranged in a plurality of corresponding second arrays and are used to store a plurality of corresponding bits for the first bits; and each of the plurality of first components includes a first array of the first arrays, a second array of the second arrays, a multiplier of the multipliers, and a first error bit detector of the first error bit detectors. The multiplier is used to perform a multiplication of a plurality of input bits and a plurality of corresponding first bits, and the first error bit detector is used to detect an error bit corresponding to one of the first bits based on the corresponding corresponding bits, wherein the detection of the error bit by the first error bit detector is synchronized with the multiplication performed by the multiplier.

2. The memory-based arithmetic system as claimed in claim 1, wherein: the arrays of the multipliers and the first error bit detectors further comprise: a plurality of data multiplexers; and each of the first groups further comprises a corresponding one of the data multiplexers; and for each of the first groups, the data multiplexers are configured to select (i) an inner product generated by the multiplier, or (ii) a preset value generated by the first error bit detector based on an output signal.

3. The in-memory computing system as claimed in claim 1, further comprising: a plurality of co-encoders; wherein: each of the first groups further comprises a corresponding co-encoder among the co-encoders; and for each of the first groups, the co-encoder is used to encode a corresponding plurality of the co-encoders based on the corresponding first bits; and for each of the first groups, the first error bit detector comprises a mutex or logic gate for receiving the first bits and the co-encoders as a plurality of inputs and thereby generating an output signal representing a first flag signal that can be declared to indicate the presence of an error bit.

4. The in-memory computing system as claimed in claim 1, further comprising: a trajectory inference data generator for generating one or more error bit trajectory inference signals based on the co-occurring bits, wherein for each of the first groups: the first error bit detector is further configured to generate an output signal representing a first flag signal that can be declared to indicate the presence of an error bit; the first array is configured with a plurality of columns and Q columns, wherein Q is a positive integer; the first groups have Q groups and the first flag signal has a corresponding Q instances; the trajectory inference data generator is configured to receive the Q instances of the first flag signal and includes: A Q:P encoder is used to encode the Q instances of the first flag signal into a P-bit signal to represent an error indicator, the error indicator being a first of one or more error bit trajectory inference signals, and P being a positive integer; and a second error bit detector is used based on the Q instances of the first flag signal to generate a second flag signal that can be declared to indicate that the error indicator points to one of the error bits, and the second flag signal being a second of one or more error bit trajectory inference signals.

5. The in-memory computing system as claimed in claim 4, further comprising: an error bit corrector for determining that one of the memory cells in one of the first arrays is a data-corrupted cell, the data-corrupted cell representing a position of the error bit based on the error index and the second flag signal.

6. A method for operating a computational system within memory, the method comprising: For each of the plurality of first groups, each comprising one of the plurality of multipliers and one of the plurality of first error bit detectors, the in-memory arithmetic system comprises a plurality of first components in a first region of a semiconductor die, the first components comprising a plurality of memory cells correspondingly used to store a plurality of units and arranged in an array, one of the plurality of multipliers and a plurality of first error bit detectors, a plurality of first components disposed in a first array and used to store a plurality of first bits of the memory cells, and a plurality of second components disposed in a second array and used to store a plurality of co-occurrence bits of the memory cells corresponding to the first bits: performing multiplication of a plurality of input bits and a plurality of first components corresponding to the first bits; and performing a first error detection of an error bit in one of the corresponding first bits based on the corresponding co-occurrence bits, the first error detection being performed synchronously with the multiplication.

7. The method of operation as described in claim 6 further includes: for each of the first groups, generating one or more error bit trajectory inference signals based on the corresponding bits.

8. The method of operation as described in claim 7, wherein: for each of the first groups, performing the first error detection comprises: generating a first flag signal that is denoted as indicating the presence of an error bit; the first array is arranged in a plurality of columns and Q columns, wherein Q is a positive integer; wherein the first groups have Q groups and the first flag signal has a corresponding Q instances; for each of the first groups, generating one or more error bit trajectory inference signals comprises: performing a Q:P encoding by encoding a P-bit signal into the Q instances of the first flag signal to represent an error index, the error index being a first of the one or more error bit trajectory inference signals, wherein P is a positive integer; and performing a second error detection to generate a second flag signal that is denoted as indicating the error index points to an error bit based on the Q instances of the first flag signal, and the second flag signal being a second of the one or more error bit trajectory inference signals.

9. A memory-based computing system, comprising: In a first region of a semiconductor die, a plurality of first components include a plurality of memory cells for storing a plurality of words, a plurality of multipliers, and a trajectory inference data generator; a plurality of the memory cells are arranged in a plurality of first arrays and used to store a plurality of first words; each of the plurality of first groups includes a first array of the first arrays, a multiplier of the multipliers, and a trajectory inference data generator, and the operation of each of the first groups is related to a plurality of the first words; the multipliers generate the first words of the first array by performing one or more (A) a plurality of input words and an associated plurality of first checksum words and (B) a plurality of weight words and an associated plurality of second checksum words; and the trajectory inference data generator generates one or more error bit trajectory inference signals based on a plurality of the selected first words.

10. The memory-based arithmetic system as claimed in claim 9, wherein for each of the first groups: the first array is arranged with a plurality of first columns and a plurality of first bars, the first characters being positioned at a plurality of intersections of the first columns and the first bars; the first characters representing a plurality of inner product characters; and for each of the first groups, the trajectory inference data generator includes: a column sum generator for generating a column sum based on the inner product characters.