In-memory computing system and operating method thereof
By integrating memory cells, multipliers and bit error detectors in the semiconductor die, performing bit error detection and multiplication operations in parallel, and using the trajectory inference data generator to quickly locate bit errors, solving the complex and time-consuming problems of the prior art bit error detection and correction, and achieving efficient bit error processing and computing system performance improvement.
Patent Information
- Application Number
- CN202411906068.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-24
AI Technical Summary
In existing semiconductor integrated circuits, bit error detection and correction become more complex and time-consuming as component size decreases and transistor density increases, affecting the efficiency of the computing system.
In a specific area of the semiconductor die, memory cells, multipliers and bit error detectors are integrated to reduce operation cycles by performing bit error detection and multiplication operations in parallel, and generate bit error trajectory inference signals through the trajectory inference data generator to quickly locate and correct bit errors.
It realizes the rapid completion of bit error detection and correction, improves the operating speed and efficiency of the computing system, reduces the out-of-die transmission volume related to bit error detection, positioning and correction, and improves overall performance.
Smart Images

Figure CN120196306A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to a compute-in-memory system and an operation method thereof in a memory. Background Art
[0002] The semiconductor integrated circuit (IC) industry manufactures various analog and digital devices to solve problems in many different fields. The development of semiconductor process technology nodes has gradually reduced the component size and tightened the pitch, resulting in a gradual increase in transistor density. ICs have become smaller. Summary of the Invention
[0003] According to one aspect of an embodiment of the present application, there is provided a compute-in-memory (CIM) system, including: in a first region of a semiconductor die, a first component includes memory cells respectively configured to store a single bit, and an array including a multiplier and a first bit error detector; a first memory cell among the memory cells is arranged in a corresponding first array and configured to store a first bit; a second memory cell among the memory cells is arranged in a corresponding second array and configured to store a parity bit corresponding to the first bit; and for a first group, each in the first group includes a corresponding one of the first array, the second array, the multiplier, and the first bit error detector, the multiplier is configured to perform a multiplication of an input bit and a corresponding one of the first bits, and the first bit error detector is configured to perform a detection of a bit error in the corresponding first bit based on the corresponding parity bit.
[0004] According to another aspect of an embodiment of the present application, there is provided a method for operating a compute-in-memory (CIM) system, the method including: for a first group, each in the first group includes a multiplier and a corresponding first bit error detector, the CIM system includes its first component in a first region of a semiconductor die, the first component includes memory cells respectively configured to store a single bit and arranged in an array, an array of multipliers, and an array of first bit error detectors, a first memory cell among the memory cells is arranged in a first array and configured to store a first bit, a second memory cell among the memory cells is arranged in a second array and configured to store a parity bit corresponding to the first bit; performing a multiplication of an input bit and a corresponding first bit; and performing a first error detection of a bit error in the corresponding first bit based on the corresponding parity bit, the performing of the first error detection is performed in parallel with the multiplication.
[0005] In yet another aspect according to an embodiment of the present application, a compute-in-memory (CIM) system is provided, including: in a first region of a semiconductor die, a first component includes memory cells configured to store words accordingly, a multiplier, and a trace inference data generator; a first memory cell among the memory cells is arranged in a first array and is configured to store a first word; for a first group, each in the first group includes a corresponding first array, a multiplier, and a trace inference data generator in the first array, and each in the first group operates in association with a corresponding first word in the first word. The multiplier is configured to generate a first word of the first array by performing one or more multiplications of (A) an input word and an associated first checksum word and (B) a corresponding weight word and an associated second checksum word, and the trace inference data generator is configured to perform the generation of one or more bit error trace inference signals based on a selected first word in the first word. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] In the drawings, one or more embodiments are illustrated by way of example and not limitation, where elements with the same reference numerals represent the same elements throughout. Unless otherwise disclosed, the drawings are not drawn to scale.
[0007] Figure 1A and Figures 1C - 1E is a functional block diagram of a system according to some embodiments.
[0008] Figure 1B is a functional block diagram of a memory arrangement according to some embodiments.
[0009] Figures 2A - 2D is a schematic diagram according to some embodiments.
[0010] Figures 3A - 3M is a schematic diagram according to some embodiments.
[0011] Figures 4A - 4G is a block diagram according to some embodiments.
[0012] Figures 5A - 5E is a flowchart of a corresponding method according to some embodiments.
[0013] Figure 6 is a functional block diagram of an electronic design automation (EDA) system according to some embodiments.
[0014] Figure 7 is a functional block diagram of an integrated circuit (IC) manufacturing system and its associated IC manufacturing process according to some embodiments. DETAILED DESCRIPTION
[0015] The following disclosure provides many different embodiments or examples for implementing the present invention. Specific embodiments or examples of components and arrangements are described below to simplify the present invention. Of course, these are merely examples and are not intended to be limiting. For example, in the following description, forming the first component above or on the second component may include embodiments in which the first component and the second component are in direct contact, and may also include embodiments in which additional components may be formed between the first component and the second component such that the first component and the second component may not be in direct contact. Additionally, the present invention may repeat reference numerals and / or letters in the various examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate a relationship between the various embodiments and / or configurations being discussed.
[0016] In addition, for ease of description, spatial relationship terms such as "below", "beneath", "lower", "above", "upper", etc. may be used herein to describe the relationship of one element or component to another element or component as shown in the figures. Except for the orientation shown in the figures, the spatial relationship terms are intended to include different orientations of the device in use or operation. The device may be positioned otherwise (rotated 90 degrees or in other orientations), and the spatial relationship descriptors used herein may be interpreted accordingly. In some embodiments, the term standard cell structure refers to a standardized building block included in various standard cell structure libraries. In some embodiments, various standard cell structures are selected from their libraries and used as components in a layout diagram representing a circuit.
[0017] In some embodiments, a compute-in-memory (CIM) system in a first memory includes a first component and an array in a first region of a first semiconductor die, the first component including memory cells each configured to store a single bit, and the array including a multiplier and a first bit error detector. The first memory cells are arranged in respective first arrays and are configured to store first bits. The second memory cells are arranged in respective second arrays and are configured to store parity bits corresponding to the first bits. For a first set, each set includes a respective one of the first array, the second array, the multiplier, and the first bit error detector: the multiplier is configured to perform a multiplication of an input bit with a corresponding bit in the first bits; and the first bit error detector is configured to perform a detection of a bit error in the corresponding first bits based on the corresponding parity bits. In some embodiments, an example of a bit error detector is configured to detect the presence of a bit error and generate a signal indicating the presence. For example, generating Figure 2AFlags such as FLG1, where the flag FLG1 is assertable to indicate the presence of a bit error. In some embodiments, the first CIM system is included as part of an artificial intelligence (AI) system. According to another method corresponding to the first CIM system, bit error detection (pre-multiplication detection) of the bit errors in the memory corresponding to the first array is performed before performing the multiplication. According to another method, bit error detection is performed using two operation cycles before the multiplication. In contrast, according to at least some embodiments of the first CIM system, a first bit error detector 210(Q-1) is included to perform bit error detection in parallel with the multiplication performed by the multiplier, which uses one operation cycle, which is one operation cycle faster than another method.
[0018] In some embodiments, the second CIM system includes a first component in a second region of a second semiconductor die, the first component including memory cells configured to store words accordingly, a multiplier, and a bit error trajectory inference data generator. The first memory cells are arranged in a first array and are configured to store a first word. For a first group, each first group includes each of the first array and the multiplier and a corresponding one of the bit error trajectory inference data generators, and each first group operates in association with a corresponding word in the first word: the multiplier is configured to generate the first word of the first array by performing one or more multiplications of (A) an input word and an associated first checksum word and (B) a corresponding weight word and an associated second checksum word, and the bit error trajectory inference data generator is configured to generate one or more bit error trajectory inference signals based on the selected words in the first word, where the bit error trajectory inference signals are used to infer the location of the bit error, e.g., a memory cell located at the intersection of an identified row and an identified column in the array. Examples of bit error trajectory inference signals according to the second CIM system include row and vector (e.g., Figure 3J Row_Sum 304J in Figure 3L , column and vector (such as Col_Sum 304L in ), etc. In some embodiments, the second CIM system is included as part of an AI system. According to another method corresponding to the second CIM system, (1) the weight array is stored in a first region of the first die, and (2) bit error detection, localization, and correction (DLC) are performed entirely by a processor and an associated random access memory (RAM) located at least on the second die. According to another method, to perform bit error DLC, a large amount of data is transferred from the weight array on the first die to the processor on the second die (off-die transfer), which results in a large amount of transfer latency, thus greatly reducing the speed of processing bit error DLC according to another method. In contrast, compared with another method, at least some embodiments of the second CIM system greatly reduce the off-die transfer volume associated with bit error detection, localization, and correction (DLC), thus achieving much faster bit error DLC than another method.
[0019] Figure 1A is a functional block diagram of a Compute-in-Memory (CIM) system 100A in a digital memory according to some embodiments.
[0020] In Figure 1A CIM system 100A includes a semiconductor die 102A that includes at least a CIM region 103A and a second region 114A. Components of the CIM region 103A (see Figures 2A - 2C , Figures 4A - 4B etc.) include: a Weight and Parity (W&P) array 104A; a multiplier 107A; an adder tree 108A; a bit error detector 110; a trace inference data generator 112; and a parity encoder 158. Region 114A includes a bit corrector 116 (see Figure 2D ). In some embodiments, region 114A is a processor and the bit error corrector 116 represents a function performed by the processor. In some embodiments, an instance of the bit error detector 110 is configured to detect the presence of a bit error and generate a signal indicating the presence. For example, generating Figure 2A flag FLG1 etc., where flag FLG1 is assertable to indicate the presence of a bit error. The trace inference data generator 112 is configured to generate one or more bit error trace inference signals for inferring the location of the bit error, such as a memory cell at the intersection of an identified row and an identified column in an array. In some embodiments, examples of the bit error trace inference signals include Figure 2C pointer signal (pointer) EPT of Figure 2C and flag signal (flag) FLG2 of
[0021] Since the bit error detector 110 and the trace inference data generator 112 are located in the same region, i.e., the CIM region 103A where other components of the CIM region 103A are located on the same die, the components of the CIM region 103A are physically closer to each other compared to the case where the bit error detector 100 and the trace inference data generator 112 are located on a different die from other components of the CIM region 103A. The increased proximity of the components of the CIM region 103A to each other helps to improve the operation speed. For example, the detector 110 can detect bit errors faster and the generator 112 can generate trace inference data faster.
[0022] Figure 1B is a functional block diagram of a W&P array 104B according to some embodiments.
[0023] In some embodiments, the W&P array 104B is Figure 1AAn example of the W&P array 104A, etc. The W&P array 104B is organized into N rows (see Figure 2A ) and Q slices 118(0)-118(Q-1), where N and Q are corresponding positive integers. Each row in the W&P array 104B has corresponding segments in the slices 118(0)-118(Q-1). Each of the slices 118(0)-118(Q-1) includes corresponding first and second sub-arrays. Regarding slice 118(0), the first sub-array is a two-dimensional (2D) weight bit array 120(0) and a one-dimensional (1D) parity bit array 122(0). Regarding slice 118(Q-1), the first sub-array is a 2D weight bit array 120(Q-1) (see Figure 2A ) and a 1D parity bit array 122(Q-1) (see Figure 2A ). In some embodiments, each of the parity bits 122(0)-122(Q-1) is a 2D array.
[0024] Figure 1C is a functional block diagram of a digital CIM system 100C according to some embodiments.
[0025] The CIM system 100C is similar to Figure 1A the CIM system 100A. For the sake of brevity, the discussion will focus on the differences between the CIM system 100C and the CIM system 100A rather than the similarities.
[0026] In Figure 1C , the CIM system 100C includes two semiconductor die, namely a first semiconductor die 102C(1) and a second semiconductor die 102C(2), while the CIM system 100A includes one die 102A. The die 102C(1) includes at least a CIM region 103C and a region 155(1). The die 102C(2) includes a region 155(2) and 114C. In some embodiments, the CIM system 100C is included as part of an AI system.
[0027] Figure 1C The CIM region 103C and the die 102C(1) of Figure 1A correspond to the CIM region 103A and the die 102A of Figure 1A . However, the parity encoder 158 and the trajectory inference data generator 112 are included in the region 155(1) of the CIM system 100C rather than in the CIM region 103A, which is different from the CIM system 100A. In some embodiments, as indicated by the imaginary line (dashed line), one or more of the parity encoder 158 or the trajectory inference data generator 112 are included in the region 155(2) of the die 102C(2) rather than in the CIM region 103C of the die 102C(1), which is another difference relative to the CIM system 100A.
[0028] Figure 1C The region 114C of Figure 1A corresponds to the region 114A of . However, the region 114C is a region of die 102C(2), rather than a region of die 102C(1). Thus, the bit corrector 116 is included as part of die 102C(2), rather than as part of die 102C(1), which is another difference relative to the CIM system 100A. In some embodiments, the region 114C is a processor, and the bit error corrector 116 represents a function executed by the processor.
[0029] Since the bit error detector 110 and the trajectory inference data generator 112 are on the same die as the CIM region 103C, the bit error detector 110 and the trajectory inference data generator 122 are physically closer to the components of the CIM region 103C than in the case where the bit error detector 100 and the trajectory inferable data generator 112 are on a different die from the CIM region 103C. The increased proximity of the bit error detector 110 and the trajectory inference data generator 112 to the components of the CIM region 103C facilitates advantages including increased operating speed, e.g., faster bit error detection by the detector 110 and faster trajectory inference data generation by the generator 112.
[0030] Figure 1D is a functional block diagram of a digital CIM system 100D according to some embodiments.
[0031] The CIM system 100D is similar to Figure 1A the CIM system 100A of . For the sake of brevity, the discussion will focus on the differences between the CIM system 100D and the CIM system 100A, rather than the similarities.
[0032] The CIM system 100D includes a semiconductor die 102D, which includes a CIM region 124D and a region 114D. Figure 1D The CIM region 124D and the die 102D of Figure 1A correspond to the CIM region 103A and the die 102A of . Figure 1D The W&P array 104B, multiplier 107B, and adder tree 108B of Figures 3A - 3L (also see Figure 1A ) correspond to the W&P array 104A, multiplier 107A, and adder tree 108A of . The trajectory inference data generator 126 (see Figures 3G - 3L ) corresponds to the trajectory inference data generator 112. However, the trajectory inference data generator 126 also includes a row sum generator 127 (see Figure 3I ) and a column sum generator 128 (see Figure 3K)。The trajectory inference data generator 126 is configured to generate a bit error trajectory inference signal for inferring the location of a bit error, such as a memory cell at the intersection of an identified row and an identified column in the array. Examples of bit error trajectory inference signals include row and vector (e.g., Row_Sum 304J in Figure 3J ), column and vector (such as Col_Sum 304L in Figure 3L ), etc. In some embodiments, the CIM system 100D is included as part of an AI system.
[0033] The CIM region 124D also includes an input data bit and checksum bit array 105 (see Figures 3A - 3C ), a product array 106 (see Figures 3G - 3H ), and a checksum generator 178 (see Figures 3H - 3L ). The checksum generator 178 also includes a row checksum generator 127 (see Figures 3A - 3C ) and a column checksum generator 128 (see Figures 3D - 3F ).
[0034] The CIM region 124D neither includes a corresponding part of the parity encoder 158 nor an error code detector 110. However, the region 114D of the die 102D includes a bit error detector, locator, and corrector 130 corresponding to the bit error detector 110 (see Figure 3M ). In some embodiments, the region 114D is a processor, and the bit error detector, locator, and corrector 130 represent corresponding functions executed by the processor.
[0035] Since the checksum generator 178 and the trajectory inference data generator 126 are located in the same region of the same die as the other components of the CIM region 124D, the components of the CIM region 124D are physically closer to each other compared to the case where the checksum generator 178 and the trajectory inference data generator 126 are located on a different die from the other components of the CIM region 124. The increased proximity of the components of the CIM region 124D to each other helps to improve the operating speed. For example, the generator 178 generates checksums faster, the generator 126 generates trajectory inference data faster, etc.
[0036] Figure 1E is a functional block diagram of a digital CIM system 100E according to some embodiments.
[0037] The CIM system 100E is similar to the Figure 1D CIM system 100D. For the sake of brevity, the discussion will focus on the differences between the CIM system 100E and the CIM system 100D rather than the similarities.
[0038] In Figure 1EIn [the figure], CIM system 100E includes two semiconductor dies, namely the first semiconductor die 102E(1) and the second semiconductor die 102E(2), while CIM system 100D includes one die 102D. Die 102E(1) at least includes CIM region 124E and region 155(3). Die 102E(2) includes region 114E, and in some embodiments, also includes region 155(4). In some embodiments, CIM system 100E is included as part of an AI system.
[0039] Figure 1E The CIM region 124E and die 102E(1) correspond to Figure 1D The CIM region 124D and die 102D. However, the checksum generator 178, including the row checksum generator 179 and the column checksum generator 180, is included in region 114D of CIM system 100E, rather than in CIM region 124D, which is a difference relative to CIM system 100D. In some embodiments, as shown by the dashed line, one or more checksum generators 178, including the row checksum generator 179 and the column checksum generator 180, are included in region 155(4) of die 102E(2), rather than in the CIM region 124E of die 102E(1), which is another difference relative to CIM system 100D.
[0040] Figure 1E Region 114E of [the figure] corresponds to Figure 1D Region 114D of [the figure]. However, region 114E is a region of die 102E(2), rather than a region of die 102E(1). Therefore, the error code detector, corrector, and locator 130 are included as part of die 102E(2), rather than as part of die 102E(1), which is another difference relative to CIM system 100D. In some embodiments, region 114E is a processor and a bit error detector, and the locator and corrector 130 represent corresponding functions executed by the processor.
[0041] Because the checksum generator 178 is located in region 155(3) and is thus on the same die as the CIM region 124E, the checksum generator 178 is physically closer to the components of the CIM region 124 compared to the case where the checksum generator 176 is on a different die from the CIM region 124E. The increased proximity of the checksum generator 178 to the components of the CIM region 124E facilitates advantages including increased operating speed, such as the generator 178 generating checksums more quickly.
[0042] Figure 2A is a schematic diagram of the CIM region 203(1) of a digital CIM system according to some embodiments.
[0043] The CIM region 203(1) is Figure 1A a part of the CIM region 103A of Figure 1C a part of the CIM region 103C of, etc. Multiple instances of the CIM region 203(1) include the CIM region 103A, the CIM region 103C, etc.
[0044] The CIM region 203(1) includes: a weight and parity bit (W&P) array 204A, which includes slices 218(0)-218(Q - 1); an enhanced multiplier (EM) array 236 of EM blocks 238(0)-238(Q - 1), each EM block including a multiplier, such as 251(Q - 1); and an adder tree 208; and where Q is a positive integer. In some embodiments, Q is a power of 2. In some embodiments, Q = 64. In some embodiments, Q is equal to a positive integer power of 2 other than 64.
[0045] On a row-by-row basis, the EM array 236 is configured to receive data bits of a given row from the W&P array 204A. Thus, the EM blocks 238(0)-238(Q - 1) are configured to receive corresponding segments of a given row. Figure 2A The EM array 236 of is also configured to receive given column input bits corresponding to the given row data bits from the W&P array 204A from the input array XIN1 and multiply them to obtain Q products PRD1(0)-PRD1(Q - 1). The adder tree 208 adds the Q products PRD1(0)-PRD1(Q - 1) to generate an output signal Out_1. The output signal Out_1 is operated on, for example, by a bit error corrector 216 (see Figure 2D ).
[0046] Taking the EM block 238(Q - 1) as an example, the EM block 238(Q - 1) includes a multiplier 251(Q - 1), a bit error detector (such as 210(Q - 1)), and a multiplexer (MUX) (254(Q - 1)). A part 242 of the W&P array 204A and the EM array 236 includes the slice 218(Q - 1) of the EM block 238(Q - 1) of the W&P array 204A and the EM array 236. Figure 2A Part 244(1) of part 242 is shown in more detail using a decomposition diagram in
[0047] In part 244(1), that is, in the decomposition diagram, the slice 218(Q - 1) includes a 2D weight array 220(Q - 1) of a bit memory cell 245 and a 1D parity array 221(Q - 1) of a bit memory cell 246. In Figure 2AIn this case, it is assumed that memory cells 245 and 246 are static random access memory (SRAM) cells. In some embodiments, memory cells 245 and 246 are a type of memory cell other than SRAM.
[0048] Memory cells 245 and 246 of slice 218(Q-1) are organized into rows and columns. For simplicity of illustration, some but not all of the signal lines for reading from or writing to slice 218(Q-1) are shown. Slice 218(Q-1) is configured to read the data bits of its single row at any given time. The selection of a given row in slice 218(Q-1) is controlled by corresponding read word lines RWL[0]-RWL[N-1], where N is a positive integer. Not only is a given row selected in slice 218(Q-1), but the same row is selected simultaneously in each of slices 218(0)-218(Q-2). The columns in slice 218(Q-1) have corresponding read bit lines RBL[0]-RBL
[12] .
[0049] Weight array 220(Q-1) is arranged with respect to lines RBL[0]-RBL
[11] . Thus, weight array 220(Q-1) is an Nx12 array, where N is a positive integer. Parity check array 221(Q-1) is arranged with respect to line RBL
[12] . Thus, parity check array 221(Q-1) is an Nx1 array. Thus, slice 218(Q-1) is an Nx(K+1) array.
[0050] For simplicity of illustration, Figure 2A it is assumed that each row in weight array 220(Q-1) stores a 12-bit word. More generally, weight array 220(Q-1) stores K-bit words, where K is a positive integer, assumed to be K = 12 in Figure 2A this case. Thus, weight array 220(Q-1) is an NxK array. In some embodiments, K is a power of 2, such as K = 8, K = 16, K = 32, etc. In some embodiments, K is a positive integer other than 8, 12, 16, or 32. Note that input XIN1 is a KxL array, where L is a positive integer.
[0051] Iteratively, on a row-by-row basis, the EM block 238(Q-1) is configured to generate an output signal PRD1(Q-1) representing the product of the multiplication of the columns of the N×K weight array 220(Q-1) with the corresponding columns of the K×L array XIN1. More specifically, iteratively, the multiplier 251(Q-1) is configured to receive the row data bits from the weight array 220(Q-1) as the multiplicand and the column input bits of the input array XIN1 as the multiplier, and multiply the multiplicand with the multiplier to obtain the product PRD1(Q-1). Thus, the multiplier 251(Q-1) is configured to receive the K data bits of a row on lines RBL[0]-RBL
[11] and the K input bits of a column of the input data XIN1, and multiply them to obtain the product PRD1(Q-1). The product PRD1(Q-1) is a single word having K+K = 2K bits. In Figure 2A the example of, the product PRD1(Q-1) has K+K = 12+12 = 24 bits. The operation of the multiplier 251(Q-1) is also discussed in Figure 4B the context of.
[0052] Taking the slice 218(Q-1) as the representation of the slices 218(0)-218(Q-1) of the W&P array 204A, when a value is stored in one of the memory cells 245 in a selected row of the slice 218(Q-1), i.e., the value on the lines RBL[0]-RBL
[11] represents a bit error, a single-bit error condition occurs in the slice 218(Q-1). When a value is stored in two of the memory cells 245 in a selected row of the slice 218(Q-1), i.e., the values on two of the lines corresponding to the lines RBL[0]-RBL
[11] represent bit errors, a double-bit error condition occurs in the slice 218(Q-1). The probability of a single-bit error condition occurring in the slice 218(Q-1) is low. The probability of a double-bit error condition occurring in the slice 218(Q-1) is much lower than the likelihood of a single-bit error condition occurring in the slice 218(Q-1). In fact, Figure 2A it is assumed that no double-bit error condition occurs in the slice 218(Q-1). Thus, Figure 2A is configured to detect and correct single-bit error conditions in the slice 218(Q-1), but is not configured to detect or correct double-bit error conditions in the slice 218(Q-1). At least some other embodiments disclosed herein are configured to detect and correct single-bit error conditions in the slice corresponding to the slice 218(Q-1), but do not detect and correct double-bit error conditions.
[0053] On a row-by-row basis, bit error detector 210(Q-1) is configured to receive K data bits on lines RBL[0]-RBL
[11] and a parity bit on line RBL
[12] . Based on the bit values of lines RBL[0]-RBL
[12] , bit error detector 210(Q-1) determines whether a bit error exists on one of lines RBL[0]-RBL
[11] and generates an output signal representing flag signal (flag) FLG1 based thereon. Flag FLG1 represents the output signal of EM block 238(Q-1) and is also provided internally to MUX 254(Q-1). It should be noted that line RBL
[12] is also provided to trajectory inference data generator 212 (see Figure 2C ).
[0054] Flag FLG1 is assertable to indicate the presence of a bit error, i.e., detector 210(Q-1) has detected a bit error in the corresponding data bit row from weight array 220(Q-1). Figure 2A Assume the following assertion states of flag FLG1: When flag FLG1 is not asserted, i.e., when FLG1 = 0, there is no error on lines RBL[0]-RBL
[11] ; when flag FLG1 is asserted, i.e., when FLG1 = 1, there is an error on one of rows RBL[0]-RBL
[11] .
[0055] Bit error detector 210(Q-1) includes an exclusive OR (XOR) gate 253(Q-1) configured to receive K+1 inputs. Generally, for any multi-input XOR gate, the output is true (or logic 1) when an odd number of inputs are true. In some embodiments, assume the opposite of the said assertion state of flag FLG1. The operation of XOR gate 253(Q-1) is also discussed in the context of Figure 4B .
[0056] According to another method corresponding to the CIM system of which CIM region 203(1) forms a part, bit error detection (pre-multiplication detection) of bit errors in the memory corresponding to weight array 220(Q-1) is performed before multiplication. According to another method, bit error detection is performed using two operation cycles before multiplication. In contrast, according to at least some embodiments, bit error detector 210(Q-1) is included in CIM region 203(1) to perform error code detection in parallel with the multiplication performed by multiplier 251(Q-1) using one operation cycle, which is one operation cycle faster than the other method.
[0057] Another method uses a Q-bit weight array corresponding to the weight array 220(Q-1). In addition, in the case of Q = 12, for each row in the weight array, another method uses 5 check bits to implement pre-multiplication detection, which imposes significant disadvantages in terms of the area consumed on the die (increased footprint), power consumption, routability of signal segments (see Figure 5C block 547 of Figure 5C ), and / or PG segments (see
[0058] 547 of
[0059] On a row-by-row basis, Figure 2A the EM array 236 of Figure 2D is configured to receive a column of bits from the input array XIN1, multiply them with the corresponding row data bits from the W&P array 204A to obtain Q products PRD1(0)-PRD1(Q-1). The adder tree 208 adds the Q products PRD1(0)-PRD1(Q-1) together to generate an output signal Out_1. The output signal Out_1 is operated on, for example, by a bit corrector 216 (see
[0060] The adder tree 208 is configured to receive the Q products PRD1(0)-PRD1(Q-1) from the EM array 236 and add them together. The adder tree 208 has J processes crs(0),..., crs(J-1) of adders 240, where J is a positive integer and J < Q. In some embodiments, the number of processes J in the adder tree 208 is related to the number Q of slices 218(0)-218(Q-1) as follows: Q is equal to 2 to the power of J, i.e., Q = 2 J。In such an embodiment, the EM array 236 generates Q product words, where the W&P array 204A has Q slices that provide Q words to the EM array 236. Accordingly, the adder tree includes J processes crs(0),..., crs(J - 1) of adders 240 and generates a single word as the output signal Out_1. Each adder 240 is configured to receive two single-word inputs. For example, in the case of Q = 64, the enhanced multiplier array 236 has J = 6 processes.
[0061] In some embodiments, each of the Q words is represented by 2K bits, such that a single word Out_1 is represented by Q * 2K = (2 J ) * 2K bits. In some such embodiments, 2K = 24, such that a single word Out_1 is represented by 1536 = (2 6 ) * 24-bit words.
[0062] Figure 2B is a schematic diagram of part 244(2) of part 242 of CIM region 203(1) of a digital CIM system according to some embodiments.
[0063] Part 244(2) includes slice 218(Q - 1) of the W&P array 204A and parity encoder 259(Q - 1). Figure 2B extends the Figure 2A example where K = 12. In some embodiments, the parity encoder 259(Q - 1) is Figure 1A an example of one of the parity encoders 158 such as
[0064] On a per-row basis, the parity encoder 259(Q - 1) is configured to generate the value of a parity bit (parity value) corresponding to the bit values of the row and write it to a corresponding one of the memory cells 246. Each memory cell 246 is also configured to be selectively written with the parity value from the parity encoder 259(Q - 1).
[0065] A given row of slice 218(Q-1) includes K = 12 instances of memory cells 245 in weight array 220(Q-1) corresponding to read bit lines RBL[0]-RBL
[11] and one instance of memory cell 246 in parity array 221(Q-1). Thus, for a given row of slice 218(Q-1), parity encoder 259(Q-1) is also configured to receive K data bits from weight array 220(Q-1) of slice 218(Q-1) located on read bit lines RBL[0]~RBL
[11] . Based on the bit values of lines RBL[0]-RBL
[11] , parity encoder 259(Q-1) generates a parity value corresponding to the given row. Then, the parity value of the given row is written into the instance of memory cell 246 included in the given row.
[0066] Parity encoder 259(Q-1) includes XOR gate 260(Q-1), which is configured to receive K inputs corresponding to the bit values on lines RBL[0]-RBL
[11] and generate an output signal representing the corresponding parity bit. The operation of XOR gate 260(Q-1) of parity encoder 259(Q-1) is also described Figure 4A in the context of.
[0067] Figure 2C is a schematic diagram of CIM region 203(2) of a digital CIM system according to some embodiments.
[0068] CIM region 203(2) is Figure 1A a part of CIM region 103A of, Figure 1C a part of CIM region 103C of, etc. Examples of multiple instances of CIM region 203(2) include CIM region 103A, CIM region 103C, etc. Trajectory inferable data generator 212 is Figure 1A , Figure 1C etc. an example of trajectory inference data generator 112. Figure 2C extends Figure 2A the example where K = 12 in.
[0069] Trajectory inference data generator 212 is configured to receive Q instances of flag FLG1 (Q flags FLG1) from EM array 236, receive Q parity bits from W&P array 204A, and generate a bit error trajectory inference signal based on this. The bit error trajectory inference signal generated by generator 212 includes a pointer signal (pointer) EPT and a flag signal (flag) FLG2. Pointer EPT and flag FLG2 are operated by bit corrector 216 (see Figure 2D ), for example. It should be noted that Q flags FLG1 correspond to Q slices 218(0)-218(Q-1) of W&P array 204A.
[0070] In Figure 2C it, the trajectory inference data generator 212 includes a Q:P encoder 262 and a slice error detector 213.
[0071] The encoder 262 is configured to receive Q instances of the flag FLG1 from the EM array 236 and generate a pointer EPT, where the pointer EPT is a P-bit word, P is a positive integer, P < Q, and Q = 2 P . In some embodiments, the encoder 262 receives Q instances of the flag FLG1 as a concatenation of Q instances of the flag FLG1. The operation of the encoder 262 is also discussed in the context of Figure 4B .
[0072] In some embodiments, the Q flags FLG1 are provided to the encoder 262 as a word having Q bits, such that Q = {f(0), f(1), …, f(i - 1), f(Q - 1)}. Extending an example of the asserted state of the flag FLG1 discussed in the context of Figure 2A , where none of the Q slices of the array W&P array 204A have bit errors, each of the bits f(0) - f(Q - 1) is set to a logical zero value. However, in the case where a given slice 218(i) has a bit error, the bit f(i) of the word having Q bits is set / asserted to a logical value of 1, i.e., f(i) = 1; and the remaining bits f(0), f(1), …, f(i - 1), f(i + 1), …, f(Q - 2), f(Q - 1) are not asserted, that is, the remaining bits are set to logical zero values.
[0073] Figure 2C The slice error detector 213 of
[0074] is configured to receive Q instances of the flag FLG1 from the EM array 236 and generate a flag signal (flag) FLG2. The flag FLG2 is assertable to indicate the presence of a slice error, i.e., the detector 213 has detected a slice error. An OR gate 263 is included in the slice error detector 213. The OR gate 263 is configured to receive Q flags FLG1 from the EM array 236 and generate the flag FLG2.
[0074] Regarding the flag FLG2, the example of the asserted state of the flag FLG1 discussed in the context of Figure 2A is extended as follows: When the flag FLG2 is not asserted, i.e., when FLG1 = 0, there is no slice error between the slices 218(0) - 218(Q - 1); when the flag FLG2 is asserted, i.e., when FLG2 = 1, there is a slice error in one of the slices 218(0) - 218(Q - 1).
[0075] Figure 2D is a schematic diagram of a digital CIM system 200D according to some embodiments.
[0076] The CIM system 200D is Figure 1A an example of the CIM system 100A Figure 1C and the CIM system 100C, etc. The CIM system 200D includes a CIM area 203(3) and a bit error corrector 216.
[0077] The CIM area 203(3) is Figure 1A an example of the CIM area 103A Figure 1C and a combination of the CIM area 103C and the area 155(1), etc. The bit error corrector 216 is Figure 1A , Figure 1C an example of the bit error corrector 116, etc.
[0078] The bit error corrector 216 is configured to receive an output signal Out_1, a pointer EPT, and a flag FLG2 from the CIM area 203(3) and operate on them according to the flowchart 290. In Figure 2D it, the flowchart 290 is shown as a decomposition diagram 244D of the bit error corrector 216. The operation of the bit error corrector 216 is also discussed in the context of Figure 4B .
[0079] The flowchart 290 includes blocks 265(1)-265(5). At block 265(1), a decision is made as to whether to assert the flag FLG2, i.e., FLG2 = 1. If the result of block 265(1) is no, i.e., FLG2 = 0, the process continues to block 265(2).
[0080] At block 265(2), the process stops because there is no chip error and thus no bit correction of the output signal Out_1 is required. If the result of block 265(1) is yes, i.e., FLG2 = 1, there is a chip error and bit correction is required, so the process continues to block 265(3).
[0081] At block 265(3), the localization of the slice error is performed. It should be recalled that: the output signal Out_1, the pointer EPT, and the flag FLG2 are generated iteratively row by row; and each iteration of the output signal Out_1, the pointer EPT, and the flag FLG2 is based on the current multiplicand, i.e., the corresponding column in the input column of XIN1, and the current multiplier, i.e., the corresponding row in the data row of the W&P array 204A. Once the flag FLG1 is asserted, the damaged row in slice 218(Q - 1) is identified as the current multiplier. Then, block 265(3) will identify the slice having the slice error. Thus, in block 265(3), the pointer EPT is checked to determine which one of bits f(0)-f(Q1) is asserted, i.e., set to the value 1. Whichever one of bits f(0)-f(Q1) is set to the value 1 indicates that a slice error exists in the corresponding one of slices 218(0)-218(Q - 1). The flow proceeds from block 265(3) to blocks 265(4) and 265(5).
[0082] A single damaged slice situation occurs when a single one of slices 218(0)-218(Q - 1) experiences a unit error condition. A double damaged slice situation occurs when two of slices 218(0)-218(Q - 1) experience the corresponding unit error conditions. The probability of a single damaged slice scenario occurring among slices 218(0)-218(Q - 1) is low. The probability of a double damaged slice scenario occurring among slices 218(0)-218(Q - 1) is substantially lower than that of the single damaged slice scenario. Thus, generally, the bit sequence of the cascaded Q flags FLG1 will have only one bit whose value is set to logic 1. At block 265(3), regardless of the total number of bits in the bit sequence of the cascaded Q flags FLG1 that are set to logic 1, each bit that is set to logic 1 also identifies the corresponding slice as having a bit error.
[0083] At block 265(4), each of slices 218(0)-218(Q - 1) in the W&P array 204A identified by block 265(3) is updated. In some embodiments, the bit error corrector 216 is configured to update each damaged slice by writing the undamaged data bit value into the memory cell 245 in the row identified in block 265(3), e.g., by copying the corresponding undamaged data bit value from a source copy or a permanent copy of the W&P array 204A. For example, the source copy or the archival copy is stored outside the CIM region containing the W&P array 204A.
[0084] At block 265(5), the value of the corrected output signal Out_1 is adjusted. In some embodiments, the output signal Out_1 is stored in a first register. In some embodiments, the bit corrector 216 is configured to multiply the current multiplicand (i.e., a corresponding column in the input column of XIN1) by the current multiplier (i.e., a corresponding row in the data row of the W&P array 204A) to form a corrected product and write the corrected product into the first register.
[0085] In some embodiments, blocks 265(4) and 265(5) are executed substantially simultaneously. In some embodiments, block 265(4) is executed before block 265(5). In some embodiments, block 265(5) is executed before block 265(4).
[0086] Figure 3A is a schematic diagram of the CIM region 324 of a digital CIM system according to some embodiments.
[0087] The CIM region 324 is Figure 1D a part of the CIM region 124D of Figure 1E a part of the CIM region 125E of, etc. Examples of multiple instances of the CIM region 324 include the CIM region 124D, the CIM region 124E, etc. The CIM region 324 is similar to Figure 2A the CIM region 203(1) of. For the sake of brevity, the discussion will focus on the differences between the CIM region 324 and the CIM region 203(1), rather than the similarities.
[0088] The CIM region 324 includes: a weight and checksum bit (W&C) array 304A including slices 318(0)-318(Q-1); a multiplier (MX) array 337 of multipliers 352(0)-352(Q-1); an adder tree 308; a C1 product array 306A; and a trajectory inference data (LID) generator 326A; and where Q is a positive integer. In some embodiments, Q is a power of 2. In some embodiments, Q = 64. In some embodiments, Q is equal to a positive power of 2 other than 64.
[0089] A portion 342 of the W&C array 304A and the MX array 337 includes the slice 318(Q-1) of the W&C array 304A and the multiplier 352(Q-1) of the MX array 337. Figure 3A A part of the portion 342 is shown in more detail using the decomposition diagram 344A. As shown in the decomposition diagram, a part of the portion 342 includes the slice 318(Q-1), the multiplier 352(Q-1), and the multiplier 364(Q-1). Although Figure 3A the multiplier 364(Q-1) is shown in, it should be noted that for simplicity of illustration, Figure 3AThe multipliers 364(0)-364(Q-2) are not shown.
[0090] Slice 318(Q-1) includes a 2D weight array 320(Q-1) and a 2D checksum (CHK) array 323(Q-1). The weight array 320(Q-1) includes a bit memory cell 349 and a CHK array 323(Q-1) consisting of a one-bit memory cell 350. In Figure 3A it is assumed that the memory cells 349 and 350 are SRAM cells. In some embodiments, the memory cells 349 and 350 are a type of memory cell other than SRAM.
[0091] On a row-by-row basis, the MX array 337 is configured to receive segments of the given row weight bits from the W&C array 304A. Thus, each multiplier 352(0)-352(Q-1) of the MX array 337 is configured to receive a segment of the corresponding weight bits of the given row. The MX array 337 is also configured to receive the input bits of the given column from the input array XIN2 (see Figure 3B ) corresponding to the bits of the given row from the W&C array 304A, and multiply them to obtain Q first products PRD2(0)-PRD2(Q-1). The first products PRD2(0)-PRD2(Q-1) are also provided to the adder tree 308A.
[0092] The adder tree 308A adds the Q first products PRD2(0)-PRD2(Q-1) to generate an output signal Out_2. The output signal Out_2 is operated on by, for example, a bit error detector, locator, and corrector unit 316 (see Figure 3M ). The adder tree 308A is similar to Figure 2A the adder tree 208. The operation of the multiplier 352(Q-1) is also discussed in the context of Figures 3E - 3G and Figure 4C .
[0093] In Figure 3A the CIM region 324, the multipliers 352(0)-352(Q-1) correspond to the respective multipliers 251(0)-251(Q-1) of the EM blocks 238(0)-238(Q-1) of Figure 2A . However, the CIM region 324 does not include the bit error detectors 210(0)-210(Q-1), nor does it include the MUXes 254(0)~254(Q-1) of the EM blocks 238(0)-238(Q-1) of Figure 2A . Instead, the CIM region 324 includes multipliers 364(0)-364(Q-1), a C1 product array 306A, and an LID generator 326A, while Figure 2A does not.
[0094] On a row-by-row basis, multiplier 364(Q-1) is configured to receive the corresponding row of weights and checksum bits from slice 318(Q-1). Multiplier 364(Q-1) is also configured to receive input bits from a given column of the input array XIN3 (see Figure 3D ) corresponding to the bits of a given row from slice 318(Q-1), and multiply them together to obtain a second product D[i][j], where D[j][j] represents a word in C1 array 306A, and i and j are non-negative integers.
[0095] On a row-by-row basis, the second product D[i][j] is cumulatively stored by multiplier 364(Q-1) and other multipliers 364(0)-364(Q-2) in C1 array 306A (see Figure 3H C1 array 306H in). The operation of multiplier 364 is also described in Figure 3H and Figure 4C in the context of.
[0096] LID generator 326A is configured to operate on C1 array 306A and generate a bit error locus inference (BELI) signal including a row sum signal Row_Sum and a column sum signal Col_Sum. The operation of LID generator 326A is also described in Figures 3I - 3M and Figures 4D - 4E in the context of. In Figure 3A , LID generator 326A is configured to receive row (i) of C1 from C1 array 306A.
[0097] For simplicity of illustration, Figure 3A assume the following: there are N rows in slice 318(Q-1), so there are N rows in each of weight array 320(Q-1) and CHK array 323(Q-1); each row in weight array 320(Q-1) stores a K = 12-bit word. Figure 3A This assumption in Figure 2A is similar to the assumption in
[0098] Figure 3A Also assume that there are N rows in CHK array 323(Q-1), and each row in CHK array 323[Q-1] stores an 11-bit word corresponding to read bit lines RBL
[12] -RBL
[22] , such that CHK array 323(Q-1) is an N×11 array. For each given row containing segments in slices 318(0)-318(Q-1), there are thus segments in each of weight array 320(Q-1) and CHK array 323(Q-1) in slices 318(0) to 318(Q-1), and the 11-bit word / segment in CHK array 323(Q-1) represents the checksum of the 12-bit word / segment in weight array 320(Q-1).
[0099] The CHK array 323(Q - 1) stores S-bit words, where S is a positive integer, assumed to be S = 11 in Figure 3A . For further generalization, the CHK array 323(Q - 1) is an NxS array. Thus, the slice 318(Q - 1) is an Nx(K + S) array. It should be noted that the input XIN3 (see Figure 3D 305D of ) is a (K + S)xZ array, where S and Z are corresponding positive integers. In some embodiments, S is a positive integer other than 11. To simplify Figure 3D the illustration of the positions in XIN3 and Figures 3B - 3C XIN2, K + S is represented by V, i.e., V = K + S, where V is a positive integer.
[0100] Regarding the multiplier 352(Q - 1), iteratively (i.e., row by row), the multiplier 352(Q - 1)i is configured to generate an output signal PRD2(Q - 1) representing the product of the row of the NxK weight array 320(Q - 1) multiplied by the corresponding column of the KxZ array XIN2. More specifically, iteratively, the multiplier 352(Q - 1) is configured to receive the row weight bits from the weight array 320(Q - 1) as the multiplicand and the column input bits of the input XIN2 as the multiplier, and multiply the multiplicand by the multiplier to obtain the product PRD2(Q - 1). The multiplier 352(Q - 1) is configured to receive a row of K data bits on lines RBL[0] - RBL
[11] and a column of K input bits of the input data XIN2 and multiply them together to obtain the product PRD2(Q - 1). The product PRD2(Q - 1) is a single word with K + K = 2K bits. In Figure 3A the example of, the product PRD2(Q - 1) has 2K = 12 + 12 = 24 bits.
[0101] Regarding the multiplier 364(Q - 1), iteratively (i.e., row by row), the multiplier 364(Q - 1) is configured to generate an output signal PRD2(Q - 2) representing a second product D[i][j], which is obtained by multiplying the row of the Nx(K + S) slice 318(Q - 1) by the corresponding column of the (K + S)×Z input array XIN3. More specifically, iteratively, the multiplier 364(Q - 1) is configured to receive a row of bits from the slice 318(Q - 2) as the multiplicand, receive a column of input bits of the input data XIN3 as the multiplier, and multiply the multiplicand by the multiplier to obtain the second product D[i][j], where the latter is shown to be provided to the C2 array 306A. The multiplier 364(Q - 1) is configured to receive a row of K + S data bits on lines RBL[0] - RBL
[22] and a column of K + S input bits of the input data XIN3 and multiply them together to obtain the second product D[i][j].
[0102] Similar to Figure 2A , as a practical matter Figure 3A it is assumed that there will be no double-bit error condition in slice 318(Q-1). Thus Figure 3A is configured to detect and correct single-bit error conditions in slice 318(Q-1), but is not configured to detect or correct double-bit error conditions in slice 318(Q-1). At least some other embodiments disclosed herein are configured to detect and correct single-bit error conditions in the slice corresponding to slice 318(Q-1), but do not detect and correct double-bit error conditions.
[0103] Although Figure 3A the discussion has been presented in terms of bit lines and the number of bits in segments of rows and columns, Figures 3B - 3M it is presented in terms of words.
[0104] Figure 3B is a schematic diagram of an input XIN2 array 305B according to some embodiments.
[0105] The input XIN2 array 305B is operated by a row checksum generator 379 to generate a row vector 386D, where the latter is appended to the XIN2 array 305B to form Figure 3D the XIN3 array 305D of Figure 1D The row checksum generator 379 is an example of the row checksum generator 179 of
[0106] The XIN2 array 305B is an FxG array, where F and G are corresponding positive integers. Each position (i, j) in the XIN2 array 305B represents a corresponding word A[i][j], where i and j are corresponding non-negative integers. Thus, the XIN2 array 305B includes positions A[0][0], …, A[F-1][G-1].
[0107] Figure 3C is a schematic diagram of a row checksum generator 379 according to some embodiments.
[0108] The row checksum generator 379 is a row checksum generator configured to generate a row checksum R_ChkSum_1. The row checksum generator 379 is Figure 1D an example of the row checksum generator 179 in
[0109] Each recursive adder 377 includes an adder 240 and a register 382. For each recursive adder 377, the adder 240 is configured to receive words from a corresponding column and receive the words in the register 382. Each instance of the register 382 is initialized to store zero. At time t = 0, the adder 240 adds the corresponding word in the 0th row of the XIN2 array 305B to the word in the corresponding register 382 (all of which were previously initialized to zero) and stores / overwrites the sum at t = 0 inside the corresponding register 382. At time t = 1, the adder 240 adds the corresponding word in the 1st row of the XIN2 array 305B to the word in the corresponding register 382 at t = 0 and stores / overwrites the corresponding register 382 with the sum at t = 1. At time t = F - 1, the adder 24 adds the corresponding word in the V - 2 row of the XIN2 array 305B to the word in the corresponding register 382 at t = F - 2 and stores or overwrites the corresponding register 382 with the sum at t = F - 1. The words at t = F - 1 in the G instances of the register 382 represent the vector R_ChkSum_1, which is appended to Figure 3B the XIN2 array 305B to form Figure 3D the input XIN3 array 305D.
[0110] Figure 3D is a schematic diagram of the input XIN3 array 305D according to some embodiments.
[0111] The XIN2 array 305B is the result of appending the R_ChkSum_1 vector 386D to the input XIN2 array 305B. Thus, the XIN3 array 305D is an (F + 1) x G array.
[0112] Figure 3E is a schematic diagram of the weight W1 array 304E according to some embodiments.
[0113] The weight W1 array 304E is Figure 3A an example of a weight array composed of parts of the weight array 320(Q - 1) of etc. The weight W1 array 304E is operated on by the column checksum generator 380 to generate Figure 3G the weight W2 array 304G and the array 305G of. The column checksum generator 380 is Figure 1D an example of the column checksum generator 180 of etc.
[0114] The weight W1 array 304E is an E x F array, where E is a positive integer. Each position (i, j) in the weight W1 array 304E represents the corresponding word B[i][j]. Thus, the weight W1 array 304E includes positions B[0][0], …, B[E - 1][F - 1].
[0115] Figure 3FSchematic diagram of column checksum generator 380 according to some embodiments.
[0116] Column checksum generator 380 is a column checksum generator configured to generate a column vector 387G representing checksum C_ChkSum_1, where the latter is appended to W1 array 304E to form Figure 3G weight W2 array 304G. Column checksum generator 380 is Figure 1D an example of column checksum generator 180 and the like. Column checksum generator 380 includes adder tree 308F. On a row-by-row basis, column checksum generator 380 is configured to generate a sum corresponding to the words in a given row of W1 array 304E and store the sum in Figure 3G the corresponding row of column C_ChkSum_1 vector 387G.
[0117] The output of column checksum generator 380 is C_chk∑(T = x), where x is a non-negative integer variable. At time t = 0, generator 380 adds the corresponding words in row 0 of W1 array 304E and stores the resulting sum as word C_chk∑(T = 0) in row 0 of checksum column vector 387G. At time t = 1, generator 380 adds the corresponding words in row 1 of W1 array 304E and stores the resulting sum as word C_chk∑(T = 1) in row 1 of checksum column vector 387G. At time t = E - 1, generator 380 adds the corresponding words in row E - 1 of W1 array 304E and stores the resulting sum as word C_chk∑(T = E - 1) in row E - 1 of checksum column vector 387G. Words C_chk∑(T = 0), …, C_chk∑(T = E - 1) together represent checksum column vector 387G, which is appended to Figure 3E weight W1 array 304E to form Figure 3G weight W2 array 304G.
[0118] Figure 3G Schematic diagram of weight W2 array 304G according to some embodiments.
[0119] W2 array 306G is the result of appending a column vector (i.e., C_ChkSum_1 vector 387G) to W1 array 304E. Thus, weight W2 array 306G is an Ex(F + 1) array.
[0120] Figure 3H Schematic diagram of product C1 array 306H according to some embodiments.
[0121] C1 array 306H is Figure 3AAn example of the C1 array 306A. The C1 array 306H is an ExG array. Each position (i, j) in the C1 array 306H represents the corresponding word C[i][j]. Thus, the C1 array 306H includes positions C[0][0], …, C[E-1][G-1].
[0122] The C1 array 306H is the product of the W2 array 304G and the XIN3 array 305D, i.e., C1 = W2 * XIN3. The C1 array 306H is an example of an array whose rows have been generated row by row by a multiplier 352, etc. The bottom row 388 of the C1 array 306H (see Figure 3A ) represents the row vector R_ChkSum_2 corresponding to the row R_ChkSum_1 vector 386D of Figure 3J . The rightmost column 389 of the C1 array 306H (see Figure 3D ) represents the column vector C_ChkSum_2 corresponding to the column C_ChkSum_1 vector 387G of Figure 3J . Figure 3G
[0123] The C1 array 306H is operated on by at least the following: a row sum generator (see Figures 3I - 3J ) for generating a row vector Row_Sum, which is used by a bit error detector, locator, and corrector (see Figure 3M ); and a column sum generator (see Figures 3K - 3L ) for generating a column vector Col_Sum, which is also used by a bit error detector, locator, and corrector (see Figure 3M ).
[0124] Figure 3I is a schematic diagram of the row sum generator 326I according to some embodiments.
[0125] The row sum generator 326I is a row checksum generator configured to generate Figure 3J the row sum 304J. The row sum generator 326I is Figure 1D Examples of row sums and generator 127. The row sum generator 326I is an array of G recursive adders 377 corresponding to the G columns of the C1 array 306H. Each recursive adder 377 is configured to generate a sum corresponding to the words in the corresponding column of the C1 array 306H by adding the words C[i][j] in the corresponding column row by row and column by column. At time t = 0, the adder 240 adds the corresponding words in the 0th row of the C1 array 306H to the words in the corresponding registers 382 (all of which were previously initialized to zero) and stores / overwrites the sum of the corresponding registers 382 at t = 0. At time t = 1, the adder 240 adds the corresponding words in the 1st row of the C1 array 306H and the words in the corresponding registers 382 at t = 0 and stores / overwrites the sum of the corresponding registers 382 at t = 1. At time t = N-1, the adder 24 adds the corresponding words in the (N-1)th row of the C1 array 326I and the words in the corresponding registers 382 at t = N-2 and stores / overwrites the sum of the corresponding registers 382 at t = N-1. The words at t = N-1 in the G instances of the register 382 represent the row vector Row_Sum, as shown in Figure 3J shown in 304J.
[0126] Figure 3J is a schematic diagram of the Row_Sum vector 304J according to some embodiments.
[0127] In Figure 3J the Row_Sum vector 304J is shown adjacent to and aligned below the C1 array 306H. In some embodiments, the Row_Sum vector 304J is a 1D array of one-bit memory cells 356.
[0128] Figure 3K is a schematic diagram of the column sum generator 328K according to some embodiments.
[0129] The column sum generator 328K is configured to generate Figure 3L the column vector sum and Col_Sum of the column checksum generator. The column sum generator 328K is Figure 1D an example of the column sum generator 128 such as. The column sum generator 328K includes an adder tree 308K(Q-1). On a row-by-row basis, the column sum generator 328K is configured to generate a sum corresponding to the words in a given row of the C1 array 306H and store the sum in Figure 3Lin the corresponding rows of the column vector Col_Sum 304L. At time t = 0, the generator 328K adds the corresponding words in the 0th row of the C1 array 306H and stores the resulting sum as the word Col∑(T = 0) in the 0th row of the column vector Col_Sum 304L. At time t = 1, the generator 328K adds the corresponding words in the 1st row of the C1 array 306H and stores the resulting sum as the word Col∑(T = 1) in the 1st row of the column vector Col_Sum 304L.... At time t = E - 1, the generator 328K adds the corresponding words in row N - 1 of the C1 array 306H and stores the resulting sum as the word Col∑(T = E - 1) in row N - 1 of the column vector Col_sum 304L. Collectively, ColΣ(T = 0), …, ColΣ(T = E - 1) represent the column vector Col_Sum 304L, in Figure 3L which is labeled as Col_Sum 304L.
[0130] Figure 3L is a schematic diagram of the Row_Sum vector 304J and the Col_Sum vector 304L according to some embodiments.
[0131] In Figure 3J it, the Col_Sum vector 304L is shown adjacent to the C1 array 306H and aligned with it on the right. Also in Figure 3J it, the Row_Sum vector 304J is shown adjacent to the C1 array 306H and aligned below it. In some embodiments, the Col_Sum vector 304L is a 1D array consisting of one-bit memory cells 357.
[0132] Figure 3M is a schematic diagram of the digital CIM system 300D according to some embodiments.
[0133] The CIM system 300D is Figure 1D an example of the CIM system 100D of Figure 1E the CIM system 100E of, etc. The CIM system 300D includes a CIM region 324 and a bit error detector, locator, and corrector (DLC) unit 316.
[0134] The CIM region 324 is Figure 1D an example of the CIM region 124D of Figure 1E a combination of the CIM region 124E and the region 114 of, etc. The DLC unit 316 is Figure 1D , Figure 1E etc., an example of the bit error detector, locator, and corrector 130.
[0135] The DLC unit 316 is configured to receive the output signals Out_2, R_ChkSum, C_ChkSum, Row_Sum, and Col_Sum from the CIM region 304M and operate on them according to the flowchart 383. In Figure 3M , the flowchart 383 is shown as the exploded view 344M of the DLC unit 316. The operation of the DLC unit 316 is also discussed in the context of Figures 4C - 4G .
[0136] The flowchart 383 includes blocks 384(1)-384(5). In some embodiments, the process advances from Figure 5D block 560 to block 384(1), as shown by the off-page connector 392.
[0137] At block 384(1), a decision is made as to whether the following conditions hold: (1) R_ChkSum = Row_Sum and (2) C_ChkSum = Col_Sum. If the result of block 384(1) is yes, the process continues to block 384(2).
[0138] At block 384(2), the process stops because there is no bit error and thus no bit error correction of the output signal Out_2 is required. However, if the result of block 384(1) is no, there is a bit error and bit error correction is required, so the process continues to block 384(3).
[0139] At block 384(3), the location of the bit error is performed. Figures 4C - 4F Examples of how to perform the location are provided. The process advances from block 384(3) to blocks 384(4) and 384(5).
[0140] At block 384(4), the bit error in the W&C array 304A identified by block 384(3) is updated. In some embodiments, the DLC unit 316 is configured to update the word with the bit error by writing an undamaged word to the memory cell 349 identified in block 384(3) as having the bit error, e.g., by copying the data bits representing the undamaged version of the word from a source copy or a permanent copy of the W&C array 304A. For example, the source copy or the permanent copy is stored outside the CIM region that contains the W&C array 304A.
[0141] At block 384(5), the value of the output signal Out_2 is corrected. In some embodiments, the output signal Out_2 is stored in a first register. In some embodiments, the DLC unit 316 is configured to multiply the current multiplicand (i.e., the corresponding column in the input column of XIN1) by the current multiplier (i.e., the corresponding row in the data row of the W&C array 304A) to form a corrected product and write the corrected product to the first register.
[0142] In some embodiments, blocks 384(4) and 384(5) are executed substantially simultaneously. In some embodiments, block 384(4) is executed before block 384(5). In some embodiments, block 384(5) is executed before block 384(4).
[0143] According to another method, which is a counterpart of the CIM system such as the CIM region 324, and the DLC unit 316 forms a part, (1) the weight array is stored in the first region of the first die, and (2) correspondingly on at least the second die, the bit error detector, positioning, and correction (DLC) are completely accomplished by the processor and the associated RAM. According to another method, in order to perform the bit error DLC, a large amount of data is transferred from the weight array on the first die to the processor on the second die (off-die transfer), which results in a large amount of transfer latency, thus greatly reducing the speed of processing the bit error DLC according to another method. In contrast, compared with other methods, at least some embodiments greatly reduce the off-die transfer amount related to bit error detection, positioning, and correction (DLC), thereby achieving a faster bit error DLC than other methods.
[0144] That is, according to at least some embodiments, the row checksum generator 379, the XIN3 array 305D, the column checksum generator 380, the W2 array 304G, the C1 array 306H, the row sum generator 326I, the row vector Row_Sum 304J, the column sum generator 328K, and the column vector Col_Sum 304L are included in the same CIM region as the W&C array 304A, the MX array 337, and the adder tree 308A, such as the CIM region 324. Compared with other methods, this increases the proximity of the arithmetic operator to storage (AOS). According to at least some embodiments, compared with another method, the proximity of (1) the storage location to (2) the arithmetic unit accessing / manipulating the storage location in the same CIM region effectively improves the AOS proximity to greatly reduce the off-die transfer amount of a part including the bit error DLC, thereby achieving a much faster DLC than other methods.
[0145] More specifically, according to at least some embodiments, the proximity of (1) the storage location (represented by the XIN3 array 305D, the W2 array 304G, the C1 array 306H, the row vector Row_Sum 304J, and the column vector Col_Sum 304L, etc.) to (2) the arithmetic unit accessing / manipulating the storage location (correspondingly represented by the row checksum generator 379, the column checksum generator 380, the row sum generator 326I, and the column sum generator 328K) in the same CIM region effectively improves the AOS proximity compared with another method to greatly reduce the off-die transfer amount of a part including the bit error DLC, and thus achieves a faster DLC compared with another method.
[0146] Figure 4A is a block diagram of a simple example of parity bit generation 466 according to some embodiments.
[0147] Parity bit generation 466 is a simplified example of how the XOR gate 260(Q - 1) of the parity encoder 259(Q - 1) etc. generates Figure 2B the parity bits stored in the parity array 221(Q - 1) as the source.
[0148] In Figure 4A the parity array 422 stores parity bits corresponding to the data bits in the weight array 420. The weight values stored in the weight array 420 are represented in decimal, and the equivalent binary representation is shown in parentheses. For example, at the intersection of the second row and the second column in the weight array 420 (position (2, 2)), the weight is represented as 3 in decimal, and the equivalent binary is represented as 11 in parentheses, that is, the position (1, 1) in the parity array 422 corresponds to the position (1, 1) in the weight array 420. Further example, the value stored in the position (1, 1) of the parity array 422 is 0, which means that the XOR operation is applied to the value 3 (1, 1) stored in the position (1, 1) of the weight array 420, that is, 1^1 = 0.
[0149] The parity check value stored in each position of the parity array 422 represents the result of applying the XOR operation to the binary representation shown in the corresponding position of the weight array 420. For example, the position (1, 1) in the parity array 422 corresponds to the position (1, 1) in the weight array 420. Further example, the value stored in the position (1, 1) of the parity array 422 is 0, which means that the XOR operation is applied to the value 3 (1, 1) stored in the position (1, 1) of the weight array 420, that is, 1^1 = 0.
[0150] Figure 4B is a block diagram of a simple example of bit error detection 468 according to some embodiments.
[0151] Bit error detection 468 is a simple example of how the bit error detector 210(Q - 1) etc. detects bit errors in the corresponding rows of the weight array 220(Q - 1).
[0152] Figure 4B includes a weight array 472, which corresponds to Figure 4A the weight array 420, with the difference that it is assumed that the weight array 472 is experiencing a bit error. More specifically, the weight value in the position (1, 1) of the weight array 472 is corrupted and wrongly shown as 2(10) instead of 3(11).
[0153] Figure 4B Also included are: Figure 4A a parity check array 422; an input array XIN1 470; a product array 474; and a bit error detection array 476. The product array 474 represents the result of multiplying the input array XIN1 470 and the weight array 472. The bit error detection array 476 represents the result of applying an XOR operation to the binary representation shown in the corresponding position of the weight array 472 and the parity check bit value in the corresponding position of the parity check array 422.
[0154] A value of 0 in a given position of the bit error detection array 476 indicates that there is no bit error in the corresponding position of the weight array 472. In contrast, a value of 1 in a given position of the bit error detection array 476 indicates that there is indeed a bit error in the corresponding position of the weight array 472.
[0155] Regarding the bit error detection array 476, for example, the position (1,1) in the bit error detection array 476 stores the result of applying an XOR operation to the value in the position (1,1) of the weight array 472 and the value in the position (1,1) of the parity check array 422. More specifically, the value stored in the position (1,1) of the bit error detection array 476 is 1, which indicates that the XOR operation has been applied to the corrupted value 2 of the binary representation (10) stored in the position (1,1) of the weight array 420 and the parity check value 1 in the position (1,1) of the parity check array 422, i.e., 1^0^0 = 1, where the parentheses (zigzag line) symbol (^) is used to indicate the XOR operation applied to the bits in the bit string 100. In contrast, if the position (1,1) in the weight array 420 is not corrupted, i.e., if the position (1,1) stores 3 (11) instead of 2 (1,0), then the position (1,1,1) of the parity check array 422 will show a parity check value of 0, i.e., 1^1^0 = 0, which indicates that there is no bit error.
[0156] Figure 4C is a block diagram of a simple example of checksum generation and array multiplication according to some embodiments.
[0157] In Figure 4C array C2 is the product of the input array XIN3 and the weight array W2, where C2 = XIN3 * W2. As described below, C2 is shown to have a bit error.
[0158] The input array XIN3 is shown as the result of appending the row vector R_ChkSum_1 to the input array XIN2, such that where the symbol is used to represent the append operator, and in this article, the text string format Indicates that B is appended to A. For example, in XIN3, XIN3_position(3,1) = 4, i.e., XIN3_position(1,1) = 1 plus XIN3_position(2,1) = 3. The third row / bottom row of XIN3 represents the row checksum R_ChkSum_1.
[0159] The weight array W2 is shown as the result of appending the column vector C_ChkSum_1 to the weight array W1 such that For example, in W2, W2_position(3,1) = 3, i.e., W2_position(1,1) = 1 plus W2_position(1,2) = 4. The third column / far right column of W2 represents the column checksum C_ChkSum_1.
[0160] Regarding the multiplication C2 = XIN3 * W2, for example, consider C2_position(2,1) = 11 in C2. The word C2_position(2,1) = 11 is the sum of (i) the product of XIN3_position(2,1) = 3 and W2_position(1,1) = 1 plus (ii) the product of XIN3_position(2,2) = 4 and W2_position(1,2) = 2.
[0161] The third row / bottom row of C2 represents the row checksum R_ChkSum_2, which corresponds to the row checksum R_ChkSum_1 in the bottom row of XIN3. The third / far right column of the C1 array 306H(Q - 1) (see Figure 3J ) represents the column vector C_ChkSumm_2 corresponding to Figure 3G the column vector C_ChkSum_1 of 387G.
[0162] To identify, locate, and correct bit errors (see Figures 4D - 4G ), Figure 4C shows that C2_position(1,1) has suffered a late-occurring bit error, i.e., a bit error that occurred after C2 was initially stored in the memory. If there were no bit errors, C2_position(1,1) would be C2_position(1,1) = 5. However, this example assumes that a potential bit error has caused C2_position(1,1) to become C2_position(1,1) = 4.
[0163] Based on the checksum word in position (3,1), a column error in C2_position(1,1) or C2_position(2,1) can be detected; see Figure 4D . The identification of a column error implicitly means that a bit error has occurred in one of the non-checksum words stored in the corresponding column. Similarly, based on the checksum word in position (3,2), a column error in C2_position(1,2) or C2_position(2,2) can be detected.
[0164] Based on the checksum word at position (3,1), a row error in C2_position (1,1) or C2_position (2,1) can be detected; see Figure 4E . Similarly, based on the checksum word at position (3,2), a row error in C2_position (1,2) or C2_position (2,2) can be detected. The identification of a row error means that a bit error has occurred in one of the non-checksum words stored in the corresponding row.
[0165] A bit error in C2 can be located at the intersection of the column identified as having a column error and the row identified as having a row error; see Figure 4F . Once located, the bit error is correctable; see Figure 4G .
[0166] Figure 4D is a block diagram of a simple example of row error location according to some embodiments.
[0167] In Figure 4D (which extends the example of Figure 4C ), a decision block 484(3)D is shown, which partially corresponds to Figure 3M 's block 384(3) etc. At block 484(3)D, it is determined verbatim whether the sum (3,j) matches the position (3,j) in R_ChkSum of C2.
[0168] To generate sum (3,1), a current column-by-column sum of the non-checksum words in the first column of C2 is performed, i.e., the incorrect C2_position (1,1) = 4 and C2_position (2,1) = 11, resulting in sum (3.1) = 15. Here, "current" is used to indicate that the sum is based on the current version of C2. Similarly, to generate sum (3,2), a column-by-column sum of the non-checksum words in the second column of C2 is performed. To generate sum (3,3), a column-by-column sum of the checksum words in the first and second rows of the third column of C2 is performed.
[0169] Sum (3,2) is determined to match (be equal to) the row checksum word at C2 position (3,2), indicating that there is no error column in the second row of C2. Sum (3,3) is determined to match (be equal to) the row checksum word at C2 position (3,3), indicating that there is no error column in the third row of C2. However, sum (3,1) is determined to not match (be not equal to) the row checksum word at C2 position (3,1), indicating that there is an error column in the first row of the C2 array.
[0170] In Figure 4EIn, sum(3, 1) = 15 represents the first mismatch because sum(3, 1) = 15 does not match the row checksum word at C2_Position(3, 1) = 16. The root cause of the first mismatch is that the word at C2_Position(1, 1) represents a delayed bit error. The first mismatch is used to locate and correct the bit error at C2_Position(1, 1) (see Figure 4D sum Figures 4F - 4G ).
[0171] Figure 4E is a block diagram of a simple example of column error localization according to some embodiments.
[0172] In Figure 4E (which extends the example of Figure 4C ), decision block 484(3)E is shown, which partially corresponds to block 384(3) etc. of Figure 3M . At block 484(3)E, it is determined verbatim whether sum(i, 3) matches the position (i, 3) in C_ChkSum of C2.
[0173] To generate sum(1, 3), the current row sum is performed on the non-checksum words in the first row of C2, i.e., the incorrect C2_Position(1, 1) = 4 and C2_Position(1, 2) = 10, resulting in sum(1, 3) = 14. Similarly, the row-by-row sum is performed on the non-checksum words in the second row of C2 to obtain sum(2, 3). The row-by-row sum is performed on the checksum words in the first and second columns of the third row of C2 to obtain sum(3, 3).
[0174] At block 484(3), sum(2, 3) is determined to match the column checksum word at C2_Position(2, 3), indicating that there is no row error in column 2 of C2. sum(3, 3) is determined to match the column checksum word at C2_Position(3, 3), indicating that there is no row error in column 3 of C2. However, sum(1, 3) is determined not to match the row checksum word at C2_Position(1, 3), indicating that there is a column error in column 1 of C2.
[0175] In Figure 4E , sum(1, 3) represents the second mismatch because sum(1, 3) = 14 does not match the column checksum word at C2_Position(1, 3) = 15. The root cause of the second mismatch is that the word at C2_Position(1, 1) represents a delayed bit error. The second mismatch is used to locate and correct the bit error at C2_Position(1, 1) (see Figures 4E - 4G ).
[0176] Figure 4F is a block diagram of a simple example of bit error localization according to some embodiments.
[0177] In Figure 4F (which extends Figures 4D - 4EIn the example of [], the position of block 484(4) is shown, which corresponds to Figure 3M Block 384(4) etc. of []. At block 484(4), the intersection point can be determined by the following: the row recognized as having a row error in the context of Figure 4D , that is, row 1 of C2; and the column identified as having a column error in the context of Figure 4E , that is, column 1 of C2. The intersection of row 1 and column 1 of C2 is C2_position(1,1). Therefore, the intersection of row 1 and column 1 of C2 locates the bit error at C2_position(1,1).
[0178] Figure 4G is a block diagram of a simple example of bit error correction according to some embodiments.
[0179] In Figure 4G (which extends the example of Figures 4D - 4F ), the correction of block 484(5) is shown, which corresponds to Figure 3M Block 384(5) etc. of []. At block 484(5), for the non-checksum position (i, j) in C2 determined at block 484(4) of Figure 4D (), the difference Δ is determined between sum(3,1) and C2 position(3,1), such that Δ = sum(3,1) - C2 position(3,1). Alternatively, the difference Δ is determined between sum(1,3) and C2 position(1,3), such that Δ = sum(1,3) - C2 position(1,3). Next, the corrected value C'[1][1] to be stored / overwritten at C2_position(1,1) is determined by adding Δ to the error value C[1][1] at C2_position(1,1), so that C'[1][1] = C[1][1] + Δ.
[0180] Figure 5A is a flowchart 500 of a method for manufacturing a memory device according to some embodiments.
[0181] According to some embodiments, the method of flowchart (flowchart) 500 can be implemented, for example, using an EDA system 600 (see Figure 6 , discussed below) and an IC manufacturing system 700 (see Figure 7 , discussed below). Examples of devices that can be manufactured according to the method of flowchart 500 include devices based on the illustrations disclosed herein.
[0182] In Figure 5A , the method of flowchart 500 includes blocks 502 - 504. At block 502, a layout diagram is generated, which includes one or more layout diagrams corresponding to the CIM system and / or die area etc. disclosed herein. According to some embodiments, for example, an EDA system 600 can be used (see Figure 6) to implement block 502. The process advances from block 502 to block 504.
[0183] At block 504, based on the layout diagram, at least one of the following is performed: (A) performing one or more photolithographic exposures, or (B) fabricating one or more photolithographic masks, or (C) fabricating one or more components in a device (e.g., semiconductor device) layer. See the discussion of the IC manufacturing system 700 below Figure 7 in
[0184] Figure 5B is a flowchart 508 of a method for operating a CIM system according to some embodiments.
[0185] Examples of CIM systems that can be operated according to the method of flowchart 508 include Figure 1A the CIM system 100A of Figure 1C the CIM system 100C of
[0186] At block 520, for the weight bits in the weight array, the corresponding parity bits are encoded by a parity encoder and stored in the parity array. An example of the weight array is Figure 2B the weight array 220(Q - 1) of Figure 2B An example of the parity array is Figure 2B the parity array 221(Q - 1) of
[0187] At block 512, the parity encoder encodes the parity bits by performing an XOR operation on the weight bits through an XOR gate. An example of the XOR gate is the XOR gate 260(Q - 1) etc. From block 512, the process continues to block 514.
[0188] In some embodiments, the execution of block 514 is performed simultaneously in time with block 510. In some embodiments, the execution of block 514 is not performed close to simultaneously with block 510.
[0189] At block 514, the multiplier receives a segment of a weight row from the weight and parity arrays and an input column from the input array, and the segment of the row and the column represent the multiplicand and the multiplier respectively. Examples of the weight and parity arrays are Figure 2A the slice 218(Q - 1) of the W&P array 204A of Figure 2Ainput arrays XIN1, etc. Examples of multipliers include multipliers 251(0)-251(Q-1) corresponding to EM blocks 238(0)-238(Q-1) of the EM array 236 corresponding to Figure 2A etc. The process proceeds from block 514 to each of blocks 516 and 518.
[0190] At block 516, the multiplicand and the multiplier are multiplied together by the multiplier to form a product. Examples of products include Q products PRD1(0)-PRD1(Q-1) generated by the EM array 236, etc. The process proceeds from block 516 to block 522 (discussed after discussing blocks 518-520).
[0191] At block 518, a bit error in the multiplicand is detected by a bit error detector, and a first flag is generated to indicate the result of the bit error detection. Examples of bit error detectors include bit error detectors 210(0)-210(Q-1) corresponding to Figure 2A EM blocks 238(0)-238(Q-1), etc., where each bit error detector generates a corresponding instance of the first error flag. Examples of the first error flag include Q instances of flag FLG1 generated correspondingly by bit error detectors 212(0)-212(Q-1), etc. Within block 518, the process proceeds to block 520.
[0192] At 520, an XOR gate performs an XOR operation on the multiplicand to generate the first error flag. Examples of XOR gates include XOR gates 253(0)-253(Q-1) corresponding to Figure 2A EM blocks 238(0)-238(Q-1), etc. An example of the XOR operation result is the state of flag FLG1 generated by XOR gate 253(Q-1), i.e., whether flag FLG1 is asserted (FLG1 = 1) to indicate a bit error, or whether it is not asserted (FLG1 = 0) to indicate no bit error, etc. From block 520, the process leaves block 518 and proceeds to block 522.
[0193] Block 518 is executed in parallel with block 516. According to another method corresponding to block 518, i.e., before performing the multiplication, bit error detection of the bits in the memory unit corresponding to the weight array 220(Q-1) is performed. Performing bit error detection before multiplication according to another method uses two operation cycles. In contrast, according to at least some embodiments, executing blocks 516 and 518 in parallel uses one operation cycle, which is one operation cycle faster than the other method.
[0194] In step 522, it is determined whether a position error is detected. An example of determining whether a position error is detected is to determine whether the XOR gate 253(Q - 1) etc. has asserted the flag FLG1, that is, whether the flag FLG1 has been set to 1. According to the determination at block 522, the process proceeds to block 524 or block 528.
[0195] If the determination at block 522 is no, that is, if no position error is detected, the process continues to block 524. At block 524, the product PRD1 is selected by the selector instead of the reference value. An example of the reference value is Figure 2A the reference REF etc. An example of the selector is Figure 2A the MUX 251(Q - 1) etc., where the MUX 251(Q - 1) is configured to receive the product PRD1(Q - 1) and the reference REF as inputs, and the flag FLG1 as the control signal, and is also configured to select the product PRD(Q - 1) when FLG = 0. From block 524, the flow proceeds to block 526, where the process stops.
[0196] If the determination at block 522 is yes, that is, if a position error is detected, the process continues to block 528. At block 528, the reference value is selected by the selector instead of the product PRD1. Extending the example of block 524, the MUX 251(Q - 1) is also configured to select the reference REF when FLG = 1. The process proceeds from block 528 to block 530.
[0197] At block 530, the trajectory inference signal is generated by the trajectory inference data generator. An example of the trajectory inference data generator is Figure 2A the trajectory inference data generator 212 etc., where the generator 212 is configured to generate an error pointer and a second error flag. An example of the error pointer is Figure 2C the pointer PRD1(Q - 1) etc. An example of the second error flag is the flag FLG2 etc. Inside block 530, the process proceeds to blocks 532 and 534.
[0198] At block 532, the Q:P encoding of Q instances of the first flag is performed by the encoder, thereby generating an error pointer. An example of Q instances of the first flag is Q instances of the flag FLG1 generated by the bit error detectors 212(0) - 212(Q - 1) of the EM blocks 238(0) - 238(Q - 1). An example of the encoder is Figure 2C the Q:P encoder 262 etc., where the encoder 262 generates an example of the error pointer, that is, the point EPT. The process proceeds from block 532 to block 538 (discussed after blocks 534 - 536).
[0199] At block 534, the chip error detector detects a chip error and generates a second error flag based on Q instances of the first flag. An example of the chip error detector isFigure 2C such as the die error detector 213, which generates an example of the second error flag, i.e., flag FLG2. Within block 534, the process proceeds to block 536.
[0200] At 536, an OR gate performs an OR operation on Q instances of the first flag to generate the second error flag. An example of the OR gate is the OR gate 263 of the trajectory inference data generator 212, etc., which generates an example of the second flag, i.e., flag FLG2. An example of the OR operation result is the state of flag FLG2 generated by the OR gate 263, i.e., whether flag FLG2 is asserted (FLG2 = 1) to indicate the presence of a die error, or whether it is not asserted (FLG2 = 0) to indicate the absence of a die error, etc. From block 536, the process leaves block 534 and proceeds to block 538.
[0201] At block 538, in response to the second flag indicating the presence of a die error, e.g., in response to flag FLG2 being asserted (FLG1 = 1) to indicate the presence of a die error, the die error is located by the bit error corrector. An example of the bit error corrector is Figure 2D the bit error corrector 216, etc. An example of bit error location is Figure 2D block 265(3) of the flowchart 290, etc. The process proceeds from block 538 to block 540.
[0202] At block 540, the bit error is corrected by the bit error corrector. An example of bit error correction is Figure 2D one or more of blocks 265(4)-265(5) of the flowchart 290, etc. From block 540, the process proceeds to block 526, where the process stops.
[0203] Figure 5C is a flowchart 543 of a method for manufacturing a CIM system according to some embodiments.
[0204] Flowchart 543 is Figure 5A an example of block 504. According to some embodiments, the method of flowchart 543 can be implemented, for example, using the IC manufacturing system 700 (see Figure 7 discussed below). Examples of digital CIM systems that can be manufactured according to the method of flowchart 543 include those based on the CIM systems disclosed herein, etc.
[0205] Flowchart 543 includes blocks 545 - 547. At block 545, in a first region of a first semiconductor die, a first structure including a first component is formed. The first component includes a memory cell configured to store a single bit accordingly, a multiplier, and a first bit error detector. Additionally, a first memory cell among the memory cells is arranged in a corresponding first array and is configured to store a first data bit. A second memory cell among the memory cells is arranged in a corresponding second array and is configured to store a parity bit corresponding to the first data bit. The first component is organized into a first group, and each first group includes a corresponding one of the first array, the second array, the multiplier, and the first bit error detector.
[0206] Regarding block 545, examples of the first structure include the structure of semiconductor devices (such as transistors), structures that facilitate coupling to transistors, etc. In some embodiments, the structure including transistors and the structure that facilitates coupling to transistors are formed in one or more first layers collectively referred to as the transistor layer. Examples of transistors include field - effect transistors (FETs), such as positive - channel metal - oxide - semiconductor (PMOS) FETs (PFETs), negative - channel metal - oxide - semiconductor (NMOS) FETs (NFETs), etc.
[0207] Examples of the structure including transistors include: an active region in a semiconductor layer; a well region surrounding a selected active region; source / drain (S / D) regions in the active region; a channel region in the active region between corresponding pairs of S / D regions; a gate structure above the corresponding active region and (optionally) a buried gate (BG) structure below the corresponding active region; or the like.
[0208] Examples of the structure that facilitates coupling to transistors include: metal - to - source / drain (MD) contacts located above the S / D regions and coupled to the S / D regions, and (optionally) corresponding buried MD (BMD) contacts located below the S / D regions and coupled to the S / D regions; metal - to - gate (MG) contacts coupled to the gate structure and (optionally) corresponding buried MG (BMG) contacts coupled to the BG structure; via - to - MD (VD) contacts coupled to the MD contacts and corresponding buried VD (BVD) contacts coupled to the BMD contacts; via - to - MG (VG) contacts coupled to the MG contacts and corresponding buried VG (BVG) contacts coupled to the BMG contacts; local interconnect (LI) structures that couple, for example, MD contacts and / or gate structures together and (optionally) buried LI (BLI) structures that couple, for example, BMD contacts and / or BG gate structures together; or the like.
[0209] Regarding block 545, examples of the first memory cell include Figure 2A memory cell 245, etc. Examples of the second memory cell include Figure 2Amemory cells 246, etc. An example of the first array is Figure 2A weight array 220(Q-1), etc. An example of the second array is Figure 2A parity check array 221(Q-1), etc. An example of the multiplier is Figure 2A multiplier 251(Q-1), etc. An example of the first bit error detector is [[ID=2 bit error detector 212(Q-1), etc. An example of the first group includes chip 218(Q-1) and EM block 238(Q-1), etc. The process advances from block 545 to block 547.
[0210] At block 547, mutual coupling is formed between the first components, resulting in at least: for each first group, the multiplier is configured to perform multiplication of the input data bits by the corresponding bits in the first data bits, and for each first group, the first bit error detector is configured to perform bit error detection of the corresponding first data bits based on an associated one of the corresponding parity bits.
[0211] Regarding block 547, examples of forming mutual coupling include forming signal segments and / or PG segments in metallization layers located respectively above and (optionally) below the transistor layer. In some embodiments, the signal segments are conductive and are configured to carry signals including input / output (I / O) signals, control signals, etc. In such embodiments, the signal segments are correspondingly coupled to VD contacts, MG contacts, (optionally) BVD contacts, (optionally) BVG contacts, etc. In some embodiments, the PG segments are conductive and are configured to be powered by the corresponding reference voltages of the power grid (PG). In such embodiments, the PG segments are correspondingly coupled to VD contacts, MG contacts, (optionally) BVD contacts, (optionally) BVG contacts, etc. For example, the first PG segment among these PG segments is configured to be powered by a first reference voltage (e.g., VDD), while the second PG segment among these PG segments is configured to be powered by a second reference voltage (e.g., VSS).
[0212] In some embodiments, regarding block 547, forming mutual coupling between the first components also results in at least for each first group (e.g., 244(1)), the first bit error detector (e.g., (210(Q-1)) being configured to perform detection of bit errors in the corresponding first data bits based on the corresponding first data bits and parity bits.
[0213] In some embodiments, regarding block 547, forming mutual coupling (e.g., block 547) between the first components also results in at least for each first group (e.g., 244(1)), the first bit error detector (e.g., 210(Q-1)) also being configured to perform detection in parallel with the multiplication performed by the multiplier (e.g., 251(N-1)).
[0214] In some embodiments, with respect to block 547, forming a mutual coupling between first components (e.g., block 547) further causes, at least for each first group (e.g., 244(1)), the CIM system to be configured to perform localization of bit errors after performing detection.
[0215] In some embodiments, forming a mutual coupling between first components (e.g., block 547) also causes, at least for each first group (e.g., 244(1)), the CIM system to be configured to perform error bit correction after performing localization.
[0216] In some embodiments, flowchart 543 further includes a first block, where in a first region (e.g., 103) of a first semiconductor die (e.g., (102A / C(1))), a second structure including a second component is formed, and the second component includes a multiplexer (e.g., 254A). In such an embodiment, each first group (e.g., 244(1)) further includes a corresponding one of the multiplexers (e.g., 254A). In such an embodiment, flowchart 543 further includes a second block, where forming a mutual coupling at least between the second component or the first components causes, at least for each first group (e.g., 244(1)), the multiplexer (e.g., 254A) to be configured to select, for example, (i) a product generated by a multiplier (e.g., 251(N - 1)) or, for example, (ii) a predefined value based on an output signal (e.g., 210(Q - 1)) generated by a first bit error detector.
[0217] In some embodiments, flowchart 543 further includes a first block, where in a first region (e.g., 103) or a second region (e.g., 155(1)) of a first semiconductor die (e.g., (102A / C(1))), or in a first region (e.g., 155(2)) of a second semiconductor die (e.g., 102C(2)), a second structure including a second component is formed, and the second component includes a parity encoder (e.g., 158). In such an embodiment, each first group (e.g., 244(1)) further includes a corresponding one of the parity encoders (e.g., 158). In such an embodiment, flowchart 543 further includes a second block, where forming a mutual coupling at least between the second component or the first components causes, at least for each first group (e.g., 244(1)), the parity encoder (e.g., 158) to be configured to encode a corresponding bit in the parity bits based on a corresponding first data bit.
[0218] In some embodiments, flow chart 543 further includes a first block, wherein in a first region (e.g., 103) of a first semiconductor die (e.g., (102A / C(1))), a third structure including a third component is formed, the third component including an exclusive - or (e.g., XOR) gate (e.g., 260(x)), and the XOR gate (e.g., 260(x)) is included as a corresponding part of a parity encoder (e.g., 158). In such an embodiment, each first group (e.g., 244(1)) further includes a corresponding one of the XOR gates (e.g., 260(x)). In such an embodiment, flow chart 543 further includes a second block, wherein a mutual coupling is formed at least among the third component, the first component, or the second component, resulting in that for each first group (e.g., 244(1)), the XOR gate (e.g., 253(x)) is configured to operate row - by - row, including receiving a row of first data bits as an input, correspondingly generating a parity bit, and storing the parity bit in a corresponding row of a second array (e.g., 221(N - 1)).
[0219] In some embodiments, flow chart 543 further includes a first block, wherein in a first region (e.g., 103) of a first semiconductor die (e.g., (102A / C(1))), a second structure including a second component is formed, the second component including an exclusive - or (e.g., XOR) gate (e.g., 253(x)), and the XOR gate (e.g., 253(x)) is included as a corresponding part of a first - bit error detector (e.g., 210(Q - 1)). In such an embodiment, each first group (e.g., 244(1)) further includes a corresponding one of the XOR gates (e.g., 253(x)). In such an embodiment, flow chart 543 further includes a second block, wherein a mutual coupling is formed at least between the second component and the first component, resulting in that for each first group (e.g., 244(1)), the XOR gate (e.g., 253(x)) is configured to receive a first data bit and a parity bit as inputs and generate an output signal representing a first flag signal (e.g., FLG1) based thereon, and the first flag signal is assertable to indicate the presence of a bit error.
[0220] In some embodiments, flow chart 543 further includes a first block, wherein in a first region (e.g., 103) or a second region (e.g., 155(1)) of a first semiconductor die (e.g., (102A / C(1))), or in a first region (e.g., 155(2)) of a second semiconductor die (e.g., 102C(2)), a second structure including a second component is formed, the second component including a trace inference data generator (e.g., 112). In such an embodiment, flow chart 543 further includes a second block, wherein a mutual coupling is formed at least between the second component and the first component, resulting in that at least the trace inference data generator (e.g., 112) is configured to generate one or more error - code trace inference signals (e.g., EPT and FLG2) in
[0221] In some embodiments, flow chart 543 further includes a first block, wherein in a first region (e.g., 103) or a second region (e.g., 155(1)) of a first semiconductor die (e.g., (102A / C(1))), a third structure including a third component is formed, the third component including a Q:P encoder (e.g., 262) and a second bit error detector (e.g., 213), and the Q:P encoder (e.g., 262) and the second bit error detector (e.g., 213) are included in a trajectory inference signal generator (e.g., 112). In such an embodiment, flow chart 543 further includes a second block, wherein couplings are formed at least between the third component, the first component, or the second component, resulting in at least: for each first group (e.g., 244(1)), the first bit error detector (e.g., 210(Q-1)) is further configured to generate an output signal representing a first flag signal (e.g., FLG1), the first flag signal being assertable to indicate the presence of a bit error; for each first group (e.g., 244(1)), the first array (e.g., 220(N-1)) is arranged in rows and Q columns, where Q is a positive integer; for each first group (e.g., 244(1)), there are Q first groups (e.g., 244(2)) and corresponding Q instances of the first flag signal (e.g., FLG1); for each first group (e.g., 244(1)), the trajectory inference data generator (e.g., 212) is configured to receive Q instances of the first flag signal (e.g., FLG1); for each first group (e.g., 244(1)), the Q:P encoder (e.g., 262) is configured to encode the Q instances of the first flag signal (e.g., FLG1) into a P-bit signal representing an error pointer (e.g., EPT), the error pointer (e.g., EPT) being the first bit error trajectory inference signal in one or more bit error trajectory inference signals, P is a positive integer, and for each first group (e.g., 244(1)), the second bit error detector (e.g., 213) generates a second flag signal (e.g., FLG2) based on the Q instances of the first flag signal (e.g., FLG1), the second flag signal (e.g., FLG2) being assertable to indicate that the error pointer (e.g., EPT) points to a bit error, and the second flag signal (e.g., FLG2) being the second bit error trajectory inference signal in one or more bit error trajectory inference signals.
[0222] In some embodiments, flow chart 543 further includes a first block, wherein in a first region (e.g., 103) of a first semiconductor die (e.g., (102A / C(1))), a fourth structure including a fourth component is formed, the fourth component including an OR gate (e.g., 263), and the OR gate (e.g., 263) is correspondingly included in the second bit error detector (e.g., 213).
[0223] In such an embodiment, the flow chart 543 further includes a second block, wherein a mutual coupling is formed between at least the fourth component, the first component, the second component, or the third component, such that for at least each first group (e.g., 244(1)), the OR gate (e.g., 263) is configured to receive Q instances of the first flag signal (e.g., FLG1) and generate an output signal representing the second flag signal (example FLG2) based thereon.
[0224] is a flow chart 550 of a method of operating a CIM system according to some embodiments.
[0225] Examples of CIM systems that can be operated according to the method of flow chart 550 include CIM system 100D of CIM system 100E of etc. Flow chart 550 includes blocks 552 - 564.
[0226] At block 552, for the input words of the input array, the row checksum generator generates the corresponding checksum and appends the latter to the input array to form an input and checksum (I&C) array. An example of the row checksum generator is row checksum generator 379 of The example of the input array is represented by XIN2 array 305B of The example of the checksum is
[0227] At block 554, for the weight words in the weight array, the column checksum generator generates the corresponding checksum and appends the latter to the weight array to form a weight and checksum (W&C) array. An example of the column checksum generator is column checksum generator 380 of The example of the weight array is represented by weight arrays 320(0) - 320(Q - 1) of The example of the checksum is represented by checksum arrays 323(0) - 323(Q - 1) of, and the checksum C_chk∑(t = x) includes
[0228] At block 556, a row segment of weights and the associated checksum (as the multiplicand) are received from the W&C array, and a column sum of inputs and the associated checksum (as the multiplier) are received from the I&C array by the multiplier. An example of the multiplier is The combination of multipliers 364(0)-364(Q-1), etc. The process advances from block 556 to block 558.
[0229] At block 558, iteratively, the multiplicand is multiplied by the multiplier to form a product row of the product array. Accordingly, an example of the product array and the product rows therein is the W1 array 304E and row (i) therein, etc. The process advances from block 558 to block 560.
[0230] At block 560, a trajectory inference signal is generated by the trajectory inference signal generator.
[0231] Examples of the trajectory inference signal include the row and vector Row_Sum 304J of the column and vector Col_Sum 304L of the row sum generator 326I of the column sum generator 328K of
[0232] At block 562, a corresponding one in the trajectory inference signal generator performs column addition row by row to form a row sum. An example of the row sum is the row and vector Row_Sum 304J of A corresponding one in the trajectory inference signal generator is an example of
[0233] At block 564, a corresponding one in the trajectory inference signal generator performs row addition row by row to form a corresponding word of the column sum. An example of the column sum is the column and vector Col_Sum 304L of A corresponding one in the trajectory inference signal generator is an example of Figure 3M the column sum generator 328K of
[0234] In some embodiments, block 564 is executed before block 562. In some embodiments, blocks 562 and 564 are executed substantially simultaneously.
[0235] Figure 5E is a flowchart 573 of a method for manufacturing a CIM system according to some embodiments.
[0236] Flowchart 573 is Figure 5A an example of block 504 of Figure 7) implemented. Examples of digital CIM systems that can be fabricated according to the method of flowchart 573 include Figure 1D CIM system 100D of Figure 1E CIM system 100E, etc.
[0237] Flowchart 573 includes blocks 575 - 577.
[0238] At block 575, in a first region of a first semiconductor die, a first structure including a first component is formed. The first component includes a memory cell configured to store a single bit, a multiplier, and a locus inference data (LID) generator, respectively. In addition, the first memory cell in the memory cells is arranged in a corresponding first array and is configured to store a first word. The first component is organized into a first group, and each first group includes a corresponding one of the first array, the multiplier, and the LID generator. Regarding block 575, examples of the first structure include the instances of the first structure discussed in the context of Figure 5C block 545 of
[0239] Examples of the first memory cell include Figure 3A memory cell 349 of Figure 3A Examples of the first array are Figure 2A product array 306A of Figure 3A Examples of the multiplier are
[0240] multiplier 251(Q - 1) of Figure 5C Examples of the LID generator are LID generator 326A, etc. An example of the first group is a group including
[0241] Figure 3A sheet 318(Q - 1), multiplier 364(Q - 1), and LID generator 326A, etc. The process advances from block 575 to block 577.
[0240] At block 577, an inter - coupling is formed between the first components, resulting in at least: for each first group, the multiplier is configured to perform one or more multiplications of (i) an input word and an associated first check - sum word and (ii) a corresponding weight word and an associated second check - sum word, and for each first group, the LID generator is configured to generate one or more LID signals based on a selected first word. Regarding block 577, examples of forming the inter - coupling include the examples of forming the inter - coupling discussed in the context of Figure 5C block 547 of
[0241] In some embodiments, flow chart 573 further includes a first block, wherein a second structure including a second component is formed in a first region (e.g., 124) or a second region (e.g., 155(3)) of a first semiconductor die (e.g., 102E(2)), or in a first region (e.g., 155(4)) of a second semiconductor die (e.g., 102E(2)), and the second component includes a row checksum generator (e.g., 379). In such an embodiment, flow chart 573 further includes a second block, wherein a mutual coupling is formed at least between the second component and one or more first components, resulting in at least the row checksum generator (e.g., 379) being configured to generate a first checksum word (e.g., the word in R_ChkSum_1) based on an input word (e.g., XIN2 305B).
[0242] In some embodiments, flow chart 573 further includes a first block, wherein a third structure including a third component is formed in the same region where the row checksum generator (e.g., 379) is located, and the third component includes a recursive adder (e.g., 379), and the recursive adder (e.g., 379) is included in the row checksum generator (e.g., 379). In such an embodiment, a second memory cell (e.g., 350) in the memory unit is arranged in a second array (e.g., XIN2 305B) and is configured to store an input word (e.g., XIN2 305B). In such an embodiment, flow chart 573 further includes a second block, wherein a mutual coupling is formed at least between the first component and the second component, further resulting in at least the second array (e.g., XIN2 305B) being arranged in a first row and one or more first columns, and the input word (e.g., XIN2 305B) is correspondingly located at the intersection of the first row and the first column (e.g., A[x][y]). In such an embodiment: the second array (e.g., 305B) is a first part of a third array (e.g., XIN3 305D); a third memory cell (e.g., 348) is arranged as a second part (e.g., 386D) of the third array (e.g., XIN3 305D), and is correspondingly configured to store a second word (the first checksum word) representing a first checksum (e.g., the word in R_ChkSum_1); the second part (e.g., 386D) of the third array (e.g., XIN3 305D) is arranged in a second column and a second row, and the first checksum word (e.g., R_ChkSum_1) is correspondingly located at the intersection of the second row and the second column. In such an embodiment, flow chart 573 further includes a second block, wherein a mutual coupling is formed at least among the third component, the first component, or the second component, resulting in at least each recursive adder (e.g., 379) being configured to generate a corresponding one of the first checksum words (e.g., the word in R_ChkSum_1) by recursively adding the input word (e.g., A[x][i]) column by column to a corresponding column in the first column.
[0243] In some embodiments, flowchart 573 further includes a first block, wherein in a first region (e.g., 124) or a second region (e.g., 155(3)) of a first semiconductor die (e.g., 102E(1)), or in a first region (e.g., 155(4)) of a second semiconductor die (e.g., 102E(2)), a second structure including a second component is formed, and the second component includes column checksum generators (e.g., 180 and 380). In such an embodiment, each first group further includes a respective one of the column checksum generators (e.g., 380). In such an embodiment, flowchart 573 further includes a second block, wherein a mutual coupling is formed between at least the second component and the first component, thereby generating at least one or each first group (e.g., 344 + 326A (e.g., Q - 1)), and the column checksum generator (e.g., 380) is configured to generate a second checksum word (e.g., C_ChkSum_1) based on a weight word (e.g., W1 304E).
[0244] In some embodiments, flowchart 573 further includes a first block, wherein in the same region where the column checksum generators (e.g., 180, 380) are located, a third structure including a third component is formed, and the third component includes an adder tree (e.g., 308F), and the adder tree (e.g., 308F) is included as a corresponding part of the column checksum generators (e.g., 180, 380). In such an embodiment: a second memory cell (e.g., 349) is arranged in a second array (e.g., W1 304E) and is configured to store weight words; the second array (e.g., W1 304E) is arranged in a first row and a first column, and the weight word (e.g., W1 304E) is correspondingly located at the intersection of the first row and the first column, and correspondingly represents a first word (e.g., B[x][y]); the second array (e.g., W1304E) is a first part of a third array (e.g., 304F); a third memory cell (e.g., 350) is arranged as a second part (e.g., 387) of the third array (e.g., 304F), and is correspondingly configured to store a second checksum word (e.g., C_ChkSum_1); the second part (e.g., 387) of the third array (e.g., 304F) is arranged in a second row and a second column, and the second checksum word (e.g., C_ChkSum_1) is correspondingly located at the intersection of the second column and the second row.
[0245] In such an embodiment, flowchart 573 further includes a second block, wherein a mutual coupling is formed between at least the third component, the second component, and the first component, resulting in at least for each first group (e.g., 344 + 326A (e.g., Q - 1)), and the adder tree (e.g., 308F) is configured to generate a second checksum word by adding a weight word (e.g., B[i][y]) to the first row of the second array (e.g., W1 304E).
[0246] In some embodiments, flowchart 573 further includes a first block, wherein in a region that is the same as the region where the trajectory inference data generator (e.g., 326I, 328K) is located, a second structure including a second component is formed, the second component includes a row and generator (e.g., 326I), and the row and generator (e.g., 326I) is included as a corresponding part of the trajectory inference data generator (e.g., 326I, 328K). In such an embodiment: a first array (e.g., Figure 3H C1 306A therein) is arranged in the first row and the first column, and the first word is correspondingly located at the intersection of the first row and the first column; the first word (e.g., Figure 3H C1 therein) represents a product word. In such an embodiment, flowchart 573 further includes a second block, wherein a mutual coupling is formed at least between the second component and the first component, resulting in that for at least each first group (e.g., 344 + 326A (e.g., Q - 1)), the row and generator (e.g., 326I) is configured to generate a row sum (e.g., Figure 3J Row_Sum therein) based on the product word.
[0247] In some embodiments, the row sum (e.g., Row_Sum) is a row vector composed of second words. In such an embodiment, flowchart 573 further includes a first block, wherein a third structure including a third component is formed, the third component includes a recursive adder (e.g., 379), and the recursive adder (e.g., 379) is included as a corresponding part of the row and generator (e.g., 326I). In such an embodiment, flowchart 573 further includes a second block, wherein a mutual coupling is formed at least among the third component, the second component, and the first component, resulting in that for each first group in the first group (e.g., 344 + 326A (Q - 1)), the recursive adder (e.g., 379) is configured to generate a corresponding second word in the row sum (e.g., Row_Sum) by recursively adding the first words in the corresponding column in the first column (e.g., Figure 3H C1 therein) column by column.
[0248] In some embodiments, the first word in a selected row (e.g., 388) in the first row represents a third checksum (e.g., Figure 3J R_ChkSum_2 therein). In such an embodiment, flowchart 573 further includes a first block, wherein in a second region (e.g., 114D) of the first semiconductor die (e.g., 102D) or a first region (e.g., 114E) of the second semiconductor die (e.g., 102E(2)), a third structure including a third component is formed, the third component includes a processor (e.g., Figure 3M114D / E, 316) in. In such an embodiment, the flowchart 573 further includes a second block, wherein a mutual coupling is formed between at least a third component, a second component, or a first component, such that at least a processor (e.g., 316) is configured to compare a row sum (e.g., 384(1)) with a corresponding checksum in a third checksum (e.g., Figure 3J R ChkSum 2) of to identify that the first row has a bit error.
[0249] In some embodiments, the flowchart 573 further includes a first block, wherein in a region same as the region where a trajectory inference data generator (e.g., 326I, 328K) is located, a second structure including a second component is formed, the second component includes a column sum generator (e.g., 328K), and the column sum generator (e.g., 328K) is included as a corresponding part of the trajectory inference data generator. In these embodiments, a first array (e.g., C1 array 306A) is arranged in a first row and a first column, and a first word is correspondingly located at the intersection of the first row and the first column; the first word represents a product word. In such an embodiment, the flowchart 573 further includes a second block, wherein a mutual coupling is formed between at least a second component or a first component, such that at least for each first group (e.g., 344 + 326A(Q - 1)), a column sum (e.g., Figure 3L Col_Sum in) is generated based on the product word (e.g., 328K).
[0250] In some embodiments, the column sum (e.g., Col_Sum) is a column vector composed of second words. In such an embodiment, the flowchart 573 further includes a first block, wherein in a region same as the region where a column sum generator (e.g., 328K(Q - 1)) is located, a third structure including a third component is formed, the third component includes an adder tree (e.g., 308K), and the adder tree (e.g., 308K) is included as a corresponding part of the column sum generator (e.g., 328K). In such an embodiment, the flowchart 573 further includes a second block, wherein a mutual coupling is formed between at least a third component, a second component, or a first component, such that for each first group in a first group (e.g., 344 + 326A (e.g., Q - 1)), the adder tree (e.g., 308K) is configured to generate a second word of the column sum (e.g., Col_Sum) row by row by adding the first words in the corresponding rows of the first row.
[0251] In some embodiments, a first word in a selected column in the first column (e.g., Figure 3L 389 in) represents a fourth checksum (e.g., Figure 3L C_ChkSum_2 of).
[0252] In such an embodiment, the flow chart 573 further includes a first block, wherein in a second region (e.g., 114D) of a first semiconductor die (e.g., 102D) or a first region (e.g., 114E) of a second semiconductor die (e.g., 102E(2)), a third structure including a third component is formed, and the third component includes a processor (e.g., Figure 3M 114D / E, 316M in
[0253] In such an embodiment, the flow chart 573 further includes a second block, wherein a mutual coupling is formed at least among the third component, the second component, or the first component, resulting in at least the processor (e.g., 316) being configured to compare the corresponding values in the column sum (e.g., 384(1)) and the fourth checksum (e.g., Figure 3L Col_Sum in
[0254] to identify that the first column has a bit error. Figure 3I - Figure 3J In some embodiments, the formation of the mutual coupling among the first components further results in at least for each first group (e.g., 344 + 326A(Q - 1)), the trace inference data generator (e.g., 326A) is further configured to perform one or more bit error trace inference signals (e.g., Figure 3L - Figure 3M Row_Sum of
[0255] Figure 6 is a functional block diagram of an electronic design automation (EDA) system 600 according to some embodiments.
[0256] In some embodiments, the EDA system 600 includes an automatic placement and routing (APR) system. In some embodiments, the EDA system 600 is a general computing device including a hardware processor 602 and a non-transitory computer-readable storage medium 604. Among other things, the storage medium 604 is further encoded with computer program code 606, i.e., a set of executable instructions. The execution of the instructions (i.e., the computer program code) 606 by the hardware processor 602 represents (at least in part) an EDA tool that implements, according to one or more embodiments (hereinafter referred to as the process and / or method), for example, a method of generating a layout diagram as disclosed herein, a method of generating a layout diagram (such as the layout diagram disclosed herein), or a method of generating a layout diagram corresponding to the device disclosed herein, etc., as part or all of the method.
[0257] The storage medium 604 stores a layout diagram 611, such as the layout diagram disclosed herein, etc.
[0258] Processor 602 is electrically coupled to computer-readable storage medium 604 via bus 608. Processor 602 is also electrically coupled to I / O interface 610 via bus 608. Network interface 612 is also electrically connected to processor 602 via bus 608. Network interface 612 is connected to network 614 such that processor 602 and computer-readable storage medium 604 can be connected to external components via network 614. Processor 602 is configured to execute computer program code 606 encoded in computer-readable storage medium 604 to make system 600 available for performing part or all of the processes and / or methods. In one or more embodiments, processor 602 is a central processing unit (CPU), a multiprocessor, a distributed processing system, an application specific integrated circuit (ASIC), and / or a suitable processing unit.
[0259] In one or more embodiments, computer-readable storage medium 604 is an electronic, magnetic, optical, electromagnetic, infrared, and / or semiconductor system (or apparatus or device). For example, computer-readable storage medium 604 includes semiconductor or solid state memory, magnetic tape, removable computer disk, random access memory (RAM), read-only memory (ROM), rigid disk, and / or optical disk. In one or more embodiments using optical disks, computer-readable storage medium 604 includes compact disk read-only memory (CD-ROM), compact disk read / write (CD-R / W), and / or digital video disk (DVD).
[0260] In one or more embodiments, storage medium 604 stores computer program code 606, which is configured to make system 600 (where such execution represents (at least in part) an EDA tool) available for performing part or all of the processes and / or methods. In one or more embodiments, storage medium 604 also stores information that facilitates performing part or all of the processes and / or methods. In one or more embodiments, storage medium 604 stores standard cell library 607 including standard cells disclosed herein. In some embodiments, storage medium 604 stores one or more layout diagrams 611.
[0261] EDA system 600 includes I / O interface 610. I / O interface 610 is coupled to external circuitry. In one or more embodiments, I / O interface 610 includes a keyboard, keypad, mouse, trackball, trackpad, touchscreen, and / or cursor direction keys for passing information and commands to processor 602.
[0262] The EDA system 600 also includes a network interface 612 coupled to the processor 602. The network interface 612 allows the system 600 to communicate with a network 614 to which one or more other computer systems are connected. The network interface 612 includes a wireless network interface, such as Bluetooth, WIFI, WIMAX, GPRS, or WCDMA; or a wired network interface, such as Ethernet, USB, or IEEE-1364. In one or more embodiments, part or all of the processes and / or methods are implemented in two or more systems 600.
[0263] The system 600 is configured to receive information through the I / O interface 610. The information received through the I / O interface 610 includes one or more of instructions, data, design rules, a standard cell library, and / or other parameters for the processor 602 to process. The information is transmitted to the processor 602 through the bus 608. The EDA system 600 is configured to receive information related to a user interface (UI) through the I / O interface 610. This information is stored as the UI 642 in the computer-readable medium 604.
[0264] In some embodiments, part or all of the processes and / or methods are implemented as a stand-alone software application executed by a processor. In some embodiments, part or all of the processes and / or methods are implemented as a software application that is part of an additional software application. In some embodiments, part or all of the processes and / or methods are implemented as a plug-in of a software application. In some embodiments, at least one of the processes and / or methods is implemented as a software application that is part of an EDA tool. In some embodiments, part or all of the processes and / or methods are implemented as a software application used by the EDA system 600. In some embodiments, tools such as those available from CADENCE DESIGN SYSTEMS, INC. are used to generate a layout including standard cells, or another suitable layout generation tool.
[0265] In some embodiments, these processes are implemented as the functions of a program stored in a non-transitory computer-readable recording medium. Examples of non-transitory computer-readable recording media include, but are not limited to, one or more of external / removable and / or internal / built-in storage or memory units, such as optical discs (such as DVDs), magnetic disks (such as hard disks), semiconductor memories (such as ROMs), RAMs, memory cards, and the like.
[0266] Figure 7 is a functional block diagram of an integrated circuit (IC) manufacturing system 700 and its associated IC manufacturing process according to some embodiments.
[0267] In some embodiments, based on Figure 5AThe layout diagram generated by the frame 502, the IC manufacturing system 700 implements Figure 5A The block 504, where the manufacturing system 700 is used to manufacture one or more of the following: (A) one or more semiconductor masks or (B) at least one component in the early semiconductor integrated circuit layer. In some embodiments, the IC manufacturing system 700 implements one or more flowcharts disclosed herein.
[0268] In Figure 7 it, the IC manufacturing system 700 includes entities that interact in the design, development, and manufacturing cycle and / or services related to manufacturing the IC device 760, such as the design house 720, the mask house 730, and the IC fabrication plant / manufacturer ("Fab") 750. The entities in the system 700 are connected by a communication network. In some embodiments, the communication network is a single network. The communication network includes wired and / or wireless communication channels. Each entity interacts with one or more other entities and provides services to and / or receives services from one or fewer other entities. In some embodiments, two or more of the design house 720, the mask house 730, and the IC fabrication plant 750 are owned by a single larger company. In some embodiments, two or more of the design house 720, the mask house 730, and the IC fabrication plant 750 coexist in a common facility and use common resources.
[0269] The design house (or design team) 720 generates the IC design layout 722. The IC design layout 722 includes various geometric patterns designed for the IC device 760. The geometric patterns correspond to the patterns of the metal, oxide, or semiconductor layers that make up the various components of the IC device 760 to be manufactured. The layers are combined to form various IC features. For example, a part of the IC design layout 722 includes various IC features, such as active regions, gate terminals, source and drain electrodes, metal wires or vias for interlayer interconnection, and openings for pads, which will be formed in a semiconductor substrate (such as a silicon wafer) and various material layers disposed on the semiconductor substrate. Depending on the context, the source / drain regions can be referred to individually or collectively as the source or drain. The design house 720 implements appropriate design procedures to form the IC design layout 722. The design process includes one or more of logic design, physical design, or placement and routing. The IC design layout 722 is presented in one or more data files with geometric pattern information. For example, the IC design layout 722 is represented in the GDSII file format or the DFII file format.
[0270] The mask chamber 730 includes data preparation 732 and mask fabrication 734. The mask chamber 730 uses the IC design layout 722 to fabricate one or more masks 735 for the respective layers of the IC device 760 to be fabricated according to the IC design layout 720. The mask chamber 730 performs mask data preparation 732, in which the IC design layout 722 is converted into a representative data file (“RDF”). The mask data preparation 732 provides the RDF to the mask fabrication 734. The mask fabrication 734 includes a mask writer. The mask writer converts the RDF into an image on a substrate, such as a mask (reticle) or a semiconductor wafer. The design layout is manipulated by the mask data preparation 732 to conform to the specific characteristics of the mask writer and / or the requirements of the IC foundry 750. In Figure 7 , the mask data preparation 732, the mask fabrication 734, and the mask 735 are shown as separate elements. In some embodiments, the mask data preparation 732 and the mask fabrication 734 are collectively referred to as mask data preparation.
[0271] In some embodiments, the mask data preparation 732 includes optical proximity correction (OPC), which uses lithography enhancement techniques to compensate for image errors, such as those that may be caused by diffraction, interference, other process effects, etc. The OPC adjusts the IC design layout 722. In some embodiments, the mask data preparation 732 further includes resolution enhancement techniques (RET), such as off-axis illumination, sub-resolution assist features, phase-shifting masks, other suitable techniques, etc. or combinations thereof. In some embodiments, inverse lithography technology (ILT) is further used, which treats OPC as an inverse imaging problem.
[0272] In some embodiments, the mask data preparation 732 includes a mask rule checker (MRC), which uses a set of mask creation rules to inspect the OPC-processed IC design layout. These rules contain certain geometric and / or connectivity limitations to ensure sufficient margins, taking into account the variability of semiconductor manufacturing processes, etc. In some embodiments, the MRC modifies the IC design layout to compensate for the limitations during mask fabrication 734, which may undo some of the modifications performed by the OPC to meet the mask creation rules.
[0273] In some embodiments, mask data preparation 732 includes lithography process checking (LPC), which simulates the processes to be implemented by the IC fabricator 750 to fabricate the IC device 760. The LPC simulates the processes based on the IC design layout 722 to fabricate a simulated fabricated device, such as the IC device 760. The process parameters in the LPC simulation can include parameters related to various processes of the IC manufacturing cycle, parameters related to the tools used to fabricate the IC, and / or other aspects of the manufacturing process. The LPC takes into account various factors, such as aerial image contrast, depth of focus ("DOF"), mask error enhancement factor ("MEEF"), other suitable factors, etc. or combinations thereof. In some embodiments, after the LPC fabricates the simulated fabricated device, if the shape of the simulated device is not close enough to meet the design rules, OPC and / or MRC are repeated to further refine the IC design layout 722.
[0274] For clarity, the above description of mask data preparation 732 has been simplified. In some embodiments, mask data preparation 732 includes additional features, such as logical operations (LOP) to modify the IC design layout according to manufacturing rules. In addition, the processes applied to the IC design layout 722 during data preparation 732 can be performed in various different orders.
[0275] After mask data preparation 732 and during mask manufacturing 734, a mask 735 or a set of masks 735 is fabricated based on the modified IC design layout. In some embodiments, based on the modified IC design layout, a pattern is formed on the mask (photomask or reticle) using an electron beam (e-beam) or a mechanism of multiple electron beams. The mask is formed by various techniques. In some embodiments, the mask is formed using binary techniques. In some embodiments, the mask pattern includes opaque regions and transparent regions. A radiation beam, such as an ultraviolet (UV) beam, used to expose an image-sensitive material layer (such as photoresist) coated on a wafer is blocked by the opaque regions and transmitted through the transparent regions. In one example, a binary mask includes a transparent substrate (such as fused quartz) and an opaque material (such as chromium) coated in the opaque regions of the mask. In another example, the mask is formed using phase shift techniques. In a phase shift mask (PSM), various features in the pattern formed on the mask are configured to have an appropriate phase difference to improve resolution and imaging quality. In various examples, the phase shift mask is an attenuated PSM or an alternating PSM. The mask generated by mask manufacturing 734 is used for various processes. For example, such a mask is used in an ion implantation process to form various doped regions in a semiconductor wafer, in an etching process to form various etched regions in a semiconductor wafer, and / or in other suitable processes.
[0276] IC manufacturer 750 is an IC manufacturing enterprise that includes one or more manufacturing facilities for manufacturing various different IC products. In some embodiments, IC manufacturer 750 is a semiconductor foundry. For example, there may be one manufacturing facility for the front-end manufacturing (front-end-of-line (FEOL) manufacturing) of multiple IC products, while a second manufacturing facility can provide back-end manufacturing (back-end-of-line (BEOL) fabrication) for the interconnect and packaging of the IC products, and a third manufacturing facility can provide other services for the foundry business.
[0277] IC manufacturer 750 uses mask 735 manufactured by mask chamber 730 and uses manufacturing tool 752 to manufacture IC device 760. Thus, IC manufacturer 750 uses IC design layout 722 at least indirectly to manufacture IC device 760. In some embodiments, semiconductor wafer 753 is manufactured by IC manufacturer 750 using mask 735 to form IC device 760. Semiconductor wafer 753 includes a silicon substrate or other suitable substrate with material layers formed thereon. The semiconductor wafer also includes one or more of various doped regions, dielectric features, multi-level interconnects, etc. (formed in subsequent manufacturing steps).
[0278] In some embodiments, a compute-in-memory (CIM) system includes: in a first region of a semiconductor die, a first component includes memory cells respectively configured to store a single bit, and an array includes a multiplier and a first bit error detector; a first memory cell among the memory cells is arranged in a corresponding first array and is configured to store a first bit; a second memory cell among the memory cells is arranged in a corresponding second array and is configured to store a parity bit corresponding to the first bit; and for a first group, each in the first group includes a corresponding one of the first array, the second array, the multiplier, and the first bit error detector, the multiplier is configured to perform multiplication of an input bit and a corresponding one of the first bits, and the first bit error detector is configured to perform detection of a bit error in the corresponding first bit based on the corresponding parity bit.
[0279] In some embodiments, for each first group, the first bit error detector is further configured to perform the detection in parallel with the multiplication performed by the multiplier.
[0280] In some embodiments, the arrays of the multiplier and the first bit error detector further include, respectively: a multiplexer; and each first group further includes a corresponding one of the multiplexers; and for each first group, the multiplexer is configured to select (i) a product generated by the multiplier or (ii) a predefined value generated by the first bit error detector based on an output signal.
[0281] In some embodiments, the CIM system further includes a parity encoder; and wherein: each first group further includes a respective one of the parity encoders in the parity encoder; and for each first group, the parity encoder is configured to encode a respective parity bit in the parity bits based on the respective first bit.
[0282] In some embodiments, for each first group, the first bit error detector includes an exclusive OR (XOR) gate configured to receive the first bit and the parity bit as inputs and generate an output signal based on the inputs to represent an assertable first flag signal indicating the presence of a bit error.
[0283] In some embodiments, the CIM system further includes: a trace inference data generator configured to generate one or more bit error trace inference signals based on the parity bits.
[0284] In some embodiments, for each first group: the first bit error detector is further configured to generate an output signal representing an assertable first flag signal to indicate the presence of a bit error; the first array is arranged in rows and Q columns, where Q is a positive integer; there are Q first groups and Q corresponding instances of the first flag signal; the trace inference data generator is configured to receive the Q instances of the first flag signal and includes: a Q:P encoder configured to encode the Q instances of the first flag signal into a P-bit signal representing an error pointer, the error pointer being the first bit error trace inference signal among one or more bit error trace inference signals, and P is a positive integer; and a second bit error detector configured to generate a second flag signal based on the Q instances of the first flag signal, the second flag signal being assertable to indicate that the error pointer points to a bit error, and the second flag signal being the second bit error trace inference signal among one or more bit error trace inference signals.
[0285] In some embodiments, the second bit error detector includes: an OR gate configured to receive the Q instances of the first flag signal and generate an output signal representing the second flag signal based on the Q instances of the first flag signal.
[0286] In some embodiments, the CIM system further includes a bit error corrector configured to determine a respective memory cell in a respective one of the memory cells in the first array as a data corrupted cell, the data corrupted cell representing the location of the bit error based on the error pointer and the second flag signal.
[0287] In some embodiments, a method of operating a computing-in-memory (CIM) system includes, for a first set, each in the first set including a multiplier and a corresponding first bit error detector, the CIM system including a first component in a first region of a semiconductor die, the first component including memory cells configured to store a single bit and arranged in an array, an array of multipliers, and a first bit error detector, a first memory cell among the memory cells being arranged in a first array and configured to store a first bit, and a second memory cell among the memory cells being arranged in a second array and configured to store a parity bit corresponding to the first bit: performing a multiplication of an input bit and a corresponding first bit among the first bits; and performing a first error detection of a bit error in the corresponding first bit based on the corresponding parity bit, the performing of the first error detection being performed in parallel with the multiplication.
[0288] In some embodiments, the method further includes, for each first set, the performing of the first error detection including: performing a logical exclusive OR (XOR) operation on the first bit and the parity bit as inputs to generate a first flag signal, the first flag signal being assertable to indicate the presence of a bit error.
[0289] In some embodiments, the method further includes, for each first set, generating one or more bit error trace inference signals based on the parity bit.
[0290] In some embodiments, for each first set, the performing of the first error detection includes: generating a first flag signal, the first flag signal being assertable to indicate the presence of a bit error; the first array being arranged in rows and Q columns, where Q is a positive integer; there being Q first sets and corresponding Q instances of the first flag signal; for each first set, the generating of one or more bit error trace inference signals includes: performing Q:P encoding by encoding the Q instances of the first flag signal into a P-bit signal representing an error pointer, the error pointer being a first bit error trace inference signal among the one or more bit error trace inference signals; and P being a positive integer; and performing a second error detection to generate a second flag signal based on the Q instances of the first flag signal, the second flag signal being assertable to indicate that the error pointer points to a bit error, and the second flag signal being a second bit error trace inference signal among the one or more bit error trace inference signals.
[0291] In some embodiments, for each first set, the performing of the second error detection includes: performing a logical OR operation on the Q instances of the first flag signal to generate an output signal representing the second flag signal based on the Q instances of the first flag signal.
[0292] In some embodiments, a Compute-in-Memory (CIM) system includes, in a first region of a semiconductor die, a first component including memory cells configured to store words accordingly, a multiplier, and a trace inference data generator; a first memory cell among the memory cells is arranged in a first array and is configured to store a first word; for a first group, each in the first group includes a respective first array, a multiplier, and a trace inference data generator in the first array, and each in the first group operates in association with a respective first word in the first word, the multiplier is configured to generate a first word of the first array by performing one or more multiplications of (A) an input word and an associated first checksum word and (B) a respective weight word and an associated second checksum word, and the trace inference data generator is configured to perform generation of one or more bit error trace inference signals based on a selected first word among the first words.
[0293] In some embodiments, for each first group: the first array is arranged in a first row and a first column, the first word is correspondingly located at the intersection of the first row and the first column; the first word represents a product word; and for each first group, the trace inference data generator includes: a row sum generator configured to generate a row sum based on the product word.
[0294] In some embodiments, for each first group, the row sum is a row vector composed of second words; and the row sum generator includes: a recursive adder corresponding to a first column in the first array, each recursive adder is configured to generate a respective one of the second words in the row sum by recursively adding the first words in the corresponding first column in the first column column by column.
[0295] In some embodiments, for each first group: the first array is arranged in a first row and a first column, the first word is correspondingly located at the intersection of the first row and the first column; the first word represents a product word; and for each first group, the trace inference data generator includes: a column sum generator configured to generate a column sum based on the product word.
[0296] In some embodiments, for each first group: the column sum is a column vector composed of second words; and the column sum generator includes: an adder tree configured to generate the second words of the column sum row by row by adding the first words in the corresponding first rows in the first row.
[0297] In some embodiments, for each first group, the trace inference data generator is further configured to perform generation of one or more bit error trace inference signals after the multiplier performs one or more multiplications.
[0298] One or more of the disclosed embodiments will readily be seen by one of ordinary skill in the art to achieve one or more of the above advantages. After reading the above specification, one of ordinary skill in the art will be able to effect various changes, substitutions of equivalents, and various other embodiments as broadly disclosed herein. Accordingly, the protection granted herein is limited only by the definitions contained in the appended claims and their equivalents.
Claims
1. A computing-in-memory system, comprising: In a first region of the semiconductor die, the first component includes memory cells respectively configured to store a single bit, and the array includes a multiplier and a first bit error detector; A first memory cell of the memory cells is arranged in a corresponding first array and configured to store a first bit; A second memory cell of the memory cells is arranged in a corresponding second array and is configured to store a parity bit corresponding to the first bit; as well as for a first group, each of the first groups includes a corresponding one of the first array, the second array, the multiplier, and the first bit error detector, The multiplier is configured to perform a multiplication of an input bit and a corresponding first one of the first bits, and The first bit error detector is configured to perform detection of a bit error in the corresponding first bit based on the corresponding parity bit.
2. The in-memory computing system of claim 1, wherein: For each first group, The first bit error detector is further configured to perform the detection in parallel with the multiplication performed by the multiplier.
3. The in-memory computing system of claim 1 , wherein: Said array of multipliers and first bit error detectors accordingly further comprises: Multiplexers; and Each first group further comprises a respective one of said multiplexers; and For each first group, The multiplexer is configured to select (i) a product generated by the multiplier or (ii) a predefined value generated by the first bit error detector based on an output signal.
4. The in-memory computing system of claim 1 , further comprising: Parity check encoder; and in: Each first group further includes a corresponding one of the parity check encoders; as well as For each first group, The parity encoder is configured to encode corresponding ones of the parity bits based on corresponding first bits.
5. The in-memory computing system of claim 1 , further comprising: The trace inference data generator is configured to generate one or more bit error trace inference signals based on the parity bits.
6. The in-memory computing system of claim 5, wherein: For each first group: The first bit error detector is further configured to generate an output signal representing a first flag signal that may be asserted to indicate the presence of a bit error; The first array is arranged in rows and Q columns, where Q is a positive integer; There are Q first groups and corresponding Q instances of the first flag signal; The trajectory inference data generator is configured to receive the Q instances of the first flag signal and comprises: A Q:P encoder configured to encode the Q instances of the first flag signal into a P-bit signal representing an error pointer, the error pointer being a first bit error trajectory inference signal of the one or more bit error trajectory inference signals, and P being a positive integer; and a second bit error detector configured to generate a second flag signal based on the Q instances of the first flag signal, the second flag signal being assertable to indicate that the error pointer points to the bit error, and the second flag signal being a second bit error trajectory inference signal among the one or more bit error trajectory inference signals.
7. A method of operating a computing system in memory, the method comprising: For a first group, each of the first group includes a multiplier and a corresponding first bit error detector, the computing system in memory includes a first component in a first region of a semiconductor die, the first component includes memory cells respectively configured to store a single bit and arranged in an array, an array of multipliers, and a first bit error detector, a first memory cell of the memory cells being arranged in a first array and configured to store a first bit, a second memory cell of the memory cells being arranged in a second array and configured to store a parity bit corresponding to the first bit; performing a multiplication of an input bit and a corresponding first one of the first bits; as well as A first error detection is performed on a bit error in a corresponding first bit based on the corresponding parity bit, the first error detection being performed in parallel with the multiplication.
8. The method according to claim 7, further comprising: For each first group, One or more bit error trajectory inference signals are generated based on the parity bits.
9. The method according to claim 8, wherein: For each first group, performing the first error detection comprises: generating a first flag signal, the first flag signal being assertable to indicate the presence of a bit error; The first array is arranged in rows and Q columns, where Q is a positive integer; There are Q first groups and corresponding Q instances of the first flag signal; For each first group, Generating one or more bit error trajectory inference signals comprises: Q:P encoding is performed by encoding the Q instances of the first flag signal into a P-bit signal representing an error pointer, the error pointer being a first bit error trajectory inference signal of the one or more bit error trajectory inference signals; and P is a positive integer; and A second error detection is performed to generate a second flag signal based on the Q instances of the first flag signal, the second flag signal being assertable to indicate that the error pointer points to the bit error, and the second flag signal being a second bit error trajectory inference signal of the one or more bit error trajectory inference signals.
10. A computing-in-memory system, comprising: In a first region of the semiconductor die, a first component includes a memory cell, a multiplier, and a trajectory inference data generator respectively configured to store a word; A first memory cell of the memory cells is arranged in a first array and configured to store a first word; for a first group, each of the first groups includes a corresponding first array of the first arrays, the multiplier, and the trajectory inference data generator, and each of the first groups operates in association with a corresponding first word of the first words, the multiplier being configured to generate the first word of the first array by performing one or more multiplications of (A) an input word and an associated first checksum word and (B) a corresponding weight word and an associated second checksum word, and The trajectory inference data generator is configured to perform generation of one or more bit error trajectory inference signals based on selected ones of the first words.