In-memory computing array, in-memory computing macro circuit and in-memory computing device
By introducing global bitline switches and collaborative optimization design methods into the in-memory computing array, the problem of mismatch between write bandwidth and computing throughput in traditional in-memory computing hardware acceleration chips is solved, and the parallelism of storage mode and computing mode is realized, and computing efficiency and storage density are improved.
Patent Information
- Application Number
- CN202510421123.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-27
AI Technical Summary
The write bandwidth and computing throughput of traditional in-memory computing hardware acceleration chips do not match, resulting in reduced computing efficiency, and the memory layout design does not match the CMOS logic design, resulting in area waste and performance bottlenecks.
By introducing global bit line switches into the in-memory computing array, parallel operation of data reading from memory cells to latch is realized, ensuring parallel operation of calculation and read and write operations. At the same time, a collaborative optimization design method is adopted to optimize the design rules of storage units and CMOS logic units to improve storage density and computing power density.
The parallelization of storage mode and computing mode is realized, eliminating data handling delays, improving computing throughput and storage density, reducing power consumption and area waste, and improving overall performance and stability.
Smart Images

Figure CN120048302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip technology, and particularly to an in-memory computing array, an in-memory computing macro circuit, and an in-memory computing device. Background Art
[0002] With the rapid development of artificial intelligence, the algorithm data and computing volume of neural networks have increased significantly. Among them, matrix-vector multiplication operations account for the vast majority of the computing requirements of generative artificial intelligence models, greatly increasing the requirement for computing parallelism. Due to the separation of storage and computing in traditional processors, there is a large amount of data transfer between the memory and the arithmetic unit, resulting in the power consumption and delay overhead of data transfer being much higher than the computing overhead, causing the "memory wall" bottleneck. The emerging in-memory computing chips are considered to be an effective technical route that is expected to break through the "memory wall" bottleneck and accelerate the computing of artificial intelligence models.
[0003] However, the current in-memory computing hardware acceleration technology still faces three challenges.
[0004] First, the write bandwidth and computing throughput rate of the current in-memory computing hardware acceleration chips do not match, and the weight write update affects the computing efficiency, and there is a conflict between the two. In the actual hardware design and software operation process, due to physical factors such as lithography limits and economic factors such as production costs, the hardware resources that can be integrated on the chip are always limited. Therefore, the current memory-computation integrated technology cannot fully deploy many large neural network models on a single chip. This means that in actual operation, the in-memory computing array can only store part of the weights of the neural network, and the weight data needs to be written and updated dynamically, as Figure 1 shown.
[0005] The traditional in-memory computing architecture does not support parallel reading / writing and computing. Therefore, when updating the data in the memory array, the computing must be paused. This design increases the time and energy overhead and reduces the overall efficiency of the in-memory computing accelerator. To solve this problem, pipeline technology and data prefetch technology can be introduced at the accelerator system level to reduce the delay caused by data transfer. However, when the data transfer time is much longer than the computing time, these technologies are difficult to effectively hide the data transmission delay, resulting in the performance bottleneck still existing.
[0006] Existing solutions mainly focus on analog in-memory computing technologies based on the current domain, voltage domain, charge domain, and time domain. These solutions have significant advantages in computing energy efficiency and computing power density. However, the computing accuracy in the analog domain is limited by factors such as device fluctuations and circuit noise, and the system stability risk is relatively high, making it difficult to be implemented. On the other hand, taking the SRAM-based in-memory computing technology as an example, in such designs, the bitlines of SRAM are usually coupled with computing units, resulting in the in-memory computing array being able to perform only computing or read / write operations at the same time. Different from the computing function that can achieve parallelism at the two-dimensional level, the weight update in the memory array still occurs in a one-dimensional row-by-row writing manner, resulting in a significant mismatch between the data writing bandwidth and the computing throughput. Therefore, when facing large neural network models, the actual computing efficiency of traditional analog in-memory computing technologies drops significantly.
[0007] Second, there are differences between the traditional memory layout design and the CMOS logic layout design. This mismatch brings additional lateral space requirements between the memory cell area and the logic area, as well as different cell heights, resulting in area waste, as Figure 2 shown. In the traditional memory-compute separation architecture, the designs of memory cells and CMOS logic cells follow different rules and optimization goals. The design of memory cells focuses on high density and repairability to meet the requirements of large-scale data storage. Taking SRAM as an example, these cells usually adopt compact design rules (Push Rule) to achieve high storage density, while allowing a certain degree of manufacturing defects because the SRAM design includes an easily implemented cell repair mechanism. This design is usually automatically generated by a dedicated compiler according to the requirements of the system architect within a certain range to form an SRAM array. In contrast, the design of CMOS logic cells pays more attention to reliability and integration. Due to the complex and diverse topological structures of logic cells, including various logic gates and flip-flops, their design must take into account a wide range of application scenarios and high uncertainty. Therefore, CMOS logic cells usually adopt more conservative design rules to ensure the generality and reliability of the design. This design rule is called Logic Rule, which allows logic cells to follow the same digital circuit design process and be plug-and-play in different digital chip designs.
[0008] However, the mismatch between these two design rules makes it difficult to achieve the best power consumption, performance, area and other metrics in the in-memory computing architecture. To solve this problem, the present invention proposes a co-optimization design method, aiming to study the design rules of memory cells and CMOS logic cells, and propose a new design method to enable these two types of cells to coexist and work together in the same architecture through customized layout design.
[0009] Third, in the design of a digital domain in-memory computing array with high-density integration, due to the tight coupling between digital computing logic units and memories involved in analog signal processing, it has a negative impact on signal and power integrity, and limits the operating speeds of both, directly affecting the performance, stability, and reliability of the entire system.
[0010] On the one hand, the memory for processing analog signals requires ensuring the accuracy of signals while only allowing slow charging and discharging of bit lines. However, the digital computing logic units tightly coupled to the memory may generate high-frequency digital signals, which can couple to the effective computing performance of the memory through various parasitic effects and may even lead to calculation errors. Summary of the Invention
[0011] In view of the above problems, the present invention provides an in-memory computing array, including: a memory cell group capable of storing data for computing; a computing unit for performing computing operations; a latch capable of storing data read from the memory cell group; and a global bit line switch connected between a global bit line and the memory cell group, wherein the global bit line can realize data interaction between the in-memory computing array and the outside, and the global bit line switch is configured to: disconnect in response to a computing operation instruction, read data from the memory cell group to the latch for performing computing operations, and close after the latch completes data latching, so that the memory cell group can read / write data through the global bit line, enabling computing operations and read / write operations to be parallel.
[0012] The in-memory computing array provided in this application realizes the parallelism of the storage mode and the computing mode through the isolation of the global bit line and the local bit line by the global bit line switch and the latching of the data used for computing by the latch. Specifically, the local bit line is responsible for connecting a smaller range of memory cells, for example, the storage cells of an in-memory computing array, while the global bit line is responsible for summarizing the signals of multiple local bit lines and performing global transmission. When entering the computing mode / responding to a computing operation instruction, the global bit line switch disconnects. At this time, the global bit line is isolated from the memory cell group and the local bit line, and the data and weights required for performing computing operations are read from the memory cell group to the latch. After the data is latched, the global bit line switch closes. If read / write operations are required, since the data / weights required for performing computing operations have been latched to the latch, the reading / writing of data to / from the memory cell group in the storage mode (also called the read / write mode) does not affect the execution of the computing task. Thus, an in-memory computing architecture is provided where the storage and computing modules are separated but closely adjacent to each other, without additional delay overhead for data transfer, and at the same time, high-speed parallelism of reading / writing and computing can be achieved.
[0013] To achieve the above-mentioned invention object, the present embodiment provides an in-memory computing macro circuit, including: an in-memory computing array group, which integrates a plurality of the in-memory computing arrays therein, and an adder tree connected to the computing units of each in-memory computing array; a decoding and driving module, configured to decode the input address information and generate corresponding driving signals to select a specific in-memory computing array group for local readout operations and control the computing of the in-memory computing array group; a read / write driving circuit, configured to control the data writing / reading of the memory cell group in the in-memory computing array group; and a control signal generation module, connected to the decoding and driving module and the read / write driving circuit, configured to generate control signals to control the decoding and driving module and the read / write driving circuit.
[0014] Optionally, the in-memory computing array group is configured to have 32*64 sub-arrays, each sub-array includes 16*4 memory cells, and every 16 memory cells cooperate with a latch to form one of the in-memory computing arrays.
[0015] Optionally, in the column direction, 32 of the sub-arrays and 1 adder tree form a DCIM block, and 64 of the DCIM blocks are used to perform multiply-accumulation operations of 1-bit inputs and 4-bit weights for 32 input channels. The 64 DCIM blocks constitute a DCIM array, supporting INT4 / 8 / 16 data format configurations.
[0016] Optionally, the in-memory computing macro circuit is configured with a 4-bit wide and 64-bit mask, each bit of which controls whether the corresponding DCIM block in the column direction is enabled. If the weights of the DCIM block do not need to be updated, the mask will cancel the write enable of the corresponding DCIM block, thereby avoiding re-write operations.
[0017] Optionally, each bit of the mask corresponds to one of the DCIM blocks. If the mask bit is 1, the corresponding DCIM block is enabled for write operations; if the mask bit is 0, the corresponding DCIM block is not enabled and the write operation is skipped.
[0018] Optionally, it further includes an in-situ decoupling circuit, including: a differential amplifier SA, a first capacitor CP1, and a second capacitor CP2. The non-inverting input terminal of the differential amplifier SA is configured to be connected to the positive signal bit line BLP, and the inverting input terminal of the differential amplifier SA is configured to be connected to the negative signal bit line BLN. The memory cell power domain VDDM is designed to be disposed between the positive signal bit line and the negative signal bit line, and the computing unit power domain VDDLO is symmetrically disposed on both sides of the memory cell power domain VDDM. The first capacitor CP1 and the second capacitor CP2 are respectively connected between the computing unit power domain VDDLO disposed on both sides of the memory cell power domain VDDM and the positive signal bit line BLP and the negative signal bit line BLN.
[0019] Optionally, the layouts of the memory cell and the computing unit are configured to be highly matched.
[0020] Optionally, the full adders of the adder tree are interleaved between the sub-arrays of the DCIM block.
[0021] To achieve the above invention objective, the present application provides a memory-in-computation device configured with the memory-in-computation array or the memory-in-computation macro circuit described above.
[0022] In summary, the memory-in-computation array and the memory-in-computation macro circuit provided by the present application eliminate the data transfer delay from the memory cell to the computing unit, enable data reading / writing and computing to be performed simultaneously, and avoid reducing the computing throughput due to waiting for weight updates. In addition, due to the optimized layout design, the memory cell and the computing unit can be tightly coupled in a smaller chip space, eliminating the space redundancy caused by the mismatch of layout rules, improving the storage density and computing power density, and significantly improving the chip performance. This embodiment also reduces the crosstalk influence between the analog and digital signals, suppresses the power supply noise, avoids the wiring congestion problem, optimizes the signal and power integrity, improves the robustness of the memory-computation integrated array under various working conditions, and ensures the stable operation of the system by setting an in-situ decoupling circuit. Description of the Drawings
[0023] Figure 1 is a schematic structural diagram of a memory-in-computation array in the prior art.
[0024] Figure 2 is a schematic layout structure diagram of a memory cell and a computing unit in a memory-in-computation array in the prior art.
[0025] Figure 3 is a schematic structural diagram of a memory-in-computation macro circuit and a memory-in-computation array provided by an embodiment of the present invention.
[0026] Figure 4It is a schematic diagram of the operation timing of the in-memory computing array provided in the embodiments of the present invention.
[0027] Figure 5 It is a schematic layout diagram of the memory cells and computing units in the in-memory computing array provided in the embodiments of the present invention.
[0028] Figure 6 It is a schematic layout diagram of the memory cells and computing units in the in-memory computing array provided in the embodiments of the present invention.
[0029] Figure 7 It is a schematic diagram of the structure of the in-situ decoupling circuit provided in the embodiments of the present invention. Specific Embodiments
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] As Figure 3 shown, this embodiment provides an in-memory computing array 100, including: a memory cell group 110 capable of storing data for computing; a computing unit 120 for performing computing operations inside the in-memory computing array; a latch 130 capable of storing the data read from the memory cell group 110; and a global bit line switch GBLSW connected between the global bit line GBL (including the positive signal bit line GBLP and the negative signal bit line GBLN) and the memory cell group 110. Among them, the global bit line GBL can realize data interaction between the in-memory computing array 100 and the outside. The global bit line switch GBLSW is configured to: disconnect in response to a computing operation instruction, read the data from the memory cell group 110 to the latch 130 for performing the computing operation, and close after the latch 130 completes data latching, so that the memory cell group 110 can read / write data through the global bit line GBL, enabling the computing operation and the read / write operation to be parallel.
[0032] The in-memory computing array provided by this embodiment realizes the parallelism of the storage mode and the computing mode through the isolation of the global bit line switch GBLSW for the global bit line and the local bit line LBL (including the positive signal bit line LBLP and the negative signal bit line LBLN), and the latching of the latch 130 for the data participating in the calculation. Specifically, the local bit line LBL is responsible for connecting a relatively small range of memory cells. For example, the memory cells of an in-memory computing array 100, while the global bit line is responsible for summarizing the signals of multiple local bit lines LBL and performing global transmission. When entering the computing mode / executing a computing operation instruction, the global bit line switch GBLSW is disconnected. At this time, the global bit line GBL is isolated from the memory cell group 110 and the local bit line. The data and weights required for the computing operation are read from the memory cell group 110 to the latch 130. In response to the latch 130 completing data latching, the global bit line switch GBLSW is closed. At this time, if a read / write operation is required, since the data / weights required for the computing operation have been latched to the latch 130, the data read / write of the memory cell group 110 in the storage mode (also called the read / write mode) does not affect the execution of the computing task. Thus, a memory and computing integrated architecture is provided where the storage and computing modules are separated but closely adjacent, without additional delay overhead for data transfer, and at the same time, high-speed parallelism of reading / writing and computing can be achieved.
[0033] The following combines Figure 4 to illustrate the operating timing of the in-memory computing array 100 provided by this embodiment.
[0034] When the in-memory computing array 100 enters the computing mode, the global bit line switch GBLSW is disconnected. Since the memory cell group 110 is isolated from the global bit line GBL, the memory cell group 110 stops / does not perform read / write operations, and the latch 130 locally reads the data required for the computing operation from the memory cell group 110. In response to the completion of local data reading, the global bit line switch GBLSW is closed. At this time, the computing unit 120 can execute the computing task and send the computing result to the adder tree for accumulation. At the same time, since the global bit line switch GBLSW is closed, the isolation between the global bit line GBL and the memory cell group 110 is released. Therefore, while the computing unit 120 is performing the computing operation, data reading / writing operations in the storage mode can be performed in parallel.
[0035] Continue to refer to Figure 3 This embodiment provides an in-memory computing macro circuit 200, including: an in-memory computing array group 210, which integrates multiple in-memory computing arrays 100 therein, and an adder tree 211 connected to the computing units 120 of each in-memory computing array;
[0036] A decoding and driving module 220, configured to decode the input address information and generate corresponding driving signals to select a specific in-memory computing array group 210 or in-memory computing array 100 for local readout operations, and to control the computing of the in-memory computing array group 210;
[0037] A read / write driving circuit 230, configured to control data writing / reading of the memory cell group 110 in the in-memory computing array group 210; and
[0038] A control signal generation module 240, connected to the decoding and driving module 220 and the read / write driving circuit 230, configured to generate control signals to control the decoding and driving module 220 and the read / write driving circuit 230.
[0039] Specifically, the in-memory computing array group 210 can be configured to have 32 * 64 sub-arrays. Each sub-array includes 16 * 4 memory cells. Every 16 memory cells cooperate with a latch to form an in-memory computing array 100. The 4 in-memory computing arrays 100 in each sub-array store a total of 16 matrices of 4-bit weights, which can be used to perform multiplication operations of 1-bit inputs and 4-bit weights. In the column direction, 32 sub-arrays and 1 adder tree (AT) form a DCIM block. There are a total of 64 DCIM blocks, which are used to perform multiply-accumulate (MAC) operations of 1-bit inputs and 4-bit weights for 32 input channels. The 64 DCIM blocks constitute a DCIM array, which supports INT4 / 8 / 16 data format configurations. Such a setting enables the architecture to flexibly process weights and input data of different bit widths to meet different precision requirements. In addition, the 64 DCIM blocks constitute a DCIM array, significantly improving the parallel processing ability of in-memory computing and being suitable for efficient inference in neural networks.
[0040] Optionally, in this embodiment, the read / write driving circuit 230 is configured with a 4-bit wide and 64-bit mask (Mask), and each bit of it controls whether the DCIM block in the corresponding column direction is enabled. When the weights of some DCIM blocks do not need to be updated, the mask cancels the write enable of these DCIM blocks, thus avoiding re-write operations.
[0041] Specifically, a 64-bit mask is mapped to 64 DCIM blocks in the column direction. Each bit in the mask corresponds to a DCIM block: if the mask bit is 1, the DCIM block is enabled for writing operations; if the mask bit is 0, the DCIM block is not enabled and the write operation is skipped.
[0042] This method allows for flexible selection of which DCIM blocks participate in writing in the column direction, thus achieving flexible configuration of the data writing positions.
[0043] The technical effects of the mask setting include:
[0044] Energy saving: It significantly reduces unnecessary write operations and lowers the overall energy consumption.
[0045] Performance improvement: It avoids unnecessary writes, reduces operation time, and improves write efficiency.
[0046] Flexibility: It supports selective update of weights in different columns, facilitating adaptation to dynamically changing workloads.
[0047] In summary, the in-memory computing macro circuit 200 of the in-memory computing array 100 provided by this embodiment, through the hierarchical design concept of DCIM blocks, sub-arrays, and adder trees, while achieving efficient parallel computing of storage mode and computing mode, realizes the efficient operation of matrix multiplication and accumulation operations.
[0048] In addition, since it supports INT4 / 8 / 16 data format configuration, the flexible data format support can adjust the precision of weights and input data according to application requirements.
[0049] Finally, the write operation is optimized through the mask function: Through a flexible 64-bit mask, unnecessary write operations are avoided, reducing power consumption and improving performance.
[0050] This design is applicable to the field of high-performance computing, especially in in-memory computing applications for neural network inference and training, with significant technical advantages.
[0051] To solve the problem of the mismatch between the layout of memory cells and computing cells in the prior art, as Figure 5 shown, the layouts of memory cells and computing cells are set to be highly matched. Such a design can make the memory cells and computing cells adjacent closely, compressing the layout area.
[0052] Furthermore, in this embodiment, as Figure 6 shown, taking the SRAM-based in-memory computing technology as an example, the height of SRAM is increased to match that of the computing unit LLNC, and the sizes of the n-well region and p-substrate region are adjusted accordingly. Such a design can not only ensure yield and performance, but also provide more space for the design and wiring of the computing unit, as well as the power / ground grid in the adder tree, which is beneficial to the performance optimization of the DCIM macro cell. The full adders (FAs) in the adder tree (AT) are staggered between the sub-arrays of the DCIM block, and the well sizes and cell heights matching the storage cells improve the area utilization rate, achieving higher integration and reliability in a smaller chip space.
[0053] Optionally, to avoid crosstalk in the digital signals of the adder tree 211 and the analog signals on signal lines such as the global bit line GBL in the in-memory computing macro circuit 200 provided in this embodiment, this embodiment provides an in-situ decoupling routing strategy. On the one hand, the digital signal lines and the analog signal lines are deployed on different metal layers to reduce their mutual interference. On the other hand, as Figure 7 shown, specifically, Figure 7 the left side of shows that in the prior art, due to its sensitivity, the memory cell power domain VDDM (which can also be called the memory unit power domain) is prone to crosstalk noise caused by the high-frequency digital signals of the computing unit power domain VDDLO, resulting in signal errors due to coupled noise interference. The right side shows that this embodiment provides an in-situ decoupling circuit, including: a differential amplifier SA, a first capacitor CP1, and a second capacitor CP2. The non-inverting input terminal of the differential amplifier SA is configured to connect to the positive signal bit line BLP, and the inverting input terminal of the differential amplifier SA is configured to connect to the negative signal bit line BLN. The memory unit power domain VDDM is designed to be disposed between the positive signal bit line and the negative signal bit line. The computing unit power domain is symmetrically disposed on both sides of the memory unit power domain VDDM. The first capacitor CP1 and the second capacitor CP2 are respectively connected between the computing unit power domain VDDM disposed on both sides of the memory unit power domain VDDM and the positive signal bit line BLP and the negative signal bit line BLN. With such a design, the coupling noise between the analog signal line and the digital signal line can be eliminated by setting decoupling capacitors in the way of differential structure routing without adding additional components.
[0054] In summary, the in-memory computing array 100 and the in-memory computing macro circuit 200 provided in this embodiment eliminate the data transfer delay from the memory unit to the computing unit, enable data reading, writing, and computing to be performed simultaneously, and avoid waiting for weight updates to reduce the computing throughput. In addition, due to the optimized layout design, the memory unit and the computing unit can be tightly coupled in a smaller chip space, eliminating the space redundancy caused by the mismatch of layout rules, improving the storage density and computing power density, and significantly improving the chip performance. This embodiment also reduces the crosstalk impact between analog and digital signals, suppresses power supply noise, avoids routing congestion problems, optimizes signal and power integrity, improves the robustness of the memory-computation integrated array under various working conditions, and ensures the stable operation of the system by setting an in-situ decoupling circuit.
[0055] So far, the technical solutions of the present invention have been described with reference to the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to the above specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. An in-memory computing array, characterized in that: include: a group of memory cells capable of storing data for computation; A computing unit, for performing computing operations; a latch capable of storing data read from the memory cell group; as well as A global bit line switch is connected between a global bit line and the memory cell group, wherein the global bit line can realize data interaction between the in-memory computing array and the outside, and the global bit line switch is configured as follows: The device is disconnected in response to a computing operation instruction, and data is read from the memory cell group to the latch for executing the computing operation, and is closed after the latch completes data latching, so that the memory cell group can read / write data through the global bit line, so that the computing operation and the read and write operations can be carried out in parallel.
2. An in-memory computing macro circuit, using the in-memory computing array according to claim 1, characterized in that: include: An in-memory computing array group, which integrates a plurality of the in-memory computing arrays and an adder tree connected to the computing units of each in-memory computing array; A decoding and driving module, used to decode the input address information and generate a corresponding driving signal to select a specific in-memory computing array group for local read operation and control the calculation of the in-memory computing array group; A read / write drive circuit, used for controlling the writing / reading of data from the memory cell group in the in-memory computing array group; as well as The control signal generating module is connected to the decoding and driving module and the read-write driving circuit, and is used to generate a control signal to control the decoding and driving module and the read-write driving circuit.
3. The in-memory computing macro circuit according to claim 2, characterized in that: The in-memory computing array group is configured to have 32*64 sub-arrays, each sub-array includes 16*4 memory cells, and every 16 memory cells cooperate with latches to form an in-memory computing array.
4. The in-memory computing macro circuit according to claim 3, characterized in that: In the column direction, 32 of the sub-arrays and 1 adder tree form a DCIM block, 64 of the DCIM blocks are used to perform multiplication and accumulation operations of 1-bit input and 4-bit weight of 32 input channels, and the 64 DCIM blocks constitute a DCIM array, supporting INT4 / 8 / 16 data format configuration.
5. The in-memory computing macro circuit according to claim 4, characterized in that: The read / write drive circuit is configured with a 4-bit wide 64-bit mask, each bit of which controls whether the DCIM block in the corresponding column direction is enabled. If the weight of the DCIM block does not need to be updated, the mask will cancel the write enable of the corresponding DCIM block, thereby avoiding re-writing operations.
6. The in-memory computing macro circuit according to claim 4, characterized in that: Each bit of the mask corresponds to a DCIM block. If the mask bit is 1, the corresponding DCIM block is enabled to enter the write operation; if the mask bit is 0, the corresponding DCIM block is not enabled and the write operation is skipped.
7. The in-memory computing macro circuit according to any one of claims 2 to 6, characterized in that: It also includes an in-situ decoupling circuit, including: a differential amplifier SA, a first capacitor CP1 and a second capacitor CP2, the non-inverting input terminal of the differential amplifier SA is configured to connect to the positive signal bit line BLP, the inverting input terminal of the differential amplifier SA is configured to connect to the negative signal bit line BLN, the memory cell power domain VDDM is designed to be set between the positive signal bit line and the negative signal bit line, the calculation unit power domain VDDLO is symmetrically set on both sides of the memory cell power domain VDDM, the first capacitor CP1 and the second capacitor CP2 are respectively connected to the calculation unit power domain VDDLO set on both sides of the memory cell power domain VDDM and between the positive signal bit line BLP and the negative signal bit line BLN.
8. The in-memory computing macro circuit according to any one of claims 4 to 6, characterized in that: The layouts of the memory unit and the computing unit are configured to be highly matched.
9. The in-memory computing macro circuit according to claim 8, characterized in that: The full adders of the adder tree are interleaved between the sub-arrays of the DCIM blocks.
10. An in-memory computing device, characterized in that: The device is provided with the in-memory computing array as claimed in claim 1 or the in-memory computing macro circuit as claimed in any one of claims 2-9.