Storage and calculation integrated chip and storage device

The Store-Compute-One chip design addresses the latency issues of von Neumann architecture by integrating storage and computation units, enhancing chip efficiency through reduced data transmission and improved computational latency.

CN120316063APending Publication Date: 2025-07-15李修录
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510482956.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The chips under the traditional von Neumann architecture have a bottleneck in computing efficiency due to the separation of storage modules and computing modules, especially in artificial intelligence and big data analysis scenarios.

Method used

The integrated storage and computing module is adopted. Through the design of the integrated storage and computing unit and the basic storage and computing array, combined with the main control module and the circuit network, the integration of data storage and computing is achieved, and the data transmission distance is shortened.

Benefits of technology

Effectively reduce computing delay, improve chip computing efficiency, enhance the computing function of RAM arrays, and improve storage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316063A_ABST
    Figure CN120316063A_ABST
Patent Text Reader

Abstract

The invention discloses a storage and calculation integrated chip and a storage device, the storage and calculation integrated chip comprises a storage and calculation integrated module formed by connecting a plurality of storage and calculation integrated units, and the storage and calculation integrated module is used for storing data and calculating; each storage and calculation integrated unit comprises a basic storage and calculation unit array formed by connecting a plurality of basic storage and calculation units, each storage and calculation integrated unit is at least connected with two other storage and calculation integrated units, each basic storage and calculation integrated unit is at least connected with two other basic storage and calculation integrated units, and each basic storage and calculation integrated unit comprises an RAM array; the main control module is used for managing calculation tasks executed by the storage and calculation integrated chip; and the circuit network is used for connecting the plurality of storage and calculation integrated units in the storage and calculation integrated module and connecting the main control module with the storage and calculation integrated module. The data is stored and calculated through the storage and calculation integrated module, so that the data transmission distance is effectively shortened, the calculation delay can be reduced, and the calculation efficiency of the chip is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chips, and particularly to a computing-in-memory chip and a storage device. Background Art

[0002] In today's information age, data processing has become the core requirement in all walks of life. From real-time image processing on smartphones, large-scale data analysis in data centers to complex decision-making systems in autonomous driving, there is an urgent need for high-performance and low-power chips.

[0003] Currently, the mainstream chip architectures on the market are still mainly based on the traditional von Neumann architecture, that is, the storage module and the computing module are separated. This chip architecture is convenient for designing the storage module and the computing module, and is also convenient for large-scale production on the production line.

[0004] However, with mainstream chips, data needs to be frequently transferred between the storage module and the computing module, resulting in a high computing latency. With the emergence of computing-intensive application scenarios such as artificial intelligence and big data analysis, the bottleneck of computing efficiency caused by the separation of the chip computing module and the storage module under the traditional von Neumann architecture has become increasingly apparent. Although the chip manufacturing process has continuously advanced, such as the application of 7-nanometer and 5-nanometer technologies, which have improved the chip performance to a certain extent, it has not fundamentally solved the computing efficiency problem caused by the separation of the storage module and the computing module. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present application provides a computing-in-memory chip and a storage device. The computing-in-memory module is used to store data and perform calculations, effectively shortening the data transmission distance, reducing the computing latency, and thus greatly improving the computing efficiency of the chip.

[0006] To solve the above problems, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a computing-in-memory chip, including: a computing-in-memory module composed of a plurality of computing-in-memory units connected together, where the computing-in-memory module is used to store data and perform calculations;

[0008] The computing-in-memory unit includes a basic computing-in-memory unit array composed of a plurality of basic computing-in-memory units connected together. Each computing-in-memory unit is connected to at least two other computing-in-memory units, and each basic computing-in-memory unit is connected to at least two other basic computing-in-memory units. The basic computing-in-memory unit includes a RAM array composed of a plurality of RAM units;

[0009] A main control module for managing the computing tasks executed by the computing-in-memory chip;

[0010] A circuit network for connecting a plurality of memory - computing units in the memory - computing integrated module and connecting the main control module to the memory - computing integrated module.

[0011] In some embodiments, the memory - computing integrated chip further includes a global cache module. The global cache module is connected to the memory - computing integrated module through the circuit network, and is used for caching the latest data received by the memory - computing integrated chip.

[0012] In some embodiments, the memory - computing unit further includes a unit cache module, a first computing module, a first control module, and a model function module;

[0013] The unit cache module, the first computing module, and the model function module are all connected to the basic memory - computing unit array, and the first control module is connected to the model function module;

[0014] The unit cache module is used for caching the latest data received by the memory - computing unit, the first computing module is used for performing calculations based on the data in the basic memory - computing unit array, the first control module is used for controlling the operation of the memory - computing unit, and the model function module is used for implementing the calculation functions required during the operation of the artificial intelligence model based on the data in the basic memory - computing unit array.

[0015] In some embodiments, the basic memory - computing unit further includes a basic unit cache module, a second computing module, and a second control module. The second computing module and the unit cache module are both connected to the RAM array, the second control module is connected to the second computing module, the basic unit cache module is used for caching the latest data received by the basic memory - computing unit, the second computing module is used for performing calculations based on the data in the RAM array, and the second control module is used for controlling the operation of the basic memory - computing unit.

[0016] In some embodiments, the second computing module includes a row decoding circuit, a column decoding circuit, a driving circuit, a read - write control circuit, a working mode control circuit, and a pulse - width modulation circuit;

[0017] The row decoding circuit is connected to each row of RAM units in the RAM array, the column decoding circuit is connected to each column of RAM units in the RAM array, the read - write control circuit is connected to the working mode control circuit, the working mode control circuit is connected to the driving circuit, the pulse - width modulation circuit is connected to the driving circuit, and the driving circuit is connected to the row decoding circuit and the column decoding circuit;

[0018] The row decoding circuit is used to control the potential of the word lines of the corresponding rows according to the first address signal, and the column decoding circuit is used to control the potential of the word lines of the corresponding columns according to the second address signal, so as to select a RAM cell in the RAM array;

[0019] The driving circuit is used to control the potentials of the word lines and bit lines in the RAM array by controlling the row decoding circuit and the column decoding circuit according to the control signal of the working mode control circuit, so as to determine the working mode of the RAM array as the target working mode;

[0020] The read-write control circuit is used to send a read signal or a write signal to the working mode control circuit according to the read-write control signal sent by the second control module;

[0021] The working mode control circuit is used to determine the working mode of the RAM array according to the read signal or the write signal. The working modes include a read mode, a write mode, and an operation mode, and the operation mode includes a multiplication operation mode;

[0022] The pulse width modulation circuit is used to generate a plurality of word line pulse signals according to the operation bit number demand signal of the driving circuit when the RAM array is in the multiplication operation mode;

[0023] The driving circuit is further used to control the potential of the corresponding word lines in the RAM array according to the word line pulse signals to control the discharge of the first target bit line when the RAM array is in the multiplication operation mode, and obtain the change amount of the voltage on the first target bit line. The change amount of the voltage on the first target bit line is the result value of the multiplication operation of the target data stored in the two RAM arrays.

[0024] In some embodiments, the operation mode further includes an addition operation mode;

[0025] The driving circuit is further used to control the potential of the corresponding word lines in the RAM array according to the control signal of the working mode control circuit to control the discharge of the second target bit line when the RAM array is in the addition operation mode, so that the voltage value of the target data storage node in the RAM array is the target voltage value, thereby obtaining the result value of the addition operation.

[0026] In some embodiments, the RAM cell includes a core storage module and a data access module. The core storage module includes a cross-coupled inverter composed of a plurality of transistors. At least two data storage nodes are formed in the cross-coupled inverter, and the data storage nodes store data through their own voltage values;

[0027] The data access module includes a plurality of transistors, a plurality of bit lines, and a plurality of word lines. The bit lines and the word lines are used to control the conduction and cutoff of the transistors in the RAM cell to change the voltage value of the data storage node, so as to read or write data.

[0028] In some embodiments, the core storage module includes a cross-coupled inverter composed of a first transistor, a second transistor, a third transistor, and a fourth transistor. Among them, the first transistor and the second transistor are connected to form a first inverter, the third transistor and the fourth transistor are connected to form a second inverter, and a first data storage node and a second data storage node are formed in the cross-coupled inverter;

[0029] The data access module includes a fifth transistor, a sixth transistor, a seventh transistor, an eighth transistor, a first bit line, a second bit line, a third bit line, a first word line, a second word line, and a source line. The drain of the fifth transistor is connected to the first bit line, the gate is connected to the first word line, and the source is connected to the second inverter; the drain of the sixth transistor is connected to the first inverter, the gate is connected to the first word line, and the source is connected to the drain of the seventh transistor and the gate of the eighth transistor to form a third data storage node; the gate of the seventh transistor is connected to the second word line, and the source is connected to the second bit line; the drain of the eighth transistor is connected to the source line, and the source is connected to the third bit line;

[0030] The first bit line, the second bit line, and the third bit line are used to access the data stored in the first data storage node, the second data storage node, or the third data storage node. The first word line is used to control the conduction or cutoff of the fifth transistor and the sixth transistor to read or write data. The second word line is used to control the conduction or cutoff of the seventh transistor to read or write data.

[0031] In some embodiments, the core storage module includes a cross-coupled inverter composed of a ninth transistor, a tenth transistor, an eleventh transistor, and a twelfth transistor. Among them, the ninth transistor and the tenth transistor are connected to form a third inverter, the tenth transistor and the eleventh transistor are connected to form a fourth inverter, and a fourth data storage node and a fifth data storage node are formed in the cross-coupled inverter;

[0032] The data access module includes a thirteenth transistor, a fourteenth transistor, a fifteenth transistor, a fourth bit line, a fifth bit line, and a third word line. The drain of the thirteenth transistor is connected to the fourth bit line, the gate is connected to the third word line, and the source is connected to the fourth inverter. The drain of the fourteenth transistor is connected to the fifth bit line, the gate is connected to the third word line, and the source is connected to the third inverter. The source of the fifteenth transistor is connected to the power supply voltage, the gate is connected to the enable signal output terminal, and the drain is connected to the sources of the ninth transistor and the eleventh transistor.

[0033] The fourth bit line and the fifth bit line are used to access the data stored in the fourth data storage node or the fifth data storage node, and the third word line is used to control the conduction or cutoff of the thirteenth transistor and the fourteenth transistor to read or write data.

[0034] In a second aspect, an embodiment of the present application provides a storage device, and the storage device includes the memory-computation integrated chip as described in the first aspect.

[0035] The present application provides a memory-computation integrated chip and a storage device. The memory-computation integrated module of the present application is used to store data and perform calculations, effectively shortening the distance of data transmission, reducing the calculation delay, and thus greatly improving the calculation efficiency of the chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 FIG. 15 is a schematic structural diagram of the first implementation manner of the memory-computation integrated chip provided by the embodiment of the present application.

[0037] Figure 2 FIG. 19 is a schematic structural diagram of the second implementation manner of the memory-computation integrated chip provided by the embodiment of the present application.

[0038] Figure 3 FIG. 23 is a schematic structural diagram of the memory-computation integrated unit provided by the embodiment of the present application.

[0039] Figure 4 FIG. 27 is a schematic structural diagram of the basic memory-computation unit provided by the embodiment of the present application.

[0040] Figure 5 FIG. 31 is a schematic structural diagram of the second calculation module provided by the embodiment of the present application.

[0041] Figure 6 FIG. 35 is a schematic circuit structural diagram of the first implementation manner of the RAM unit provided by the embodiment of the present application.

[0042] Figure 7 FIG. 39 is a schematic circuit structural diagram of the second implementation manner of the RAM unit provided by the embodiment of the present application.

[0043] Figure 8It is a schematic structural diagram of a storage device provided by an embodiment of the present application. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts belong to the scope of protection of the present application.

[0045] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.

[0046] The present application provides a computing-in-memory chip and a storage device. The computing-in-memory module is used to store data and perform calculations, effectively shortening the distance of data transmission, reducing calculation latency, and thus greatly improving the calculation efficiency of the chip.

[0047] Next, the computing-in-memory chip provided by the present application will be specifically described in conjunction with the accompanying drawings.

[0048] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of the first implementation manner of the computing-in-memory chip provided by an embodiment of the present application. As Figure 1 shown, in some implementation manners, the computing-in-memory chip 1 includes a computing-in-memory module 10 composed of a plurality of computing-in-memory units 11 connected, a main control module 20, and a circuit network. Figure 1 Only a part of the circuit network connecting each computing-in-memory unit 11 and the main control module 20 is shown in

[0049] but this should not be construed as a limitation to the present application.

[0050] Please refer to Figure 2 ,Figure 2 This is a schematic structural diagram of the second implementation mode of the memory - computing integrated chip provided by the embodiments of the present application. As Figure 2 shown, in some embodiments, each memory - computing integrated unit 11 is connected to at least two other memory - computing integrated units 11.

[0051] Optionally, the memory - computing integrated units 11 are arranged and connected in an array. Each row of memory - computing integrated units 11 includes a plurality of serially connected memory - computing integrated units 11, and each column of memory - computing integrated units 11 includes a plurality of serially connected memory - computing integrated units 11.

[0052] In some embodiments, the memory - computing integrated unit 11 is used to store the weight parameter data of the artificial intelligence model.

[0053] Optionally, the weight parameter data is a weight matrix, a weight vector, or a single weight value, etc.

[0054] In some embodiments, in the memory - computing integrated unit 11, calculation operations such as accumulation, matrix multiplication, or vector dot - product are performed on the weight parameter data.

[0055] In some embodiments, the memory - computing integrated chip 1 further includes a global cache module 30. The global cache module 30 is connected to the memory - computing integrated module 10 through a circuit network. The global cache module 30 is used to cache the latest data received by the memory - computing integrated chip 10.

[0056] Please refer to Figure 3 , Figure 3 This is a schematic structural diagram of the memory - computing integrated unit provided by the embodiments of the present application. As Figure 3 shown, in some embodiments, the memory - computing integrated unit 11 includes a basic memory - computing unit array 110 composed of a plurality of basic memory - computing units 111 connected. The memory - computing integrated unit 11 includes a unit cache module 112, a first computing module 113, a first control module 114, and a model function module 115. The unit cache module 112, the first computing module 113, and the model function module 115 are all connected to the basic memory - computing unit array 110, and the first control module 114 is connected to the model function module 115. The unit cache module 112 is used to cache the latest data received by the memory - computing integrated unit 11, the first computing module 113 is used to perform calculations based on the data in the basic memory - computing unit array 110, the first control module 114 is used to control the operation of the memory - computing integrated unit 11, and the model function module 115 is used to implement the calculation functions required during the operation of the artificial intelligence model based on the data in the basic memory - computing unit array 110.

[0057] In some embodiments, the basic memory - computing unit array 110 can implement the first - step calculation operation, the first computing module 113 can implement the second - step calculation operation, and the model function module 115 can implement the third - step calculation operation.

[0058] Exemplarily, the model function module includes a pooling calculation module and an activation calculation module.

[0059] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the basic memory and computing unit provided by the embodiments of the present application. As Figure 4 shown, in some embodiments, the basic memory and computing unit 111 further includes a second computing module 1112, a basic unit cache module 1113, and a second control module 1114. The second computing module 1112 and the unit cache module 1113 are both connected to the RAM array 1111, and the second control module 1114 is connected to the second computing module 1112. The basic unit cache module 1113 is used to cache the latest data received by the basic memory and computing unit 111, the second computing module 1112 is used to perform calculations based on the data in the RAM array 1111, and the second control module 1114 is used to control the operation of the basic memory and computing unit 111.

[0060] In some embodiments, the second control module 1114 is connected to the RAM array 1111.

[0061] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of the second computing module provided by the embodiments of the present application. As Figure 5 shown, in some embodiments, the second computing module 1112 includes a row decoding circuit 11121, a column decoding circuit 11122, a driving circuit 11123, a read / write control circuit 11124, a working mode control circuit 11125, and a pulse width modulation circuit 11126.

[0062] In some embodiments, the row decoding circuit is connected to each row of RAM cells in the RAM array, the column decoding circuit is connected to each column of RAM cells in the RAM array, the read / write control circuit is connected to the working mode control circuit, the working mode control circuit is connected to the driving circuit, the pulse width modulation circuit is connected to the driving circuit, and the driving circuit is connected to the row decoding circuit and the column decoding circuit.

[0063] In some embodiments, the row decoding circuit is used to control the potential of the word lines of the corresponding row according to the first address signal, and the column decoding circuit is used to control the potential of the word lines of the corresponding column according to the second address signal, so as to select a RAM cell in the RAM array. The driving circuit is used to control the potentials of the word lines and bit lines in the RAM array according to the control signal of the operating mode control circuit, so as to determine the operating mode of the RAM array as the target operating mode. The read-write control circuit is used to send a read signal or a write signal to the operating mode control circuit according to the read-write control signal sent by the second control module. The operating mode control circuit is used to determine the operating mode of the RAM array according to the read signal or the write signal, and the operating modes include a read mode, a write mode, and an arithmetic mode. Among them, the arithmetic mode includes a multiplication arithmetic mode. The pulse width modulation circuit is used to generate a plurality of word line pulse signals according to the operation bit number demand signal of the driving circuit when the RAM array is in the multiplication arithmetic mode. The driving circuit is further used to control the potential of the corresponding word lines in the RAM array according to the word line pulse signals to control the discharge of the first target bit line when the RAM array is in the multiplication arithmetic mode, and obtain the change amount of the voltage on the first target bit line, and the change amount of the voltage on the first target bit line is the result value of the multiplication operation of the target data stored in the two RAM arrays.

[0064] In some embodiments, when the RAM array is in the multiplication arithmetic mode, the driving circuit is used to store the first target data in a column of RAM cells in the RAM array, and then the pulse width modulation circuit generates a plurality of word line pulse signals according to the operation bit number demand signal of the driving circuit. The driving circuit reads the second target data stored in another column of RAM cells in the RAM array, and the data signals of each bit of the second weight vector are serially input into the driving circuit. The driving circuit controls the potential of the word lines of the corresponding RAM cells in the RAM array according to the AND result of the data signal of each bit of the second weight vector and the corresponding word line pulse signal, so as to control the discharge of the first target bit line and obtain the change amount of the voltage on the first target bit line. The change amount of the voltage on the first target bit line is the result value of the multiplication operation of the two target data.

[0065] Optionally, the target data may be a weight vector.

[0066] In some embodiments, each word line pulse signal is composed of signals of multiple sub-stages. The operation time of each sub-stage of each word line pulse signal conforms to from low to high: 2 i t0 = 2 i-1 T s1 = 2 i-2 T s2 = 2 i-2 T s2 = … = 2 0 T si , where Tsi represents the operation time of the i-th sub-stage, and t0 represents the least significant bit pulse width.

[0067] In some embodiments, the operation mode further includes an addition operation mode. The driving circuit is further configured to control the potential of the corresponding word line in the RAM array according to the control signal of the working mode control circuit to control the discharge of the second target bit line when the RAM array is in the addition operation mode, so that the voltage value of the target data storage node in the RAM array is the target voltage value, thereby obtaining the result value of the addition operation.

[0068] By the above method, the RAM array can calculate according to the stored data, and can perform addition calculation and multiplication calculation, so as to realize the multiply-accumulate calculation of the matrix.

[0069] In some embodiments, the RAM cell includes a core storage module and a data access module. The core storage module includes a cross-coupled inverter composed of multiple transistors. At least two data storage nodes are formed in the cross-coupled inverter, and the data storage nodes store data through their own voltage values. Among them, at least two data storage nodes can store at least three numerical combinations.

[0070] In some embodiments, the data access module includes multiple transistors, multiple bit lines and multiple word lines. The bit lines and word lines are used to control the on and off of the transistors in the RAM cell to change the voltage value of the data storage node, so as to read or write data.

[0071] Please refer to Figure 6 , Figure 6 which is a schematic circuit diagram of the first embodiment of the RAM cell provided by the embodiment of the present application. As Figure 6As shown, in some embodiments, the core storage module includes a cross-coupled inverter A1 composed of a first transistor M5, a second transistor M6, a third transistor M7, and a fourth transistor M8. Among them, the first transistor M5 and the second transistor M6 are connected to form a first inverter A11, the third transistor M7 and the fourth transistor M8 are connected to form a second inverter A12, and a first data storage node Vx and a second data storage node Vy are formed in the cross-coupled inverter A1. The data access module includes a fifth transistor M1, a sixth transistor M2, a seventh transistor M3, an eighth transistor M4, a first bit line BL, a second bit line BLB, a third bit line BLR, a first word line WL1, a second word line WL2, and a source line SL. The drain of the fifth transistor M1 is connected to the first bit line BL, the gate is connected to the first word line WL1, and the source is connected to the second inverter A12. The drain of the sixth transistor M2 is connected to the first inverter A11, the gate is connected to the first word line WL1, and the source is connected to the drain of the seventh transistor M3 and the gate of the eighth transistor M4 to form a third data storage node Vz. The gate of the seventh transistor M3 is connected to the second word line WL2, and the source is connected to the second bit line BLB. The drain of the eighth transistor M4 is connected to the source line SL, and the source is connected to the third bit line BLR. The first bit line BL, the second bit line BLB, and the third bit line BLR are used to access the data stored in the first data storage node Vx, the second data storage node Vy, or the third data storage node Vz. The first word line WL1 is used to control the conduction or cutoff of the fifth transistor M1 and the sixth transistor M2 to read or write data, and the second word line WL2 is used to control the conduction or cutoff of the seventh transistor M3 to read or write data.

[0072] As Figure 6 shown, exemplarily, the first transistor M5 is a PMOS transistor, and the second transistor M6 is an NMOS transistor. The source of the first transistor M5 is connected to the power supply voltage Vdd, the gate is connected to the gate of the second transistor M6 and the drain of the sixth transistor M2, and the drain is connected to the drain of the second transistor M6. The source of the second transistor M6 is grounded. The third transistor M7 is a PMOS transistor, and the fourth transistor M8 is an NMOS transistor. The source of the third transistor M7 is connected to the power supply voltage Vdd, the gate is connected to the gate of the fourth transistor M8 and the source of the fifth transistor M1, and the drain is connected to the drain of the fourth transistor M8. The source of the fourth transistor M8 is grounded.

[0073] Exemplarily, the fifth transistor M1, the sixth transistor M2, the seventh transistor M3, and the eighth transistor M4 are all NMOS transistors.

[0074] Exemplarily, the first data storage node Vx is located between the drain of the first transistor M5 and the drain of the second transistor M6, and the voltage at the first data storage node Vx can remain unchanged when the RAM cell A is powered on. The second data storage node Vy is located between the drain of the third transistor M7 and the drain of the eighth transistor M8, and the voltage at the second data storage node Vy can remain unchanged when the RAM cell A is powered on.

[0075] In some embodiments, the operating modes are divided into a read mode, a write mode, and an arithmetic mode in the normal operation mode, and a read mode, a write mode, and an arithmetic mode in the enhanced operation mode. In the normal operation mode, the RAM cell stores SRAM-type stored data through the first data storage node and the second data storage node. In the enhanced operation mode, the RAM cell stores SRAM-type stored data through the first data storage node and the second data storage node, and stores DRAM-type stored data through the third data storage node. At this time, the numerical combinations that the RAM cell can store include "011", "010", "101", and "100", a total of 4 numerical combinations.

[0076] In some embodiments, when the driving circuit controls the RAM cell A to be in a holding state by controlling the row decoding circuit and the column decoding circuit according to the control signal of the operation mode control circuit, the first word line WL1 is pulled down to a low level, while the first bit line BL and the second bit line BLB are pre-charged to a high level, and the fifth transistor M1 and the sixth transistor M2 are turned off, so that the voltage values of the first data storage node and the second data storage node remain unchanged, thereby storing data.

[0077] Optionally, the voltage values of the first data storage node Vx and the second data storage node Vy are one high and one low, and the two data storage nodes are used together to store 1 bit of data.

[0078] Optionally, the voltage values of the first data storage node Vx and the second data storage node Vy are one high and one low, and the two data storage nodes are used together to store 2 bits of data.

[0079] Optionally, the voltage values of the first data storage node Vx and the second data storage node Vy are opposite.

[0080] In some embodiments, in the normal operation mode, when the driving circuit controls the RAM cell A to be in the read mode by controlling the row decoding circuit and the column decoding circuit according to the control signal of the operation mode control circuit, the first word line WL1 and the second word line WL2 are pre-charged to a high level, the fifth transistor M1 and the sixth transistor M2 are turned on, the first bit line BL is discharged to the ground through the first inverter A11, while the second bit line BLB remains at a high level, and there is a voltage difference between the first bit line BL and the second bit line BLB. This voltage difference is captured by the amplifier circuit in the driving circuit and then amplified and output, thereby completing the read operation.

[0081] In some embodiments, in the normal operation mode, when the driving circuit controls the RAM cell A to be in the write mode by controlling the row decoding circuit and the column decoding circuit according to the control signal of the operation mode control circuit, the first word line WL1 and the second word line WL2 are pre-charged to a high level, the fifth transistor M1 and the first transistor M5 are turned on, and data is written by inputting a write control signal to the first bit line BL and the second bit line BLB. For example, if you want to write the data "0" into the RAM cell, the first bit line BL is connected to a low level, and the second bit line BLB is connected to a high level. Before the write operation starts, the data stored in the first data storage node Vx is "1", and the data stored in the second data storage node Vy is "0". After the write operation starts, the fifth transistor M1 and the first transistor M5 are turned on. In the cross-coupled inverter A1, the voltage values of the first data storage node Vx and the second data storage node Vy will flip, so the data stored in the two data storage nodes will flip. Then the data stored in the first data storage node Vx is "0", and the data stored in the second data storage node Vy is "1", and the data is written.

[0082] In some embodiments, in the normal operation mode, the source line SL and the third bit line BLR are kept at 0V. This can ensure that no current flows through the eighth transistor M4 regardless of the voltage value at the third data storage node Vz.

[0083] In some embodiments, in the enhanced operation mode, the RAM cell A can store one more bit of data through the third data storage node Vz. At this time, the first word line WL1 and the second word line WL2 are both pre-charged to a high level to turn on the seventh transistor M3 and the eighth transistor M4. In the write mode, the data write signal changes the voltage value of the third data storage node Vz through the third bit line BLB and the seventh transistor M3, thereby writing the data. During the write operation, the source line is kept at 0V, and the third bit line BLR is also discharged to 0V.

[0084] In some embodiments, in the read mode of the enhanced operation mode, the source line is pulled high to a high level, and the data stored in the third data storage node Vz is read by sensing the cumulative value of the voltage change on the third bit line during discharge.

[0085] Optionally, the seventh transistor M3 and the eighth transistor M4 form an embedded DRAM cell, while the other transistors in the RAM cell A form an SRAM cell. The embedded DRAM cell can store independent data in a dynamic manner, while static data is stored in the SRAM cell.

[0086] It can be understood that since the third data storage node Vz is in the read or write circuit of the SRAM data, the DRAM data stored on the third data storage node Vz will be destroyed during the read or write operation of the SRAM data stored on the first data storage node Vx and the second data storage node Vy. In some embodiments, this problem can be circumvented by a specific data writing mode. For example, in a specific data writing mode, the SRAM data is written first and then the DRAM data is written, and the DRAM data is read first and then the SRAM data is read. In this way, two bits of data can be stored in the RAM cell, improving the storage efficiency.

[0087] Exemplarily, in a deep learning network, the first data storage node Vx and the second data storage node Vy can be used to store weight parameter data, while the third data storage node Vz is used to store input activation parameter data.

[0088] Please refer to Figure 7 , Figure 7 is a schematic circuit diagram of the second embodiment of the RAM cell provided by the embodiment of the present application. As Figure 7As shown, in some embodiments, the core storage module includes a cross-coupled inverter A2 composed of a ninth transistor M9, a tenth transistor M10, an eleventh transistor M11, and a twelfth transistor M12. Among them, the ninth transistor M9 and the tenth transistor M10 are connected to form a third inverter A21, the tenth transistor M10 and the eleventh transistor M11 are connected to form a fourth inverter A22, and a fourth data storage node Q and a fifth data storage node Qb are formed in the cross-coupled inverter A2. The data access module includes a thirteenth transistor M13, a fourteenth transistor M14, a fifteenth transistor M15, a fourth bit line BL1, a fifth bit line BLB1, and a third word line WL3. The drain of the thirteenth transistor M13 is connected to the fourth bit line BL1, the gate is connected to the third word line WL3, and the source is connected to the fourth inverter A22. The drain of the fourteenth transistor M14 is connected to the fifth bit line BLB1, the gate is connected to the third word line WL3, and the source is connected to the third inverter A21. The source of the fifteenth transistor M15 is connected to the power supply voltage, the gate is connected to the enable signal output terminal ENb, and the drain is connected to the source of the ninth transistor M9 and the source of the eleventh transistor M11. The fourth bit line BL1 and the fifth bit line BLB1 are used to access the data stored in the fourth data storage node Q or the fifth data storage node Qb. The third word line WL3 is used to control the conduction or cutoff of the thirteenth transistor M13 and the fourteenth transistor M14 to read or write data.

[0089] In some embodiments, the ninth transistor M9 and the eleventh transistor M11 are PMOS transistors, and the tenth transistor M10, the eleventh transistor M11, the twelfth transistor M12, the thirteenth transistor M13, the fourteenth transistor M14, and the fifteenth transistor M15 are all NMOS transistors.

[0090] In some embodiments, in the general operation mode, the enable signal output from the enable signal output terminal ENb turns on the fifteenth transistor M15. At this time, the working process of the RAM unit of the second embodiment can refer to the working process of the RAM unit of the first embodiment above.

[0091] In some embodiments, in the enhanced operation mode, the enable signal output from the enable signal output terminal ENb turns off the fifteenth transistor M15. At this time, since the cross-coupled inverter A1 is no longer connected to the power supply voltage, the positive feedback of the cross-coupled inverter A1 is weakened. In this way, three different numerical combinations can be stored in the fourth data storage node Q and the fifth data storage node Qb, and each circuit state corresponds to one of the numerical combinations. The three different numerical combinations are (1, 0), (0, 1), and (0, 0).

[0092] The method of writing the numerical combination (1, 0) or (0, 1) refers to the working process of the RAM unit of the first embodiment above.

[0093] Specifically, when writing the numerical combination (0, 0), the third word line WL3 is pulled high to a high level, and the fourth bit line BL1 and the fifth bit line BLB1 are pulled to 0V. Both the fourth data storage node Q and the fifth data storage node Qb are discharged, thereby turning off the tenth transistor M10 and the twelfth transistor M12. Since the fifteenth transistor M15 is turned off, the fourth data storage node Q and the fifth data storage node Qb are not connected to the power supply voltage. Therefore, the voltage values of both the fourth data storage node Q and the fifth data storage node Qb are 0.

[0094] It should be noted that in the prior art, it is impossible to store the numerical combination (0, 0) because there is a strong positive feedback between two cross-coupled inverters, and the strong positive feedback will always make the voltage values of the two data storage nodes one high and one low. Through the above method, the RAM unit can store three numerical combinations, thereby improving the storage efficiency.

[0095] The embodiment of the present application further provides a control method for a memory-computation integrated chip, which is applied to the main control module of the memory-computation integrated chip as described above. The control method of the memory-computation integrated chip includes steps S100 to step S300.

[0096] Step S100: Obtain the usage information of the memory-computation integrated chip.

[0097] Among them, the usage information is used to represent the usage of the artificial intelligence model deployed on the memory-computation integrated chip.

[0098] Exemplarily, the usage information is to perform natural language processing, image processing, collaborative filtering, or autonomous driving sensor data processing, etc. using the artificial intelligence model. Among them, collaborative filtering refers to matching users and commodities in an intelligent recommendation system to recommend commodities to users.

[0099] Step S200: Determine the target data storage mode according to the usage information and the amount of data to be processed.

[0100] Among them, the target data storage mode is to dynamically select the general operation mode or the enhanced operation mode for storage and calculation, and adopt the corresponding target data storage structure in the general operation mode or the enhanced operation mode.

[0101] In some embodiments, the RAM unit is used to store the weight parameter data of the artificial intelligence model. The weight parameter data may include a weight parameter matrix and an intermediate activation value.

[0102] In some embodiments, when the RAM unit adopts the circuit structure of the first embodiment as Figure 6 shown, step S200 includes steps S201 to S202.

[0103] Step S201: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for natural language processing, it is determined that the first target data storage structure adopted in the enhanced operation mode is to store the numerical bits and sign bits separately.

[0104] When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for natural language processing, the weight parameter matrix needs to be accessed frequently, and multiplication and addition calculations are performed on the weight parameter matrix. At this time, the intermediate activation values of the artificial intelligence model are mostly sparse binary data, and the binary data is 0 or 1.

[0105] In some embodiments, when the RAM unit adopts the circuit structure of the first embodiment as Figure 6 shown, the third data storage node is used to store the numerical bits, and the first data storage node and the second data storage node are used to store the sign bits.

[0106] As described above, since the third data storage node Vz is in the read or write circuit of the SRAM data, the DRAM data stored on the third data storage node Vz will be destroyed during the read or write operation of the SRAM data stored on the first data storage node Vx and the second data storage node Vy. In some embodiments, this problem can be avoided by a specific data writing mode. For example, in the specific data writing mode, the SRAM data is written first and then the DRAM data is written, and the DRAM data is read first and then the SRAM data is read. At this time, the numerical bits are stored as DRAM data and the sign bits are stored as SRAM data.

[0107] Step S202: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for natural language processing, it is determined that the second target data storage structure adopted in the general operation mode is to use the first data storage node and the second data storage node to store the intermediate activation values.

[0108] Among them, the first data storage node and the second data storage node are jointly used to store 1 bit of the intermediate activation value.

[0109] In some embodiments, when the RAM unit adopts the circuit structure of the first embodiment as Figure 6 shown, step S200 includes steps S203 to S204.

[0110] Step S203: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for image processing, it is determined that the first target data storage structure adopted in the enhanced operation mode is to hierarchically store the gray values of the image.

[0111] At this time, the RAM unit is also used to store the image data to be processed in the artificial intelligence model.

[0112] In some embodiments, the third data storage node is used to store the first-level gray value, and the first data storage node and the second data storage node are used to store the second-level gray value.

[0113] Exemplarily, the 2-bit data that a RAM unit can store are 00, 01, 10, and 11, and the gray values represented by 00, 01, 10, and 11 increase gradually. Among them, the data of the first bit from left to right is the first-level gray value, and the data of the second bit is the second-level gray value. For example, in "01", "0" is the first-level gray value and "1" is the second-level gray value.

[0114] In some embodiments, when the requirement for processing accuracy is not high, only the first-level gray value in the RAM unit needs to be read for processing. When the requirement for processing accuracy is high, all the data of the two-level gray values in the RAM unit can be read. In this way, the calculation speed can be increased when the requirement for processing accuracy is not high.

[0115] Step S204: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for image processing, it is determined that the second target data storage structure adopted in the general operation mode is to use the first data storage node and the second data storage node to store filter coefficients.

[0116] Exemplarily, the filter coefficient is the enable flag of the convolution kernel and is single-bit data.

[0117] In some embodiments, when the RAM unit adopts the circuit structure of the first embodiment as shown in Figure 6 Step S200 includes steps S205 to S206.

[0118] Step S205: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for collaborative filtering, it is determined that the first target data storage structure adopted in the enhanced operation mode is to use the third data storage node to store the sign bit and use the first data storage node and the second data storage node to store the numerical bit.

[0119] When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for collaborative filtering, the in-memory computing chip needs to store and calculate large-scale sparse matrices, so it is necessary to efficiently store non-zero elements. In addition, the calculation process involves frequent row and column scans and logical judgments, and non-zero element positioning is required. Therefore, using the third data storage node to store the sign bit and using the first data storage node and the second data storage node to store the numerical bit can facilitate reading the sign bit first when reading data to determine whether the stored element is non-zero.

[0120] Step S206: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for collaborative filtering and the RAM unit is used to store matrix index pointer data, it is determined that the general operation mode is adopted for the RAM unit, and the second target data storage structure adopted in the general operation mode is to use the first data storage node and the second data storage node to store matrix index pointer data.

[0121] In some embodiments, when the RAM unit adopts the circuit structure of the first embodiment as shown in Figure 6 Figure 1, step S200 includes steps S207 to S208.

[0122] Step S207: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for autonomous driving sensor data processing, it is determined that the first target data storage structure adopted in the enhanced operation mode is to use the first data storage node, the second data storage node, and the third data storage node to store matrix elements.

[0123] Step S208: When it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for autonomous driving sensor data processing, it is determined that the second target data storage structure adopted in the general operation mode is to use the first data storage node and the second data storage node to store sensor fault flags.

[0124] When the artificial intelligence model is used for autonomous driving sensor data processing, a large amount of sensor data needs to be fused. Therefore, it is necessary to determine whether the sensor is faulty according to a large number of sensor fault flags, and then the sensor data can be read and calculated. The sensor fault flag is generally single-bit Boolean data. Therefore, in the general operation mode, the first data storage node and the second data storage node can be used to store sensor fault flags to reduce energy consumption.

[0125] Step S300: Control the in-memory computing module to store data and perform calculations using the target data storage mode.

[0126] In some embodiments, the general operation mode or the enhanced operation mode in the target data storage mode is adopted for storage and calculation based on the type of current calculation.

[0127] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for natural language processing and the type of current calculation is to calculate the weight parameter matrix, the enhanced operation mode is adopted for storage and calculation.

[0128] In some embodiments, when the enhanced operation mode is selected for storage and calculation, the third data storage node is used to store the numerical bits, and the first data storage node and the second data storage node are used to store the sign bits. When writing data, the sign bits are written first and then the numerical bits. When reading data, the numerical bits are read first for the first-step calculation, and then the sign bits are read for the second-step calculation.

[0129] Optionally, the sign bit is 0 or 1.

[0130] Optionally, 0 can be used to indicate that the data in the RAM cell is non-negative, and 1 can be used to indicate that the data in the RAM cell is negative.

[0131] Optionally, 1 can be used to indicate that the data in the RAM cell is non-negative, and 0 can be used to indicate that the data in the RAM cell is negative.

[0132] In some embodiments, when reading data, the numerical bits are read first for the first-step multiply-accumulate calculation, and then the sign bits are read for the second-step calculation to adjust the sign of the operation result. In this way, the calculation efficiency can be improved without damaging the numerical bits.

[0133] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for natural language processing and the current calculation type is calculating intermediate activation values, the general operation mode is adopted for storage and calculation. In this way, the power consumption of the chip can be reduced.

[0134] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for image processing and the current calculation type is image calculation, only the first-level gray value stored in the third data storage node is read according to the calculation accuracy requirement, or all the second-level gray value data in the RAM cell is read.

[0135] Exemplarily, when performing edge detection calculation, the calculation accuracy requirement is only 1-bit gray value data. At this time, only the first-level gray value stored in the third data storage node is read and calculated. When performing image convolution calculation, the calculation accuracy requirement is 2-bit gray value data. At this time, the first-level gray value stored in the third data storage node is read first, and then the second-level gray value stored in the first data storage node and the second data storage node is read and calculated.

[0136] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for image processing and the current calculation type is filter coefficient calculation, the filter coefficients stored in the first data storage node and the second data storage node are read and calculated.

[0137] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for collaborative filtering and the current type of calculation is matrix calculation, first read the sign bit stored in the third data storage node. When it is determined according to the sign bit that the data in the RAM unit is non-zero, then read the numerical bits stored in the first data storage node and the second data storage node and perform calculations.

[0138] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for collaborative filtering and the current type of calculation is to select elements in the matrix, read the matrix index pointer data stored in the first data storage node and the second data storage node.

[0139] In some embodiments, the matrix index pointer data can be multi-bit, and one RAM unit can be used to store 1 bit of matrix index pointer data. After reading the matrix index pointer data stored in multiple RAM units, the enhanced operation mode can be switched to and the matrix elements in the corresponding RAM units can be read and calculated.

[0140] In some embodiments, when it is determined according to the usage information that the artificial intelligence model deployed on the in-memory computing chip is used for autonomous driving sensor data processing and the current type of calculation is sensor fault detection, read the sensor fault flags stored in the first data storage node and the second data storage node. Then, when it is determined according to the sensor fault flags that the sensors for acquiring data are working properly, the enhanced operation mode can be switched to and the sensor data in the RAM unit can be read.

[0141] In some embodiments, when the RAM unit adopts the circuit structure of the second embodiment as Figure 7 shown, the data can be read simultaneously and calculated.

[0142] In summary, the in-memory computing chip provided by the embodiments of the present application has the following advantages:

[0143] 1. By using the in-memory computing module to store data and perform calculations, the distance of data transmission is effectively shortened, the calculation latency can be reduced, and thus the calculation efficiency of the chip can be greatly improved.

[0144] 2. By means of the operation modes including the multiplication operation mode and the addition operation mode, the RAM array can perform calculations according to the stored data, and addition calculations and multiplication calculations can be performed, so as to realize the multiply-accumulate calculation of the matrix, enhancing the calculation function of the RAM array.

[0145] 3. By means of the RAM unit including the embedded DRAM unit and the SRAM unit, two-bit data can be stored in the RAM unit, improving the storage efficiency.

[0146] 4. Through the second embodiment of the RAM unit, the RAM unit can store three numerical combinations, thereby improving the storage efficiency.

[0147] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the storage device provided by the embodiment of the present application. As Figure 8 shown, the present application also provides a storage device 2, and the storage device 2 includes the above-mentioned computing-in-memory chip 1.

[0148] In some embodiments, the storage device is used to deploy an artificial intelligence model.

[0149] In summary, the present application provides a computing-in-memory chip and a storage device. The computing-in-memory chip includes: a computing-in-memory module formed by connecting a plurality of computing-in-memory units, and the computing-in-memory module is used to store data and perform calculations; the computing-in-memory unit includes a basic computing-in-memory unit array formed by connecting a plurality of basic computing-in-memory units, each computing-in-memory unit is connected to at least two other computing-in-memory units, each basic computing-in-memory unit is connected to at least two other basic computing-in-memory units, and the basic computing-in-memory unit includes a RAM array composed of a plurality of RAM units; a main control module for managing the calculation tasks executed by the computing-in-memory chip; a circuit network for connecting a plurality of computing-in-memory units in the computing-in-memory module and connecting the main control module to the computing-in-memory module. By using the computing-in-memory module to store data and perform calculations, the present application effectively shortens the distance of data transmission, can reduce the calculation latency, and thus greatly improves the calculation efficiency of the chip.

[0150] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A processing-in-memory chip, characterized in that, Comprising: A memory - computing integrated module composed of multiple memory - computing integrated units connected together, where the memory - computing integrated module is used to store data and perform calculations; The memory - computing integrated unit includes a basic memory - computing unit array composed of multiple basic memory - computing units connected together. Each memory - computing integrated unit is connected to at least two other memory - computing integrated units, and each basic memory - computing unit is connected to at least two other basic memory - computing units. The basic memory - computing unit includes a RAM array composed of multiple RAM units; A main control module for managing the computing tasks executed by the memory - computing integrated chip; A circuit network that is used to connect multiple memory - computing integrated units in the memory - computing integrated module and connect the main control module to the memory - computing integrated module.

2. The memory - computing integrated chip according to claim 1, wherein: The memory - computing integrated chip further includes a global cache module. The global cache module is connected to the memory - computing integrated module through the circuit network, and the global cache module is used to cache the latest data received by the memory - computing integrated chip.

3. The memory - computing integrated chip according to claim 1, wherein: The memory - computing integrated unit further includes a unit cache module, a first computing module, a first control module, and a model function module; The unit cache module, the first computing module, and the model function module are all connected to the basic memory - computing unit array, and the first control module is connected to the model function module; The unit cache module is used to cache the latest data received by the memory - computing integrated unit, the first computing module is used to perform calculations based on the data in the basic memory - computing unit array, the first control module is used to control the operation of the memory - computing integrated unit, and the model function module is used to implement the computing functions required during the operation of the artificial intelligence model based on the data in the basic memory - computing unit array.

4. The memory - computing integrated chip according to claim 1, wherein: The basic memory - computing unit further includes a basic unit cache module, a second computing module, and a second control module. The second computing module and the unit cache module are both connected to the RAM array, the second control module is connected to the second computing module, the basic unit cache module is used to cache the latest data received by the basic memory - computing unit, the second computing module is used to perform calculations based on the data in the RAM array, and the second control module is used to control the operation of the basic memory - computing unit.

5. The memory - computing integrated chip according to claim 4, wherein: The second computing module includes a row decoding circuit, a column decoding circuit, a driving circuit, a read - write control circuit, a working mode control circuit, and a pulse width modulation circuit; The row decoding circuit is connected to each row of RAM units in the RAM array, the column decoding circuit is connected to each column of RAM units in the RAM array, the read - write control circuit is connected to the working mode control circuit, the working mode control circuit is connected to the driving circuit, the pulse width modulation circuit is connected to the driving circuit, and the driving circuit is connected to the row decoding circuit and the column decoding circuit; The row decoding circuit is used to control the potential of the word line of the corresponding row according to the first address signal, and the column decoding circuit is used to control the potential of the word line of the corresponding column according to the second address signal, so as to select a RAM cell in the RAM array; The driving circuit is used to control the potentials of the word lines and bit lines in the RAM array by controlling the row decoding circuit and the column decoding circuit according to the control signal of the working mode control circuit, so as to determine the working mode of the RAM array as the target working mode; The read-write control circuit is used to send a read signal or a write signal to the working mode control circuit according to the read-write control signal sent by the second control module; The working mode control circuit is used to determine the working mode of the RAM array according to the read signal or the write signal. The working modes include a read mode, a write mode, and an arithmetic mode, and the arithmetic mode includes a multiplication arithmetic mode; The pulse width modulation circuit is used to generate a plurality of word line pulse signals according to the operation bit number demand signal of the driving circuit when the RAM array is in the multiplication arithmetic mode; The driving circuit is further used to control the potential of the corresponding word line in the RAM array according to the word line pulse signal to control the discharge of the first target bit line when the RAM array is in the multiplication arithmetic mode, so as to obtain the change amount of the voltage on the first target bit line. The change amount of the voltage on the first target bit line is the result value of the multiplication operation of the target data stored in the two RAM arrays.

6. The in-memory computing chip according to claim 5, wherein The arithmetic mode further includes an addition arithmetic mode; The driving circuit is further used to control the potential of the corresponding word line in the RAM array according to the control signal of the working mode control circuit to control the discharge of the second target bit line when the RAM array is in the addition arithmetic mode, so that the voltage value of the target data storage node in the RAM array is the target voltage value, thereby obtaining the result value of the addition operation.

7. The in-memory computing chip according to claim 1, wherein The RAM cell includes a core storage module and a data access module. The core storage module includes a cross-coupled inverter composed of a plurality of transistors. At least two data storage nodes are formed in the cross-coupled inverter, and the data storage nodes store data through their own voltage values; The data access module includes a plurality of transistors, a plurality of bit lines, and a plurality of word lines. The bit lines and the word lines are used to control the conduction and cutoff of the transistors in the RAM cell to change the voltage values of the data storage nodes, so as to read or write data.

8. The in-memory computing chip according to claim 7, wherein The core storage module includes a cross-coupled inverter composed of a first transistor, a second transistor, a third transistor, and a fourth transistor. Among them, the first transistor and the second transistor are connected to form a first inverter, the third transistor and the fourth transistor are connected to form a second inverter, and a first data storage node and a second data storage node are formed in the cross-coupled inverter; The data access module includes a fifth transistor, a sixth transistor, a seventh transistor, and an eighth transistor 、 a first bit line, a second bit line, a third bit line, a first word line, a second word line, and a source line. The drain of the fifth transistor is connected to the first bit line, the gate is connected to the first word line, and the source is connected to the second inverter; the drain of the sixth transistor is connected to the first inverter, the gate is connected to the first word line, and the source is connected to the drain of the seventh transistor and the gate of the eighth transistor to form a third data storage node; the gate of the seventh transistor is connected to the second word line, and the source is connected to the second bit line; the drain of the eighth transistor is connected to the source line, and the source is connected to the third bit line; The first bit line, the second bit line, and the third bit line are used to access the data stored in the first data storage node, the second data storage node, or the third data storage node. The first word line is used to control the conduction or cutoff of the fifth transistor and the sixth transistor to read or write data. The second word line is used to control the conduction or cutoff of the seventh transistor to read or write data.

9. The computing-in-memory chip according to claim 7, wherein The core storage module includes a cross-coupled inverter composed of a ninth transistor, a tenth transistor, an eleventh transistor, and a twelfth transistor. Among them, the ninth transistor and the tenth transistor are connected to form a third inverter, the tenth transistor and the eleventh transistor are connected to form a fourth inverter, and a fourth data storage node and a fifth data storage node are formed in the cross-coupled inverter; The data access module includes a thirteenth transistor, a fourteenth transistor, a fifteenth transistor, a fourth bit line, a fifth bit line, and a third word line. The drain of the thirteenth transistor is connected to the fourth bit line, the gate is connected to the third word line, and the source is connected to the fourth inverter; the drain of the fourteenth transistor is connected to the fifth bit line, the gate is connected to the third word line, and the source is connected to the third inverter; the source of the fifteenth transistor is connected to the power supply voltage, the gate is connected to the enable signal output terminal, and the drain is connected to the source of the ninth transistor and the source of the eleventh transistor; The fourth bit line and the fifth bit line are used to access the data stored in the fourth data storage node or the fifth data storage node. The third word line is used to control the conduction or cutoff of the thirteenth transistor and the fourteenth transistor to read or write data.

10. A storage device, characterized in that, The storage device includes the computing-in-memory chip according to any one of claims 1 to 9.

Citation Information

Cited By

  • Storage and calculation integrated high-speed external interface design method for storage unit buffer

    CN122111915A