Simulation in-memory computing macro architecture and data processing method and device
By introducing input buffers and dynamic buffers into the analog in-memory computing macro architecture, and combining vector operations with capacitor compensation arrays, the problem of insufficient driving capability of analog in-memory computing macros is solved, improving running speed and computing power, and achieving efficient data processing.
Patent Information
- Application Number
- CN202511280738.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2026-01-13
AI Technical Summary
Simulated in-memory computation macros suffer from insufficient driving capability in large-scale parallel computing tasks, affecting their running speed and computing power.
It adopts an analog in-memory computing macro architecture, including an input buffer, several levels of analog in-memory computing core arrays and a capacitor compensation array. The input buffer provides initial drive, the dynamic buffer continues to provide drive for the input features, and the capacitor compensation array performs vector operations. The operation results are converted into digital signals by an analog-to-digital converter.
It improves the running speed and computing power of simulated in-memory computing macros, alleviates the problem of insufficient driving capability, and improves the overall performance and energy efficiency of the system.
Smart Images

Figure CN121328640A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent computing, and in particular to a simulated in-memory computing macro architecture, a data processing method and device. BACKGROUND
[0002] An artificial intelligence accelerator has characteristics such as high dense operation and high parallelism. The memory wall is a key factor limiting the high throughput and high energy efficiency ratio of the artificial intelligence accelerator. In the related art, the in-memory computing architecture is an effective solution to break through the memory wall. However, the simulated in-memory computing macro usually integrates a large-scale storage computing unit matrix to achieve efficient in-memory computing. In order to enhance the driving capability, a driver is usually integrated at the input end of the simulated in-memory computing macro. However, for the input units far away, even if there is a driver, the driving capability may still be insufficient. This deficiency limits the running speed and computing power of the simulated in-memory computing macro, and further affects its performance in large-scale parallel computing tasks. SUMMARY
[0003] The present application aims to at least partly solve one of the problems in the prior art.
[0004] To this end, the present application aims to provide an efficient simulated in-memory computing macro architecture, a data processing method and device.
[0005] In order to achieve the above technical purpose, the technical solutions adopted by the embodiments of the present application include the following aspects:
[0006] On the one hand, the embodiments of the present application provide a simulated in-memory computing macro architecture, comprising: an input buffer for providing initial driving for input features; a plurality of levels of simulated in-memory computing core arrays, a dynamic buffer being arranged between adjacent two levels of simulated in-memory computing core arrays, the dynamic buffer being used for transmitting the input features and providing driving for the input features; a primary simulated in-memory computing core array for receiving input features through the input buffer; a non-primary simulated in-memory computing core array for receiving input features through the simulated in-memory computing core array of the previous level and the previous dynamic buffer; a capacitance compensation array connected with each level of the simulated in-memory computing core array, the capacitance compensation array being used for performing vector operation on the multi-bit weight and the input features in the corresponding level of the simulated in-memory computing core array, and converting the operation result into a digital signal through an analog-to-digital converter. The input buffer provides initial driving for the input features, and the dynamic buffer continues to provide driving for the input features, which is conducive to alleviating the problem of insufficient driving capability in the simulated in-memory computing macro, and improving the running speed and computing power of the simulated in-memory computing macro.
[0007] In addition, the simulated in-memory computing macro architecture according to the above embodiments of the present application can also have the following additional technical features:
[0008] The analog in-memory computing macro architecture of the embodiment of the present application, the dynamic buffer includes a source follower and a level shifter, and the level shifter is used to lift the voltage of the input feature.
[0009] In an embodiment of the present application, the analog in-memory computing macro architecture further includes: a digital-to-analog converter used to convert the digital input feature into a voltage input feature and transmit the voltage input feature to the input buffer.
[0010] In an embodiment of the present application, the analog in-memory computing core array includes a multi-bit weight unit, each bit weight unit is used to receive an input feature corresponding to a bit, and the capacitance compensation array is used to perform vector operation on the multi-bit weight and input feature in the corresponding stage of the analog in-memory computing core array, including: the capacitance compensation array is used to perform an analog vector product operation on the multi-bit weight and multi-bit input feature in the corresponding stage.
[0011] In an embodiment of the present application, the weight unit in each current stage of the analog in-memory computing core array transmits the input feature to the weight unit of the same bit in the next stage of the analog in-memory computing core array through the dynamic buffer corresponding to the bit.
[0012] In another aspect, the embodiment of the present application proposes a data processing method applied to the above-mentioned analog in-memory computing macro architecture, the method comprising:
[0013] Obtaining an input feature;
[0014] Transmitting the input feature to each stage of the analog in-memory computing core array in turn, performing vector operation through each stage of the capacitance compensation array to obtain the operation result of each stage;
[0015] Simultaneously receiving the next input feature as a new input feature, and returning to the step of obtaining the input feature.
[0016] Further, the data processing method of the embodiment of the present application, the transmitting the input feature to each stage of the analog in-memory computing core array in turn, performing vector operation through each stage of the capacitance compensation array to obtain the operation result of each stage, comprises:
[0017] If the current stage of the capacitance compensation array obtains the operation result of the current stage, at the same time, the current stage of the analog in-memory computing core array transmits the input feature of the current stage to the next stage of the analog in-memory computing core array through the dynamic buffer, until the input feature of the current stage is transmitted to the last stage of the analog in-memory computing core array.
[0018] Further, the data processing method of the embodiment of the present application, the receiving the next input feature as a new input feature, comprises:
[0019] The second-stage operation result is obtained through a second-stage capacitor compensation array, and meanwhile, the input buffer receives the next input feature as a new input feature.
[0020] In another aspect, an embodiment of the present application provides an artificial intelligence accelerator, comprising the analog in-memory computing macro architecture.
[0021] In another aspect, an embodiment of the present application provides a data processing apparatus, comprising the artificial intelligence accelerator.
[0022] The present application provides an analog in-memory computing macro architecture, a data processing method and apparatus, wherein the analog in-memory computing macro architecture comprises: an input buffer configured to provide initial driving for an input feature; a plurality of levels of analog in-memory computing core arrays, wherein a dynamic buffer is arranged between two adjacent levels of analog in-memory computing core arrays, and the dynamic buffer is configured to transmit and provide driving for the input feature; a primary analog in-memory computing core array configured to receive the input feature through the input buffer; a non-primary analog in-memory computing core array configured to receive the input feature through a previous level of analog in-memory computing core array and a previous dynamic buffer; and a capacitor compensation array connected to each level of the analog in-memory computing core array, wherein the capacitor compensation array is configured to perform vector operation on a plurality of weight bits and the input feature in the corresponding level of analog in-memory computing core array, and convert the operation result into a digital signal through an analog-to-digital converter. The input buffer provides initial driving for the input feature, and the dynamic buffer continues to provide driving for the input feature, which is conducive to alleviating the problem of insufficient driving capability in the analog in-memory computing macro, and improving the running speed and computing power of the analog in-memory computing macro. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the drawings in the following introduction are only for the convenience of clearly describing part of the embodiments of the technical solutions in the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0024] Figure 1 The structural schematic diagram of an embodiment of the analog in-memory computing macro architecture provided by the present application;
[0025] Figure 2 The flowchart of an embodiment of the data processing method provided by the present application. DETAILED DESCRIPTION
[0026] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0027] Currently, AI accelerators are characterized by high computational density and high parallelism. The memory wall is a key factor limiting the high throughput and energy efficiency of AI accelerators. At present, in-memory computing architecture and systolic dataflow are two effective solutions to overcome the memory wall limitation.
[0028] Systolic dataflow is a high-efficiency computing architecture based on regularized data flow. Its core idea is to achieve rhythmic, synchronous data transmission and parallel computation through locally interconnected arrays of Processing Elements (PEs). This architecture mimics the "contraction-relaxation" rhythm of a biological heart, allowing data to flow between adjacent PEs along pre-defined paths while simultaneously completing computational tasks. Its typical structure consists of multiple PEs arranged in a one-dimensional or two-dimensional topology. Each PE communicates directly only with its neighboring units, and its operation is synchronized with a global clock, forming a highly pipelined computational flow.
[0029] Analog in-memory computing is a novel computing paradigm that breaks through the memory wall. Its core idea is to perform computational operations directly within memory cells using analog signals, eliminating frequent data transfer between the processor and memory, thereby significantly improving energy efficiency and computing density. This technology utilizes the physical characteristics of storage devices (such as resistance, capacitance, and current / voltage relationships) to achieve multiply-accumulate (MAC) operations, making it particularly suitable for data-intensive tasks such as neural networks and signal processing. Compared to digital in-memory computing, analog computing offers high energy efficiency and a simpler structure in low-precision computing environments. Analog in-memory computing is more suitable for energy-efficient edge computing and neuromorphic computing applications.
[0030] Therefore, Systolic Dataflow and Analog In-Memory Computing (AIMC) are two key technologies for overcoming the memory wall bottleneck and significantly improving system energy efficiency and throughput. However, these two technologies still face some key problems and barriers in their respective application scenarios.
[0031] 1. Power Consumption Challenges of Pulsating Data Streams: In pulsating data stream architectures, each processing unit needs to buffer the current input data (such as input activations or weights) for the next level of data transfer. This buffer and the corresponding control logic are easily implemented in digital processing units, making pulsating data streams widely used in digital AI accelerators. However, in low-precision applications (such as edge computing), the power consumption of digital machine learning accelerators has become a critical issue that urgently needs to be addressed. Traditional digital implementations are inadequate for low-power requirements, limiting their further application in scenarios with extremely high energy efficiency demands, such as edge computing.
[0032] 2. Driving Capability Bottleneck of Simulated In-Memory Computation: Simulated in-memory computation macros typically integrate large-scale matrices of memory computing units to achieve efficient in-memory computation. To enhance driving capability, a driver is usually integrated at the input end of the simulated in-memory computation macro. However, for input units that are far apart, even with a driver, insufficient driving capability may still be encountered. This deficiency limits the running speed and computing power of simulated in-memory computation macros, thus affecting their performance in large-scale parallel computing tasks.
[0033] In light of the aforementioned issues, combining pulsating data streams with analog in-memory computing holds promise for breaking down the key barriers between the two technologies and achieving complementary advantages. The efficient data flow mechanism of pulsating data streams can provide optimized data scheduling support for analog in-memory computing, improving data transmission efficiency and alleviating the problem of insufficient driving capabilities in analog in-memory computing macros. Simultaneously, the low-power characteristics of analog in-memory computing can effectively reduce power consumption in the pulsating data stream architecture, making it more suitable for low-precision, low-power edge computing applications. This combination not only fully leverages the synergistic effect of both technologies but also provides a more efficient solution for low-power applications such as edge computing, powerfully promoting the practical application and development of related technologies.
[0034] The following describes in detail, with reference to the accompanying drawings, the simulated in-memory computing macroarchitecture proposed according to an embodiment of the present invention. See also Figure 1 The present invention provides a simulated in-memory computing macroarchitecture including:
[0035] An input buffer is used to provide initial drive for input features;
[0036] The system comprises several levels of analog in-memory computing core arrays, with a dynamic buffer between adjacent levels. The dynamic buffer is used to transmit the input features and provide a drive for the input features. The primary analog in-memory computing core array is used to receive the input features via the input buffer. The non-primary analog in-memory computing core array is used to receive the input features via the previous level analog in-memory computing core array and the preceding dynamic buffer.
[0037] A capacitor compensation array is connected to the analog in-memory computing core array of each level. The capacitor compensation array is used to perform vector operations on the multi-bit weights and input features in the analog in-memory computing core array of the corresponding level, and convert the operation results into digital signals through an analog-to-digital converter.
[0038] An input buffer is used to receive input features and provide initial drive for them, ensuring that subsequent computing units can effectively receive and process signals. The input features in this application can be analog features, such as voltage input features converted from digital signals. The number of stages in the analog in-memory computing core array in this application can be set according to actual needs, and this application does not impose a specific limitation. Figure 1 The array consists of a 64-level analog in-memory computing core array. A dynamic buffer provides the drive for the transfer of input features between the analog in-memory computing core arrays. The capacitor-compensated array is used to perform vector operations on multi-bit weights and input features in the corresponding levels of the analog in-memory computing core arrays connected to it.
[0039] The analog in-memory computing macroarchitecture of this invention includes a dynamic buffer comprising a source follower and a level shifter, wherein the level shifter is used to boost the voltage of the input feature.
[0040] The source follower in this application can be of any form, such as Figure 1 In this system, transistors are used. Level shifters are implemented using capacitors and resistors.
[0041] In one embodiment of the present invention, the analog in-memory computing macroarchitecture further includes: a digital-to-analog converter, the digital-to-analog converter being used to convert digital input characteristics into voltage input characteristics and transmit the voltage input characteristics to the input buffer.
[0042] In this application, the input features transmitted and processed in the input buffer, the multi-stage analog in-memory computing core array, and the capacitor compensation array are all in the form of voltage.
[0043] In one embodiment of the present invention, the analog in-memory computing core array includes a multi-bit weight unit, each weight unit being used to receive the input feature of the corresponding bit, and the capacitor compensation array being used to perform vector operations on the multi-bit weights and input features in the corresponding level of the analog in-memory computing core array, including: the capacitor compensation array being used to perform analog vector product operations on the multi-bit weights and multi-bit input features in the corresponding level.
[0044] The number of bits in the weight units of the simulated in-memory computing core array in this application can be set according to actual needs, such as... Figure 1 In this analog in-memory computing core array, there are 64-bit weight units, and it also receives 64-bit input features.
[0045] In one embodiment of the present invention, each weight unit in the current level analog in-memory computing core array transmits the input feature to the weight unit of the same bit in the next level analog in-memory computing core array through the dynamic buffer of the corresponding bit.
[0046] On the other hand, embodiments of the present invention propose a data processing method applied to the aforementioned simulated in-memory computing macroarchitecture, referring to... Figure 2 As shown, the method includes:
[0047] Step S100: Obtain input features;
[0048] Step S200: The input features are sequentially transmitted to each level of analog in-memory computing core array, and vector operations are performed through each level of capacitor compensation array to obtain the operation results at each level;
[0049] Step S300: Simultaneously receive the next input feature as a new input feature, and return to the step of obtaining the input feature.
[0050] In this application, the data undergoes vector operations on one hand, and on the other hand, while the vector operations are being performed, the input features are being passed along the analog memory in-processing core array at each level; at the same time, new input features are being received.
[0051] Furthermore, in the data processing method of this embodiment of the invention, the step of sequentially transmitting the input features to each level of analog in-memory computing core array, and performing vector operations through each level of capacitor compensation array to obtain the operation results at each level includes:
[0052] If the current-level capacitor compensation array obtains the current-level operation result, the current-level analog in-memory computing core array will simultaneously transmit the current-level input features to the next-level analog in-memory computing core array through a dynamic buffer, until the current-level input features are transmitted to the last-level analog in-memory computing core array.
[0053] Furthermore, in the data processing method of this embodiment of the invention, receiving the next input feature as a new input feature includes:
[0054] The second-stage operation result is obtained through the second-stage capacitor compensation array. At the same time, the input buffer receives the next input feature as a new input feature.
[0055] The following detailed description of the simulated in-memory computing macroarchitecture and data processing method provided in this application uses specific embodiments:
[0056] Simulated in-memory computing macros with pulsating data streams represent a promising computing architecture that holds the promise of overcoming existing technological bottlenecks, more effectively breaking through strong storage limitations, and significantly improving system energy efficiency and computing power. However, at the technical implementation level, the following key issues remain:
[0057] 1. Buffer design issues in traditional pulsating data streams
[0058] In traditional pulsating dataflow architectures, buffers are typically integrated within the processing unit to temporarily store data and pass it to the next processing unit. Analog buffers are mostly built based on source-level followers, which suffer from two major problems: high static power consumption and significant swing decay. Static power consumption significantly reduces the system's energy efficiency ratio, while swing decay affects computational accuracy, thereby reducing the overall system performance.
[0059] 2. Data processing timing optimization issues
[0060] In analog in-memory computation macros with pulsating data streams, precise control of data processing timing is crucial for improving system performance. Efficient buffer integration is particularly important for charge-based analog in-memory computation macros, directly impacting data transfer efficiency, overall computing power, and throughput. Therefore, optimizing buffer integration to achieve efficient data transfer is a pressing technical challenge.
[0061] Figure 1 A schematic diagram of a high-performance, charge-based analog-in-memory computing macroarchitecture based on pulsating dataflow is shown. Structurally, this computing macroarchitecture mainly includes an analog input buffer, an analog computing-in-memory array, a multi-bit capacitor compensation array, and an analog dynamic buffer equipped with a level shifter. The analog input buffer, located at the input, provides initial driving force for the input analog voltage to ensure that subsequent computing units can effectively receive and process signals. Each stage of the analog dynamic buffer consists of a source follower and a level shifter. This analog dynamic buffer alleviates the driving requirements of the corresponding computing units on the input activation features, thereby reducing the requirement for initial input driving strength and contributing to higher system throughput. The capacitor compensation array performs vector product operations on the input features and multi-bit weights, and is connected to an analog-to-digital converter (ADC). The ADC converts the analog vector product result into a digital signal.
[0062] From a workflow perspective, the voltage-to-digital converter (DAC) first converts the input digital feature information into a voltage signal. An input buffer then enhances the driving capability of the input voltage signal to meet subsequent processing requirements. This input signal is then processed with the first-stage multi-bit weights. Each stage's capacitor compensation array is responsible for accumulating the current multi-bit vector operation results. The vector product result is then passed to the corresponding stage's ADC, which converts it into a digital signal, thus deriving the vector product of the current input feature and a set of weights. Additionally, the weights within the in-memory computation macro are pre-trained and pre-written in software. Simultaneously, the analog input signal is laterally passed to the next stage's analog dynamic buffer, providing input activation features for the next set of multi-bit weights. Furthermore, the analog dynamic buffer enhances the driving capability of this stage, ensuring stable performance at higher operating frequencies. Compared to ordinary analog buffers, this dynamic buffer integrates clock control functions (such as...). Figure 1 As shown), this effectively avoids static power consumption. Simultaneously, the analog buffer also includes a level shifter (such as...). Figure 1 As shown, the second stage (as shown) is used to boost the output voltage to maintain its stability and ensure computational accuracy. The analog input signal of the second stage is calculated with the stage's weights. The result is converted into a voltage through a supplementary capacitor array, and then converted into a digital signal by an ADC. This digital signal represents the result of a multi-bit vector product operation. At this point, the first stage receives a new input signal, while the second stage passes the processed input signal to the next stage's computational unit. Therefore, the analog input signal sequentially passes through each stage's dynamic buffer before being passed to the next stage for vector product operations. The vector product result of each stage is converted into a corresponding voltage through a capacitor compensation array, and this voltage is then converted into a corresponding digital signal by an ADC.
[0063] This design combines pulsating data streams with charge-based analog in-memory computing to achieve efficient data processing and computation. It fully leverages the low-power advantage of analog in-memory computing while ensuring orderly data flow and efficient transmission through the efficient data scheduling mechanism of pulsating data streams. This combination not only improves the overall system performance but also provides strong support for low-power, high-efficiency computing tasks, making it particularly suitable for energy-intensive applications such as edge computing.
[0064] Compared with the prior art, the present invention has significant advantages in the following two key aspects:
[0065] 1. High-performance, charge-type analog in-memory computing macro
[0066] In traditional charge-based analog in-memory computing macros, relying solely on input drivers is insufficient to meet the demands of large-scale computing units, resulting in limited overall system throughput and computing power. This invention effectively solves this critical problem by introducing a pulsating data flow mechanism. By setting analog buffers between stages, optimized data scheduling support is provided for analog in-memory computing, significantly improving data transmission efficiency and alleviating the problem of insufficient driving capability in analog in-memory computing macros. Therefore, system throughput and computing power are significantly improved, enabling more efficient fulfillment of intelligent computing needs.
[0067] 2. Power consumption optimization and amplitude preservation of analog buffers
[0068] The static power consumption of traditional analog buffers reduces the overall system energy efficiency, while amplitude attenuation also affects the accuracy of in-memory computation and overall prediction accuracy. This invention employs a dynamic analog buffer, effectively suppressing static power consumption, improving system energy efficiency, and enhancing computing power. Furthermore, the integrated level shifter effectively compensates for amplitude attenuation, ensuring the system's computational accuracy.
[0069] In summary, compared with traditional solutions, the technology of this invention can bring higher energy efficiency and computing power to analog in-memory computing (CIM) systems, significantly improving the overall performance of the system. The introduction of this technology will provide stronger support for edge-side artificial intelligence computing and promote the efficient development of related applications.
[0070] This invention provides an integrated implementation of pulsating data streams. For charge-based analog in-memory computation macros, the key technical challenge lies in achieving energy-efficient data transmission across multiple computational units.
[0071] This invention provides a high-efficiency analog buffer design. Analog buffer design is crucial and challenging for multi-level computing arrays. A high-efficiency analog buffer improves system computing power while ensuring system energy efficiency.
[0072] This invention provides an analog buffer design based on a level shifter. Level shifters are crucial for buffers used in analog in-memory computations. They effectively suppress the swing attenuation caused by voltage buffers, ensuring the overall macro computational performance of in-memory computations.
[0073] In this invention, the input activation data is converted into an analog voltage signal by a digital-analog converter (DAC). This voltage data signal flows to the next stage under the action of an interstage buffer, and the driving capability of the system is also improved. A key feature of this invention is its use of a pulsed data flow charge-type analog in-memory computing macroarchitecture.
[0074] This invention provides an extension of the output data distribution of the in-memory computing core. Based on a level-conversion analog dynamic buffer, it enhances the inter-stage driving capability of the analog in-memory computing macro and the system computing power while suppressing voltage swing attenuation, which is another major feature of this invention.
[0075] On the other hand, embodiments of the present invention provide an artificial intelligence accelerator, including the above-described analog in-memory computing macroarchitecture.
[0076] It is evident that the content of the above-described architecture embodiments is applicable to this accelerator embodiment. The specific functions implemented in this accelerator embodiment are the same as those in the above-described architecture embodiments, and the beneficial effects achieved are also the same as those achieved in the above-described architecture embodiments.
[0077] On the other hand, embodiments of the present invention provide a data processing apparatus, including the aforementioned artificial intelligence accelerator.
[0078] It is evident that the content of the above-described architecture embodiments is applicable to this device embodiment. The specific functions implemented in this device embodiment are the same as those in the above-described architecture embodiments, and the beneficial effects achieved are also the same as those achieved in the above-described architecture embodiments.
[0079] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.
[0080] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.
[0081] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0082] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable programs for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, a program execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can retrieve and execute a program from or in conjunction with such a program execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit a program for use by or in conjunction with a program execution system, apparatus, or device.
[0083] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0084] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0085] In the foregoing description of this specification, references to terms such as "one embodiment," "another embodiment," or "some embodiments" indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0086] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0087] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.
Claims
1. A simulated in-memory computing macroarchitecture, characterized in that, The simulated in-memory computing macroarchitecture includes: An input buffer is used to provide initial drive for input features; The system comprises several levels of analog in-memory computing core arrays, with a dynamic buffer between adjacent levels. The dynamic buffer is used to transmit the input features and provide a drive for the input features. The primary analog in-memory computing core array is used to receive the input features via the input buffer. The non-primary analog in-memory computing core array is used to receive the input features via the previous level analog in-memory computing core array and the preceding dynamic buffer. A capacitor compensation array is connected to the analog in-memory computing core array of each level. The capacitor compensation array is used to perform vector operations on the multi-bit weights and input features in the analog in-memory computing core array of the corresponding level, and convert the operation results into digital signals through an analog-to-digital converter.
2. The analog in-memory computing macroarchitecture according to claim 1, characterized in that, The dynamic buffer includes a source follower and a level shifter, the level shifter being used to boost the voltage of the input characteristic.
3. The analog in-memory computing macroarchitecture according to claim 1, characterized in that, The analog in-memory computing macroarchitecture further includes a digital-to-analog converter (DAC) for converting digital input characteristics into voltage input characteristics and transmitting the voltage input characteristics to the input buffer.
4. The analog in-memory computing macroarchitecture according to claim 1, characterized in that, The analog in-memory computing core array includes multiple weight units, each weight unit is used to receive the input feature of the corresponding bit, and the capacitor compensation array is used to perform vector operations on the multiple weights and input features in the corresponding level of the analog in-memory computing core array, including: the capacitor compensation array is used to perform analog vector product operations on the multiple weights and multiple input features in the corresponding level.
5. The analog in-memory computing macroarchitecture according to claim 2, characterized in that, Each weight unit in the current level of the analog in-memory computing core array transmits the input features to the weight unit of the same bit in the next level of the analog in-memory computing core array through the dynamic buffer of the corresponding bit.
6. A data processing method, characterized in that, Applied to any one of the analog in-memory computing macroarchitectures as described in claims 1 to 5, the method comprises: Obtain input features; The input features are sequentially transmitted to each level of analog in-memory computing core array, and vector operations are performed through each level of capacitor compensation array to obtain the operation results at each level; Simultaneously, the next input feature is received as a new input feature, and the process returns to the step of obtaining the input feature.
7. The data processing method according to claim 6, characterized in that, The process of sequentially transmitting the input features to various levels of analog in-memory computing core arrays, performing vector operations through various levels of capacitance compensation arrays, and obtaining the operation results at each level includes: If the current-level capacitor compensation array obtains the current-level operation result, the current-level analog in-memory computing core array will simultaneously transmit the current-level input features to the next-level analog in-memory computing core array through a dynamic buffer, until the current-level input features are transmitted to the last-level analog in-memory computing core array.
8. The data processing method according to claim 6, characterized in that, The receiving of the next input feature as a new input feature includes: The second-level operation result is obtained through the second-level capacitor compensation array. At the same time, the input buffer receives the next input feature as a new input feature.
9. An artificial intelligence accelerator, characterized in that, The artificial intelligence accelerator includes an analog in-memory computing macroarchitecture as described in any one of claims 1 to 5.
10. A data processing apparatus, characterized in that, The data processing device includes the artificial intelligence accelerator as described in claim 9.