Neural network accelerator and neural network acceleration method based on storage-computation integration

CN122840142APending Publication Date: 2026-09-29SUZHOU KUANWEN ELECTRONICS SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611307788.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明创造实施例提供的一种基于存算一体的神经网络加速器和神经网络加速方法,至少解决相关技术中基于存算一体的神经网络加速器的算法映射效率低、通用性差的问题

Benefits of technology

[0015]本发明创造实施例提供的基于存算一体的神经网络加速器,将存算集群拆分为小簇(多个存算宏单元独立配置使能开关)与大簇(多个存算宏单元共享一个使能开关),所有使能开关均接入调度子系统,可实现分级启停控制,根据神经网络不同层的输入通道规模动态匹配算力资源;算力链路为存算计算集群、内部加法树、第一激活量化模块、外部加法树、第二激活量化模块,形成两级串行运算处理链路,可支持尺寸大于集群容量的网络层,拓展了可加速的网络范围;通过算力分级适配、存储分时复用、调度分层自动化的协同设计,一方面从算力、存储、调度三个维度适配不同规模、不同结构的神经网络,另一方面通过自动化的权重加载、数据流调度、算力匹配,有助于降低算法映射的调度开销,提升存算资源利用率。能够解决相关技术中基于存算一体的神经网络加速器的算法映射效率低、通用性差的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840142A_ABST
    Figure CN122840142A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of neural networks, and provides a neural network accelerator based on a memory-computing integrated system and a neural network acceleration method, which comprise a memory-computing computing cluster, an internal addition tree, a first activation quantization module, an external addition tree and a second activation quantization module which are sequentially connected, the memory-computing computing cluster comprises a small cluster cluster and a large cluster cluster, the small cluster cluster comprises a plurality of first memory-computing macro units each of which is configured with an enabling switch, the large cluster cluster comprises a plurality of second memory-computing macro units which share one enabling switch, and each enabling switch is connected to a scheduling subsystem; a weight update manager which is connected with a weight storage and the memory-computing computing cluster; a feature storage which is connected with the memory-computing computing cluster and the external addition tree; and a scheduling subsystem which is connected with the control ends of the above parts. The application solves the problems of low algorithm mapping efficiency and poor universality of the neural network accelerator based on the memory-computing integrated system in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of neural network technology, and in particular to a neural network accelerator and neural network acceleration method based on in-memory computing. Background Technology

[0002] The demands on computing power, power consumption, and real-time performance for edge deep learning inference continue to increase. Traditional von Neumann architecture suffers from severe memory walls, and frequent data transfers lead to high power consumption and low computing power bottlenecks.

[0003] Compute-in-memory (CIM) performs multiplication and accumulation operations within the storage array, significantly reducing data migration overhead and becoming a mainstream solution for edge AI acceleration. Various CIM macros have implemented basic operations such as convolution and fully connected operations. However, CIM-based neural network accelerators suffer from the following problems: a single CIM macro cannot handle the entire neural network computation process, and the fixed configuration of the CIM unit makes it difficult to dynamically allocate computing resources, resulting in low algorithm mapping efficiency; there is no corresponding computation scheduling for special network structures such as ResNet skip connections, leading to poor versatility. Summary of the Invention

[0004] The present invention provides a neural network accelerator and neural network acceleration method based on in-memory computing, which at least solves the problems of low algorithm mapping efficiency and poor versatility of neural network accelerators based on in-memory computing in related technologies.

[0005] This invention provides an in-memory computing-based neural network accelerator, comprising: a computing power subsystem, including an in-memory computing cluster, an internal addition tree, a first activation quantization module, an external addition tree, and a second activation quantization module connected in sequence; wherein the in-memory computing cluster includes a small cluster and a large cluster; the small cluster includes multiple first in-memory macro units, each configured with an enable switch; the large cluster includes multiple second in-memory macro units, all sharing an enable switch; and each enable switch is connected to a scheduling subsystem; a storage subsystem, including a weight memory, a feature memory, and a weight update manager; the weight update manager is connected to the weight memory and the in-memory computing cluster, and the feature memory is connected to the in-memory computing cluster and the external addition tree; and a scheduling subsystem, connected to the control terminals of each part of the computing power subsystem and the storage subsystem, for managing the start / stop and data flow of the corresponding parts.

[0006] As an optional solution, the aforementioned neural network accelerator is configured with three states: system idle, system update, and system computation. In the system idle state, the scheduling subsystem continuously monitors weight update signals and computation start signals. When the scheduling subsystem detects the weight update signal, it enters the system update state. The weight update manager reads weight data from the weight memory and writes it to the in-memory computing cluster. After writing, it sends an update completion signal back to the scheduling subsystem. When the scheduling subsystem detects both the update completion signal and the computation start signal, it enters the system computation state. The scheduling subsystem controls the feature memory to load input features into the in-memory computing cluster, performs neural network inference computation, and outputs the computation results.

[0007] As an optional approach, the scheduling subsystem controls the feature memory to load input features into the in-memory computing cluster, executes neural network inference calculations, and outputs the calculation results. This includes: initialization; calculating the size of the current layer input features and the size of the in-memory computing cluster; if the size of the current layer input features is less than or equal to the size of the in-memory computing cluster, the feature memory loads the current layer input features into the in-memory computing cluster, executes neural network inference calculations, and outputs the calculation results; if the size of the current layer input features is greater than the size of the in-memory computing cluster, the current layer input features are loaded into the in-memory computing cluster in batches, and the calculation results corresponding to each batch are accumulated through the external addition tree and then output, wherein the calculation results corresponding to each batch are temporarily stored in the feature memory.

[0008] As an optional approach, if the size of the current layer input features is greater than the size of the in-memory computing cluster, the quotient of the size of the current layer input features divided by the size of the in-memory computing cluster is taken as the total number of batches. If there is a remainder, the quotient is incremented by one to take the total number of batches.

[0009] As an optional solution, the aforementioned scheduling subsystem incorporates a sub-scheduling state machine, which includes sequentially switching sub-idle states, input feature loading states, computation states, and output feature storage states. The sub-idle state is used for initialization. The input feature loading state is used to control the feature memory to load the current layer input features into the in-memory computing cluster. The computation state is used to manage the in-memory computing cluster and perform neural network inference operations. The output feature storage state is used to control the final computation result to be stored back into the feature memory.

[0010] As an optional solution, an instruction first-in-first-out (FIFO) cache unit is also included; the instruction FIFO cache unit is connected to the scheduling subsystem and provides configuration information for the neural network layer, which is used to generate the weight update signal and the computation start signal.

[0011] As an optional solution, the aforementioned small cluster includes 9 first in-memory macro units, and the aforementioned large cluster includes 27 second in-memory macro units. The capacity of the aforementioned first in-memory macro units and the aforementioned second in-memory macro units is 64×64 bits, and both use bit serial transmission. The aforementioned weight memory is 256KB static random access memory, equipped with a 32-bit read / write interface for the CPU and a 64-bit weight loading interface for the aforementioned in-memory computing cluster. The aforementioned feature memory is 16KB static random access memory, configured with 8 sets of mutually exclusive enabled access interfaces, of which 3 sets of access interfaces are used for input feature access, 2 sets of access interfaces are used for intermediate calculation result access, and 3 sets of access interfaces are used for output feature access.

[0012] As an optional approach, the aforementioned weight update manager pre-stores the base address of the weight memory and the base address of the in-memory computing cluster. When the aforementioned scheduling subsystem detects a weight update signal, the aforementioned weight update manager generates a weight read address sequence based on the aforementioned weight memory base address, which is used to read weight data from the aforementioned weight memory with a bit width of 64 bits. The aforementioned weight update manager generates a weight write address sequence based on the aforementioned in-memory computing cluster base address, which is used to write weight data to the aforementioned in-memory computing cluster with a bit width of 64 bits.

[0013] As an optional solution, the connection between the feature memory and the external addition tree includes: a first path for transmitting the intermediate results of historical batches temporarily stored in the feature memory to the external addition tree as an accumulation input; and a second path for transmitting the accumulation results output by the external addition tree back to the feature memory for storage.

[0014] This invention also provides a neural network acceleration method based on any of the aforementioned neural network accelerators. The method includes the following steps: decoding the configuration instructions of the current neural network layer to obtain configuration information; based on the configuration information, reading the corresponding weight data from the weight memory through the weight update manager and writing it into the in-memory computing cluster; reading input feature data from the feature memory and sending it into the in-memory computing cluster to output the convolution operation result; performing accumulation, activation, pooling, and quantization processing on the convolution operation results of different channels to obtain the output feature data of the current batch; writing the output feature data into the feature memory as the input feature of the next neural network layer; repeating the above steps until the inference computation of all neural network layers is completed.

[0015] The in-memory computing-based neural network accelerator provided in this invention splits the in-memory computing cluster into small clusters (multiple in-memory computing macrounits with independently configured enable switches) and large clusters (multiple in-memory computing macrounits sharing one enable switch). All enable switches are connected to a scheduling subsystem, enabling hierarchical start / stop control and dynamic matching of computing resources based on the input channel size of different layers of the neural network. The computing power link consists of an in-memory computing cluster, an internal addition tree, a first activation quantization module, an external addition tree, and a second activation quantization module, forming a two-level serial computation processing link that can support network layers larger than the cluster capacity, expanding the range of networks that can be accelerated. Through a collaborative design of hierarchical computing power adaptation, time-sharing storage reuse, and hierarchical automated scheduling, it adapts to neural networks of different sizes and structures from three dimensions: computing power, storage, and scheduling. Furthermore, through automated weight loading, data flow scheduling, and computing power matching, it helps reduce the scheduling overhead of algorithm mapping and improve the utilization rate of in-memory computing resources. This solves the problems of low algorithm mapping efficiency and poor versatility in related technologies for in-memory computing-based neural network accelerators. Attached Figure Description

[0016] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other embodiments based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of a module of a neural network accelerator based on in-memory computing in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram illustrating the working principle of a neural network accelerator based on in-memory computing in an embodiment of the present invention.

[0019] Figure 3 This is a flowchart illustrating the steps of a neural network acceleration method in an embodiment of the present invention. Detailed Implementation

[0020] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0021] Compute-in-memory (CIM) performs multiplication and accumulation operations within the storage array, significantly reducing data migration overhead and becoming a mainstream solution for edge AI acceleration. Various CIM macros have implemented basic operations such as convolution and fully connected operations. However, CIM-based neural network accelerators suffer from the following problems: a single CIM macro cannot handle the entire neural network computation process, and the fixed configuration of the CIM unit makes it difficult to dynamically allocate computing resources, resulting in low algorithm mapping efficiency; there is no corresponding computation scheduling for special network structures such as ResNet skip connections, leading to poor versatility.

[0022] Therefore, please refer to Figure 1 As shown in the figure, the present invention provides a neural network accelerator based on in-memory computing, including a computing subsystem, a storage subsystem and a scheduling subsystem.

[0023] The computing power subsystem includes a storage computing cluster 11, an internal addition tree 12, a first activation quantization module 13, an external addition tree 14, and a second activation quantization module 15 connected in sequence. The storage computing cluster 11 includes a small cluster and a large cluster. The small cluster includes multiple first storage computing macro units, each of which is configured with an enable switch. The large cluster includes multiple second storage computing macro units, all of which share an enable switch. Each enable switch is connected to the scheduling subsystem 19.

[0024] The storage subsystem includes a weight memory 16, a feature memory 17, and a weight update manager 18. The weight update manager 18 is connected to the weight memory 16 and the in-memory computing cluster 11, respectively. The feature memory 17 is connected to the in-memory computing cluster 11 and the external addition tree 14, respectively.

[0025] The scheduling subsystem 19 is connected to the control terminals of each part of the above-mentioned computing power subsystem and storage subsystem, and is used to manage the start-up, shutdown and data flow of the corresponding parts.

[0026] This invention provides an AI accelerator architecture suitable for edge-side neural network acceleration. Its core mainly includes a computing subsystem, a storage subsystem, and a scheduling subsystem. This architecture is adaptable to in-memory computing macros with reconfigurable bit widths and dynamically adjusts the number of in-memory macros participating in the computation and the computation data flow according to the size of different input neural networks. This architecture can be widely used in the design of AI accelerators based on in-memory macros, improving the mapping efficiency of neural network algorithms and expanding the application scenarios of in-memory computing.

[0027] Regarding in-memory computing clusters, the first in-memory macro unit of a small cluster and the second in-memory macro unit of a large cluster are both in-memory macro units, but they differ in functional division and enable switch control.

[0028] Taking a computing cluster containing 4×9 in-memory macros as an example, it is mainly responsible for the multiplication and accumulation calculations of the neural network. It consists of 4×9 in-memory macros, each of which is 64x64=4Kb in size. The data transmission method is bit serial, and a single macro can receive input from a maximum of 64 input channels. This will be further explained in this embodiment.

[0029] To facilitate power consumption optimization, the entire Cluster (in-memory computing cluster) is divided into small and large clusters. The small cluster consists of 9 Macros (9 First In-Memory Macro Units), capable of processing the initial layer of a neural network with 3 input channels. Each of the 9 First In-Memory Macro Units has an independent enable switch for fine-grained control. The large cluster comprises 27 CIM Macros (27 Second In-Memory Macro Units), sharing a single enable switch. When the number of input channels exceeds 64, both the small and large clusters are activated simultaneously for collaborative processing, maximizing computing power.

[0030] In addition to in-memory computing based on the in-memory computing cluster, the computing subsystem also performs computations according to the structure of a neural network. Taking a typical convolutional neural network as an example, it usually includes the following layers: convolutional layer (first convolutional layer), activation layer, pooling layer, quantization layer, convolutional layer (second convolutional layer), etc. When accelerating the convolutional neural network based on the accelerator provided in this embodiment, each module also follows this structure. First, the in-memory computing cluster is used to complete the convolution operation with a large computational scale. Second, since convolution is a 4-dimensional convolution, the convolution results of different channels need to be added together. At this time, an addition tree is needed to accumulate the convolution results of different channels. The accumulated data is sent to the activation layer for activation, for example, data less than 0 is directly set to 0. The activated data is sent to the pooling module for pooling, which mainly reduces the computational scale. The pooling result is sent to quantization, for example, reconstructing the 23-bit computation result into a commonly used 8-bit or 16-bit computational bit width. Otherwise, the bit width will become longer and longer, which is not conducive to the utilization of computing resources. In practice, pooling operations can be merged into convolution operations, that is, pooling operations are completed during the convolution process. Therefore, the adder provided in this embodiment does not specifically limit the pooling module.

[0031] In other words, the adder provided in this embodiment adds the results of the convolution kernels through the internal adder tree to initially form output features; activates and quantizes the calculation results of the current layer through the first activation quantization module; accumulates the quantization results of the current layer and the quantization results of the previous layer or the remaining parts that were not calculated in one go through the external adder tree; and activates and quantizes the calculation results of the external adder tree through the second activation quantization module.

[0032] The adder provided in this embodiment additionally provides an external adder tree and a second activation quantization module as external operators compared to related technologies. When the edge-side in-memory neural network processor (Neural Network Processor, AI Accelerator, or NPU) cannot complete the computation of an entire layer, or when the computation results of two adjacent layers need to be added, such as in the "skip connection" in ResNet, the previous computation result of the NPU will be temporarily stored in the feature memory and processed together with the current computation result in the external operators. Another reason for additionally setting up external operators is to facilitate the subsequent pipelined operation of internal and external operators, which is suitable for multi-layer network scenarios.

[0033] For example, the weight memory 16 uses a 256KB SRAM and provides three sets of interfaces: wm_load and wm_store signals for CPU read and write, both with a data width of 32 bits; and wm_cim_load signal for in-memory computing, with a data width of 64 bits, which mainly controls the weight writing of the in-memory computing cluster.

[0034] For example, the feature memory uses a 16KB SRAM. Input features, output features, and intermediate results share this feature memory. The feature memory provides eight sets of interfaces, each of which cannot be enabled simultaneously. Specifically, the input feature `ifm` is allocated three sets of interfaces, the output feature `ofm` is allocated three sets of interfaces, and the intermediate result `orb` is allocated two sets of interfaces. The intermediate results are primarily used to handle cases where computation cannot be completed in one round. This will be further explained in this embodiment.

[0035] The weight update manager primarily receives weight update signals. Based on the number of input / output feature channels, the maximum capacity of the in-memory macrocells and the cluster, it retrieves data from the weight memory and writes the weight data to the in-memory computing cluster. `wis_Max_ifC`, `wis_Max_ofC`, `ref_cim_cx_addr_base`, and `ref_wm_addr_base` are the pre-prepared feature configurations and addresses of the in-memory and weight memory. When the weight update signal `weight_update_start` is high, the weight update manager enters the weight loading state. It is pulled high on the next rising edge of the clock (`wm_cim_load_en`), at which point it begins reading 64 bits of weight data from the weight memory.

[0036] The capacity of an in-memory computing cluster is equal to the capacity of a single in-memory macrocell multiplied by the number of in-memory macrocells. For example, if a single in-memory macrocell can perform a 64×64 bit (4kb) convolution operation, then the capacity of the in-memory computing cluster is equal to 64. 64 Number of computational macrounits. Taking image features as an example, the size of the input features is equal to the size of a single channel (length). high) Number of channels. All input features are stored in the feature memory. During computation, the weight manager calculates the size of the current layer's input features and the computational size of the in-memory computing cluster. If the size of the current layer's input features is less than or equal to the computational size of the in-memory computing cluster, all input features of the current layer are retrieved from the feature memory and written to the in-memory computing cluster. If the size of the current layer's input features is greater than the computational size of the in-memory computing cluster, the size of the input features is divided by the size of the in-memory computing cluster, and then written to the in-memory computing cluster in batches. That is, after each batch of computation is completed, the input features of the current layer that have not yet been computed are reloaded and rewritten to the in-memory computing cluster for data refresh.

[0037] The adder provided in this embodiment uses a hierarchical scheduler to control the state of the NPU system. The entire system includes three states: system idle, system update, and system computation. Among them, the system computation state can be divided into four sub-states: sub-idle, input feature loading, computation, and output feature storage.

[0038] When the NPU is in an idle state, the system checks whether the weight update signal and the computation start signal are activated. If the weight update signal is detected, the system state switches to the system update state. At this time, the weight update manager (weight controller) begins transferring data from the weight storage to the in-memory computing cluster. After the transfer is complete, an update completion signal is issued, and the system state switches to the system computation state.

[0039] The system computation state includes a subsystem scheduler with four states: sub_idle, sub_ifmloader, sub_compute, and sub_ofmstore. The subsystem scheduler is initially initialized in the sub_idle state. When an input feature loading signal is detected, it switches to the sub_ifmloader state and loads the input feature data from the feature memory. After loading, it switches to sub_compute, at which point the input features are sent to the in-memory computing cluster for inference computation. If the size of a layer is smaller than the size supported by the in-memory computing cluster, the scheduler switches to the sub_ofmstore state and stores the computation result (output feature) in the feature memory before starting computation for the next layer. If the size of a layer is larger than the range supported by the in-memory computing cluster, the computation result is temporarily stored in the feature memory, and the system scheduler switches back to the system update state. The weight update manager updates the remaining weight data for that layer, and the computation controller (computation manager) also loads the remaining feature data into the in-memory computing cluster. Subsequently, the calculation results of these remaining data are sent to the addition layer (external addition tree) in the external arithmetic unit (external operator) and added to the result previously stored in the feature memory. The result after addition is processed sequentially through the ReLU layer and the quantization layer (second activated quantization module). After the external arithmetic unit completes the calculation, the final result is stored in the feature memory.

[0040] For example, Figure 2 This is a schematic diagram of the working principle of the adder provided in this embodiment, which includes three stages: decoding configuration instructions, loading the weights of the neural network, and starting calculation to complete inference.

[0041] In the first stage, the AI ​​accelerator decodes configuration information from the instruction FIFO. These instructions contain configuration information such as data precision, operator type, and input / output features for each layer of the neural network. After the computation of each layer is completed, a layer completion flag is sent to the instruction controller. Then, the upper layer of the system sends the configuration information for the next layer into the FIFO, and reads the instructions for the next layer from the FIFO for decoding, until the computation of all layers is completed.

[0042] In the second phase, the weight controller, also known as the weight refresh manager, reads weight data from the weight storage and writes it to the in-memory computing cluster (CIM Cluster). The amount of weight data required depends on the size of the in-memory computing cluster and the scale of each network layer. For example, if the in-memory computing cluster consists of 36 macrocells, each 4Kb, then the total weight data for the cluster is 144Kb. When the required weight data exceeds the capacity of the in-memory computing cluster, the weight controller will continue to transfer the remaining weight data to the cluster after the previous calculation is completed.

[0043] In the third phase, the Compute Manager, part of the scheduling subsystem, first enables or disables relevant computing units based on the configuration information of the current layer. Then, the Compute Manager reads feature data from the feature memory and begins managing the data streams generated by the enabled computing units during computation. After computation is complete, the output feature data is written to the Route Memory and used as input features for the next layer.

[0044] The accelerator provided in this invention splits the in-memory computing cluster into small clusters (multiple in-memory macrounits with independently configured enable switches) and large clusters (multiple in-memory macrounits sharing one enable switch). All enable switches are connected to the scheduling subsystem, enabling hierarchical start / stop control and dynamic matching of computing resources based on the input channel size of different layers of the neural network. The computing power link consists of an in-memory computing cluster, an internal addition tree, a first activation quantization module, an external addition tree, and a second activation quantization module, forming a two-level serial computation processing link that can support network layers larger than the cluster capacity, expanding the range of networks that can be accelerated. Through the collaborative design of hierarchical computing power adaptation, time-sharing storage reuse, and hierarchical automated scheduling, it adapts to neural networks of different sizes and structures from three dimensions: computing power, storage, and scheduling. Furthermore, through automated weight loading, data flow scheduling, and computing power matching, it helps reduce the scheduling overhead of algorithm mapping and improve the utilization rate of in-memory computing resources. This solves the problems of low algorithm mapping efficiency and poor versatility in related technologies based on in-memory computing neural network accelerators.

[0045] As an optional approach, the aforementioned neural network accelerator is configured with three states: system idle, system update, and system computation.

[0046] When the system is idle, the aforementioned scheduling subsystem detects weight update signals and calculates start signals in real time.

[0047] When the scheduling subsystem detects the weight update signal, it enters the system update state. The weight update manager reads the weight data from the weight memory and writes it into the storage and computing cluster. After writing, it sends an update completion signal back to the scheduling subsystem.

[0048] When the scheduling subsystem detects the update completion signal and the calculation start signal, it enters the system calculation state. The scheduling subsystem controls the feature memory to load input features into the in-memory computing cluster, performs neural network inference calculations, and outputs the calculation results.

[0049] This scheduling architecture decouples the weight refresh and inference computation processes through hardware-level closed-loop state management. It is compatible with various neural networks of different sizes, improving the accelerator's versatility. It also simplifies the upper-layer software control logic, standardizes data flow timing, reduces algorithm deployment costs, and improves the utilization rate of in-memory computing resources, effectively addressing the pain point of low algorithm mapping efficiency in in-memory computing accelerators.

[0050] As an optional approach, the scheduling subsystem controls the feature memory to load input features into the in-memory computing cluster, executes neural network inference calculations, and outputs the calculation results. This includes: initialization; calculating the size of the current layer input features and the size of the in-memory computing cluster; if the size of the current layer input features is less than or equal to the size of the in-memory computing cluster, the feature memory loads the current layer input features into the in-memory computing cluster, executes neural network inference calculations, and outputs the calculation results; if the size of the current layer input features is greater than the size of the in-memory computing cluster, the current layer input features are loaded into the in-memory computing cluster in batches, and the calculation results corresponding to each batch are accumulated through the external addition tree and then output, wherein the calculation results corresponding to each batch are temporarily stored in the feature memory.

[0051] By automatically comparing features with cluster size in hardware, differentiating single-batch / multi-batch computing processes, caching intermediate results in feature memory, and combining them with external addition trees, the system overcomes the capacity limitations of in-memory computing clusters, is compatible with neural networks of various sizes, and solves the problem of poor versatility. On the other hand, it pushes the logic of block division, caching, and result merging down to the hardware scheduling layer, reducing the amount of upper-layer software mapping development, improving algorithm mapping efficiency, and optimizing the utilization of computing power and storage resources.

[0052] As an optional approach, if the size of the current layer input features is greater than the size of the in-memory computing cluster, the quotient of the size of the current layer input features divided by the size of the in-memory computing cluster is taken as the total number of batches. If there is a remainder, the quotient is incremented by one to take the total number of batches.

[0053] By using a batch calculation rule with remainder rounding up, it can fully cover input features of any size and break through the cluster capacity limit. On the other hand, it pushes the batch number calculation logic down to the hardware, eliminating the need for manual block calculation in software, simplifying the multi-batch inference scheduling process, effectively improving the mapping efficiency of neural network algorithms, and ensuring the completeness and accuracy of inference calculation results.

[0054] As an optional solution, the aforementioned scheduling subsystem incorporates a sub-scheduling state machine, which includes sequentially switching sub-idle states, input feature loading states, calculation states, and output feature storage states.

[0055] The aforementioned idle state is used for initialization. The aforementioned input feature loading state is used to control the aforementioned feature memory to load the current layer's input features into the aforementioned in-memory computing cluster. The aforementioned computing state is used to manage the aforementioned in-memory computing cluster and execute neural network inference operations. The aforementioned output feature storage state is used to control the final calculation results to be stored back into the aforementioned feature memory.

[0056] By using a four-layer fixed-process sub-scheduling state machine, the hardware of the single-layer neural network inference process is standardized. A single set of state logic is compatible with networks of all sizes and various network structures, improving versatility. It also enables time-sharing wake-up of hardware modules, optimizing power consumption and resource utilization on the device side.

[0057] As an optional solution, an instruction first-in-first-out (FIFO) cache unit is also included.

[0058] The aforementioned instruction FIFO buffer unit is connected to the aforementioned scheduling subsystem. The aforementioned instruction FIFO buffer unit provides configuration information for the neural network layer and is used to generate the aforementioned weight update signal and the aforementioned computation start signal.

[0059] An instruction FIFO cache unit is added and two types of core trigger signals are generated based on the layer configuration. This can cache multi-layer network configurations, adapt to continuous inference and various network structures, and improve the accelerator's versatility. It also realizes instruction pre-storage and decoupling of software and hardware scheduling, simplifying hardware control logic.

[0060] As an optional solution, the aforementioned small cluster includes 9 first in-memory macro units, and the aforementioned large cluster includes 27 second in-memory macro units. The capacity of the aforementioned first in-memory macro units and the aforementioned second in-memory macro units is 64×64 bits, and both use bit serial transmission to transmit data.

[0061] The aforementioned weight memory is a 256KB static random access memory, equipped with a 32-bit read / write interface for the CPU and a 64-bit weight loading interface for the aforementioned in-memory computing cluster.

[0062] The aforementioned feature memory is a 16KB static random access memory, configured with 8 sets of mutually exclusive access interfaces. Among them, 3 sets of access interfaces are used for input feature access, 2 sets of access interfaces are used for intermediate calculation result access, and 3 sets of access interfaces are used for output feature access.

[0063] The tiered, large- and small-cluster computing power clusters cover the entire network channel. Dedicated intermediate result interfaces support block computation and residual networks. Large-capacity on-chip weight storage adapts to multiple model specifications and is compatible with various neural networks. Unified bit-serial macrocells simplify data flow configuration. Feature memory is functionally divided into dedicated interfaces, reducing software timing adjustments and block logic development workload, and lowering the difficulty of network deployment and debugging. High-bit-width storage and loading interfaces shorten weight refresh time and improve inference efficiency. Tiered enabling computing power units and shared feature storage reduce power consumption and hardware area, adapting to low-power AI acceleration scenarios on the edge.

[0064] As an optional approach, the aforementioned weight update manager pre-stores the weight memory base address and the storage computing cluster base address.

[0065] When the scheduling subsystem detects a weight update signal, the weight update manager generates a weight read address sequence based on the weight memory base address, which is used to read weight data from the weight memory with a 64-bit width; the weight update manager generates a weight write address sequence based on the in-memory computing cluster base address, which is used to write weight data to the in-memory computing cluster with a 64-bit width.

[0066] By pre-stored dual-base addresses and automatically generating 64-bit read / write address sequences in hardware, on the one hand, the address logic is decoupled from the network scale, making it compatible with various neural networks and multi-model switching; on the other hand, the weight address allocation and data distribution logic are pushed down to the hardware, saving the manual address development and debugging work in software, improving the efficiency of algorithm mapping, and at the same time, the high-bit-width batch transmission shortens the weight loading latency and optimizes the overall inference performance of the accelerator.

[0067] As an alternative, the connection between the aforementioned feature memory and the aforementioned external addition tree includes a first path and a second path.

[0068] The first path is used to transfer the intermediate results of historical batches temporarily stored in the aforementioned feature memory to the aforementioned external addition tree as the accumulation input.

[0069] The second path is used to send the accumulated result output by the external addition tree back to the feature memory for storage.

[0070] By establishing a bidirectional transmission path that separates reading and writing between the feature memory and the external addition tree, a closed-loop processing link for intermediate result caching, accumulation, and storage back is constructed. On the one hand, this breaks through the capacity limitations of the storage-computing cluster and is compatible with ultra-large network layers and special network structures of residual types. On the other hand, by solidifying the data interaction logic of block accumulation in hardware, the workload of software mapping development and debugging can be reduced. At the same time, existing storage resources are reused, the read and write paths are separated, and the stability of hardware operation is optimized.

[0071] like Figure 3 As shown, the present invention also provides a neural network acceleration method based on any of the above-mentioned neural network accelerators, the method including steps S301 to S306.

[0072] Step S301: Decode the configuration instructions of the current neural network layer to obtain configuration information.

[0073] Step S302: Based on the above configuration information, the corresponding weight data is read from the weight storage through the weight update manager and written to the storage computing cluster.

[0074] Step S303: Read the input feature data from the feature memory and send it to the in-memory computing cluster, and output the convolution operation result.

[0075] Step S304: The convolution operation results of different channels are accumulated, activated, pooled, and quantized to obtain the output feature data of the current batch.

[0076] Step S305: Write the above output feature data into the above feature memory as input features for the next neural network layer.

[0077] Step S306: Repeat the above steps until the inference calculations of all neural network layers are completed.

[0078] The aforementioned neural network acceleration method relies on the accelerator hardware architecture described above. Through a standardized, closed-loop, cyclical six-layer inference process, it adapts to neural networks of various sizes and structures. It solidifies the underlying hardware scheduling logic into a standard process, simplifying software development and debugging of algorithm mapping, improving mapping efficiency, and simultaneously aligning with the native computational flow of neural networks, fully reusing on-chip storage resources, and optimizing overall inference speed and edge power consumption. This addresses the problems of low algorithm mapping efficiency and poor versatility in related technologies based on in-memory computing neural network accelerators.

[0079] The present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.

[0080] The present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the method of the embodiments of the present invention.

[0081] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method of this invention.

[0082] Computer programs for implementing the methods of embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0083] In the context of embodiments of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0084] It should be noted that the term "comprising" and its variations used in the embodiments of this invention are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of this invention are illustrative and not restrictive, and those skilled in the art should understand that unless explicitly indicated otherwise in the context, they should be understood as "one or more". The descriptions of terms such as "first", "second", etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of indicated technical features.

[0085] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this invention are subject to strict compliance with relevant laws, regulations, and regulatory requirements in their collection, storage, use, processing, transmission, provision, and disclosure, and adhere to the principles of legality, legitimacy, necessity, and good faith. The acquisition of relevant information and data is premised on the user's explicit consent or other legitimate reasons, and a clear and convenient authorization management approach is provided to the user, allowing the user to independently choose to consent, withdraw consent, or refuse to provide relevant information. For functions that rely on user information, if the user does not authorize or withdraws authorization, the corresponding technical function cannot be implemented, and the technical solution of this invention is not applicable in this scenario.

[0086] The steps described in the method embodiments provided by the present invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this respect.

[0087] The term "embodiment" in this specification refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily imply the same embodiment, nor does it imply independence or alternativeity from other embodiments. The various embodiments in this specification are described in a related manner, with reference to each other for similar or identical parts. In particular, for apparatus, device, and system embodiments, since they are substantially similar to method embodiments, the description is relatively simple, and relevant details are referred to in the description of the method embodiments.

[0088] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A neural network accelerator based on in-memory computing, characterized in that, include: The computing power subsystem includes a memory computing cluster, an internal addition tree, a first activation quantization module, an external addition tree, and a second activation quantization module connected in sequence. The memory computing cluster includes a small cluster and a large cluster. The small cluster includes multiple first memory computing macro units, each of which is configured with an enable switch. The large cluster includes multiple second memory computing macro units, all of which share an enable switch. Each enable switch is connected to the scheduling subsystem. The storage subsystem includes a weight memory, a feature memory, and a weight update manager. The weight update manager is connected to the weight memory and the in-memory computing cluster, respectively. The feature memory is connected to the in-memory computing cluster and the external addition tree, respectively. The scheduling subsystem is connected to the control terminals of each part of the computing subsystem and the storage subsystem, and is used to manage the start-up, shutdown and data flow of the corresponding parts.

2. The neural network accelerator according to claim 1, characterized in that, The neural network accelerator is configured with three states: system idle, system update, and system computation. When the system is idle, the scheduling subsystem detects weight update signals and calculates start signals in real time. When the scheduling subsystem detects the weight update signal, it enters the system update state. The weight update manager reads weight data from the weight memory and writes it into the storage and computing cluster. After writing is completed, it sends an update completion signal back to the scheduling subsystem. When the scheduling subsystem detects the update completion signal and the computation start signal, it enters the system computation state. The scheduling subsystem controls the feature memory to load input features into the in-memory computing cluster, performs neural network inference computation, and outputs the computation results.

3. The neural network accelerator according to claim 2, characterized in that, The scheduling subsystem controls the feature memory to load input features into the in-memory computing cluster, performs neural network inference calculations, and outputs the calculation results, including: Perform initialization; Calculate the scale of the current layer's input features and the scale of the in-memory computing cluster; If the size of the current layer input features is less than or equal to the size of the in-memory computing cluster, the current layer input features are loaded from the feature memory into the in-memory computing cluster, neural network inference calculation is performed, and the calculation result is output. If the size of the current layer input features is greater than the size of the in-memory computing cluster, the current layer input features are loaded into the in-memory computing cluster in batches. The calculation results corresponding to each batch are accumulated by the external addition tree and then output. The calculation results corresponding to each batch are temporarily stored in the feature memory.

4. The neural network accelerator according to claim 3, characterized in that, If the size of the current layer input feature is greater than the size of the in-memory computing cluster, the quotient of the size of the current layer input feature divided by the size of the in-memory computing cluster is taken as the total number of batches. If there is a remainder, the quotient is incremented by one and taken as the total number of batches.

5. The neural network accelerator according to claim 3, characterized in that, The scheduling subsystem has a built-in sub-scheduling state machine, which includes a sequentially switching sub-idle state, an input feature loading state, a calculation state, and an output feature storage state. The sub-idle state is used for initialization; The input feature loading status is used to control the feature memory to load the current layer input features into the in-memory computing cluster; The computing state is used to manage the in-memory computing cluster and execute neural network inference operations; The output feature storage state is used to control the final calculation result to be stored back into the feature memory.

6. The neural network accelerator according to claim 2, characterized in that, It also includes an instruction first-in-first-out (FIFO) cache unit; The instruction FIFO cache unit is connected to the scheduling subsystem. The instruction FIFO cache unit provides configuration information for the neural network layer and is used to generate the weight update signal and the computation start signal.

7. The neural network accelerator according to claim 1, characterized in that, The small cluster includes 9 first in-memory macro units, and the large cluster includes 27 second in-memory macro units. The capacity of both the first and second in-memory macro units is 64×64 bits, and both transmit data in a bit-serial manner. The weight memory is a 256KB static random access memory, equipped with a 32-bit read / write interface for the CPU and a 64-bit weight loading interface for the in-memory computing cluster. The feature memory is a 16KB static random access memory, configured with 8 sets of mutually exclusive access interfaces, of which 3 sets of access interfaces are used for input feature access, 2 sets of access interfaces are used for intermediate calculation result access, and 3 sets of access interfaces are used for output feature access.

8. The neural network accelerator according to claim 7, characterized in that, The weight update manager pre-stores the weight memory base address and the storage computing cluster base address; When the scheduling subsystem detects a weight update signal, the weight update manager generates a weight read address sequence based on the weight memory base address, for reading weight data from the weight memory with a 64-bit width; the weight update manager generates a weight write address sequence based on the in-memory computing cluster base address, for writing weight data to the in-memory computing cluster with a 64-bit width.

9. The neural network accelerator according to claim 1, characterized in that, The connection between the feature memory and the external addition tree includes: The first path is used to transmit the intermediate results of historical batches temporarily stored in the feature memory to the external addition tree as an accumulation input; The second path is used to send the accumulated result output by the external addition tree back to the feature memory for storage.

10. A method for accelerating neural networks, characterized in that, Acceleration based on any one of claims 1-9, the method comprising the following steps: The configuration instructions for the current neural network layer are decoded to obtain the configuration information. Based on the configuration information, the corresponding weight data is read from the weight storage through the weight update manager and written to the storage-computing cluster; The system reads input feature data from the feature memory and sends it to the in-memory computing cluster, then outputs the convolution operation result. The convolution results from different channels are accumulated, activated, pooled, and quantized to obtain the output feature data of the current batch. The output feature data is written into the feature memory and used as the input features of the next neural network layer. Repeat the above steps until the inference calculations for all neural network layers are completed.