Priori knowledge-based heterogeneous in-memory computing network-on-chip architecture optimization method
By building a joint simulation platform and a three-stage search strategy to optimize the heterogeneous in-memory computing on-chip network architecture, the problems of large design space and low simulation efficiency of heterogeneous network chips are solved, and efficient hardware configuration search and optimization are achieved.
Patent Information
- Application Number
- CN202510809677.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-10-03
AI Technical Summary
Existing heterogeneous in-memory computing network chip architecture designs face huge design space challenges. The lack of dedicated simulators makes it impossible to efficiently evaluate hardware configurations, and existing frameworks cannot fully explore the optimal configuration, resulting in inaccurate simulation results and low search efficiency.
A joint simulation platform for heterogeneous in-memory computing network-on-chip is constructed. A three-stage search strategy based on prior knowledge and a preset simulated annealing algorithm are used to optimize the search process. The force-directed algorithm is combined to optimize the core layout and determine the target architecture configuration of the heterogeneous in-memory computing network-on-chip.
It improves the search efficiency and simulation accuracy of heterogeneous in-memory computing on-chip network architecture, optimizes computing delay time, power consumption and area, and significantly improves the performance of heterogeneous network chips.
Smart Images

Figure CN120745540A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of integrated circuit chip design, and in particular to a method for optimizing heterogeneous in-memory computing on-chip network architecture based on prior knowledge. Background Art
[0002] In recent years, convolutional neural networks (CNNs) have demonstrated remarkable capabilities in fields such as object detection and face recognition. However, as CNN models grow in complexity, traditional von Neumann architectures, such as CPUs and GPUs, face significant challenges due to the large amount of data movement between memory and computing units. This data movement consumes over 80% of the system's energy and exacerbates the "memory wall" problem. To address these challenges, the Processing-in-Memory (PIM) architecture has emerged. This architecture performs matrix-vector multiplication (MVM) directly within the memory array, eliminating data movement overhead and improving energy efficiency.
[0003] As the number of neural network parameters increases, single-core in-memory computing accelerators face computational limitations and require frequent weight data refreshes, resulting in significant overhead. Current solutions typically involve designing homogeneous or heterogeneous in-memory computing network-on-chip (NoC) accelerators to store all weights on-chip. However, the performance of homogeneous in-memory computing network-on-chip architectures is highly dependent on the type of in-memory computing performed by a single core and the size of the array, limiting their scalability and flexibility. Digital in-memory computing architectures based on SRAM exhibit low latency by integrating logic computations, while analog in-memory computing architectures based on RRAM offer higher computational density but require additional analog-to-digital conversion overhead. Heterogeneous in-memory computing network-on-chip architectures combine different in-memory computing types and core sizes, enabling greater scalability and flexibility to meet stringent power, performance, and area (PPA) requirements. However, the main challenge in designing heterogeneous in-memory computing network-on-chip architectures is the large design space required, making manual design nearly impossible. Therefore, efficient design space exploration methods are needed.
[0004] The time required for design space exploration depends on the simulation speed and search strategy for each architecture. From a simulation perspective, most existing multicore simulators, such as DNN+NeuroSIM v2.0 and MNSIM 2.0, primarily target homogeneous in-memory computing architectures and do not support simulation of heterogeneous in-memory computing architectures. Furthermore, existing simulators primarily focus on core-level metric simulation and lack accurate inter-core communication models in network chips. Furthermore, general-purpose network chip simulators, such as Booksim 2.0, cannot accurately simulate data dependencies between neural network layers, resulting in inaccurate latency simulation results.
[0005] From the perspective of architecture search, current architecture search frameworks are unable to efficiently evaluate various hardware configurations due to the lack of simulators tailored for heterogeneous in-memory network chip architectures. For example, the PIM-HLS framework used in the paper by Y. Zhu, Z. Zhu, G. Dai, F. Tu, H. Sun, K.-T. Cheng, H. Yang, and Y. Wang, “Pim-hls: An automatic hardware generation tool for heterogeneous processing-in-memory-based neural network accelerators,” and the framework used in the paper by H. Sun, T. Xie, Z. Zhu, G. Dai, H. Yang, and Y. Wang, “Minimizing communication conflicts in network-on-chip based processing-in-memory architecture,” for in-memory computing based on on-chip networks mainly focus on workload distribution and mapping schemes rather than hardware configuration search. Other frameworks, such as Gibbon and AIG-CIM, primarily search for homogeneous configurations or a limited range of heterogeneous configurations, failing to fully explore the broader design space encompassing a variety of in-memory compute types, core sizes, and array configurations. Therefore, a dedicated search framework is urgently needed to find the optimal configuration for heterogeneous in-memory compute-based network chip architectures. Summary of the Invention
[0006] In response to one of the defects in the prior art, the purpose of this application is to provide a method for optimizing heterogeneous in-memory computing on-chip network architecture based on prior knowledge.
[0007] In a first aspect of the present application, a method for optimizing heterogeneous in-memory computing network-on-chip architecture based on prior knowledge is provided, comprising:
[0008] Constructing a joint simulation platform for heterogeneous in-memory computing network-on-chip, and using the joint simulation platform to perform joint pipeline simulation of computing and communication delays on the heterogeneous in-memory computing network-on-chip;
[0009] A three-stage search strategy based on prior knowledge is used to search within the heterogeneous architecture space of the heterogeneous in-memory computing network-on-chip, and a preset lookup table and a preset simulated annealing algorithm are used to optimize the search process to determine a target architecture configuration of the heterogeneous in-memory computing network-on-chip;
[0010] According to the target architecture configuration of the heterogeneous in-memory computing network on chip, a force-directed algorithm is used to optimize the core layout of the heterogeneous in-memory computing network on chip, and an optimized heterogeneous in-memory computing network on chip architecture is determined.
[0011] Optionally, the joint simulation platform includes a heterogeneous in-memory computing core simulator and an on-chip network data transmission simulator, wherein the heterogeneous in-memory computing core simulator is used to simulate the computing behavior of multiple types of heterogeneous in-memory computing cores, and the on-chip network data transmission simulator is used to model the data relationships and inter-layer communication behaviors of neural networks.
[0012] Optionally, the method adopts a three-stage search strategy based on prior knowledge to search within the heterogeneous architecture space of the heterogeneous in-memory computing network-on-chip, and adopts a preset lookup table and a preset simulated annealing algorithm to optimize the search process to determine the target architecture configuration of the heterogeneous in-memory computing network-on-chip, including:
[0013] performing a homogeneous architecture search at the neural network layer according to preset target indicators to determine an initial homogeneous architecture configuration corresponding to the preset target indicators, wherein the preset target indicators represent computational latency, power consumption, and area targets of the heterogeneous in-memory computing core configuration;
[0014] On the initial homogeneous architecture configuration, according to the preset target indicators, perform an inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer, optimize the architecture configuration of the heterogeneous in-memory computing core corresponding to each neural network layer, and determine the inter-layer heterogeneous in-memory computing core configuration;
[0015] According to the preset target indicators, an intra-layer heterogeneous architecture search is performed on the heterogeneous in-memory computing core corresponding to each neural network layer, and the heterogeneous in-memory computing core corresponding to each neural network layer is personalized configured to determine the target architecture configuration of the heterogeneous in-memory computing network on chip.
[0016] Optionally, based on the initial homogeneous architecture configuration, performing an inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer according to the preset target indicator, optimizing the architecture configuration of the heterogeneous in-memory computing core corresponding to each neural network layer, and determining the inter-layer heterogeneous in-memory computing core configuration includes:
[0017] According to the computing delay time, power consumption and area targets of the heterogeneous in-memory computing core configuration, an inter-layer heterogeneous architecture search is performed between the heterogeneous in-memory computing cores corresponding to each of the neural network layers to determine the core type, number of PEs and array size of the heterogeneous in-memory computing corresponding to each of the neural network layers corresponding to the computing delay time, power consumption and area targets of the heterogeneous in-memory computing core configuration, and determine the inter-layer heterogeneous in-memory computing core configuration.
[0018] Optionally, performing an intra-layer heterogeneous architecture search on the heterogeneous in-memory computing core corresponding to each neural network layer according to the preset target indicator, performing personalized configuration on the heterogeneous in-memory computing core corresponding to each neural network layer, and determining a target architecture configuration of the heterogeneous in-memory computing network-on-chip includes:
[0019] Determining the total capacity of the heterogeneous in-memory computing cores corresponding to each neural network layer based on the inter-layer heterogeneous in-memory computing core configuration;
[0020] Based on the preset weight storage target of each neural network and the total capacity of the heterogeneous in-memory computing cores corresponding to each neural network layer, the number of heterogeneous in-memory computing cores corresponding to each neural network layer is determined, and the target architectural configuration of the heterogeneous in-memory computing network-on-chip is determined.
[0021] Optionally, determining the number of heterogeneous in-memory computing cores corresponding to each neural network layer and determining the target architecture configuration of the heterogeneous in-memory computing network-on-chip based on a preset weight storage target of each neural network and a total capacity of heterogeneous in-memory computing cores corresponding to each neural network layer includes:
[0022] If the preset weight storage target is greater than the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer, adding a preset number of cores to the heterogeneous in-memory computing corresponding to the neural network layer;
[0023] If the preset weight storage target is smaller than the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer and the difference between the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer and the preset weight storage target is greater than the minimum core capacity of the heterogeneous in-memory computing corresponding to the neural network layer, remove the minimum core of the heterogeneous in-memory computing corresponding to the neural network layer.
[0024] Optionally, the method adopts a three-stage search strategy based on prior knowledge to search within the heterogeneous architecture space of the heterogeneous in-memory computing network-on-chip, and adopts a preset lookup table and a preset simulated annealing algorithm to optimize the search process to determine the target architecture configuration of the heterogeneous in-memory computing network-on-chip, further comprising:
[0025] According to the preset target indicator, calling the preset lookup table to determine the heterogeneous in-memory computing core configuration corresponding to the preset target indicator;
[0026] According to the heterogeneous in-memory computing core configuration corresponding to the preset target indicator and the preset fusion indicator, a predefined search operator is used to optimize the heterogeneous in-memory computing core configuration.
[0027] Optionally, the preset fusion index represents a fusion index of computing delay time, power consumption and area targets of the computing core configuration in the heterogeneous memory;
[0028] The predefined search operators include a search operator for adjusting the type of the heterogeneous in-memory computing core, a search operator for adjusting the number of PEs, a search operator for adjusting the array size, a search operator for adjusting the mapping order of the heterogeneous in-memory computing core, and a search operator for exchanging the position of the heterogeneous in-memory computing core corresponding to the neural network layer.
[0029] Optionally, the method for constructing the preset lookup table includes:
[0030] Using the joint simulation platform to simulate the core configurations of all heterogeneous in-memory computing, and determining the target indicator corresponding to each core configuration of the heterogeneous in-memory computing;
[0031] The preset lookup table is constructed according to all the heterogeneous in-memory computing core configurations and the target indicators corresponding to each of the heterogeneous in-memory computing core configurations.
[0032] Optionally, the step of determining an optimized heterogeneous in-memory computing network-on-chip architecture based on the target architecture configuration of the heterogeneous in-memory computing network-on-chip and optimizing the core layout of the heterogeneous in-memory computing network-on-chip using a force-directed algorithm includes:
[0033] Determining the interaction force between the cores of the heterogeneous in-memory computing according to the topological connection of the heterogeneous in-memory computing network-on-chip;
[0034] According to the preset interaction force, the position of the core of the heterogeneous in-memory computing and the interaction force between the cores of the heterogeneous in-memory computing are adjusted to determine the optimized heterogeneous in-memory computing on-chip network architecture.
[0035] The simulation and search method of the heterogeneous in-memory computing on-chip network based on prior knowledge in this application adopts a three-stage search strategy based on prior knowledge to search for hardware configuration, calls a preset lookup table and uses a preset simulated annealing algorithm to optimize the search process, quickly and effectively searches for the most optimal configuration that meets the requirements, improves search efficiency, and uses a force-directed algorithm to further optimize communication delay and layout area, thereby improving the calculation delay time of the heterogeneous in-memory computing on-chip network architecture as well as the optimization degree of power consumption and area.
[0036] Other technical effects brought about by the additional features will be further explained in the corresponding embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0038] Figure 1 The present invention is a flowchart showing a method for simulating and searching a heterogeneous in-memory computing network-on-chip based on prior knowledge according to an exemplary embodiment.
[0039] Figure 2 The figure is a schematic diagram of a prototype structure of a heterogeneous in-memory computing network-on-chip architecture according to an exemplary embodiment.
[0040] Figure 3 The figure is a simulation flow chart of a heterogeneous in-memory computing network-on-chip according to an exemplary embodiment.
[0041] Figure 4 The figure is a search flow chart of a heterogeneous in-memory computing network-on-chip according to an exemplary embodiment.
[0042] Figure 5 The figure is a flowchart showing an optimization search process according to an exemplary embodiment. DETAILED DESCRIPTION
[0043] The present application is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that those skilled in the art may make several variations and improvements without departing from the scope of the present application. These all fall within the scope of protection of the present application.
[0044] The existing technology lacks simulators specifically tailored for heterogeneous in-memory computing network chip architectures, making it impossible to evaluate various hardware configurations. Furthermore, existing heterogeneous in-memory computing architectures based on on-chip networks cannot fully search for optimal configurations of heterogeneous in-memory computing network chip architectures. Based on the above issues, embodiments of the present application provide a method for optimizing heterogeneous in-memory computing network-on-chip architectures based on prior knowledge to address these existing issues.
[0045] Figure 1 The present invention is a flowchart showing a method for simulating and searching a heterogeneous in-memory computing network-on-chip based on prior knowledge according to an exemplary embodiment.
[0046] Reference Figure 1 As shown, an embodiment of the present application is a simulation and search method for a heterogeneous in-memory computing network-on-chip based on prior knowledge, including S11 to S13.
[0047] S11, build a joint simulation platform for heterogeneous in-memory computing on-chip networks, and use the joint simulation platform to perform joint pipeline simulation of computing and communication delays on the heterogeneous in-memory computing on-chip networks.
[0048] Specifically, the heterogeneous in-memory computing network-on-chip includes multiple in-memory computing cores, and the neural network is mapped to the in-memory computing cores of the heterogeneous in-memory computing network-on-chip to complete the calculation.
[0049] S12, adopts a three-stage search strategy based on prior knowledge to search in the heterogeneous architecture space of the heterogeneous in-memory computing on-chip network, and uses a preset look-up table (LUT) and a preset simulated annealing algorithm to optimize the search process to determine the target architecture configuration of the heterogeneous in-memory computing on-chip network.
[0050] S13, according to the target architecture configuration of the heterogeneous in-memory computing network-on-chip, and using a force-directed algorithm to optimize the core layout of the heterogeneous in-memory computing network-on-chip, determine the optimized heterogeneous in-memory computing network-on-chip architecture.
[0051] The above-mentioned embodiments of the present application construct a joint simulation platform for heterogeneous in-memory computing on-chip networks, realize the simulation of heterogeneous in-memory computing on-chip networks, improve the simulation accuracy, and facilitate the evaluation of hardware configurations based on heterogeneous in-memory computing on-chip networks; adopt a three-stage search strategy based on prior knowledge to search for hardware configurations, call a preset lookup table and use a preset simulated annealing algorithm to optimize the search process, quickly and effectively search for the most optimal in-memory computing on-chip network architecture configuration that meets the requirements, and improve search efficiency; adopt a force-directed algorithm to further optimize the communication delay and layout area of the heterogeneous in-memory computing on-chip network architecture, and improve the communication delay, power consumption and area optimization degree of the heterogeneous in-memory computing on-chip network architecture.
[0052] In some specific embodiments of the present application, the joint simulation platform includes a heterogeneous in-memory computing core simulator and an on-chip network data transmission simulator. The heterogeneous in-memory computing core simulator is used to simulate the computing behavior of various types of heterogeneous in-memory computing cores, and the on-chip network data transmission simulator is used to model the data relationships and inter-layer communication behaviors of neural networks.
[0053] Illustratively, the heterogeneous in-memory computing core simulator can simulate the behavior of heterogeneous in-memory computing cores of SRAM and RRAM.
[0054] In order to build a joint simulation platform for heterogeneous in-memory computing on-chip networks, in some specific embodiments of the present application, for S11, a joint simulation platform for heterogeneous in-memory computing on-chip networks is built, and the joint simulation platform is used to perform a joint pipeline simulation of computing and communication delays on the heterogeneous in-memory computing on-chip networks, including:
[0055] Build a heterogeneous in-memory computing core simulator and an on-chip network data transmission simulator, and determine a joint simulation platform;
[0056] The heterogeneous in-memory computing core simulator is used to simulate the computing behavior of the heterogeneous in-memory computing core of the heterogeneous in-memory computing on-chip network. The on-chip network data transmission simulator is used to model the data relationship and inter-layer communication behavior of the neural network to simulate the heterogeneous in-memory computing on-chip network.
[0057] The above-mentioned embodiments of the present application use a heterogeneous in-memory computing core simulator to simulate the computing behavior of various types of heterogeneous in-memory computing cores, support hardware configuration evaluation, use an on-chip network data transmission simulator to overcome the problem of inaccurate delay simulation caused by existing simulators not considering the dependencies between neural network layers, and use a joint simulation platform to perform joint pipeline simulation of computing and communication delays on the heterogeneous in-memory computing on-chip network, thereby improving the prediction accuracy of the computing delay time of the heterogeneous in-memory computing on-chip network architecture.
[0058] The heterogeneous architecture space of the heterogeneous in-memory computing network on chip is searched to improve the search efficiency. In some specific implementations of the present application, for S12, a three-stage search strategy based on prior knowledge is adopted to search in the heterogeneous architecture space of the heterogeneous in-memory computing network on chip, and a preset lookup table and a preset simulated annealing algorithm are used to optimize the search process and determine the target architecture configuration of the heterogeneous in-memory computing network on chip. S121 to S123 can be used.
[0059] The three-stage search strategy based on prior knowledge of this embodiment includes: in the first stage, determining the initial homogeneous architecture configuration corresponding to the preset target indicator; in the second stage, based on the initial homogeneous architecture configuration, performing inter-layer heterogeneous architecture search on the corresponding heterogeneous in-memory computing cores between each neural network layer; in the third stage, based on the inter-layer heterogeneous architecture search, performing intra-layer heterogeneous architecture search on the corresponding heterogeneous in-memory computing cores within each neural network layer.
[0060] The details are as follows:
[0061] S121, according to the preset target indicators, perform a homogeneous architecture search in the neural network layer to determine the initial homogeneous architecture configuration corresponding to the preset target indicators.
[0062] Specifically, the preset target indicators represent the computing delay time, power consumption and area targets of the computing core configuration in the heterogeneous memory.
[0063] The initial homogeneous architecture configuration corresponding to the preset target indicators is an initial architecture benchmark that is close to the computing latency, power consumption and area targets of the heterogeneous in-memory computing core configuration.
[0064] S122. Based on the initial homogeneous architecture configuration, according to the preset target indicators, perform inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer, optimize the architecture configuration of the heterogeneous in-memory computing core corresponding to each neural network layer, and determine the inter-layer heterogeneous in-memory computing core configuration.
[0065] Specifically, based on the initial homogeneous architecture configuration, inter-layer heterogeneous architecture search is performed between the corresponding heterogeneous in-memory computing cores between neural network layers. According to the computing requirements of each neural network layer, heterogeneous in-memory computing cores of different types or sizes are used to accurately adapt to the diverse computing requirements of each neural network layer.
[0066] S123, according to the preset target indicators, perform intra-layer heterogeneous architecture search on the corresponding heterogeneous in-memory computing cores in each neural network layer, perform personalized configuration on the corresponding heterogeneous in-memory computing cores in each neural network layer, and determine the target architecture configuration of the heterogeneous in-memory computing network on chip.
[0067] Specifically, based on the inter-layer heterogeneous architecture search, an intra-layer heterogeneous architecture search is performed between the corresponding heterogeneous in-memory computing cores in a single-layer neural network layer, and the heterogeneous in-memory computing cores are personalized configured to realize the configuration of heterogeneous in-memory computing cores of different specifications in the same neural network layer.
[0068] In the above steps S121 to S123, the heterogeneous in-memory computing core configuration is ensured during each search process to meet the weight storage requirements of the neural network after each operation, thereby avoiding the generation of unreasonable configuration.
[0069] The above-mentioned embodiments of the present application adopt a three-stage search strategy based on prior knowledge to effectively solve the problem of difficulty in searching the huge design space of heterogeneous architectures and improve the search efficiency of hardware configurations.
[0070] When performing an inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer, the architectural configuration of the heterogeneous in-memory computing cores is optimized. In some specific embodiments of the present application, S122, based on the initial homogeneous architecture configuration, performs an inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer according to a preset target indicator, optimizes the architectural configuration of the heterogeneous in-memory computing cores corresponding to each neural network layer, and determines the inter-layer heterogeneous in-memory computing core configuration. The following method can be used:
[0071] According to the calculation delay time, power consumption and area targets of the heterogeneous in-memory computing core configuration, an inter-layer heterogeneous architecture search is performed between the heterogeneous in-memory computing cores corresponding to each neural network layer to determine the core type, PE number and array size of the heterogeneous in-memory computing corresponding to each neural network layer corresponding to the calculation delay time, power consumption and area targets of the heterogeneous in-memory computing core configuration, and determine the inter-layer heterogeneous in-memory computing core configuration.
[0072] The above-mentioned embodiment of the present application realizes that the configuration of the heterogeneous in-memory computing cores corresponding to the neural network layer is as close to the preset target indicators as possible by performing inter-layer heterogeneous search on the heterogeneous in-memory computing cores.
[0073] When performing an intra-layer heterogeneous architecture search between the corresponding heterogeneous in-memory computing cores in each neural network layer, optimizing the architectural configuration of the neural network layer, in some specific embodiments of the present application, S123, performing an intra-layer heterogeneous architecture search on the corresponding heterogeneous in-memory computing cores in each neural network layer according to a preset target indicator, performing personalized configuration on the corresponding heterogeneous in-memory computing cores in each neural network layer, and determining the target architectural configuration of the heterogeneous in-memory computing network-on-chip, may include:
[0074] According to the configuration of heterogeneous in-memory computing cores between layers, the total capacity of heterogeneous in-memory computing cores corresponding to each neural network layer is determined.
[0075] Based on the preset weight storage target of each neural network and the total capacity of the heterogeneous in-memory computing cores corresponding to each neural network layer, the number of heterogeneous in-memory computing cores corresponding to each neural network layer is determined, and the target architecture configuration of the heterogeneous in-memory computing on-chip network is determined.
[0076] For example, if the preset weight storage target is greater than the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer, a preset number of cores are added to the heterogeneous in-memory computing corresponding to the neural network layer to meet the storage requirements.
[0077] Exemplarily, if the preset weight storage target is smaller than the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer and the difference between the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer and the preset weight storage target is greater than the minimum core capacity of the heterogeneous in-memory computing corresponding to the neural network layer, the minimum core of the heterogeneous in-memory computing corresponding to the neural network layer is removed to prevent resource redundancy.
[0078] The above-mentioned embodiments of the present application perform intra-layer heterogeneous search on the heterogeneous in-memory computing cores corresponding to the neural network layer, personalize the configuration of the heterogeneous in-memory computing cores corresponding to the neural network layer, and timely supplement or remove the heterogeneous in-memory computing cores to meet storage requirements and prevent resource redundancy.
[0079] In order to optimize the search process, in some specific embodiments of the present application, for S12, a three-stage search strategy based on prior knowledge is adopted to search in the heterogeneous architecture space of the heterogeneous in-memory computing network on chip, and a preset lookup table and a preset simulated annealing algorithm are used to optimize the search process and determine the target architecture configuration of the heterogeneous in-memory computing network on chip, which also includes S124 to S125.
[0080] S124 , calling a preset lookup table according to the preset target indicator to determine the heterogeneous in-memory computing core configuration corresponding to the preset target indicator.
[0081] Specifically, the preset lookup table includes all possible heterogeneous in-memory computing core configurations and target indicators corresponding to each heterogeneous in-memory computing core configuration, that is, corresponding area, power consumption and computing delay time targets.
[0082] In the above embodiment of the present application, step S124 represents the step of calling a preset lookup table to optimize the search process.
[0083] S125 , optimizing the core configuration of the heterogeneous in-memory computing using a predefined search operator according to the heterogeneous in-memory computing core configuration corresponding to the preset target indicator and the preset fusion indicator.
[0084] Specifically, the preset fusion index represents a fusion index of the computing delay time, power consumption and area of the computing core configuration in the heterogeneous memory, and the preset fusion index serves as a feedback index for optimizing the computing core configuration in the heterogeneous memory.
[0085] The predefined search operators include search operators for adjusting the type of heterogeneous in-memory computing cores, search operators for adjusting the number of PEs, search operators for adjusting the array size, search operators for adjusting the mapping order of heterogeneous in-memory computing cores, and search operators for exchanging the positions of heterogeneous in-memory computing cores corresponding to neural network layers.
[0086] In the above embodiment of the present application, step S125 represents a step of optimizing the search process using a preset simulated annealing algorithm (SA).
[0087] In the above-mentioned embodiment of the present application, a preset lookup table is called to search for the heterogeneous in-memory computing core configuration corresponding to the preset target indicator, thereby avoiding repeated calculations and repeated simulations during the search process, achieving simulation acceleration, and improving search efficiency. The preset simulated annealing algorithm is an algorithm specifically suitable for accelerating the search of heterogeneous architectures. By using the fusion index of the calculation delay time, power consumption and area targets of the core configuration of the comprehensive heterogeneous in-memory computing as a feedback indicator and defining a dedicated search operator, the search process is guided to efficiently tend towards a more optimal heterogeneous configuration, thereby optimizing the search process.
[0088] In some specific implementations of the present application, a method for constructing a preset lookup table includes:
[0089] A joint simulation platform is used to simulate the core configurations of all heterogeneous in-memory computing and determine the target indicators corresponding to each core configuration of heterogeneous in-memory computing.
[0090] Specifically, the area, power consumption, and computing delay time corresponding to each core configuration of heterogeneous in-memory computing are determined as corresponding target indicators.
[0091] A preset lookup table is constructed based on all heterogeneous in-memory computing core configurations and the target indicators corresponding to each heterogeneous in-memory computing core configuration.
[0092] Specifically, all the above-mentioned heterogeneous in-memory computing core configurations and the area, power consumption and computing delay time corresponding to each heterogeneous in-memory computing core configuration are stored in a lookup table, and the lookup table data is called during the search process to quickly obtain the target indicators corresponding to the configuration.
[0093] In the above-mentioned embodiments of the present application, possible core configurations of heterogeneous in-memory computing are simulated in advance and the results are stored. The lookup table data is directly called during the search process to avoid repeated calculations and repeated simulations during the search process, thereby improving simulation efficiency.
[0094] To further optimize the heterogeneous in-memory computing network-on-chip architecture, in some specific implementations of the present application, for S13, according to the target architecture configuration of the heterogeneous in-memory computing network-on-chip, and using a force-directed algorithm to optimize the core layout of the heterogeneous in-memory computing network-on-chip, the optimized heterogeneous in-memory computing network-on-chip architecture is determined, and S131 to S132 can be used.
[0095] S131, determining the interaction force between the heterogeneous in-memory computing cores according to the topological connection of the heterogeneous in-memory computing network-on-chip.
[0096] S132, adjusting the positions of the heterogeneous in-memory computing cores and the interaction forces between the heterogeneous in-memory computing cores according to the preset interaction forces, and determining an optimized heterogeneous in-memory computing on-chip network architecture.
[0097] Specifically, steps S131 to S132 compare each core of the on-chip network to a particle based on the force-directed algorithm. According to the gravitational repulsion model between particles, two particles will generate gravitational force on each other when they are far away, and repulsive force on each other when they are close. The interaction force between particles will move with the direction of the force.
[0098] By setting the interaction force between cores, iteratively adjusting the position of the cores, and then changing the connection relationship between the cores, the data transmission path is adjusted, the layout area and communication delay time of the heterogeneous in-memory computing on-chip network are minimized, and the optimal layout of the heterogeneous in-memory computing on-chip network is achieved, forming a compact and efficient optimized heterogeneous in-memory computing on-chip network architecture.
[0099] The above-mentioned embodiments of the present application optimize the position of the core of the heterogeneous in-memory computing on-chip network based on the force-directed algorithm, minimize the layout area and communication delay of the heterogeneous in-memory computing on-chip network, and form a compact and efficient heterogeneous in-memory computing on-chip network layout solution.
[0100] In some specific embodiments of the present application, in the early stage of searching for heterogeneous in-memory computing on-chip networks, that is, in the first and second stages of searching for heterogeneous in-memory computing on-chip networks using the three-stage search strategy based on prior knowledge, a higher frequency is used to simulate the on-chip network to determine the data transmission mode; in the later stage of searching for heterogeneous in-memory computing on-chip networks, that is, in the third stage, the simulation frequency of the on-chip network is dynamically and gradually reduced to shorten the simulation time.
[0101] Specifically, in the first and second stages of the on-chip network search for heterogeneous in-memory computing, the neural network layer is adjusted simultaneously to map to the core configuration of the heterogeneous in-memory computing. The data transmission channel of the on-chip network changes dramatically, which requires a higher frequency simulation of the on-chip network.
[0102] In the third stage of the heterogeneous in-memory computing on-chip network search, the core configuration of the heterogeneous in-memory computing on-chip network gradually stabilizes, and only a certain core configuration mapped within the neural network layer is adjusted, with minor changes. Therefore, reducing the simulation frequency can accelerate the search process and improve simulation efficiency.
[0103] The preferred features of the above embodiments can be used alone in any embodiment, or in any combination without conflict. In addition, parts not described in detail in the embodiments can be implemented using existing technologies.
[0104] The following further illustrates the present application in conjunction with specific application examples / comparative examples in order to better understand the above technical solutions of the present application. It should be understood that the following are merely partial examples and are not intended to limit the present application.
[0105] In order to cope with the simulation of workload distribution, inter-core data transfer, and neural network data interaction, it is necessary to first define a prototype of the on-chip network architecture based on heterogeneous in-memory computing.
[0106] Figure 2 The figure is a schematic diagram of a prototype structure of a heterogeneous in-memory computing network-on-chip architecture according to an exemplary embodiment.
[0107] Reference Figure 2 As shown in (a), a prototype of a heterogeneous in-memory computing network-on-chip architecture includes a heterogeneous in-memory computing-based heterogeneous in-memory computing network-on-chip accelerator, a heterogeneous in-memory computing-based heterogeneous in-memory computing network-on-chip core, and a heterogeneous in-memory computing array. Figure 2 As shown in (b), it also includes PE peripheral circuits, refer to Figure 2 As shown in (c), it also includes a weight and activation value partitioning module.
[0108] Specifically, refer to Figure 2 As shown in (c), the weights and activation values are divided according to the core and PE peripheral circuits of the heterogeneous on-chip network based on heterogeneous in-memory computing, and the data transmission volume between the cores is calculated.
[0109] The computational process for neural network layers, such as convolutional or linear layers, involves convolution and activation matrix multiplication. When the activation matrix size exceeds the core or PE capacity within a core, the weight matrix must be partitioned into multiple blocks based on the core or PE capacity and mapped to multiple cores or PE units within the core. Simultaneously, the activation values are transferred to the corresponding cores or PEs based on the computational dependencies of the original matrix multiplication to complete the matrix multiplication after the blocks are partitioned.
[0110] When calculating the entire neural network, the output of the previous layer serves as the input to the next layer, that is, the activation value of the next layer. After each layer is divided according to the above method, the starting core and the ending core position of the data transmission are determined.
[0111] The amount of activation value data that needs to be transmitted between cores is determined based on the result of matrix multiplication of each block after segmentation.
[0112] Since the capacity difference between cores prevents the load balancing method of homogeneous systems from being directly applied to heterogeneous architectures, a backtracking algorithm is used to implement workload distribution and mapping of heterogeneous cores. Each core is simulated to obtain the preset target indicators, namely, calculation delay time, power consumption and area targets.
[0113] Figure 3 The figure is a simulation flow chart of a heterogeneous in-memory computing network-on-chip according to an exemplary embodiment.
[0114] Reference Figure 3 As shown, the present application proposes a heterogeneous in-memory computing on-chip network architecture based on prior knowledge. As a simulator, the simulation process of the optimized heterogeneous in-memory computing on-chip network architecture includes three processes: computing core simulation, on-chip network simulation and data interaction.
[0115] The compute core simulation uses neural network parameters (CNN parameters) and heterogeneous architectures to perform load slicing and mapping, calculating data transfer volume and corresponding node injection rates. This mapping information and injection rates are then exported to the on-chip network simulation via data exchange, yielding path latency and on-chip network performance parameters. The latency co-simulation model is combined with heterogeneous power and area models, and the resulting data is used to evaluate and output the architecture's power consumption, area, and latency.
[0116] The NoC simulation module uses NoC parameters to generate the NoC structure, derives data flow scheduling based on mapping information, and decomposes paths into unit routing operations. The traffic simulation module simulates routing and calculates path delays. The power / area model evaluates the power and area of the NoC, and feeds the power and area of the NoC and path delays back to the compute core simulation.
[0117] Reference Figure 3 As shown, based on the configuration of neural network and optimized heterogeneous in-memory computing on-chip network architecture, workload mapping and inter-core data transmission path are generated, and the computational delay time, power consumption and area results of the interconnection circuit are obtained through on-chip network simulation, and co-simulation is performed with the simulation results of the heterogeneous in-memory computing core.
[0118] Figure 4 The figure is a search flow chart of a heterogeneous in-memory computing network-on-chip according to an exemplary embodiment.
[0119] Reference Figure 4 As shown, a heterogeneous in-memory computing on-chip network architecture based on prior knowledge proposed in this application is used for search.
[0120] Specifically, the neural network architecture is first input. Based on the hardware architecture parameters, the neural network is segmented by input channel. The segmented load is mapped onto the hardware architecture, and the data flow scheduling is simultaneously obtained. Each time the circuit metrics of the new architecture are obtained, they are compared with the target to determine whether to update the hardware architecture and proceed to the next search.
[0121] The use of heterogeneous architecture can be regarded as the characteristics of diverse combinations of homogeneous architectures. Prior knowledge is introduced to reduce the scale of the search space, and the search complexity is reduced through a three-stage search strategy.
[0122] In the first stage, based on the pre-calculated core computing latency, power consumption and area (PPA) change trends, a homogeneous architecture configuration that is closest to the target PPA indicator is quickly located as the starting point.
[0123] In the second stage, based on the homogeneous architecture configuration, we explored the inter-layer heterogeneous architecture configuration between different neural network layers. According to the specific computing requirements of different layers, by adjusting the type of heterogeneous in-memory computing, the number of PEs and array size, the core configuration of each layer was gradually optimized according to the target indicator PPA trend chart to approach or meet the target indicator PPA requirements.
[0124] If inter-layer heterogeneous configurations still fail to meet requirements, we move on to the third phase, employing intra-layer heterogeneous search. This allows each core within the same network layer to be independently configured, further refining performance requirements. To ensure the effectiveness of intra-layer search configurations, fine-tuning checks and adjustments are performed after each operation to avoid situations where weights exceed core capacity or become excessively redundant.
[0125] Furthermore, a customized simulated annealing algorithm for heterogeneous architecture search based on prior knowledge of heterogeneous in-memory computing network-on-chip architectures is proposed. This algorithm utilizes predefined search operators, such as those for adjusting the core type of heterogeneous in-memory computing, the number of physical processing units (PEs), the array size, the mapping order of heterogeneous in-memory computing cores, and the position of neural network cores, to achieve fast and efficient search. The simulated annealing algorithm uses a fusion of metrics (FoM)—computational latency, power consumption, and area—as feedback at each search step to efficiently guide the search process toward a more optimal heterogeneous in-memory computing network-on-chip architecture configuration. The simulated annealing algorithm also achieves a balance between search efficiency and accuracy by gradually reducing the number of operations, allowing the search to converge as it approaches the target configuration.
[0126] Finally, the optimized heterogeneous in-memory computing on-chip network architecture configuration parameters are output.
[0127] Figure 5The figure is a flowchart showing an optimization search process according to an exemplary embodiment.
[0128] Reference Figure 5 In the process of searching using a heterogeneous in-memory computing on-chip network architecture provided by this application, since each step in the search process needs to be simulated, after multiple iterations, the simulation time becomes unacceptable.
[0129] Each simulation will re-simulate all cores of the current heterogeneous architecture to obtain the target index, which significantly reduces the overall search speed. This embodiment uses a lookup table (LUT) for simulation acceleration.
[0130] Reference Figure 5 As shown in the figure, the red arrows represent the original search and simulation process, and the green arrows represent the process accelerated by the lookup table. All possible heterogeneous in-memory computing core configurations are pre-simulated, and the target indicator results are stored to build the lookup table.
[0131] During the search process, the lookup table is called, and pre-computed results are reloaded and accumulated, eliminating redundant core simulations. After accelerating core simulation, on-chip network simulation becomes the most time-consuming part of the search process. To further optimize this, the frequency of on-chip network simulation is reduced during the search process, as inter-core data transfer patterns gradually stabilize during the search.
[0132] Finally, a force-directed algorithm is used to adjust the core layout of the heterogeneous in-memory computing network on chip. According to the preset interaction force, the position of the core is adjusted, and the interaction force between the cores in the topological connection is adjusted to minimize the layout area and delay, ensuring the compact arrangement of the heterogeneous cores.
[0133] This application proposes a method for optimizing heterogeneous in-memory computing on-chip network architecture based on prior knowledge. Compared with the method of searching heterogeneous in-memory computing on-chip networks without using a preset lookup table and reducing the frequency, the search time of the method proposed in this application can be reduced by more than 2 times. In the application scenario of the ResNet-18 neural network, compared with the homogeneous architecture, the convergence index (FoM) can be reduced by up to 37.41%, significantly improving the computing delay time, power consumption and area optimization of heterogeneous network chips.
[0134] In the embodiments of the present application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0135] The above describes the specific embodiments of the present application. It should be understood that the present application is not limited to the specific embodiments described above, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the substantive content of the present application. The above preferred features may be used in any combination as long as they do not conflict with each other.
Claims
1. A method for optimizing heterogeneous in-memory computing network-on-chip architecture based on prior knowledge, characterized in that: include: Constructing a joint simulation platform for heterogeneous in-memory computing network-on-chip, and using the joint simulation platform to perform joint pipeline simulation of computing and communication delays on the heterogeneous in-memory computing network-on-chip; A three-stage search strategy based on prior knowledge is used to search within the heterogeneous architecture space of the heterogeneous in-memory computing network-on-chip, and a preset lookup table and a preset simulated annealing algorithm are used to optimize the search process to determine a target architecture configuration of the heterogeneous in-memory computing network-on-chip; According to the target architecture configuration of the heterogeneous in-memory computing network on chip, a force-directed algorithm is used to optimize the core layout of the heterogeneous in-memory computing network on chip, and an optimized heterogeneous in-memory computing network on chip architecture is determined.
2. The method according to claim 1, characterized in that The joint simulation platform includes a heterogeneous in-memory computing core simulator and an on-chip network data transmission simulator. The heterogeneous in-memory computing core simulator is used to simulate the computing behavior of various types of heterogeneous in-memory computing cores, and the on-chip network data transmission simulator is used to model the data relationships and inter-layer communication behaviors of neural networks.
3. The method according to claim 1, characterized in that The method adopts a three-stage search strategy based on prior knowledge to search within the heterogeneous architecture space of the heterogeneous in-memory computing network-on-chip, and uses a preset lookup table and a preset simulated annealing algorithm to optimize the search process and determine the target architecture configuration of the heterogeneous in-memory computing network-on-chip, including: performing a homogeneous architecture search at the neural network layer according to preset target indicators to determine an initial homogeneous architecture configuration corresponding to the preset target indicators, wherein the preset target indicators represent computational latency, power consumption, and area targets of the heterogeneous in-memory computing core configuration; On the initial homogeneous architecture configuration, according to the preset target indicators, perform an inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer, optimize the architecture configuration of the heterogeneous in-memory computing core corresponding to each neural network layer, and determine the inter-layer heterogeneous in-memory computing core configuration; According to the preset target indicators, an intra-layer heterogeneous architecture search is performed on the heterogeneous in-memory computing core corresponding to each neural network layer, and the heterogeneous in-memory computing core corresponding to each neural network layer is personalized configured to determine the target architecture configuration of the heterogeneous in-memory computing network on chip.
4. The method according to claim 3, characterized in that The method further comprises: performing an inter-layer heterogeneous architecture search between the heterogeneous in-memory computing cores corresponding to each neural network layer based on the preset target indicators on the initial homogeneous architecture configuration, optimizing the architecture configuration of the heterogeneous in-memory computing cores corresponding to each neural network layer, and determining the inter-layer heterogeneous in-memory computing core configuration, including: According to the computing delay time, power consumption and area targets of the heterogeneous in-memory computing core configuration, an inter-layer heterogeneous architecture search is performed between the heterogeneous in-memory computing cores corresponding to each of the neural network layers to determine the core type, number of PEs and array size of the heterogeneous in-memory computing corresponding to each of the neural network layers corresponding to the computing delay time, power consumption and area targets of the heterogeneous in-memory computing core configuration, and determine the inter-layer heterogeneous in-memory computing core configuration.
5. The method according to claim 4, characterized in that The method further comprises: performing an intra-layer heterogeneous architecture search on the heterogeneous in-memory computing core corresponding to each neural network layer according to the preset target indicator, performing personalized configuration on the heterogeneous in-memory computing core corresponding to each neural network layer, and determining a target architecture configuration of the heterogeneous in-memory computing network-on-chip, including: Determining the total capacity of the heterogeneous in-memory computing cores corresponding to each neural network layer based on the inter-layer heterogeneous in-memory computing core configuration; Based on the preset weight storage target of each neural network and the total capacity of the heterogeneous in-memory computing cores corresponding to each neural network layer, the number of heterogeneous in-memory computing cores corresponding to each neural network layer is determined, and the target architectural configuration of the heterogeneous in-memory computing network-on-chip is determined.
6. The method according to claim 5, characterized in that The step of determining the number of heterogeneous in-memory computing cores corresponding to each neural network layer and determining the target architecture configuration of the heterogeneous in-memory computing network-on-chip based on a preset weight storage target of each neural network and a total capacity of heterogeneous in-memory computing cores corresponding to each neural network layer comprises: If the preset weight storage target is greater than the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer, adding a preset number of cores to the heterogeneous in-memory computing corresponding to the neural network layer; If the preset weight storage target is smaller than the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer and the difference between the total capacity of the heterogeneous in-memory computing cores corresponding to the neural network layer and the preset weight storage target is greater than the minimum core capacity of the heterogeneous in-memory computing corresponding to the neural network layer, remove the minimum core of the heterogeneous in-memory computing corresponding to the neural network layer.
7. The method according to claim 3, characterized in that The method adopts a three-stage search strategy based on prior knowledge to search within the heterogeneous architecture space of the heterogeneous in-memory computing network-on-chip, and uses a preset lookup table and a preset simulated annealing algorithm to optimize the search process to determine the target architecture configuration of the heterogeneous in-memory computing network-on-chip, further comprising: According to the preset target indicator, calling the preset lookup table to determine the heterogeneous in-memory computing core configuration corresponding to the preset target indicator; According to the heterogeneous in-memory computing core configuration corresponding to the preset target indicator and the preset fusion indicator, a predefined search operator is used to optimize the heterogeneous in-memory computing core configuration.
8. The method according to claim 7, characterized in that The preset fusion index represents a fusion index of the computing delay time, power consumption and area target of the computing core configuration in the heterogeneous memory; The predefined search operators include a search operator for adjusting the type of the heterogeneous in-memory computing core, a search operator for adjusting the number of PEs, a search operator for adjusting the array size, a search operator for adjusting the mapping order of the heterogeneous in-memory computing core, and a search operator for exchanging the position of the heterogeneous in-memory computing core corresponding to the neural network layer.
9. The method according to claim 7, characterized in that The method for constructing the preset lookup table includes: Using the joint simulation platform to simulate the core configurations of all heterogeneous in-memory computing, and determining the target indicator corresponding to each core configuration of the heterogeneous in-memory computing; The preset lookup table is constructed according to all the heterogeneous in-memory computing core configurations and the target indicators corresponding to each of the heterogeneous in-memory computing core configurations.
10. The method according to claim 1, characterized in that The step of configuring the target architecture of the heterogeneous in-memory computing network-on-chip and optimizing the core layout of the heterogeneous in-memory computing network-on-chip using a force-directed algorithm to determine an optimized heterogeneous in-memory computing network-on-chip architecture includes: Determining the interaction force between the cores of the heterogeneous in-memory computing according to the topological connection of the heterogeneous in-memory computing network-on-chip; According to the preset interaction force, the position of the core of the heterogeneous in-memory computing and the interaction force between the cores of the heterogeneous in-memory computing are adjusted to determine the optimized heterogeneous in-memory computing on-chip network architecture.