Software and hardware collaborative optimization method of hybrid in-memory architecture
By constructing a context-aware predictive agent model and a multi-objective optimization framework, the problem of hardware and software co-optimization of hybrid in-memory computing architecture was solved, achieving efficient and accurate architecture optimization of in-memory computing chips and improving the execution efficiency and adaptability of AI algorithms.
Patent Information
- Application Number
- CN202511071004.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Existing technologies struggle to effectively optimize the complexities and fine-grained parameters of hybrid in-memory computing architectures through hardware-software co-optimization, particularly when adapting to diverse artificial intelligence algorithms such as Transformer and Mamba, where they lack refined multi-objective optimization capabilities.
We construct a context-aware predictive agent model, fine-tune the in-memory computing architecture through a multi-objective optimization framework, extract algorithm features and combine them with hardware parameters, use self-attention mechanism and parallel prediction network to predict PPA index, and embed a multi-objective evolutionary algorithm to search for Pareto optimal configuration.
It significantly improves the overall performance of in-memory computing chips when executing target AI algorithms, enhances execution efficiency and adaptability, provides diverse hardware optimization options, meets area and power consumption requirements, and supports the development of efficient compilers.
Smart Images

Figure CN120973728A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer architecture and artificial intelligence, and in particular to a hardware and software co-optimization method for hybrid in-memory architecture. Background Technology
[0002] With the rapid development of artificial intelligence, especially deep learning, large language models (such as Transformer), and emerging complex models (such as Mamba), unprecedented challenges have been placed on the computing power, energy efficiency, and bandwidth of computing hardware. Traditional von Neumann architectures face bottlenecks in handling large-scale AI computations due to limitations imposed by the "memory wall" and "power wall." In-Memory Computing (IMC) technology, by integrating computing units directly near or inside storage units, reduces latency and power consumption caused by data movement and is considered one of the effective ways to address the challenges of AI computing.
[0003] In-memory computing chip architecture is complex, containing numerous configurable parameters and microarchitectural options, such as the type, number, and precision configuration of in-memory computing units, the pipeline design of neural network processors, the interconnect topology of multi-core systems, and memory hierarchy interfaces. Different combinations of these parameters create an extremely broad design space. Furthermore, different AI algorithms have varying hardware architecture requirements; their computational graph characteristics, tensor space dimensions, sparsity, and the proportion of mixed-precision operations all differ, leading to significant variations in the adaptability and execution performance of specific in-memory computing architecture configurations for different AI algorithms.
[0004] Therefore, how to scientifically and efficiently explore the design space of in-memory computing architectures has become a key technical problem that urgently needs to be solved in the field of in-memory computing chip design. Existing technologies have already made some explorations in this area. For example, in order to accelerate design space exploration (DSE), researchers have begun to use novel generative models to directly produce excellent hardware designs. Wang et al. proposed a DSE method based on a diffusion model (Y.Wang et al., "DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-DrivenOptimization," Proceedings of the 56th Annual IEEE / ACM International Symposium on Microarchitecture (MICRO'23)). This method learns from existing data distributions and generates high-performance DNN accelerator configurations through iterative denoising.
[0005] However, such methods still have shortcomings: First, they mainly focus on generating macroscopic parameters for general-purpose DNN accelerators, without delving into the microscopic level of the hybrid in-memory computing architecture that this invention focuses on, such as the precise ratio of mixed-signal computing units and the specific parameters like ADC / DAC accuracy. The complex interactions between these parameters are difficult for general-purpose generative models to capture. Second, this method lacks a clear mechanism to systematically utilize the "contextual features" of AI algorithms (such as sparsity, operation type distribution, etc.) as conditional inputs to guide the generation process, thus limiting its ability to perform fine-grained optimization for specific algorithm categories. Finally, directly generating a configuration set that satisfies multi-objective Pareto optimality is a challenge for diffusion models, while this invention embeds a surrogate model into the framework of traditional multi-objective optimization algorithms, enabling more direct multi-objective trade-offs.
[0006] On the other hand, in the field of hardware-software co-design, Geng et al. proposed a co-exploration framework called Gibbon (Y.Geng et al., "Gibbon: An efficient co-exploration framework of nnmodel and processing-in-memory architecture," 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA).), which can simultaneously search for neural network (NN) model structures and in-memory processing (PIM) hardware architectures.
[0007] Nevertheless, this technology also has its inherent limitations: First, its problem is set as "cooperative search," which involves changing both the algorithm and the hardware simultaneously. This has its advantages when searching for entirely new solutions, but it cannot solve the problem of how to deeply optimize and adapt the AI algorithm (such as Transformer) to the hardware architecture when the algorithm is relatively fixed. The latter is precisely the scenario that this invention focuses on. Second, in order to make the huge "algorithm-hardware" joint search space controllable, cooperative search frameworks usually need to simplify the parameter space of the hardware, making it difficult to perform fine-grained parameter configuration and exploration of the microstructure of in-memory computation as this invention does. Third, this method aims to find one or a few optimal "algorithm-hardware" matching pairs, while this invention aims to output a set of Pareto optimal hardware architecture solutions, providing designers with multiple options to weigh different PPA objectives when meeting specific algorithm requirements.
[0008] In summary, existing technologies or generative models are ill-suited to the complexity and fine-grained parameters of hybrid in-memory computing architectures, or they rely on simplified hardware models for "algorithm-hardware" collaborative search. They generally lack the ability proposed in this invention to perform fine-grained multi-objective optimization of in-memory computing architectures for specific algorithm contexts. Summary of the Invention
[0009] This invention aims to address the design and optimization challenges posed by the vast design space and complex parameter configuration of existing in-memory computing architectures when adapting to diverse artificial intelligence (AI) algorithms, particularly complex large models such as Transformer and Mamba. It provides a hardware-software co-optimization algorithm for hybrid in-memory computing architectures. The core of this method lies in constructing and utilizing a data-driven predictive proxy model, and then fine-tuning the key configurable parameters and microstructures within the in-memory computing architecture through a multi-objective optimization framework. Specifically, this includes: jointly characterizing the target AI algorithm (such as Transformer, Mamba, etc.) and the in-memory computing architecture, extracting contextual features such as computation graph features, tensor space dimension, sparsity, and the proportion of mixed-precision operations, and parameterizing the in-memory computing architecture's in-memory computing unit configuration, neural network processor pipeline, multi-core interconnect topology, and memory hierarchy interface; constructing an offline benchmark dataset containing feasible domain information of the architecture and corresponding performance, power consumption, and area (PPA) metrics; developing and training a high-fidelity predictive surrogate model with context awareness capabilities. This model uses an embedding method to process discrete architecture parameters, employs an encoder with a self-attention mechanism to learn feature associations, and achieves joint prediction of multi-dimensional PPA metrics through a parallel prediction network. A feasible domain constraint learning mechanism is introduced during training; and embedding the trained surrogate model into a multi-objective evolutionary algorithm, with energy efficiency ratio, average computing power utilization, and model execution latency as optimization objectives, while simultaneously satisfying chip area efficiency and power consumption constraints, to search for the Pareto optimal set of in-memory computing architecture configurations. This invention can significantly improve the execution performance and adaptability of AI algorithms on in-memory computing chips, provide a scientific architecture optimization scheme for the design of high-performance in-memory computing chips, and provide hardware basis for subsequent compiler development.
[0010] The optimized architecture, which has been verified in practice, is the final output of this invention and provides a hardware basis for subsequent compiler development. The compiler will, based on the optimized architecture characteristics, including the ratio and capability of pure digital in-memory computing or mixed-precision analog-digital in-memory computing units, the specific parameters of the static random access memory pipeline within the neural network processor, multi-core heterogeneous topology (if applicable), and the optimized data path between static random access memory and three-dimensional stacked dynamic random access memory, achieve efficient and accurate mapping from AI models to in-memory computing unit hardware instructions.
[0011] The present invention is achieved by at least one of the following technical solutions.
[0012] A hardware-software co-optimization method for a hybrid in-memory architecture includes the following steps:
[0013] S1. Extract the context features of the target AI algorithm to form the AI algorithm context feature vector, and perform parameterized description of the configurable parameters and microstructure of the in-memory computing architecture to obtain multiple sets of in-memory computing architecture parameter configuration vectors, forming an in-memory computing architecture parameter set.
[0014] S2. Based on the in-memory computing architecture parameter set in step S1, construct an offline benchmark dataset. The offline benchmark dataset includes multiple sets of in-memory computing architecture parameter configurations sampled from the in-memory computing architecture parameter set, corresponding AI algorithm context feature vectors, PPA indicators obtained through hardware simulation, and architecture feasible domain information.
[0015] S3. Based on the offline benchmark dataset, train a predictive agent model with context awareness. The predictive agent model can predict the corresponding in-memory computing metrics according to the input in-memory computing architecture parameter configuration vector and AI algorithm context feature vector. The training process includes using an embedding method to process discrete architecture parameters, using an encoder with a self-attention mechanism to learn the complex relationships between features, using a parallel prediction network to achieve joint prediction of multi-dimensional PPA metrics, and introducing a feasible domain constraint learning mechanism. While supervising the learning of PPA relationships between feasible architecture configurations, the mechanism introduces an adversarial negative sampling mechanism to discriminately learn infeasible configurations.
[0016] S4. The trained predictive agent model is used as an evaluation function and embedded into a multi-objective optimization algorithm framework;
[0017] S5. Define multiple optimization objectives for the in-memory computing architecture. The optimization objectives are based on the PPA index and meet chip area efficiency and power consumption constraints.
[0018] S6. Run the multi-objective optimization algorithm, use the predictive surrogate model to quickly evaluate the PPA index of the candidate in-memory computing architecture configuration, search and output a set of Pareto optimal in-memory computing architecture parameter configuration schemes that meet the preset optimization objectives.
[0019] Furthermore, the contextual features include computation graph features, tensor space dimension, sparsity, and the proportion of mixed-precision operations.
[0020] Furthermore, the configurable parameters of the in-memory computing architecture include in-memory computing unit configuration parameters, neural network processor pipeline parameters, multi-core interconnect topology, and storage hierarchy interface parameters.
[0021] Furthermore, the pipeline parameters of the neural network processor include the organization of its internal static random access memory and the number of pipeline stages; the multi-core interconnect topology includes the connection method and bandwidth between different cores; and the memory hierarchy interface parameters include the optimized data path characteristics and prefetch logic between static random access memory and three-dimensional stacked dynamic random access memory.
[0022] Furthermore, the offline benchmark dataset distinguishes between feasible architecture configuration points and infeasible architecture configuration points. Feasible architecture configuration points refer to configurations that can successfully compile, map, and execute AI algorithms and whose PPA metrics are within a preset reasonable range. Infeasible architecture configuration points refer to configurations that cannot meet the basic requirements due to physical constraints, compilation errors, or extremely poor performance.
[0023] Furthermore, in step S3, the feasible domain constraint learning mechanism aims to enable the predictive agent model to minimize the prediction error at feasible architecture configuration points through supervised learning, and to punish the tendency of infeasible architecture configuration points to produce good PPA index predictions through discriminative learning, so as to ensure that the predictive agent model has accurate architecture evaluation capabilities.
[0024] The adversarial negative sampling mechanism is used to enhance the predictive agent model's ability to identify infeasible configurations and its robustness in prediction.
[0025] Furthermore, the structure of the predictive agent model includes: mapping discrete in-memory computing architecture parameter configuration vectors into continuous vectors through an embedding layer; processing the continuous vectors using an encoder layer with a self-attention mechanism to generate a unified architecture feature representation; fusing the unified architecture feature representation with the AI algorithm context feature vector; inputting the fused features into a parallel multi-head prediction network and aggregating the outputs of each prediction head through a top-level attention mechanism or direct output method to generate a multi-dimensional PPA index prediction vector.
[0026] A system for implementing the aforementioned hardware-software co-optimization method for a hybrid in-memory architecture includes:
[0027] The feature extraction and parameterization module is used to extract the context feature vector of the target AI algorithm and to provide a parameterized description of the configurable parameters and microstructure of the in-memory computing architecture.
[0028] The dataset construction module is used to build an offline benchmark dataset. The dataset contains multiple sets of in-memory computing architecture parameter configurations, corresponding AI algorithm context feature vectors, and PPA indicators obtained through hardware simulation, and records architecture feasibility domain information.
[0029] The proxy model training module is used to train a predictive proxy model with context awareness capabilities. The proxy model can predict the corresponding PPA index based on the input in-memory computing architecture parameter configuration and AI algorithm context feature vector.
[0030] The multi-objective optimization module is used to embed the trained predictive agent model as an evaluation function into the multi-objective optimization algorithm framework. Based on multiple predefined optimization objectives based on PPA indicators, the multi-objective optimization algorithm is run to quickly evaluate the PPA indicators of candidate in-memory computing architecture configurations using the predictive agent model, and to search for and output a set of Pareto optimal in-memory computing architecture parameter configuration schemes.
[0031] A computer device according to the present invention includes a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, which, when executed by the processor, causes the processor to implement the method described herein.
[0032] The present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the method described herein.
[0033] The present invention has the following advantages and effects compared with the prior art:
[0034] 1) By performing more refined joint feature representation (including computation graph features, mixed precision ratio, multi-core interconnect topology, etc.) of AI algorithms (such as Transformer, Mamba, etc.) and in-memory computing architecture, and building a proxy model with context-aware capabilities, it is possible to more accurately capture the PPA index of in-memory computing architecture parameters under different AI algorithm scenarios, thereby achieving fine optimization for the adaptability of specific AI algorithms.
[0035] 2) High-fidelity predictive surrogate models (using embedding, self-attention encoders, and parallel prediction networks) are used to replace time-consuming hardware simulations for PPA evaluation. A feasible region constraint learning mechanism (including adversarial negative sampling) is introduced, which significantly accelerates the optimization process of multi-objective optimization algorithms, improves the accuracy of model predictions and the ability to identify feasible regions, and makes it possible to find excellent architecture configurations in a broad design space with complex constraints.
[0036] 3) By adopting a multi-objective optimization framework, it can simultaneously optimize multiple conflicting PPA indicators (energy efficiency ratio, performance, area efficiency, and delay), and take into account actual power consumption and area efficiency constraints to obtain a set of Pareto optimal solutions, providing designers with diverse trade-off options.
[0037] 4) By finely optimizing the key configurable parameters and microstructures within the in-memory computing architecture, the overall performance of the in-memory computing chip when executing target AI algorithms can be significantly improved, such as higher energy efficiency, computing power utilization, lower latency, and meeting area and power consumption requirements.
[0038] 5) The optimized architecture configuration output by this method has been verified in practice and can directly provide clear hardware basis for subsequent compiler development. It guides the compiler to achieve efficient and accurate mapping of AI models to hardware instructions based on the optimized architecture characteristics (such as mixed precision unit ratio, neural network processor pipeline parameters, multi-core topology, memory path, etc.), thereby forming a closed loop of software and hardware co-design. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of the architecture of the predictive agent model in the embodiment.
[0040] Figure 2 This is a flowchart of a hardware and software co-optimization method for a hybrid in-memory architecture.
[0041] Figure 3 This is a flowchart of a multi-objective optimization method according to an embodiment. Detailed Implementation
[0042] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0043] like Figures 1-3 As shown in this embodiment, a hardware-software co-optimization method for a hybrid in-memory architecture includes the following steps:
[0044] S1. For the target AI algorithm (such as typical models like Transformer and Mamba), extract its model context features, including computation graph features (such as operation type distribution and dependency complexity), tensor space dimension, sparsity, and the proportion of mixed-precision operations, to form an AI algorithm context feature vector a. Simultaneously, parameterize the configurable parameters and microstructure of the in-memory computing architecture to obtain multiple sets of in-memory computing architecture parameter configuration vectors. These multiple sets of in-memory computing architecture parameter configuration vectors form an in-memory computing architecture parameter set. The configurable parameters of the in-memory computing architecture include in-memory computing unit configuration parameters (such as the ratio and capability of pure digital in-memory computing or analog-digital hybrid in-memory computing mixed-precision units), neural network processor pipeline parameters (such as the specific parameters of the internal static random access memory pipeline), multi-core interconnect topology (such as the connection method and bandwidth between different cores), and storage hierarchy interface parameters (such as the optimized data path characteristics and prefetch logic between static random access memory and three-dimensional stacked dynamic random access memory).
[0045] As one example, in the AI algorithm and in-memory computing architecture feature characterization stage, feature extraction is first performed on the target AI algorithm. Taking the Transformer model as an example, the feature set of the target AI algorithm is defined as a = (c1, c2, ..., c...). i ), where each feature c i This represents a quantifiable feature of the AI algorithm. Taking the Transformer model as an example, the key features selected include c1 to c6, where c1 is the computation graph feature (such as operation type distribution and critical path length), c2 is the number of model parameters, c3 is the ratio of attention layer to feedforward network layer, c4 is the average tensor dimension, c5 is the proportion of mixed-precision operations (such as the ratio of FP16 and INT8 operations), and c6 is the average sparsity. Simultaneously, the configurable parameters of the in-memory computing architecture are parameterized, and the in-memory computing architecture parameter set is defined as x = (p1, p2, ..., p6). The selected key in-memory computing architecture parameters include p1 to p6, where p1 is the number of core computing units in the mixed-precision in-memory computing unit, p2 is the conversion precision of the analog-to-digital converter or digital-to-analog converter, p3 is the number of stages in the neural network processor pipeline, p4 is the number of in-memory computing cores, p5 is the link bandwidth of the on-chip network, and p6 is the size of the prefetch buffer of the storage interface from static random access memory to the three-dimensional stacked dynamic random access memory hierarchy. The entire design space X is the set of all possible combinations of in-memory computing architecture parameters.
[0046] S2. Based on the AI algorithm features and in-memory computing architecture parameters defined in step S1, construct an offline benchmark dataset D through systematic and precise hardware simulation experiments or actual hardware tests. Each entry in dataset D contains an in-memory computing architecture configuration x and its corresponding PPA (Power Performance, Performance, Area Efficiency) index y. This dataset also contains design points D that have been verified as feasible. feasible (i.e., the set of feasible architectural configuration points with effective performance metrics y) and infeasible design points D. infeasible (i.e., the set of infeasible architecture configuration points, without effective performance metrics), that is, D = D feasible ∪D infeasible A feasible architecture configuration point refers to an architecture configuration point that can successfully compile, map, and execute AI algorithms, and whose PPA (Power, Performance, and Area) metrics are within a preset reasonable range. An infeasible architecture configuration point is an architecture configuration point that cannot meet the basic requirements due to physical constraints, compilation errors, or extremely poor performance.
[0047] In the stage of building the offline benchmark dataset, a large-scale experiment was conducted using a precise hardware simulation platform (such as a Cycle-accurate simulator) to build the offline benchmark dataset D for a series of AI algorithm context feature vectors a and in-memory computing architecture parameter configurations x.
[0048] Dataset D is further divided into a set of feasible architecture configuration points D feasible and the set of infeasible architecture configuration points D infeasible The feasible architecture configuration point refers to a configuration that can successfully compile, map, and execute the AI algorithm and whose PPA (Programme-Performance Ratio) is within a preset reasonable range. The infeasible architecture configuration point refers to a configuration that cannot meet the basic requirements due to physical constraints, compilation errors, or extremely poor performance. That is, D = D feasible ∪D infeasible =={(x1,y1),(x2,y2),…,((x M ,y M ))∪{x′1,x′2,…,x′ N}. M and N are the number of feasible and infeasible architecture configuration points, respectively, where y M It is a multi-dimensional PPA indicator vector that includes energy efficiency ratio, performance (such as throughput or execution time), area efficiency, etc. M It is the parameter vector of the Mth feasible architecture configuration point in the dataset, x′ N It is the parameter vector of the Nth infeasible architecture configuration point in the dataset.
[0049] S3. Train a context-aware predictive agent model f using the offline benchmark dataset constructed in step S2. θ(x), where x represents the in-memory computing architecture configuration and θ represents the parameters of the proxy model. The function of this proxy model is to predict the corresponding PPA metric vector based on the input in-memory computing architecture configuration parameters x and the AI algorithm context feature vector a. (e.g., delay), that is The process of this proxy model is as follows:
[0050] The input in-memory computing architecture configuration x is mapped into a continuous vector through the embedding layer, and the continuous parameters of the continuous vector can be processed through the linear layer.
[0051] The encoding layer with self-attention mechanism is used to process continuous vectors, learn complex relationships between features, and generate a unified architecture feature representation E. arch ; Represent architectural features as E arch The fused feature E is obtained by fusing it with the AI algorithm's context feature vector 'a' (e.g., concatenating them and passing them through a fully connected layer). fused .
[0052] Fusion feature E fused The input is a parallel multi-head prediction network, where each head is responsible for predicting one PPA dimension (e.g., one head predicts energy efficiency ratio, one head predicts latency, and one head predicts area efficiency), ultimately generating a multi-dimensional PPA prediction vector. The set of feasible data points, i.e., feasible architecture configuration points, is D. feasible The above method minimizes the prediction vector of the PPA index through supervised learning. The error between the actual PPA metric y and the actual PPA metric y (e.g., mean squared error); simultaneously, an adversarial negative sampling mechanism is introduced to include infeasible configuration points, i.e., the set of infeasible architecture configuration points D. infeasible Discriminative learning is performed using negative samples. By introducing a penalty term into the loss function, the tendency of the predictive agent model to produce good PPA (Professional Performance Average) predictions for these infeasible points is suppressed, ensuring that the model has accurate architecture evaluation capabilities and the ability to identify feasible regions.
[0053] During training, the surrogate model aims to minimize the difference between predicted and actual performance metrics, effectively distinguish between feasible and infeasible configurations, and avoid overly optimistic estimates for unseen design points. Its training process incorporates a feasible domain constraint learning mechanism. This mechanism aims to enable the predictive surrogate model to minimize prediction errors at feasible architectural configurations through supervised learning, while treating infeasible architectural configurations as negative samples and penalizing their tendency to produce good PPA (Professional Performance Area) predictions through discriminative learning, thus ensuring the model possesses accurate architectural evaluation capabilities. The adversarial negative sampling mechanism enhances the model's ability to identify infeasible configurations and improves its prediction robustness.
[0054] loss function The design incorporates supervised learning loss and feasible region discrimination loss, and is not limited to the set of feasible architecture configuration points D. feasible The standard supervised learning loss term also includes an additional regularization term to adjust the model's behavior on two special classes of data points:
[0055] Among them, the loss function of the standard supervised learning loss term The definition is shown in formula (2):
[0056]
[0057] in, It is in the set of feasible architecture configuration points D feasible The supervised learning loss, expressed in the norm form of the mean squared error, is used to minimize the model's dependence on the input x. i Predictive performance index f θ (x i ) and the true performance index vector y i The differences between them. It is a conservative term. Opt(f) θ ) represents an attempt to find a model f that makes the current learning model f θ (x) predicts the approximately stochastic optimization process for the optimal performance index. This is the negative sample configuration generated by the process. By minimizing the negative value of this term, the model is encouraged to predict higher costs for these potentially overestimated points, thereby suppressing overly optimistic estimates of unknown regions. α is a hyperparameter and satisfies α>0.
[0058] Final loss function As shown in formula (3):
[0059]
[0060] Where β is a hyperparameter satisfying β>0, Used to integrate the set of infeasible architecture configuration points D infeasible Information. By using a method similar to formula (2), the model is encouraged to consider these infeasible configurations x′. i The predicted higher cost value f θ (x′ i This is used to penalize the model for its tendency to predict good performance index values for infeasible configuration points.
[0061] S4, in the predictive agent model f θ (x) After training, it is embedded into a multi-objective optimization algorithm framework and used as an evaluation function. The multi-objective optimization algorithm used is a decomposition-based multi-objective evolutionary algorithm. The optimization problem is defined as: finding a set of Pareto optimal in-memory computing architecture parameter configurations x. *This allows for a predefined set of optimization objectives O = {obj1, obj2, ..., obj}. k At the same time, obj achieves optimality. k For the k-th target, as shown in formula (4):
[0062] x * ∈ParetoOptimalSet(obj1(f θ (x,a)),obj2(f θ (x,a)),…,obj k (f θ (x,a))) (4)
[0063] Where ParetoOptimalSet(.) represents the Pareto optimal set, which contains a series of optimal solutions for which no other objective can be improved without sacrificing at least one objective. 'a' is the feature vector of the target AI algorithm, and 'f' is the optimal solution. θ (x, a) represents the performance index vector predicted by the surrogate model for a given in-memory computation configuration x and the feature vector a of the target AI algorithm. The function obj k (·) Calculate the value of the k-th optimization objective from the predicted performance index vector.
[0064] As an example, taking Transformer inference as an example, several optimization objectives are defined for the in-memory computing architecture. These optimization objectives are based on the PPA metric and satisfy chip area efficiency and power consumption constraints. Chip area efficiency refers to the effective computing power or performance provided per unit area, and power consumption constraints refer to the upper limit of the total power consumption of the chip under typical workloads.
[0065] The optimization objectives can be set as follows: obj1 maximizes the energy efficiency ratio, obj2 maximizes the average computational utilization, and obj3 minimizes the model execution latency. Simultaneously, chip area and power consumption constraints must be met: Area(x) ≤ γ, Power(x) ≤ μ, Feasible(x) = 1. γ is the area threshold, Area(.) is the chip area with in-memory computing architecture parameters configured as x, Power(.) is the chip's power performance per kilobyte (PPA) value with in-memory computing architecture parameters configured as x, and Feasible(.) indicates whether the current in-memory computing architecture parameter configuration x is feasible, with 1 indicating feasibility and 0 indicating infeasibility.
[0066] In the iterative process of the multi-objective optimization algorithm, the performance evaluation of candidate architecture configurations is quickly completed by the surrogate model, thus significantly accelerating the optimization process. The specific objective function is shown in formula (1):
[0067]
[0068] S5. The multi-objective optimization algorithm iteratively generates a population of candidate in-memory computing architecture configurations. It then uses the predictive surrogate model to quickly evaluate the PPA (Performance-to-Average Performance) of these candidate configurations, searching for and outputting a set of Pareto-optimal in-memory computing architecture parameter configurations that satisfy the preset optimization objectives. In each generation, the performance index of each individual in the population is determined by the surrogate model f. θ (x) Fast evaluation. The algorithm performs selection, crossover, and mutation operations based on the decomposed scalar optimization subproblems and neighborhood information, gradually converging towards the Pareto front.
[0069] S6. After the multi-objective optimization algorithm finishes running, output a set of Pareto optimal in-memory computing architecture configuration schemes. Let q represent the Pareto optimal in-memory computing architecture configuration. This set of configurations represents multiple optimal configurations that balance different optimization objectives. Representative configurations are selected and thoroughly validated using a precise hardware simulation platform to confirm their performance metrics under real-world operating conditions and verify whether they meet the preset core technical indicators. Further performance attribution analysis is conducted to quantify the specific contribution of the optimized configuration to reducing the computational bottleneck of the target AI model (such as the Transformer) and to evaluate the performance robustness of the optimized architecture under different AI workloads and potential process perturbations.
[0070] Validated optimized architecture configurations provide a solid hardware basis for subsequent compiler development. Compiler designers can fully leverage these optimized architectural features, such as: the explicit ratio and computational power of pure digital in-memory computation or mixed-precision analog-digital in-memory computation units; the specific parameters of the static random access memory pipeline within the neural network processor (such as depth, width, and access mode); the established multi-core heterogeneous topology and inter-core communication mechanisms; and the optimized data paths and prefetch strategies between static random access memory and three-dimensional stacked dynamic random access memory. Based on this precise hardware information, compilers can develop more efficient mapping strategies, instruction scheduling algorithms, and data arrangement schemes, thereby achieving efficient and accurate mapping of AI models (such as Transformer, Mamba, etc.) to target in-memory computation hardware instructions, fully realizing the potential of the optimized architecture.
[0071] This embodiment also provides a system for implementing a hardware-software co-optimization algorithm for a hybrid in-memory architecture, including:
[0072] The feature extraction and parameterization module is used to extract features from target AI algorithms (including Transformer, Mamba, etc.). The extracted features include computation graph features, tensor space dimension, sparsity, and the proportion of mixed precision operations, forming an AI algorithm context feature vector. It also provides a parameterized description of the configurable parameters and microstructures of the in-memory computing architecture. The in-memory computing architecture parameters include in-memory computing unit configuration parameters, neural network processor pipeline parameters, multi-core interconnect topology, and storage hierarchy interface parameters, forming an in-memory computing architecture parameter set.
[0073] The dataset construction module is used to build an offline benchmark dataset through large-scale simulation experiments. The dataset contains multiple sets of in-memory computing architecture parameter configurations, corresponding AI algorithm context feature vectors, and PPA indicators obtained through accurate hardware simulation, and records architecture feasibility domain information.
[0074] The proxy model training module is used to train a context-aware predictive proxy model based on the offline benchmark dataset. The proxy model can predict the corresponding PPA index according to the input in-memory computing architecture parameter configuration and AI algorithm context feature vector.
[0075] The multi-objective optimization module is used to embed the trained predictive agent model as an evaluation function into the multi-objective optimization algorithm framework. Based on multiple predefined optimization objectives based on PPA indicators, the multi-objective optimization algorithm is run to quickly evaluate the PPA indicators of candidate in-memory computing architecture configurations using the predictive agent model, and to search for and output a set of Pareto optimal in-memory computing architecture parameter configuration schemes.
[0076] The above description of the preferred embodiments is quite specific and detailed, but it merely illustrates one feasible implementation of the present invention and is not intended to limit the scope of the invention. It should be noted that those skilled in the art can add several modifications or improvements based on these preferred embodiments within the framework of the present invention, but these are all within the protection scope of the present invention. The protection scope of the present invention should be determined by the appended claims.
Claims
1. A hardware-software co-optimization method for a hybrid in-memory architecture, characterized in that, Includes the following steps: S1. Extract the context features of the target AI algorithm to form the AI algorithm context feature vector, and perform parameterized description of the configurable parameters and microstructure of the in-memory computing architecture to obtain multiple sets of in-memory computing architecture parameter configuration vectors, forming an in-memory computing architecture parameter set. S2. Based on the in-memory computing architecture parameter set in step S1, construct an offline benchmark dataset. The offline benchmark dataset includes multiple sets of in-memory computing architecture parameter configurations sampled from the in-memory computing architecture parameter set, corresponding AI algorithm context feature vectors, PPA indicators obtained through hardware simulation, and architecture feasible domain information. S3. Based on the offline benchmark dataset, train a predictive agent model with context awareness. The predictive agent model can predict the corresponding in-memory computing metrics according to the input in-memory computing architecture parameter configuration vector and AI algorithm context feature vector. The training process includes using an embedding method to process discrete architecture parameters, using an encoder with a self-attention mechanism to learn the complex relationships between features, using a parallel prediction network to achieve joint prediction of multi-dimensional PPA metrics, and introducing a feasible domain constraint learning mechanism. While supervising the learning of PPA relationships between feasible architecture configurations, the mechanism introduces an adversarial negative sampling mechanism to discriminately learn infeasible configurations. S4. The trained predictive agent model is used as an evaluation function and embedded into a multi-objective optimization algorithm framework; S5. Define multiple optimization objectives for the in-memory computing architecture. The optimization objectives are based on the PPA index and meet chip area efficiency and power consumption constraints. S6. Run the multi-objective optimization algorithm, use the predictive surrogate model to quickly evaluate the PPA index of the candidate in-memory computing architecture configuration, search and output a set of Pareto optimal in-memory computing architecture parameter configuration schemes that meet the preset optimization objectives.
2. The hardware and software co-optimization method for a hybrid in-memory architecture according to claim 1, characterized in that, The contextual features include computation graph features, tensor space dimension, sparsity, and the proportion of mixed precision operations.
3. The hardware and software co-optimization method for a hybrid in-memory architecture according to claim 1, characterized in that, The configurable parameters of the in-memory computing architecture include in-memory computing unit configuration parameters, neural network processor pipeline parameters, multi-core interconnect topology, and storage hierarchy interface parameters.
4. The hardware and software co-optimization method for a hybrid in-memory architecture according to claim 3, characterized in that, Neural network processor pipeline parameters include the organization of its internal static random access memory and the number of pipeline stages; multi-core interconnect topology includes the connection methods and bandwidth between different cores; storage hierarchy interface parameters include the optimized data path characteristics and prefetch logic between static random access memory and three-dimensional stacked dynamic random access memory.
5. The hardware and software co-optimization method for a hybrid in-memory architecture according to claim 1, characterized in that, The offline benchmark dataset distinguishes between feasible architecture configuration points and infeasible architecture configuration points. Feasible architecture configuration points refer to configurations that can successfully compile, map, and execute AI algorithms and whose PPA indicators are within a preset reasonable range. Infeasible architecture configuration points refer to configurations that cannot meet the basic requirements due to physical constraints, compilation errors, or extremely poor performance.
6. The hardware-software co-optimization method for a hybrid in-memory architecture according to claim 1, characterized in that, In step S3, the feasible domain constraint learning mechanism aims to enable the predictive agent model to minimize the prediction error at feasible architecture configuration points through supervised learning, and to punish the tendency of infeasible architecture configuration points to produce good PPA index predictions through discriminative learning, so as to ensure that the predictive agent model has accurate architecture evaluation capabilities. The adversarial negative sampling mechanism is used to enhance the predictive agent model's ability to identify infeasible configurations and its robustness in prediction.
7. The hardware and software co-optimization method for a hybrid in-memory architecture according to claim 1, characterized in that, The structure of the predictive agent model includes: mapping discrete in-memory computing architecture parameter configuration vectors into continuous vectors through an embedding layer; processing the continuous vectors with an encoder layer containing a self-attention mechanism to generate a unified architecture feature representation; fusing the unified architecture feature representation with the AI algorithm context feature vector; inputting the fused features into a parallel multi-head prediction network and aggregating the outputs of each prediction head through a top-level attention mechanism or direct output method to generate a multi-dimensional PPA index prediction vector.
8. A system for implementing the hardware-software co-optimization method for a hybrid in-memory architecture as described in claim 1, characterized in that, include: The feature extraction and parameterization module is used to extract the context feature vector of the target AI algorithm and to provide a parameterized description of the configurable parameters and microstructure of the in-memory computing architecture. The dataset construction module is used to build an offline benchmark dataset. The dataset contains multiple sets of in-memory computing architecture parameter configurations, corresponding AI algorithm context feature vectors, and PPA indicators obtained through hardware simulation, and records architecture feasibility domain information. The proxy model training module is used to train a predictive proxy model with context awareness capabilities. The proxy model can predict the corresponding PPA index based on the input in-memory computing architecture parameter configuration and AI algorithm context feature vector. The multi-objective optimization module is used to embed the trained predictive agent model as an evaluation function into the multi-objective optimization algorithm framework. Based on multiple predefined optimization objectives based on PPA indicators, the multi-objective optimization algorithm is run to quickly evaluate the PPA indicators of candidate in-memory computing architecture configurations using the predictive agent model, and to search for and output a set of Pareto optimal in-memory computing architecture parameter configuration schemes.
9. A computer device comprising a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, characterized in that: When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Parameter optimization method and device, compiling method and device, electronic device and medium
CN118551820A
In-memory computing system
CN118626408A
Monolithic three-dimensional integration implementation method and device for hybrid precision in-memory computing architecture
CN120086485A
Cited By
Model performance automatic optimization method for artificial intelligence chip
CN121835423A