General modeling, ai processor architecture evaluation method, device, equipment and medium

CN117349638BActive Publication Date: 2026-09-29SHANGHAI SUIYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311498500.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2026-09-29
Estimated Expiration
2043-11-10

AI Technical Summary

Technical Problem

由于AI处理器的门阵列总量,随着半导体工艺的发展呈现指数型增长,其仿真代价也越来越大,耗时也越来越长

Benefits of technology

[0027]本发明实施例的技术方案通过使用通用建模方法,可以构建得到适配复杂问题场景的回归模型,快速实现该复杂问题场景的性能预测。同时,本发明实施例的技术方案可以泛化应用于硬件架构性能评估,硬件架构性能预测,基于硬件架构性能的处理器分等级筛片,以及基于硬件架构性能预测的软件系统性能调优等领域。该方法具有模型泛化能力强,应用场景广泛,建模成本低,训练代价小,推理精度高以及速度快的特点,在工程实践中可以充分发挥能效,解决其他方法无法解决的工程难题,将流片风险降到最低,将芯片良率提升到更高,为系统性能平稳提供有力保障。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117349638B_ABST
    Figure CN117349638B_ABST
Patent Text Reader

Abstract

The application discloses a general modeling, AI processor architecture performance evaluation method and device, equipment and medium, the general modeling method comprises: obtaining a general neural network framework, the general neural network framework comprises an adjustable number of input, intermediate and output nodes, the intermediate nodes form a hierarchical architecture with an adjustable number of layers, and two nodes are connected through a composite network; according to the input parameters and the output parameters in the problem scene, the model reconstruction parameters are determined, and the general neural network framework is reconstructed according to the model reconstruction parameters to obtain an adaptive model; according to the sample set matched with the problem scene, the weight parameters and the weight distribution vector of each composite network in the adaptive model are adjusted to obtain a regression model. The technical scheme of the embodiment of the application provides a modeling method with strong generalization ability, wide application scene and low cost, and the performance of the AI processor architecture can be accurately predicted, health diagnosed and optimized before the AI processor is actually taped out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a general modeling method, apparatus, device, and medium for evaluating the performance of AI processor architecture. Background Technology

[0002] As large and complex systems in the semiconductor field, the performance of artificial intelligence (AI) processors cannot be fully described by limited simulation samples. Because the total gate array size of AI processors increases exponentially with the development of semiconductor technology, the simulation cost and time consumption are also increasing. This makes it increasingly difficult to accurately describe system performance through the traditional method of increasing the simulation sample size.

[0003] With increasingly complex processor architectures and the exponential growth in the number of hardware components and interconnections within a single AI processor, it is becoming increasingly important to conduct a thorough and comprehensive analysis of the AI ​​processor architecture performance before tape-out.

[0004] Therefore, how to train a performance evaluation model for an AI processor architecture with minimal simulation cost, and achieve accurate prediction of specific performance regions in an infinite performance space, is an important problem that needs to be solved. Summary of the Invention

[0005] This invention provides a general modeling method, apparatus, device, and medium for evaluating the performance of AI processor architectures. It offers a modeling method with strong generalization ability, wide application scenarios, and low cost, enabling accurate prediction, health diagnosis, and optimization of AI processor architecture performance before the actual tape-out of the AI ​​processor.

[0006] According to one aspect of the present invention, a general modeling method is provided, comprising:

[0007] Obtain a pre-built general neural network framework, which includes an adjustable number of input nodes, intermediate nodes, and output nodes. The intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer, and the output end of the intermediate node is connected to the intermediate node of the next layer. The nodes are connected to each other through a composite network, which includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network.

[0008] Based on the input and output parameters in the problem scenario, determine the model reconstruction parameters, and reconstruct the general neural network framework according to the model reconstruction parameters to obtain the adapted model;

[0009] Based on a sample set that matches the problem scenario, the weight parameters and weight allocation vectors of each composite network in the adaptation model are adjusted to obtain a regression model that describes the input-output relationship in the problem scenario.

[0010] According to another aspect of the present invention, a performance evaluation method for an AI processor architecture is also provided, comprising:

[0011] In the performance evaluation scenario of AI processor architecture, the healthy core distribution scheme, quality control scheme, burst length and read / write type of AI processor are set as input parameters, and transmission latency and bandwidth are set as output parameters.

[0012] Based on the set input parameters and output parameters, a performance evaluation model for describing the input-output relationship in a performance evaluation scenario of an AI processor architecture is constructed using a general modeling method as described in any embodiment of the present invention.

[0013] The target AI processor architecture to be evaluated is obtained. After obtaining the target input parameters of the target AI processor architecture, the parameter features that match the target input parameters are input into the performance evaluation model to obtain the predicted values ​​of the target output parameters that match the target AI processor architecture.

[0014] According to another aspect of the present invention, a universal modeling apparatus is also provided, comprising:

[0015] The general framework acquisition module is used to acquire a pre-built general neural network framework. The general neural network framework includes an adjustable number of input nodes, intermediate nodes, and output nodes. The intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer, and the output end of the intermediate node is connected to the intermediate node of the subsequent layer. The nodes are connected to each other through a composite network. The composite network includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network.

[0016] The adaptation model reconstruction module is used to determine the model reconstruction parameters based on the input and output parameters in the problem scenario, and to reconstruct the general neural network framework based on the model reconstruction parameters to obtain the adaptation model;

[0017] The regression model generation module is used to adjust the weight parameters and weight allocation vectors of each composite network in the adaptation model based on a sample set that matches the problem scenario, so as to obtain a regression model that describes the input-output relationship in the problem scenario.

[0018] According to another aspect of the present invention, a performance evaluation apparatus for an AI processor architecture is also provided, comprising:

[0019] The input / output parameter setting module is used to set the health core distribution scheme, quality control scheme, burst length and read / write type of the AI ​​processor as input parameters, and transmission latency and bandwidth as output parameters in the performance evaluation scenario of AI processor architecture.

[0020] The performance evaluation model construction module is used to construct a performance evaluation model for describing the input-output relationship in the performance evaluation scenario of AI processor architecture, based on the set input parameters and output parameters and using a general modeling method as described in any embodiment of the present invention.

[0021] The output parameter prediction module is used to obtain the target AI processor architecture to be evaluated. After obtaining the target input parameters of the target AI processor architecture, it inputs the parameter features that match the target input parameters into the performance evaluation model to obtain the predicted values ​​of the target output parameters that match the target AI processor architecture.

[0022] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a general modeling method as described in any embodiment of the present invention, or to perform a performance evaluation method for an AI processor architecture as described in any embodiment of the present invention.

[0026] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the general modeling method described in any embodiment of the present invention, or to implement the performance evaluation method of the AI ​​processor architecture as described in any embodiment of the present invention.

[0027] The technical solution of this invention, through the use of a general modeling method, can construct a regression model adapted to complex problem scenarios, enabling rapid performance prediction for such scenarios. Furthermore, the technical solution of this invention can be generalized to areas such as hardware architecture performance evaluation, hardware architecture performance prediction, processor grading based on hardware architecture performance, and software system performance tuning based on hardware architecture performance prediction. This method features strong model generalization ability, wide application scenarios, low modeling cost, low training cost, high inference accuracy, and high speed. In engineering practice, it can fully leverage energy efficiency, solve engineering problems that other methods cannot address, minimize tape-out risks, maximize chip yield, and provide strong assurance for stable system performance.

[0028] Meanwhile, the novel hardware architecture prediction method based on artificial intelligence, generalized from the technical solutions of this invention, can train a fast-converging performance evaluation model for a specific hardware architecture with minimal simulation cost, reducing exploration time for subsequent performance evaluation and providing directional guidance.

[0029] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of a general modeling method provided according to Embodiment 1 of the present invention;

[0032] Figure 2 This is a schematic diagram comparing the structure of the general neural network architecture applicable to the embodiments of the present invention and the traditional neural network model;

[0033] Figure 3 This is a flowchart of another general modeling method provided in Embodiment 2 of the present invention;

[0034] Figure 4 This is a flowchart of another general modeling method provided in Embodiment 3 of the present invention;

[0035] Figure 5 This is a flowchart of a performance evaluation method for an AI processor architecture provided in Embodiment 4 of the present invention;

[0036] Figure 6 This is a schematic diagram illustrating a specific performance evaluation process for an AI processor architecture applicable to an embodiment of the present invention;

[0037] Figure 7 This is the correspondence between the maximum average error of the performance evaluation method for the AI ​​processor architecture to which this invention is applicable and the proportion of the training set;

[0038] Figure 8 This is a schematic diagram of the histogram of the maximum error of the performance evaluation method for the AI ​​processor architecture to which this invention is applicable when the proportion of the training dataset is 30%.

[0039] Figure 9 This is a structural diagram of a universal modeling device provided in Embodiment 5 of the present invention;

[0040] Figure 10 This is a structural diagram of a performance evaluation device for an AI processor architecture provided in Embodiment Six of the present invention;

[0041] Figure 11 This is a schematic diagram of the structure of an electronic device that implements the general modeling method or the performance evaluation method of the AI ​​processor architecture in the embodiments of the present invention. Detailed Implementation

[0042] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0043] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0044] Example 1

[0045] Figure 1This is a flowchart of a general modeling method provided in Embodiment 1 of the present invention. This embodiment is applicable to modeling complex problem scenarios (typically, performance evaluation scenarios of complex hardware architectures). The method can be executed by a general modeling device, which can be implemented in hardware and / or software and is generally configured in electronic devices with data processing capabilities, such as various types of terminals or servers. Figure 1 As shown, the method includes:

[0046] S110. Obtain a pre-built general neural network framework.

[0047] In this embodiment, to address performance evaluation scenarios involving complex hardware architectures, typically AI processor architectures, a novel general-purpose neural network framework is creatively designed. Specifically, in... Figure 2 The diagram shows a structural comparison between the general neural network architecture applicable to the embodiments of the present invention and the traditional neural network model.

[0048] like Figure 2 As shown, unlike traditional neural network models where model elements are fixed and connected by a chain structure, the general neural network framework in each embodiment of this invention includes an adjustable number of input nodes, intermediate nodes, and output nodes. Figure 2 The example above limits the number of nodes and is not intended to limit the actual general neural network framework. Intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. Figure 2 The example provided above defines an optional hierarchical architecture and is not intended to limit a general neural network framework. The input of an intermediate node is connected to both the input node and the intermediate nodes of the preceding layer, while the output of an intermediate node is connected to the intermediate nodes of the following layer. All nodes are connected through a composite network, which includes multiple types of neural networks. Figure 2 The example provided is a composite network consisting of four types of neural networks (this example is not intended to limit a general neural network framework) and a weight allocation vector used to adjust the weights of each type of neural network.

[0049] Furthermore, to improve the expressive power of the model, this embodiment of the invention increases the connections between input and output nodes, as well as the number of intermediate nodes, thereby transforming the original chain-like network into a graph-structured network. Secondly, the original single neural network is changed into a composite structure of multiple different types of networks. For example... Figure 2As shown, the dashed arrows represent composite networks obtained by connecting multiple neural networks in parallel. The left side shows a traditional multilayer perceptron structure. This structure uses different types of parameters as input, making it difficult to distinguish the influence of different parameters on the result. Furthermore, its simple structure cannot fit complex relationships. The right side shows an improved network structure with three intermediate nodes. It can be seen that each input node is connected to an intermediate node, allowing for independent fitting of the relationship between each input and output. The number of intermediate nodes can be freely changed according to the complexity of the problem; the more nodes, the more connections, and the more complex the relationships that can be described. On the other hand… Figure 2 The connections between nodes form a composite network, which is obtained by weighted superposition of multiple neural networks. This allows it to fit different relationships, thereby further improving the expressive power of the model.

[0050] Furthermore, in combination Figure 2 The core elements of the aforementioned novel general neural network framework are described in detail.

[0051] 1. InputNode - Input node: Used to store initial parameter settings or features.

[0052] 2. OutputNode – Output node: Used to store the model's prediction results to obtain the simulation or calculation results of the experiment.

[0053] 3. InterNode – Intermediate Node: Used to store the feature matrix of intermediate steps. The input of each intermediate node is the input node and the previous intermediate node, and the output is to the subsequent intermediate node and the output node.

[0054] 4. Network Set – A composite network: A collection of multiple neural networks (NNs) of different types. The connections between nodes are composite connections calculated by weighting and summing the individual neural networks in the composite network.

[0055] 5. α – Weight Vector: Each composite connection between nodes corresponds to a weight vector. If there are N types of neural networks in the composite network, then the weight vector is represented as [α1, α2, ... α]. N ], where α1+α2+…+α N =1, and α i ∈[0,1], i=1,2,…N, representing the weights involved in each type of neural network.

[0056] Optionally, the type of neural network in the composite network includes at least two of the following:

[0057] Fully connected networks, convolutional neural networks, pooling networks, recurrent neural networks, long short-term memory networks, and multilayer perceptrons; or

[0058] Optionally, the type of neural network in the composite network includes at least two of the following:

[0059] Nonlinear neural networks, linear neural networks, and 0 networks used to represent networks where there are no connections between any two nodes.

[0060] In this embodiment, multiple types of neural networks can be embedded in a composite network. The types of neural networks can be categorized using two dimensions. One dimension is the structural type of the neural network, such as a fully connected network or a convolutional neural network. Different structural types of neural networks exhibit varying predictive performance in different application scenarios. By introducing neural networks of various structural types into a composite network, the optimal connection methods between pairs of nodes can be extensively explored in various application scenarios, ensuring the general neural network framework's ability to solve various complex problems.

[0061] Another dimension for classification is the type of input-output relationship in the neural network, such as non-linear relationships, linear relationships, and non-existent relationships. It is understandable that, in the general neural network framework of this invention, the connection between each pair of nodes is established to the greatest extent possible in the form of a graph structure. However, when facing complex problem scenarios, the input and output relationships between these pairs of nodes are often complex, for example, combinations of linear, non-linear, partially linear, and partially non-linear relationships, and even in some extreme cases, the two may have no relationship at all. By introducing neural networks with different input-output types into a composite network, the complex input-output relationships between pairs of nodes can be more accurately uncovered.

[0062] S120. Based on the input and output parameters in the problem scenario, determine the model reconstruction parameters, and reconstruct the general neural network framework according to the model reconstruction parameters to obtain the adapted model.

[0063] In this context, the input parameters in the problem scenario can be understood as various parameters that can be obtained in the performance evaluation scenario of complex hardware architectures, such as the hardware and software characteristic parameters or hardware and software configuration parameters of the complex hardware architecture. Examples include processor type, number of processors, processor layout, processor computing power, and read / write type. The output parameters in the problem scenario can be understood as the important performance parameters that need to be predicted in the performance evaluation scenario of complex hardware architectures, such as computation time, communication time, bandwidth, failure rate, or failure recovery time.

[0064] Since the number of nodes and node architecture of the general neural network framework provided in this embodiment of the invention are variable, it is necessary to first determine the values ​​of the above parameters (i.e., model reconstruction parameters) according to the actual problem scenario, and then, based on the above model reconstruction parameters, finally uniquely determine the adaptive model that matches the problem scenario.

[0065] The model reconstruction parameters may include: the number of input nodes, intermediate nodes, and output nodes; the hierarchical structure of each intermediate node, i.e., the total number of levels and the number of intermediate nodes in each level; furthermore, the model reconstruction parameters may also include the specific types of neural networks in the composite network.

[0066] S130. Based on the sample set that matches the problem scenario, adjust the weight parameters and weight allocation vectors of each composite network in the adaptation model to obtain a regression model that describes the input-output relationship in the problem scenario.

[0067] After constructing a fitted model with a fixed structure, the training process for each model parameter in the fitted model can be implemented based on a sample set. Each sample in the sample set is labeled with the mapping relationship between the set input parameters and the set output parameters. The trained model parameters include: the weight coefficients of each layer in each type of neural network within each composite network, and the weight allocation vector in each composite network.

[0068] After the training termination condition is met, a regression model describing the input-output relationship in the problem scenario can be obtained. Then, when new input parameters for the same problem scenario are input into this regression model, the model predicts and outputs the corresponding output parameters to meet the user's actual needs.

[0069] The technical solution of this invention, through the use of a general modeling method, can construct a regression model adapted to complex problem scenarios, enabling rapid performance prediction for such scenarios. Furthermore, the technical solution of this invention can be generalized to hardware architecture performance evaluation, hardware architecture performance prediction, processor grading based on hardware architecture performance, and software system performance tuning based on hardware architecture performance prediction. This method features strong model generalization ability, wide application scenarios, low modeling cost, low training cost, high inference accuracy, and high speed. It demonstrates energy efficiency in engineering practice, solves engineering problems that other methods cannot address, minimizes tape-out risks, maximizes chip yield, and provides strong assurance for stable system performance.

[0070] Example 2

[0071] Figure 3This is a flowchart of a general modeling method provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiments. In this embodiment, the operation of "determining model reconstruction parameters according to input parameters and output parameters in the problem scenario" is specified.

[0072] Correspondingly, such as Figure 3 As shown, the method may include:

[0073] S310. Obtain a pre-built general neural network framework.

[0074] The general neural network framework includes an adjustable number of input nodes, intermediate nodes, and output nodes. Intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer, and the output end of the intermediate node is connected to the intermediate node of the subsequent layer. Nodes are connected to each other through a composite network. The composite network includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network.

[0075] S320. Determine the first number of input nodes based on the number of input parameters in the problem scenario, and determine the second number of output nodes based on the number of output parameters in the problem scenario.

[0076] In this embodiment, the actual parameter settings and expected simulation results required for the problem scenario can be determined based on the actual modeling scenario, and these can be used as input parameters and output parameters, respectively. Consequently, the number of input parameters and output parameters can be determined accordingly.

[0077] Correspondingly, the number of input parameters can be determined as the number of input nodes, and the number of output parameters can be determined as the number of output nodes.

[0078] S330. Based on the first quantity, determine the third quantity of intermediate nodes, and based on the third quantity, determine the hierarchical structure of the intermediate nodes.

[0079] Understandably, the more input parameters there are, the more complex the relationship between the input and output parameters may become. Therefore, it is advisable to set the number of intermediate nodes to be proportional to the number of input parameters; that is, the more input parameters there are, the more intermediate nodes there will be, in order to describe more complex relationships using intermediate nodes. In a specific example, the number of intermediate nodes can be the same as the number of input parameters, or it can be an integer multiple of the number of input parameters. This embodiment does not impose any restrictions on this.

[0080] After determining the number of intermediate nodes, the hierarchical architecture of the intermediate nodes can be determined based on the number of intermediate nodes. The hierarchical architecture includes the number of layers and the number of intermediate nodes in each layer.

[0081] In a specific example, the hierarchical architecture of intermediate nodes can be set separately for different numbers of intermediate nodes. For example, when there are 2 intermediate nodes, a single-level structure containing 2 intermediate nodes can be set, and when there are 3 intermediate nodes, a two-level architecture can be set, with the first level containing 2 intermediate nodes and the second level containing 1 intermediate node.

[0082] S340. Based on the number of input parameters and / or the number of output parameters in the problem scenario, determine the target type of the neural network contained in the composite network, and at least one hyperparameter contained in the neural network of each target type.

[0083] In this embodiment, in addition to fixing the type of neural network contained in the composite network, the target type of the neural network contained in the composite network can be determined according to the number of input parameters in the problem scenario, or according to the number of output parameters in the problem scenario, or according to the sum of the number of input parameters and output parameters in the problem scenario. This embodiment does not impose any restrictions on this.

[0084] In a specific example, the more input parameters and / or output parameters there are in the problem scenario, the more target types of the neural networks included in the composite network will be.

[0085] Furthermore, after selecting the target type of neural network, it is necessary to initialize and set the hyperparameters of each target type of neural network accordingly. For example, the initial weight parameters and bias values ​​in the neural network.

[0086] S350. Reconstruct the general neural network framework according to the model reconstruction parameters to obtain the adapted model.

[0087] S360. Based on the sample set that matches the problem scenario, adjust the weight parameters and weight allocation vectors of each composite network in the adaptation model to obtain a regression model that describes the input-output relationship in the problem scenario.

[0088] This invention proposes a novel general neural network framework that can dynamically build adaptive models. In practical applications, it can not only adaptively change the number of input nodes, intermediate nodes, and output nodes according to the problem scenario, but also change the number and types of neural networks participating in the operation in the composite network. Therefore, the modeling process has high flexibility and can be applied to modeling various physical systems.

[0089] Example 3

[0090] Figure 4 This is a flowchart of another general modeling method provided in Embodiment 3 of the present invention. This embodiment is a refinement based on the above embodiments. In this embodiment, the operation of "adjusting the weight parameters and weight allocation vectors of each composite network in the adaptation model according to the sample set matching the problem scenario to obtain a regression model for describing the input-output relationship in the problem scenario" is specified.

[0091] Correspondingly, such as Figure 4 As shown, the method may specifically include:

[0092] S410. Obtain a pre-built general neural network framework.

[0093] The general neural network framework includes an adjustable number of input nodes, intermediate nodes, and output nodes. Intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer, and the output end of the intermediate node is connected to the intermediate node of the subsequent layer. Nodes are connected to each other through a composite network. The composite network includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network.

[0094] S420. Based on the input and output parameters in the problem scenario, determine the model reconstruction parameters, and reconstruct the general neural network framework according to the model reconstruction parameters to obtain the adapted model.

[0095] S430. Initialize the various hyperparameters included in the adaptation model.

[0096] The hyperparameters in the adaptation model may include: the ratio of training set to test set, the total number of training iterations, learning rate, regularization term, weight decay rate, and batch size. Additionally, they may include: the weight parameters of each composite network and the update frequency of the weight allocation vector.

[0097] S440. After dividing the sample set into a training set and a test set, use the training set to train the adapted model for at least one round to obtain an intermediate model and obtain the training error during the training process.

[0098] S450. After obtaining the test error of the adapted model on the test set, check whether the training error and / or test error meet the expected error conditions: if yes, proceed to S460; otherwise, proceed to S470.

[0099] The training error is the difference between the predicted value of the training sample and the labeled value of the training sample by the adaptive model, and the test error is the difference between the predicted value of the test sample and the labeled value of the test sample by the adaptive model.

[0100] S460. The intermediate model is determined to be a regression model.

[0101] S470. Based on the error type of training error and / or testing error, readjust the initial values ​​of the reconstruction parameters or hyperparameters of the adapted model, and then return to execute S440.

[0102] In an optional implementation of this embodiment, readjusting the initial values ​​of the reconstruction parameters or hyperparameters of the adapted model based on the error type of the training error and / or testing error may include:

[0103] If the training error does not change smoothly, the initial values ​​of the first type of hyperparameters are readjusted; if the difference between the training error and the test error exceeds the first difference threshold, the initial values ​​of the second type of hyperparameters are readjusted; if the training error exceeds the second difference threshold, at least one objective reconstruction parameter in the adapted model is readjusted, and a new adapted model is obtained.

[0104] Specifically, if the training error does not change smoothly during at least one round of model training, it is necessary to adjust the weight parameters and update frequency of the weight allocation vectors of each composite network, as well as hyperparameters such as the learning rate. If the difference between the training error and the test error is large, it indicates overfitting, and in this case, it is necessary to adjust hyperparameters such as the regularization term and the weight decay rate. If the training error remains large and difficult to reduce during at least one round of model training, it indicates that the complexity of the reconstructed model is mismatched with the problem scenario. In this case, it is necessary to adjust the number of intermediate nodes and the type and number of neural networks included in the composite network.

[0105] The technical solution of this invention, by readjusting the initial values ​​of the reconstruction parameters or hyperparameters of the adaptation model during the training process according to the error type of the training error and / or test error, enables the training of the adaptation model to converge quickly, achieving the effect of low modeling cost and low training cost, and meeting the modeling needs in real-world scenarios.

[0106] Example 4

[0107] Figure 5 This is a flowchart illustrating a performance evaluation method for an AI processor architecture according to Embodiment 4 of the present invention. This embodiment is applicable to situations where a performance evaluation scenario for an AI processor architecture is modeled, and the performance evaluation model obtained from the modeling is used to evaluate the performance of a given AI processor architecture. This method can be executed by an AI processor architecture performance evaluation device, which can be implemented in hardware and / or software and is generally configured in electronic devices with data processing capabilities, such as various types of terminals or servers. Figure 5As shown, the method includes:

[0108] S510: In the performance evaluation scenario of AI processor architecture, the healthy core distribution scheme, quality control scheme, burst length and read / write type of AI processor are set as input parameters, and transmission latency and bandwidth are set as output parameters.

[0109] S520. Based on the set input parameters and output parameters, a performance evaluation model for describing the input-output relationship in the performance evaluation scenario of AI processor architecture is constructed using the general modeling method described in any embodiment of the present invention.

[0110] S530: Obtain the target AI processor architecture to be evaluated, and after obtaining the target input parameters of the target AI processor architecture, input the parameter features that match the target input parameters into the performance evaluation model to obtain the predicted values ​​of the target output parameters that match the target AI processor architecture.

[0111] One successful engineering application of this invention is the performance prediction of AI processor hardware architecture. Through the aforementioned modeling, simulation, and inference processes, the goal of accurately predicting performance is achieved.

[0112] Specifically, in order to better adapt to the computation of neural network training and inference, AI processor architectures typically deploy a large number of computing cores. Due to the unavoidable manufacturing defects in semiconductor processes, computing cores within the processor may randomly become unusable (Harvest) due to manufacturing defects. This results in different distributions of healthy cores on the processor among different AI processors, leading to differences in the basic performance of the hardware architecture. By setting different quality control (QoS) schemes, it is possible to ensure that as many high-level chips as possible are obtained under various Harvest conditions, and to achieve stable basic performance of the hardware architecture of chips of the same level, thereby maximizing the benefits of the product series.

[0113] This embodiment selects the specific engineering problem-solving scenario mentioned above, uses deep learning methods to train the model, and achieves performance prediction through inference.

[0114] (I) Input feature extraction and encoding preprocessing.

[0115] First, feature extraction is performed on four input parameters for the problem to be solved: 1) Health core distribution scheme of AI processor; 2) QoS scheme; 3) Burst length (BL); 4) Read / write (RW) type: including read type, write type and read type plus write type.

[0116] Preprocessing is performed on four types of input parameters, and the processing strategies are shown in Table 1:

[0117] Table 1

[0118]

[0119] In this embodiment, based on various parameter settings and considering the significant differences in the encoding methods of each parameter, the input parameters are separated and different encoders are used for feature extraction. The features of each sample include the health core distribution scheme, QoS scheme, BL scheme, and RW type. Therefore, a total of 4 encoders are required for encoding.

[0120] (II) Output Feature Extraction and Preprocessing

[0121] The outputs of the performance evaluation model mainly include 1) transmission delay and 2) bandwidth. In this embodiment, in addition to establishing a performance evaluation model with two output parameters, different performance evaluation models can be established for different output parameters to make the performance evaluation results more accurate. Depending on the importance of the output parameters, two performance evaluation models can be established in the order of bandwidth model and transmission delay model, respectively. This embodiment does not impose any restrictions on this.

[0122] (III) Establishment of deep learning models (i.e., adaptation models)

[0123] Because the strategy for computing cores is highly complex, this embodiment adopts the following two approaches: Figure 6 The specific performance evaluation process of the AI ​​processor architecture shown involves establishing a performance evaluation model.

[0124] 1) The performance evaluation model can automatically distinguish the importance of different inputs, thereby fully extracting the features of different inputs.

[0125] 2) The performance evaluation model itself should be complex enough to learn the relationship between inputs and outputs.

[0126] Specifically, such as Figure 6 As shown, the performance evaluation model contains a total of four intermediate nodes. Each intermediate node takes the results of four InputNodes as input. In addition, subsequent intermediate nodes take all the preceding nodes as input. For ease of description, the i-th intermediate node is denoted as InterNode. i For example, InterNode0's input includes the outputs of 4 InputNodes, InterNode1's input includes the outputs of InputNodes plus the output of InterNode0, while InterNode2 includes 6 inputs, and so on.

[0127] In this embodiment, each composite network contains a total of five types of neural networks. The first three are nonlinear networks (containing activation functions) with different numbers of layers or neurons, used to fit nonlinear relationships. The fourth is a linear network (without activation functions), used to fit linear relationships. The fifth is a zero network, representing that there are no connections between nodes.

[0128] With the above settings, this performance evaluation model can not only learn various input features, but also learn the relationship between various features and output. For example, if a certain input has a high correlation with the result, the proportion of 0 operations in the output of that node is very low. Conversely, if the correlation is very low, the proportion of 0 operations in the output of that node is very high. At the same time, if a certain input and output mainly exhibit a non-linear relationship, the output of that node mainly exhibits a weighted sum of the first three types of networks. Conversely, if the linear relationship is strong, the output mainly exhibits a linear network.

[0129] Based on the forward propagation principle, the data sequentially passes through InterNode0, InterNode1, InterNode2, and InterNode3, and is then output through InterNode3, representing the simulation results to be fitted, such as bandwidth. This performance evaluation model can provide prediction results within seconds, significantly reducing simulation time costs.

[0130] (iv) Error Analysis

[0131] The performance evaluation model mainly consists of two types of parameters: the first is the weight allocation vector α of each composite network, and the second is the weight parameter w of each neural network within each composite network. Therefore, the objective function LOSS that this model needs to optimize is:

[0132]

[0133] Where y is the sample labeled value (true value), and f(x,w,α) is the model output value (predicted value), which is found through backpropagation. * With α * To minimize the objective function, where σ i (i = 1, 2) are manually set hyperparameters.

[0134] In this experiment, the sample set included a total of 3072 data samples. These samples were divided into training and test sets according to a certain ratio. The training set was trained 12 times. To increase the reliability of the experiment, three random number seeds (0, 1, and 2) were used to repeat the experiment three times. To analyze the performance evaluation of the model's fit to the data, the proportion of the training set in the total sample set was gradually changed from 10% to 90%, while the remaining sample set was used as the test set. To quantify the results, the predicted value p was calculated for each of the N tests. i Compared with the true value y i The maximum mean absolute error M and the minimum mean absolute error m are expressed as follows.

[0135]

[0136]

[0137] Since the maximum mean error M represents the worst prediction result of the model in actual industrial production and has more important significance, the maximum mean absolute error is subsequently used to measure the performance of the model.

[0138] Experimental results are as follows Figure 7 and Figure 8 As shown. Figure 7 This describes the correspondence between the maximum average error of the performance evaluation method for the AI ​​processor architecture to which this invention is applicable and the proportion of the training set. For example... Figure 7 As shown, as the proportion gradually increases, the maximum average error gradually decreases. When the proportion reaches 30%, the maximum error has gradually decreased and stabilized. Figure 8 This is a histogram showing the maximum error of the performance evaluation method for the AI ​​processor architecture applicable to this embodiment of the invention when the training dataset accounts for 30%. At this point, the maximum average error is 2.98 Gbyte / s, the error distribution is below 10, and with a theoretical peak bandwidth of 192 GB / s, the overall prediction error is less than 5%, fully demonstrating the accuracy and effectiveness of the model.

[0139] The technical solution of this invention uses minimal simulation overhead to train a performance evaluation model for the AI ​​processor architecture, enabling accurate prediction of specific performance regions within an infinite performance space. This allows for a thorough and complete analysis of the AI ​​processor architecture performance before tape-out, achieving accurate prediction, health diagnosis, and optimization of the AI ​​processor architecture performance.

[0140] Furthermore, in subsequent engineering practices, by continuously increasing the number of effective training samples and input feature types, the stability and accuracy of the model can be further improved, the initial configuration scheme of the model can be continuously refined and optimized, and various engineering problems can be continuously solved, thus empowering and enhancing engineering practices.

[0141] Example 5

[0142] Figure 9 This is a schematic diagram of the structure of a general modeling device provided in Embodiment 5 of the present invention. Figure 9 As shown, the device includes: a general framework acquisition module 910, an adaptation model reconstruction module 920, and a regression model generation module 930, wherein:

[0143] The general framework acquisition module 910 is used to acquire a pre-built general neural network framework. The general neural network framework includes an adjustable number of input nodes, intermediate nodes, and output nodes. The intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer. The output end of the intermediate node is connected to the intermediate node of the subsequent layer. The nodes are connected to each other through a composite network. The composite network includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network.

[0144] The adaptation model reconstruction module 920 is used to determine the model reconstruction parameters based on the input and output parameters in the problem scenario, and to reconstruct the general neural network framework based on the model reconstruction parameters to obtain the adaptation model.

[0145] The regression model generation module 930 is used to adjust the weight parameters and weight allocation vectors of each composite network in the adaptation model based on a sample set that matches the problem scenario, so as to obtain a regression model that describes the input-output relationship in the problem scenario.

[0146] The technical solution of this invention, through the use of a general modeling method, can construct a regression model adapted to complex problem scenarios, enabling rapid performance prediction for such scenarios. Furthermore, the technical solution of this invention can be generalized to hardware architecture performance evaluation, hardware architecture performance prediction, processor grading based on hardware architecture performance, and software system performance tuning based on hardware architecture performance prediction. This method features strong model generalization ability, wide application scenarios, low modeling cost, low training cost, high inference accuracy, and high speed. It demonstrates energy efficiency in engineering practice, solves engineering problems that other methods cannot address, minimizes tape-out risks, maximizes chip yield, and provides strong assurance for stable system performance.

[0147] Based on the above embodiments, the adaptive model reconstruction module 920 can be specifically used for:

[0148] Based on the number of input parameters in the problem scenario, determine the first number of input nodes, and based on the number of output parameters in the problem scenario, determine the second number of output nodes.

[0149] Based on the first quantity, determine the third quantity of intermediate nodes, and based on the third quantity, determine the hierarchical structure of the intermediate nodes, wherein the hierarchical structure includes the number of layers and the number of intermediate nodes included in each layer.

[0150] Based on the above embodiments, the adaptive model reconstruction module 920 can be further specifically used for:

[0151] Based on the number of input parameters and / or output parameters in the problem scenario, determine the target type of the neural network included in the composite network, and at least one hyperparameter included in the neural network of each target type.

[0152] Based on the above embodiments, the type of neural network in the composite network may include at least two of the following:

[0153] Fully connected networks, convolutional neural networks, pooling networks, recurrent neural networks, long short-term memory networks, and multilayer perceptrons; or

[0154] The type of neural network in the composite network may include at least two of the following:

[0155] Nonlinear neural networks, linear neural networks, and 0 networks used to represent networks where there are no connections between any two nodes.

[0156] Based on the above embodiments, the regression model generation module 930 may specifically include:

[0157] The hyperparameter initialization unit is used to initialize the various hyperparameters included in the adaptation model;

[0158] The model training unit is used to divide the sample set into a training set and a test set, use the training set to train the adapted model at least once to obtain an intermediate model, and obtain the training error during the training process.

[0159] The error verification unit is used to detect whether the training error and / or test error meet the expected error conditions after obtaining the test error of the adapted model on the test set.

[0160] The first error processing unit is used to determine the intermediate model as a regression model when the detected training error and / or test error meet the expected error conditions.

[0161] The second error processing unit is used to readjust the initial values ​​of the reconstruction parameters or hyperparameters of the adaptation model according to the error type of the training error and / or test error when the detected training error and / or test error do not meet the expected error conditions. Then, it returns to execute the operation of dividing the sample set into training set and test set, and using the training set to train the adaptation model at least once, until the end of training conditions are met.

[0162] Based on the above embodiments, the second error processing unit can be specifically used for:

[0163] If the training error does not change smoothly, the initial values ​​of the first type of hyperparameters are readjusted.

[0164] If the difference between the training error and the test error exceeds the first difference threshold, the initial values ​​of the second type of hyperparameters are readjusted.

[0165] If the training error exceeds the second difference threshold, at least one objective reconstruction parameter in the adapted model is readjusted, and a new adapted model is obtained.

[0166] The general modeling apparatus provided in the embodiments of the present invention can execute the general modeling method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0167] Example 6

[0168] Figure 10 This is a schematic diagram of the structure of a performance evaluation device for an AI processor architecture provided in Embodiment Six of the present invention. Figure 10 As shown, the device includes: an input / output parameter setting module 1010, a performance evaluation model construction module 1020, and an output parameter prediction module 1030, wherein:

[0169] The input / output parameter setting module 1010 is used to set the health core distribution scheme, quality control scheme, burst length and read / write type of the AI ​​processor as input parameters and transmission latency and bandwidth as output parameters in the performance evaluation scenario of AI processor architecture.

[0170] The performance evaluation model construction module 1020 is used to construct a performance evaluation model for describing the input-output relationship in the performance evaluation scenario of AI processor architecture based on the set input parameters and output parameters, using a general modeling method as described in any embodiment of the present invention.

[0171] The output parameter prediction module 1030 is used to obtain the target AI processor architecture to be evaluated, and after obtaining the target input parameters of the target AI processor architecture, input the parameter features that match the target input parameters into the performance evaluation model to obtain the predicted values ​​of the target output parameters that match the target AI processor architecture.

[0172] The technical solution of this invention uses minimal simulation overhead to train a performance evaluation model for the AI ​​processor architecture, enabling accurate prediction of specific performance regions within an infinite performance space. This allows for a thorough and complete analysis of the AI ​​processor architecture performance before tape-out, achieving accurate prediction, health diagnosis, and optimization of the AI ​​processor architecture performance.

[0173] The performance evaluation device for AI processor architecture provided in this embodiment of the invention can execute the performance evaluation method for AI processor architecture provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0174] Example 7

[0175] Figure 11 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0176] like Figure 11 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0177] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0178] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as general modeling methods or performance evaluation methods for AI processor architectures.

[0179] In some embodiments, the general modeling method, or the performance evaluation method for AI processor architecture, may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the general modeling method or the performance evaluation method for AI processor architecture described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the general modeling method or the performance evaluation method for AI processor architecture by any other suitable means (e.g., by means of firmware).

[0180] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0181] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0182] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0183] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0184] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0185] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0186] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0187] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A performance evaluation method for an artificial intelligence (AI) processor architecture, characterized in that, include: In the performance evaluation scenario of AI processor architecture, the healthy core distribution scheme, quality control scheme, burst length and read / write type of AI processor are set as input parameters, and transmission latency and bandwidth are set as output parameters. Based on the set input parameters and output parameters, a performance evaluation model is constructed using a general modeling method to describe the input-output relationship in the performance evaluation scenario of AI processor architecture. The target AI processor architecture to be evaluated is obtained. After obtaining the target input parameters of the target AI processor architecture, the parameter features that match the target input parameters are input into the performance evaluation model to obtain the predicted values ​​of the target output parameters that match the target AI processor architecture. The general modeling method includes: Obtain a pre-built general neural network framework, which includes an adjustable number of input nodes, intermediate nodes, and output nodes. The intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer, and the output end of the intermediate node is connected to the intermediate node of the next layer. The nodes are connected to each other through a composite network, which includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network. Based on the input and output parameters in the problem scenario, determine the model reconstruction parameters, and reconstruct the general neural network framework according to the model reconstruction parameters to obtain the adapted model; Based on a sample set that matches the problem scenario, the weight parameters and weight allocation vectors of each composite network in the adaptation model are adjusted to obtain a regression model that describes the input-output relationship in the problem scenario.

2. The method according to claim 1, characterized in that, Based on the input and output parameters in the problem scenario, determine the model reconstruction parameters, including: Based on the number of input parameters in the problem scenario, determine the first number of input nodes, and based on the number of output parameters in the problem scenario, determine the second number of output nodes. Based on the first quantity, determine the third quantity of intermediate nodes, and based on the third quantity, determine the hierarchical structure of the intermediate nodes, wherein the hierarchical structure includes the number of layers and the number of intermediate nodes included in each layer.

3. The method according to claim 1 or 2, characterized in that, Based on the input and output parameters in the problem scenario, the model reconstruction parameters are determined, including: Based on the number of input parameters and / or output parameters in the problem scenario, determine the target type of the neural network included in the composite network, and at least one hyperparameter included in the neural network of each target type.

4. The method according to claim 1, characterized in that, The types of neural networks in the composite network include at least two of the following: Fully connected networks, convolutional neural networks, pooling networks, recurrent neural networks, long short-term memory networks, and multilayer perceptrons; or The types of neural networks in the composite network include at least two of the following: Nonlinear neural networks, linear neural networks, and 0 networks used to represent networks where there are no connections between any two nodes.

5. The method according to claim 1, characterized in that, Based on a sample set matching the problem scenario, the weight parameters and weight allocation vectors of each composite network in the adaptation model are adjusted to obtain a regression model that describes the input-output relationship in the problem scenario, including: Initialize the hyperparameters included in the adaptation model; After dividing the sample set into a training set and a test set, the training set is used to train the adaptation model for at least one round to obtain an intermediate model and to obtain the training error during the training process. After obtaining the test error of the adapted model on the test set, check whether the training error and / or test error meet the expected error conditions. If so, the intermediate model is identified as a regression model; otherwise, based on the error type of the training error and / or test error, the initial values ​​of the reconstruction parameters or hyperparameters of the adaptation model are readjusted, and the process is repeated to divide the sample set into a training set and a test set, and then use the training set to train the adaptation model for at least one round until the end of training conditions are met.

6. The method according to claim 5, characterized in that, Based on the error type of training error and / or testing error, the initial values ​​of the reconstruction parameters or hyperparameters of the adapted model are readjusted, including: If the training error does not change smoothly, the initial values ​​of the first type of hyperparameters are readjusted. If the difference between the training error and the test error exceeds the first difference threshold, the initial values ​​of the second type of hyperparameters are readjusted. If the training error exceeds the second difference threshold, at least one objective reconstruction parameter in the adapted model is readjusted, and a new adapted model is obtained.

7. A performance evaluation device for an artificial intelligence (AI) processor architecture, characterized in that, include: The input / output parameter setting module is used to set the health core distribution scheme, quality control scheme, burst length and read / write type of the AI ​​processor as input parameters, and transmission latency and bandwidth as output parameters in the performance evaluation scenario of AI processor architecture. The performance evaluation model construction module is used to construct a performance evaluation model that describes the input-output relationship in the performance evaluation scenario of AI processor architecture based on the set input parameters and output parameters and through a general modeling method. The output parameter prediction module is used to obtain the target AI processor architecture to be evaluated, and after obtaining the target input parameters of the target AI processor architecture, input the parameter features that match the target input parameters into the performance evaluation model to obtain the predicted values ​​of the target output parameters that match the target AI processor architecture. The performance evaluation model construction module is further used for: Obtain a pre-built general neural network framework, which includes an adjustable number of input nodes, intermediate nodes, and output nodes. The intermediate nodes are used to form a hierarchical architecture with an adjustable number of layers. The input end of the intermediate node is connected to the input node and the intermediate node of the previous layer, and the output end of the intermediate node is connected to the intermediate node of the next layer. The nodes are connected to each other through a composite network, which includes multiple types of neural networks and a weight allocation vector for adjusting the weights of each type of neural network. Based on the input and output parameters in the problem scenario, determine the model reconstruction parameters, and reconstruct the general neural network framework according to the model reconstruction parameters to obtain the adapted model; Based on a sample set that matches the problem scenario, the weight parameters and weight allocation vectors of each composite network in the adaptation model are adjusted to obtain a regression model that describes the input-output relationship in the problem scenario.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the performance evaluation method of the AI ​​processor architecture as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the performance evaluation method for the AI ​​processor architecture as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Hybrid neural network algorithm-based performance assessment method used for complex industrial product

    CN106096723A

  • Image detection method and apparatus, and electronic device

    WO2022193132A1