Method, device, electronic device, medium and program product for generating architecture information

CN122840133APending Publication Date: 2026-09-29艾酷软件技术(上海)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510333046.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

单目标搜索提供的是满足单一目标的解决方案,即便该方案对于某个特定目标或特定目标聚合而言是最优的,但对于其余某些目标来说,可能是不可接受的

Benefits of technology

[0021]本申请实施例中,在对设计空间探索的过程中,由于可以针对每个网络层探索得到至少两个候选架构信息,其中,每个候选架构信息可以作为对应网络层的一种解决方案,即在进行设计空间探索的过程中可以针对每个网络层实现多目标探索,该过程中,由于可以针对每个网络层实现多目标探索,从而有利于实现多目标优化,进而可以在空间探索的过程中,有效的平衡推理能耗、推理延迟和面积等存在冲突性质的目标,以改善基于所述N个第一架构信息,生成的目标架构信息质量,这样,相对于单一目标探索而言,有利于提高推理加速器的综合性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840133A_ABST
    Figure CN122840133A_ABST
Patent Text Reader

Abstract

The application discloses a kind of architecture information generation method, device, electronic equipment, medium and program product, belong to inference accelerator technical field, the architecture information generation method includes: obtaining the model information of network model and the hardware architecture information of inference accelerator;Joint exploration space is constructed based on model information and hardware architecture information, and joint exploration space includes multiple candidate architecture information;N first architecture information corresponding to N network layers is generated based on joint exploration space, and first architecture information includes at least two candidate architecture information;Target architecture information is generated based on the N first architecture information, wherein the target architecture information includes the architecture corresponding to each network layer in the N network layer in the inference accelerator, and the mapping relationship between the network layer and the corresponding architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of inference accelerator technology, specifically to a method, apparatus, electronic device, medium, and program product for generating architectural information. Background Technology

[0002] In related technologies, running increasingly complex Artificial Intelligence (AI) models in embedded systems is quite challenging due to resource and power consumption limitations of edge devices. To meet the various constraints of edge inference devices, a common solution is to tailor-make dedicated hardware accelerators for specific AI algorithms, thereby achieving efficient and high-throughput mapping. For AI algorithms to run efficiently on dedicated hardware architectures, a scheduling method from algorithm to hardware, known as mapping, needs to be defined. Mapping defines the time and space scheduling of multiplication and accumulation operations at each layer of the network. Each mapping affects the access patterns to architectural-level buffers, thus impacting inference power consumption and latency. Each mapping also requires specific buffer capacities and several spatial component instances, which translates into different chip area footprints.

[0003] However, most design space exploration (DSE) frameworks in related technologies perform single-objective exploration (Energy / Latency / Area). If multiple metrics need to be satisfied simultaneously, this approach addresses this by aggregating multiple objectives into a single optimization metric for single-objective exploration, such as calculating the product of energy and latency. Single-objective search provides a solution that satisfies only one objective. Even if the solution is optimal for a specific objective or a specific aggregation of objectives, it may be unacceptable for others. For example, when latency is used as the search objective, the solution with the minimum latency is returned, but the corresponding area is too large to meet the area constraints of the inference accelerator. Conversely, when area is used as the search objective, it may lead to solutions with excessively high energy or latency. Therefore, the architectural information obtained using the DSE approach in related technologies can easily lead to problems such as excessive energy consumption or excessive latency in hardware accelerators during AI model execution. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, electronic device, medium, and program product for generating architecture information, which can improve the overall performance of inference accelerators.

[0005] In a first aspect, embodiments of this application provide a method for generating architecture information, including:

[0006] Obtain model information of the network model and hardware architecture information of the inference accelerator. The network model includes N network layers. The model information includes N dimension information corresponding one-to-one with the N network layers. N is an integer greater than 1. The dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers. The hardware architecture information includes processing unit information and storage architecture information.

[0007] A joint exploration space is constructed based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, wherein the candidate architecture information includes: a candidate architecture corresponding to a network layer, and the mapping relationship between the network layer and the corresponding candidate architecture;

[0008] Based on the joint exploration space, N first architecture information corresponding one-to-one with the N network layers are generated, wherein the first architecture information includes at least two candidate architecture information;

[0009] Based on the N first architecture information, target architecture information is generated, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

[0010] Secondly, embodiments of this application provide an apparatus for generating architecture information, including:

[0011] The acquisition module is used to acquire model information of the network model and hardware architecture information of the inference accelerator. The network model includes N network layers, the model information includes N dimension information corresponding one-to-one with the N network layers, where N is an integer greater than 1, the dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers, and the hardware architecture information includes processing unit information and storage architecture information.

[0012] A construction module is used to construct a joint exploration space based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, wherein the candidate architecture information includes: a candidate architecture corresponding to a network layer, and the mapping relationship between the network layer and the corresponding candidate architecture;

[0013] The generation module is used to generate N first architecture information corresponding one-to-one with the N network layers based on the joint exploration space, wherein the first architecture information includes at least two candidate architecture information;

[0014] The generation module is further configured to generate target architecture information based on the N first architecture information, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

[0015] Thirdly, embodiments of this application provide a first chip in which a network model is deployed, wherein the network model is a network model deployed based on target architecture information, and the target architecture information is generated based on the architecture information generation method described in the first aspect.

[0016] Fourthly, embodiments of this application provide an accelerator in which a network model is deployed, wherein the network model is a network model deployed based on target architecture information, and the target architecture information is generated based on the architecture information generation method described in the first aspect.

[0017] Fifthly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0018] In a sixth aspect, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0019] In a seventh aspect, embodiments of this application provide a second chip, the second chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method described in the first aspect.

[0020] Eighthly, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method described in the first aspect.

[0021] In this embodiment of the application, during the exploration of the design space, at least two candidate architecture information can be obtained for each network layer. Each candidate architecture information can serve as a solution for the corresponding network layer. That is, during the exploration of the design space, multi-objective exploration can be carried out for each network layer. Since multi-objective exploration can be carried out for each network layer, it is beneficial to achieve multi-objective optimization. In this way, during the space exploration, conflicting objectives such as inference energy consumption, inference latency, and area can be effectively balanced to improve the quality of the target architecture information generated based on the N first architecture information. Thus, compared with single-objective exploration, it is beneficial to improve the overall performance of the inference accelerator. Attached Figure Description

[0022] Figure 1 This is one of the flowcharts illustrating a method for generating architecture information provided in an embodiment of this application;

[0023] Figure 2 This is a schematic diagram of the hardware architecture of the inference accelerator in the embodiments of this application;

[0024] Figure 3 This is a schematic diagram of the storage hierarchy in the hardware architecture of the inference accelerator in this application embodiment;

[0025] Figure 4 This is a schematic diagram of the processing unit array in the hardware architecture of the inference accelerator in this application embodiment;

[0026] Figure 5 This is a schematic diagram of a mapping method for the CONV operator on the hardware architecture in the embodiments of this application;

[0027] Figure 6 This is a flowchart illustrating the method for generating architecture information based on the DSE framework in an embodiment of this application.

[0028] Figure 7 This is a flowchart illustrating the Mapper stage in an embodiment of this application;

[0029] Figure 8 yes Figure 7 A schematic diagram of the joint exploration space;

[0030] Figure 9 This is a schematic diagram comparing the search results of single-target search and multi-target search in an embodiment of this application;

[0031] Figure 10 This is a schematic diagram of the genetic algorithm flow in the Mapper stage of this application embodiment;

[0032] Figure 11This is a flowchart illustrating the Negotiation stage in an embodiment of this application;

[0033] Figure 12 This is a flowchart illustrating the genetic algorithm during the Negotiation stage in an embodiment of this application;

[0034] Figure 13 This is a schematic diagram illustrating the configuration of a point in the ArraySize subspace in an embodiment of this application;

[0035] Figure 14 This is a schematic diagram showing the result of factor number allocation and factor decomposition of the values ​​of the seven dimensions of the convolutional layer in the embodiments of this application;

[0036] Figure 15 This is a schematic diagram of the configuration result of a point in the IndexFactor subspace in an embodiment of this application;

[0037] Figure 16 This is a schematic diagram of the configuration result of a point in the LoopAllocation subspace in an embodiment of this application;

[0038] Figure 17 This is a schematic diagram of the initial space and joint exploration space in the embodiments of this application;

[0039] Figure 18 This is a schematic diagram of the structure of an architecture information generation device provided in an embodiment of this application;

[0040] Figure 19 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application;

[0041] Figure 20 A schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application. Detailed Implementation

[0042] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0043] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0044] The following description, in conjunction with the accompanying drawings, details a method, apparatus, electronic device, medium, and program product for generating architecture information provided in this application, through specific embodiments and application scenarios.

[0045] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for generating architecture information according to an embodiment of this application. The method for generating architecture information includes the following steps:

[0046] Step 101: Obtain the model information of the network model and the hardware architecture information of the inference accelerator. The network model includes N network layers, the model information includes N dimension information corresponding to the N network layers, the dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers, and the hardware architecture information includes processing unit information and storage architecture information.

[0047] Step 102: Construct a joint exploration space based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, wherein the candidate architecture information includes: a candidate architecture corresponding to a network layer, and the mapping relationship between the network layer and the corresponding candidate architecture;

[0048] Step 103: Generate N first architecture information corresponding one-to-one with the N network layers based on the joint exploration space, wherein the first architecture information includes at least two candidate architecture information;

[0049] Step 104: Based on the N first architecture information, generate target architecture information, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

[0050] The aforementioned inference accelerator can be various inference devices or hardware accelerators used to deploy AI models. For example, in some embodiments of this application, the inference accelerator is a Neural Network Processor (NPU) accelerator. Please refer to [link to relevant documentation]. Figure 2 This is a schematic diagram of the hardware architecture of an NPU accelerator. For ease of understanding, the following explanation uses the NPU accelerator as an example to further illustrate the method for generating architecture information provided in this application embodiment.

[0051] Understandably, the hardware architecture of an inference accelerator typically includes processing units and a storage architecture; please refer to [link to relevant documentation]. Figure 3 This is a schematic diagram of the storage hierarchy in the storage architecture of the inference accelerator in this embodiment of the application. The storage architecture includes four storage levels from the inside out: L0, L1, L2, and L3. The storage space closer to the innermost storage level is smaller, and the response speed is faster; correspondingly, the storage space closer to the outermost storage level is larger, and the response speed is slower. Please refer to... Figure 3 In some embodiments of this application, the data types that the storage architecture needs to store during operation typically include: input data (Input, I), output data (Output, O), and weight data (Weight, W). The data types allowed to be stored at each storage level can be pre-defined. For example, storage level L0 stores weight data W, storage level L1 stores output data O, storage level L2 stores input data I and output data O, and storage level L3 stores weight data W, input data I, and output data O. Accordingly, the above storage architecture information may include: the storage levels in the storage architecture, the size of the storage space at each storage level, and the data types stored at each storage level. It should be noted that... Figure 3 This is just one example of this application. In fact, the number of storage layers can also be other numbers, such as 3 layers, 5 layers, etc., which can be determined by the actual hardware structure of the inference accelerator. In addition, the data types pre-stored in each layer can also be set according to actual needs.

[0052] Please see Figure 4 This is a schematic diagram of the processing unit array of the hardware architecture of the inference accelerator in an embodiment of this application, wherein each multiply-accumulate (MAC) unit in the processing unit array is considered as a processing unit. See also... Figure 4In some embodiments of this application, the processing unit array includes three spatial dimensions: D1, D2, and D3, and each spatial dimension includes multiple processing units. Based on this, the aforementioned processing unit information may include the number of spatial dimensions in the processing unit array, the number of processing units included in each spatial dimension, and the total number of processing units included in the processing unit array. It should be noted that... Figure 4 This is just one example of the present application. In fact, the spatial dimension of the processing unit array can also be other dimensions, such as two-dimensional, one-dimensional, etc. The number of processing units included in each dimension of the processing unit array and the total number of processing units included in the processing unit array can also be determined according to the actual hardware structure of the inference accelerator.

[0053] The network model described above can be various types of AI models, such as intelligent question answering models, image recognition models, object detection models, etc. Accordingly, the above N network layers can include common network layers such as convolutional (CONV) layers, matrix multiplication (MatMul) layers, and fully connected (FC) layers.

[0054] In related technologies, computationally intensive operators such as the CONV operator, MatMul operator, and FC operator can generally be uniformly described as a 7-dimensional nested cyclic operation expression. The dimension values ​​of different operators can be different. For example, the operation expression for the CONV operator is:

[0055]

[0056]

[0057] The CONV operator includes seven dimensional parameters: B, K, C, OY, OX, FY, and FX. Here, B represents the number of input network samples, K represents the output channels, C represents the number of input network samples, OY and OX represent the spatial dimensions of the feature map, and FX and FY represent the spatial dimensions of the weights. Accordingly, the dimension value of B is 4, the dimension value of K is 64, the dimension value of C is 64, the dimension value of OY is 56, the dimension value of OX is 56, the dimension value of FY is 3, and the dimension value of FX is 3. It should be noted that the above dimensional values ​​are only an example of this application; the dimensional values ​​of each dimension are usually different for different types of operators.

[0058] The aforementioned joint exploration space can be a joint space composed of multiple subspaces. For example, a subspace can be constructed based on the processing unit information, a subspace can be constructed based on the factor decomposition result of the dimension value, a subspace can be constructed based on the ranking result of the factor decomposition result, and a subspace can be constructed based on the storage architecture information. The four constructed subspaces can then be combined to form a joint exploration space.

[0059] It is understood that within the joint exploration space, each subspace may contain multiple searchable candidate solutions. For example, the subspace corresponding to the processing unit information may include multiple processing architecture solutions, and the subspace corresponding to the factorization results may include multiple factorization result solutions. Based on this, the aforementioned candidate architecture information includes combination information formed by combining one candidate solution from each subspace of the joint exploration space. Thus, by arranging and combining the candidate solutions in each subspace, the multiple candidate architecture information can be obtained.

[0060] The mapping relationship between the aforementioned network layers and their corresponding candidate architectures is used to indicate the order of operations of operators in the corresponding network layers, as well as the correspondence between operators and hardware. For example, please refer to... Figure 5 This is a mapping method for the CONV operator on the hardware architecture. Taking the CONV operator as an example, the above-mentioned mapping method for the CONV operator on the hardware architecture is as follows:

[0061]

[0062]

[0063] The aforementioned CONV operator's operational expression is decomposed into 12 processing loops. For ease of understanding, the CONV operator will be used as an example below to further explain the method in the embodiments of this application.

[0064] The above-mentioned generation of N first architecture information corresponding one-to-one with the N network layers based on the joint exploration space may include: sampling multiple candidate architecture information corresponding to the first network layer from the joint exploration space, and evaluating the sampled multiple candidate architecture information to determine the power, performance, and area (PPA) of the convolution operator in the first network layer when running on the inference accelerator based on the architecture indicated in each candidate architecture information; then, searching from the joint exploration space based on the PPA of the inference accelerator to find the top k candidate architecture information that can give the inference accelerator a better PPA, and using the searched k candidate architecture information as the first architecture information of the first network layer, wherein the first network layer is any one of the N network layers, and k is an integer greater than 1. It is understood that the embodiments of this application only use the generation process of the first architecture information of the first network layer as an example to explain the generation process of the first architecture information of each network layer of the network model. In fact, each network layer of the network model can generate corresponding first architecture information according to the same method as the generation method of the first architecture information of the first network layer.

[0065] It should be noted that the number of candidate architecture information included in the above N first architecture information can be the same or different. That is, the value of k can be the same or different in different first architecture information.

[0066] In this implementation, during the design space exploration process, at least two candidate architecture information can be obtained for each network layer. Each candidate architecture information can serve as a solution for the corresponding network layer, meaning that multi-objective exploration can be achieved for each network layer during the design space exploration process. This multi-objective exploration facilitates multi-objective optimization, effectively balancing conflicting objectives such as inference energy consumption, inference latency, and area during space exploration. This improves the quality of the target architecture information generated based on the N first architecture information, thus enhancing the overall performance of the inference accelerator compared to single-objective exploration.

[0067] Optionally, generating N first architecture information points corresponding one-to-one with the N network layers based on the joint exploration space includes:

[0068] Based on the joint exploration space, the multi-objective genetic algorithm NSGA-II is used for iterative selection to obtain the first architecture information corresponding to the first network layer. The first network layer is any one of the N network layers. The parent population selected in the first iteration includes P candidate architecture information sampled from the joint exploration space. Each candidate architecture information is an individual in the parent population selected in the first iteration. P is an integer greater than 2.

[0069] The parent population of the i-th selection in the iterative selection is the (i-1)-th offspring population obtained from the (i-1)-th selection, where i is an integer greater than 1;

[0070] The first architecture information corresponding to the first network layer includes: the top k individuals with better PPA results in the offspring population obtained from the last selection in the iterative selection, where k is an integer greater than 1.

[0071] The above-mentioned iterative selection using the multi-objective genetic algorithm NSGA-II based on the joint exploration space to obtain the first architecture information corresponding to the first network layer may include the following steps:

[0072] K candidate architecture information samples obtained from the joint exploration space are used as the initialization population, wherein each candidate architecture information in the initialization population is an individual in the initialization population;

[0073] Each individual in the initial population is verified, and illegal individuals are filtered out to obtain a population excluding illegal individuals. This population excluding illegal individuals is then determined as the parent population selected in the first step. Illegal individuals are those whose hardware architecture information does not match. For example, if the architecture parameters of a candidate architecture recorded in an individual exceed the architecture parameters in the hardware architecture information, that individual is determined to be illegal. Another example is if an individual lacks a mapping relationship between some network layers and their corresponding candidate architectures. It should be noted that this embodiment only illustrates illegal individuals. In fact, various filtering conditions for illegal individuals can be set in advance as needed, and all individuals in the initial population can be verified and illegal individuals filtered based on these filtering conditions.

[0074] The first selection in the iterative selection process involves continuously selecting different parents from the parent population to mate with, generating new individuals. A cost estimation model is used to estimate the cost of each new individual, yielding a Producer-Per-Action (PPA) result. New individuals with superior PPA results are selected. Simultaneously, some individuals with superior PPA results are selected from the parent population as immigrant individuals. The cluster consisting of these new individuals with superior PPA results and the immigrant individuals is called the first offspring population. The generation of new individuals can include combining genes from different individuals in the parent population, or it can include mutating genes in the parent population to generate new individuals.

[0075] Then, the second selection in the iterative selection is performed. Specifically, the first offspring population is used as the parent population, and different parents are continuously selected for mating to generate new individuals. The cost is estimated for each new individual based on the Cost Model to obtain the PPA result of each new individual. New individuals with better PPA results can be selected. At the same time, some individuals with better PPA results can be selected from the parent population selected in the second selection as immigrant individuals. The cluster composed of the selected new individuals with better PPA results and the immigrant individuals is called the second offspring population.

[0076] Thus, the process is iterated continuously using the above method until the iteration termination condition is met. The iteration termination condition can be that the number of iterations exceeds a preset threshold, or it can be that the PPA target is met. The preset threshold can be determined empirically, for example, it could be 50 iterations, 100 iterations, etc. Meeting the PPA target means that the PPA results of all individuals in the offspring population after a certain iteration meet the preset PPA target.

[0077] Then, the top k individuals with better PPA results from the descendant population obtained by the last selection in the iterative selection are determined as the first architecture information corresponding to the first network layer.

[0078] In some embodiments of this application, to meet the needs of multi-objective optimization and large-scale spatial search, the embodiments of this application modify the NSGA-II (Non-dominated Sorting Genetic Algorithm II) algorithm to implement multi-objective search. Specifically, the genetic algorithm selects from the joint exploration space through multiple generations to output a Pareto front result that meets the requirements. For details, please refer to... Figure 10In some embodiments of this application, the specific process of obtaining the first architecture information corresponding to the first network layer by iterative selection using the multi-objective genetic algorithm NSGA-II based on the joint exploration space may include the following steps:

[0079] Initialize the population, obtain multiple candidate architecture information sampled from the joint exploration space, and use the sampled multiple candidate architecture information as the initial population, wherein each candidate architecture information in the initial population is an individual in the initial population;

[0080] Verify the population by verifying all individuals in the initial population and filtering out illegal individuals in the initial population to obtain a population that does not include illegal individuals. This population that does not include illegal individuals is then determined as the parent population selected in the first step.

[0081] Normalized fitness is achieved by normalizing the parent population selected in the first round.

[0082] Entering the evolutionary cycle, that is, performing the first selection in the iterative selection;

[0083] Parental selection (NSGA-II+E dominance) is a step in which different parents can be continuously selected from the parental population selected in the first selection to mate with and generate new individuals.

[0084] Record optimal fitness and diversity, that is, record the optimal fitness and diversity of each new individual;

[0085] Determine if the migration interval has been reached;

[0086] If so, introduce immigrant individuals, that is, select some individuals with better PPA results from the first selected parent population as immigrant individuals. Then, the cluster composed of the new individuals with better PPA results and the immigrant individuals is called the first offspring population. Then, execute the step "adjust parameters (adaptive) reinforcement learning supervision".

[0087] If not, the cluster of new individuals with better PPA results is called the first offspring population, and then the step "adjusting parameters (adaptive) reinforcement learning supervision" is directly executed.

[0088] Preserve the Pareto frontier results, i.e., preserve the first descendant population;

[0089] How do I determine if the iteration count or PPA target has been reached?

[0090] If not, return to the "Verify Population" step and proceed to the next iteration;

[0091] If so, then the process ends, and the top k individuals with better PPA results from the last selected offspring population are taken as the first architecture information corresponding to the first network layer.

[0092] The Pareto front results include several individuals that meet the criteria. After decoding the chromosome of each individual, a corresponding Arch and Mapping configuration can be obtained. From this Mapping, a MinimalArch is derived, which calculates the minimum hardware resources actually required under the current layer and mapping configuration. For example, suppose that under the mapping configuration, the L0 buffer only needs to store 1B of weights. In the Max Architecture, the L0 buffer size is set to 2B. Then, the MinimalArch needs to update the L0 buffer size to 1B and select the buffer larger than 1B and closest to 1B from the memory repository as the L0 buffer. After calculating the MinimalArch, the mapping information needs to be fine-tuned and updated based on the MinimalArch results. Then, the current Layer, MinimalArch, and Mapping information are re-input into the CostModel for evaluation, resulting in the updated PPA results. Therefore, the final Pareto output of the Mapper stage consists of several solutions. Each solution group recorded MinimalArch, Mapping, and PPA information. An example of one output solution is shown below:

[0093]

[0094]

[0095]

[0096] It is understood that the embodiments of this application only illustrate the generation process of the first architecture information of the first network layer to explain the generation process of the first architecture information of each network layer of the network model. In fact, each network layer of the network model can generate its corresponding first architecture information according to the same method used to generate the first architecture information of the first network layer.

[0097] In this implementation, by constructing a joint exploration space and combining it with the modified NSGA-II algorithm, multi-objective optimization is effectively achieved to meet the actual constraints of multiple objectives. Furthermore, a set of Pareto solutions is output for each layer of the network, effectively addressing the shortcomings of single-objective search and improving the overall performance of the inference accelerator.

[0098] Optionally, the step of iteratively selecting based on the joint exploration space using the multi-objective genetic algorithm NSGA-II to obtain the first architecture information corresponding to the first network layer includes:

[0099] The joint exploration space is sampled multiple times using the NSGA-II to obtain multiple sampled architecture information, wherein the sampled architecture information is the architecture information corresponding to the first network layer;

[0100] Based on the cost estimation model, the cost of the multiple sampling architecture information is estimated to obtain multiple power, performance and area PPA results that correspond one-to-one with the multiple sampling architecture information.

[0101] Based on the multiple PPA results, the multiple sampling architecture information is filtered to obtain the P candidate architecture information, wherein the P candidate architecture information are the top P sampling architecture information with better PPA results among the multiple sampling architecture information;

[0102] Using the P candidate architecture information as the parent population for the first selection, iterative selection is performed based on NSGA-II to obtain the first architecture information corresponding to the first network layer.

[0103] It is understood that the specific implementation process of using the P candidate architecture information as the parent population for the first selection, and performing iterative selection based on NSGA-II to obtain the first architecture information corresponding to the first network layer is the same as in the above embodiment. To avoid repetition, it will not be described again here.

[0104] In this implementation, the cost estimation model is used to successfully estimate the PPA results of each sampled architecture information, and then the P sampled architecture information with better PPA results is selected from multiple sampled architecture information. Then, the P candidate architecture information is used as the parent population for the first selection for iterative selection. In this process, since the P candidate architecture information in the parent population of the first selection is the P sampled architecture information with better PPA results selected from multiple sampled architecture information, it is beneficial to improve the quality of the parent population determined for the first selection, and thus to improve the quality of the target architecture information obtained after iterative selection.

[0105] Optionally, generating N first architecture information points corresponding one-to-one with the N network layers based on the joint exploration space includes:

[0106] Identification information is set for each of the N network layers. Among the N network layers, the identification information of network layers with the same operator type is the same, and the identification information of network layers with different operator types is different.

[0107] When generating the first architecture information corresponding to the second network layer, and when the N network layers include other network layers with the same identification information as the second network layer, the first architecture information corresponding to the second network layer is determined to be the first architecture information corresponding to the other network layers, wherein the second network layer is any one of the N network layers.

[0108] Specifically, each layer in the network model is labeled, and identical layers use the same label. For example... Figure 6 As shown, the network model consists of four layers. If layer 1 and layer 2 are identical Conv layers, both can be labeled as 1. Thus, the entire network model actually only has three different layers that need to be analyzed layer by layer. For identical layers, only the analysis results of the layers with the same label need to be copied, which greatly reduces the exploration time of the entire DSE framework. Figure 6 In the illustrated embodiment, during the generation of four first architecture information, only three first architecture information corresponding to layer0, layer1, and layer3 need to be generated, and the first architecture information of layer1 is used as the first architecture information of layer2, thus realizing the generation process of four first architecture information.

[0109] In this embodiment, by generating the first architecture information corresponding to the second network layer, and when the N network layers include other network layers with the same identification information as the second network layer, the first architecture information corresponding to the second network layer is determined to be the first architecture information corresponding to the other network layers. In this way, the number of network layers that need to be analyzed in the process of generating the first architecture information can be reduced, thereby reducing the time required for the entire DSE framework exploration process.

[0110] In related technologies, the DSE framework explores the network layer by layer. Therefore, for each layer of a Deep Neural Network (DNN), a joint "Architecture-Mapping" solution is obtained. The Architecture in the "Architecture-Mapping" solution searched by the DSE framework may differ between different layers of the same network. Although mapping can vary layer by layer because the NPU can be configured at runtime to support specific mapping methods, architecture is a static property that must be determined at design time. However, the DSE frameworks in related technologies search for "Architecture-Mapping" solutions for single layers, without outputting a unified Architecture structure for the entire network, thus failing to meet practical application requirements.

[0111] Optionally, generating target architecture information based on the N first architecture information includes:

[0112] A negotiation space is constructed based on the N first architecture information, wherein each point in the negotiation space includes one of the candidate architecture information from each of the N first architecture information;

[0113] Target architecture information is generated based on the negotiation space.

[0114] The aforementioned negotiation space can be a joint space composed of N subspaces, wherein each of the N subspaces corresponds one-to-one with the N first architecture information, and at least two candidate architecture information from each first architecture information form a subspace. A point in the aforementioned negotiation space can include combined information formed by sampling one candidate architecture information from each of the N subspaces, that is, a point in the negotiation space can include N candidate architecture information corresponding one-to-one with the N network sides, and each point in the negotiation space can serve as a spatial exploration solution for the network model.

[0115] The above-mentioned generation of target architecture information based on the negotiation space may include: sampling multiple points from the negotiation space by means of sampling, and evaluating the multiple points obtained by sampling to determine the power, performance, area (PPA) of the network model when running on the inference accelerator based on the architecture indicated in each point; then searching from the joint exploration space according to the PPA of the inference accelerator to find points that can enable the inference accelerator to have a better PPA; and the architecture information recorded in the searched points can be determined as the target architecture information.

[0116] In this implementation, a negotiation space is constructed based on the N first architecture information. This allows for further exploration of the constructed negotiation space to obtain the target architecture information. Compared to directly obtaining the target architecture information from the joint exploration space, this approach helps to narrow the exploration scope and thus improves the efficiency of space exploration.

[0117] Optionally, generating target architecture information based on the negotiation space includes:

[0118] Based on the negotiation space, NSGA-II is used for iterative selection to obtain the target individual. The parent population of the first selection in the iterative selection includes multiple points sampled from the negotiation space, each point being an individual from the parent population of the first selection. The parent population of the j-th selection in the iterative selection is the (j-1)-th descendant population obtained from the (j-1)-th selection, where j is an integer greater than 1. The target individual is the individual with the best PPA result among the descendant population obtained from the last selection in the iterative selection.

[0119] The target architecture information is generated based on the target individual.

[0120] The above-mentioned iterative selection using NSGA-II based on the negotiation space to obtain the target individual may include the following steps:

[0121] Multiple points sampled from the negotiation space are obtained, and the multiple points sampled are used as an initialization population, wherein each point in the initialization population is an individual in the initialization population;

[0122] The first selection in the iterative selection process involves continuously selecting different parents from the first-selection parent population for mating to generate new individuals. Cost estimation is performed on each new individual based on the Cost Model to obtain the PPA result for each new individual. New individuals with better PPA results can be selected. Simultaneously, some individuals with better PPA results can be selected from the first-selection parent population as immigrant individuals. The cluster consisting of the selected new individuals with better PPA results and the immigrant individuals is called the first offspring population. The generation process of new individuals can include combining the genes of different individuals in the parent population to form new individuals, or it can include mutating the genes of individuals in the parent population to generate new individuals.

[0123] Then, the second selection in the iterative selection is performed. Specifically, the first offspring population is used as the parent population, and different parents are continuously selected for mating to generate new individuals. The cost is estimated for each new individual based on the Cost Model to obtain the PPA result of each new individual. New individuals with better PPA results can be selected. At the same time, some individuals with better PPA results can be selected from the parent population selected in the second selection as immigrant individuals. The cluster composed of the selected new individuals with better PPA results and the immigrant individuals is called the second offspring population.

[0124] Thus, the process is iterated continuously using the above method until the iteration termination condition is met. The iteration termination condition can be that the number of iterations exceeds a preset threshold, or it can be that the PPA target is met. The preset threshold can be determined empirically, for example, it could be 50 iterations, 100 iterations, etc. Meeting the PPA target means that the PPA results of all individuals in the offspring population after a certain iteration meet the preset PPA target.

[0125] Then, the individual with the best PPA result in the offspring population obtained from the last selection in the iterative selection is determined as the target individual.

[0126] Please see Figure 12 The flowchart of the genetic algorithm in this embodiment mainly includes the following steps: loading mapper output, wherein the mapper output includes N first architecture information; initializing the population; fitness evaluation; non-dominated sorting; crowding calculation; parent selection; crossover and mutation; population evaluation; outputting a new population; determining whether the number of iterations has reached the preset number, or determining whether the PPA objective is met; if not, returning to the "fitness evaluation" step for the next round of iteration; if yes, saving the Pareto result.

[0127] Since the Negotiation phase also employs a multi-objective optimization algorithm, the algorithm's output is a set of Pareto solutions for the entire network model. Each solution contains information such as... Figure 11As shown, there is a Negotiated_Arch, the mapping for each layer, and the total PPA of the entire network. Specifically, the PPA results for each layer can be calculated separately based on the CostModel mentioned above. Then, the PPA results for each layer of the network model are summed to obtain the total PPA. Since layer 1 and layer 2 are identical, only layer 1 is evaluated during the Negotiation stage, and layer 2 can directly use the mapping and PPA results of layer 1. In addition, although layer 1 and layer 2 have the same shape, it is possible to choose different mappings for each layer. This is because of the combined effect of two mappings; each mapping optimizes a different objective, which may still achieve a better trade-off than simply replicating a single mapping on two identical layers. Therefore, in some other embodiments of this application, the PPA results of layer 1 and layer 2 can also be calculated separately during the Negotiation stage.

[0128] In this implementation, a negotiation mechanism is used to consider all layers of the model, achieving the generation of a unified and deterministic NPU architecture and mapping for each layer for a single model, thus satisfying the goal of end-to-end multi-objective optimization. This solution can meet the practical needs of designing a dedicated NPU architecture that conforms to PPA constraints for a specific network and is applicable to edge computing scenarios such as mobile devices.

[0129] Optionally, generating the target architecture information based on the target individual includes:

[0130] The target individual is parsed to obtain N second architecture information that correspond one-to-one with the N network layers. The second architecture information includes: the architecture of the corresponding network layer, and the mapping relationship between the network layer and the corresponding architecture.

[0131] The target architecture is determined based on the N architectures included in the N second architecture information, wherein each architecture parameter in the target architecture is the maximum value among the corresponding N architecture parameters in the N architectures;

[0132] Based on the N mapping relationships included in the N second architecture information, a mapping relationship between the N network layers and the target architecture is generated to obtain N target mapping relationships, wherein the target architecture information includes the target architecture and the N target mapping relationships.

[0133] In the execution of the genetic algorithm, after each individual is selected from the space, a Negotiation phase is performed to obtain the PPA results for all layers of the entire model. The Negotiation phase decodes the corresponding MinimalArch and Mapping from the pareto solution set of each layer based on the individual's chromosome results. Because the entire model requires a unified NPU architecture, and the hardware resource requirements recorded in the MinimalArch of each layer only meet the needs of that layer, the maximum value of the MinimalArch for all layers is taken to obtain the NegotiatedArch, which satisfies the hardware resources of all layers. Furthermore, the Mapping for each layer is obtained in the Mapper phase; when the architecture changes from MinimalArch to NegotiatedArch, the corresponding Mapping for each layer also needs to be updated. Then, all layers of the model, the NegotiatedArch, and the updated mapping are sent to CostModel for evaluation to obtain the PPA results for the entire network.

[0134] The architecture parameters described above can primarily include logical parameters and memory parameters. The logical parameters may include the MAC array dimensions (MAC Array Dims) and the number of MAC units (MAC Number). The MAC array dimensions refer to the dimensions of the MAC array. For example, please refer to [link to relevant documentation]. Figure 4 MAC ArrayDims is 3, and correspondingly, MAC Number is the number of MAC addresses. The Memory parameter mentioned above can include: the maximum storage capacity parameter for each tier in the storage architecture; for example, please refer to [link to relevant documentation]. Figure 3 ,exist Figure 3 In the illustrated embodiment, the memory parameters may include the following parameters: L0 Buffer Max Size, L1 Buffer Max Size, L2 Buffer Max Size, and L3 Buffer Max Size.

[0135] It is understandable that, since each of the above N architectures includes the following architecture parameters: MAC ArrayDims, MAC Number, L0 Buffer Max Size, L1 Buffer Max Size, L2 Buffer Max Size, and L3 Buffer Max Size, the maximum value among the N MAC ArrayDims parameters in the N architectures can be determined as the MAC ArrayDims parameter of the target architecture; the maximum value among the N L0 Buffer Max Size parameters in the N architectures can be determined as the L0 Buffer Max Size parameter of the target architecture; the maximum value among the N L1 Buffer Max Size parameters in the N architectures can be determined as the L1 Buffer Max Size parameter of the target architecture; the maximum value among the N L2 Buffer Max Size parameters in the N architectures can be determined as the L2 Buffer Max Size parameter of the target architecture; and the maximum value among the N L3 Buffer Max Size parameters in the N architectures can be determined as the L3 Buffer Max Size parameter of the target architecture. In this way, all the architecture parameters of the target architecture can be obtained.

[0136] The above-mentioned generation of the mapping relationship between the N network layers and the target architecture based on the N mapping relationships included in the N second architecture information can refer to: replacing the architecture in each of the above second architecture information with the target architecture, and replacing the mapping relationship between the architecture and the corresponding network layer recorded in the second architecture information with the mapping relationship between the target architecture and the corresponding network layer, thereby obtaining N target mapping relationships.

[0137] In this implementation, by making each architecture parameter in the target architecture the maximum value among the corresponding N architecture parameters in the N architectures, it can be ensured that the generated target architecture can meet the hardware requirements of each network layer of the network model.

[0138] Optionally, the joint exploration space includes an array size subspace, an indicator factor subspace, a cyclic sorting subspace, and a cyclic allocation subspace. The array size subspace includes multiple processing architecture schemes that can be constructed based on the processing unit information. The indicator factor subspace includes multiple factor decomposition schemes for factoring the dimensional information. The cyclic sorting subspace includes multiple sorting schemes for sorting each factor decomposition scheme. The cyclic allocation subspace includes multiple storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes.

[0139] The candidate architecture information includes: a processing architecture scheme in the array size subspace, a factorization scheme in the index factor subspace, a sorting scheme in the cyclic sorting subspace, and a storage architecture scheme in the cyclic allocation subspace.

[0140] The aforementioned joint exploration space includes the following four subspaces:

[0141] Subspace0 (first dimension): Represents the ArraySize subspace, recording all possible configurations of the processing unit array in the hardware architecture. Given computing power constraints and the dimensions of the processing unit array, an ArraySize subspace can be constructed. Each point in this subspace represents a configuration result and also one of the aforementioned processing architecture schemes. For example, please refer to... Figure 13 This application provides a processing architecture scheme in which the MAC array of the hardware architecture can be configured with three dimensions, and the number of MACs in each dimension is 32. It should be noted that... Figure 13 This is merely one example of a processing architecture scheme in this application; in fact, a large number of other different processing architecture schemes can be explored.

[0142] Subspace1 (the second dimension of the space): Represents the index factor subspace, recording the possible results of factorization for each dimension of the operator. The size of this subspace is limited by the values ​​of each dimension of the convolution operator and the total number of loops required by the mapping. Assume the values ​​of each dimension of the convolution operator are as follows: B=4, K=64, C=64, OY=56, OX=56, FY=3, FX=3. The mapping requires a total of 12 loops. Factorization of each dimension's values ​​yields a minimum of 1 factor and a maximum of the number of prime factors. For example, with OY=56, the number of factors is 1, meaning no factorization is performed, resulting in the loop (OY, 56). Alternatively, a sufficient prime factorization can be performed: 56 = 7 * 2 * 2 * 2, resulting in a maximum of 4 loops: (OY, 7), (OY, 2), (OY, 2), (OY, 2). In summary, there are limitations on the minimum (1) and maximum (number of prime factors) of the number of factors that can be decomposed into each dimension. Therefore, the problem becomes how to distribute 12 loops across 7 dimensions, which is a combinatorial problem in mathematics. All possible combinations constitute the size of the IndexFactor subspace. Below, a point in this subspace corresponds to one factorization scenario, also representing one of the factorization schemes described above. Please refer to... Figure 14 , Figure 14 This application provides a factorization scheme in which the number of loops corresponding to B is 1, meaning that the dimension value 4 of B is not decomposed. The number of loops corresponding to K is 2, therefore, the dimension value 64 of K can be decomposed into (K, 8) and (K, 8), that is, 64 is decomposed into 8×8. The number of loops corresponding to C is 2, therefore, the dimension value 64 of C can be decomposed into (C, 16) and (C, 4), that is, 64 is decomposed into 16×4. The number of loops corresponding to OY is 3, therefore, the dimension value 56 of OY can be decomposed into (OY, 7), (OY, 4), and (OY, 2), that is, 56 is decomposed into 7×4×2. The number of loops corresponding to OX is 2, therefore, the dimension value 56 of OX can be decomposed into (OX, 28) and (OX, 2), that is, 56 is decomposed into 28×2. The number of loops corresponding to FY is 1, meaning that the dimension value 3 of FY is not decomposed. The number of loops corresponding to FX is 1, meaning that the dimension value 3 of FX is not decomposed.

[0143] Subspace2 (the third dimension of the space): Represents the LoopPermutation subspace, which records the order of the factors. See also... Figure 15Mapping includes Spatial Mapping and Temporal Mapping. Spatial Mapping refers to the spatial parallelism of the operator. For example, if the number of MACs in the hardware architecture is configured as {D1:32, D2:32, D3:32}, it can be mapped to (K, 16) in the D1 dimension, meaning that each cycle can compute K=16 in parallel. Temporal Mapping refers to the fact that due to the limited computing power of the hardware architecture, the entire Conv operator needs to be executed multiple times in time to complete the computation. The loops decomposed in the IndexFactor subspace are divided into Spatial loops and Temporal loops. In the LoopPermutation subspace, the corresponding loop sorting needs to be determined according to the number of spatial loops and temporal loops set in the constraint. If 12 loops are set in the constraint file, with 7 of them being temporal loops, the remaining 5 loops are distributed in the dimensions of the 3D MAC array as D1:2, D2:2, D3:1. The problem is transformed into dividing 12 loops into 4 parts, each containing 7, 2, 2, and 1 loops, with the loops within each part being internally ordered. The allocation and ordering of all loops constitute the size of the subspace. Furthermore, mapping restrictions for loops can be specified in the constraint file according to requirements. For example, it can be stipulated that loops with the character 'FX' cannot be mapped on dimension D1, and loops with the character 'FY' cannot be mapped on dimension D2. The content of the constraint file can be set according to actual needs. Please refer to [link to relevant documentation]. Figure 15 This is one possible sorting scheme, where the Spatial Mapping includes: D1 corresponding to loop 1 and loop 0, D2 corresponding to loop 4 and loop 2, and D3 corresponding to loop 3. The Temporal Mapping includes: Temporal corresponding to loop 5, loop 6, loop 7, loop 8, loop 1, loop 11, loop 10, and loop 9. It should be noted that... Figure 15 This is merely one example of a sorting scheme in this application; in fact, a large number of other different sorting schemes can be explored.

[0144] Subspace3 (the fourth dimension of space): Represents the loop allocation subspace, recording the allocation of factors across different storage levels. Due to limitations in the data types stored at different storage levels in the hardware architecture, [the following is omitted as it is not directly related to the subspace description]. Figure 16 For example, Weight can only be stored at L0 and L3, Input can only be stored at L2 and L3, and Output can only be stored at L1, L2, and L3. Therefore, to address the uneven distribution of different data types, the memory level to which a Loop can be mapped must be determined based on the dimness of each Loop. Considering the space requirements, the memory levels to be mapped for all spatial loops can be predefined, so this subspace only explores the allocation of temporal loops across various memory levels. Please see [link to relevant documentation]. Figure 16 This application provides a storage architecture scheme using 7 temporary loops as an example. For the weight, only 2 memory levels are available. The 7 sorted loops need to be allocated to these 2 levels. One allocation method is: level 0 corresponds to loop 5; level 1 corresponds to loops 6, 7, 8, 1, 11, 10, and 9. The allocation method is similar for Input and Output, as detailed below. Figure 16 As shown. The size of this subspace is determined by all possible allocations of temporal loops across the three data types. It's important to note that these temporal loops are a result of the Looppermutation subspace output; that is, the order of the temporal loops is already determined, only the allocation differs. From... Figure 16 As can be seen, although the number of memory levels differs among the three data types, resulting in different allocation outcomes, all loops are allocated from the inside out, from low level to high level. It should be noted that... Figure 16 This is merely one example of a storage architecture scheme in this application; in fact, a large number of other different storage architecture schemes can be explored.

[0145] Understandably, the entire joint exploration space consists of four subspaces, and each point in the space represents a way of configuring and mapping architectural hardware.

[0146] In this implementation, by incorporating the search for architectural parameters into the joint exploration space, the optimal architectural configuration and mapping configuration for a single layer can be explored simultaneously, thereby improving the quality of the generated target architectural information. Furthermore, in related technologies, most frameworks only support uniform mapping when constructing the mapping space, failing to support non-uniform mapping, resulting in an incomplete mapping space and suboptimal search results. Therefore, in this embodiment, by constructing a cyclic allocation subspace within the joint exploration space, different data types can be processed separately, effectively supporting non-uniform mapping configurations.

[0147] In related technologies, the search space constructed by the DSE framework is extremely large, which poses a certain challenge to the search. Research has revealed that many design flaws exist within this space (e.g., some mapping methods lead to storage capacity exceeding limits), but most DSE frameworks fail to pre-emptively remove these flawed points after constructing the space. This results in two negative consequences: first, the lack of space pruning leads to an excessively large space, causing excessively long convergence times for the search algorithm; second, the search algorithm retrieves many unreasonable results from the space, and discarding these results after inspection and restarting the search significantly reduces search efficiency.

[0148] Optionally, before constructing the joint exploration space based on the model information and the hardware architecture information, the method further includes:

[0149] Obtain a constraint file, wherein the constraint file includes at least one of the following constraint information: first constraint information, second constraint information, third constraint information, and fourth constraint information; the first constraint information is used to constrain the number of processing units in the processing architecture scheme, the second constraint information is used to constrain the number of processing loops in the factorization scheme, the third constraint information is used to constrain the mapping relationship between the processing units and the processing loops, and the fourth constraint information is used to constrain the storage limit of each storage level and the data storage type of each storage level in the storage architecture scheme.

[0150] The construction of the joint exploration space based on the model information and the hardware architecture information includes:

[0151] An initial exploration space is constructed based on the model information and the hardware architecture information. The initial exploration space includes an initial array size subspace, an initial indicator factor subspace, an initial cyclic sorting subspace, and an initial cyclic allocation subspace. The initial array size subspace includes all processing architecture schemes that can be constructed based on the processing unit information. The initial indicator factor subspace includes all factor decomposition schemes that factorize the dimensional information. The initial cyclic sorting subspace includes all sorting schemes that sort each factor decomposition scheme. The initial cyclic allocation subspace includes all storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes.

[0152] Based on the constraint file, the initial exploration space is spatially pruned to obtain the constructed joint exploration space, wherein the candidate architecture information matches the constraint file.

[0153] Please see Figure 13 In some embodiments of this application, the first constraint information mentioned above may include: the minimum number of Macs (Min Mac Num) is 64, the maximum number of Macs (Man Mac Num) is 12288, and the number of dimensions (Dim Array Num) is 3. Thus, under the constraint of the first constraint, a possible processing architecture scheme is: D1 is 32, D2 is 32, and D3 is 32.

[0154] The number of processing units in the above processing architecture is the same as the number of loops. For example, please refer to [link to example]. Figure 14 , Figure 14 The processing architecture scheme in the above describes a total of 12 processing units. In practical implementation, the number of processing units in the processing architecture scheme can be constrained according to actual needs using the aforementioned second constraint information.

[0155] The aforementioned third constraint information can be used to constrain the mapping relationship between the processing unit and the processing loop. This mapping relationship can include relationships where the processing unit can be mapped by dimension parameters, and relationships where the processing unit cannot be mapped by dimension parameters. For example, please refer to [link to relevant documentation]. Figure 15 In some embodiments of this application, the third constraint information may include: D1 Disallow Dim: FX, and D2 Disallow Dim: FY, wherein D1 Disallow Dim: FX indicates that a loop with FX cannot be mapped in the D1 dimension; and D2 Disallow Dim: FY indicates that a loop with FY cannot be mapped in the D2 dimension.

[0156] It should be noted that, in addition to constraining the storage limit and data storage type of each storage tier in the aforementioned storage architecture scheme, the fourth constraint information can also constrain other storage-related information. For example, it can constrain the number of storage tiers that each type of data can correspond to based on the fourth constraint information. For example, please refer to [link to relevant documentation]. Figure 16 In some embodiments of this application, the fourth constraint information may include: W level: 2; I level: 2; O level: 3; where W level: 2 indicates that the number of storage levels corresponding to data of data type W is 2, such as... Figure 3 As shown, data of data type W corresponds to storage levels Level 0 and Level 1, meaning data of data type W can only be stored in Level 0 and Level 1. Correspondingly, data of data type I corresponds to storage levels Level 0 and Level 1. Data of data type O corresponds to storage levels Level 0, Level 1, and Level 2.

[0157] Please refer to Table 1 below, which is an example of the constraint file for some embodiments of this application:

[0158] Table 1:

[0159]

[0160] In Table 1, the constraint variable Spatial loops num corresponds to the first constraint information; the constraint variable Temporal loops num corresponds to the second constraint information; the constraint variable Disallowedspatial loop dim corresponds to the third constraint information; and the constraint variables User defined loops and Maxmemory sizes correspond to the fourth constraint information.

[0161] It should be noted that the descriptions of the first constraint information, the second constraint information, the third constraint information, and the fourth constraint information in the above embodiments are only some examples in the embodiments of this application. In fact, users can set the specific content of the first constraint information, the second constraint information, the third constraint information, and the fourth constraint information as needed.

[0162] Matching the candidate architecture information with the constraint file means that the candidate architecture information satisfies the content constrained by all the constraint information in the constraint file.

[0163] Specifically, please see Figure 17 (a), Figure 17(a) An initial exploration space can be constructed based on the above model information and the hardware architecture information. Since the initial exploration space is very large and due to the mutual constraints of the architecture and mapping, the configuration of many points in the space is unreasonable. Based on this, in this embodiment of the application, the initial exploration space can be trimmed by the constraint file to remove unreasonable points in the initial exploration space, thereby reducing the size of the space.

[0164] In this embodiment, by setting constraint information in the constraint file, after constructing the initial exploration space, all points in the initial exploration space can be further checked based on the constraint file, and points that do not conform to the constraint file can be removed. The space after clipping is much smaller than the original space; the specific reduction ratio and the specific size of the operator depend on the constraint file set. Please refer to... Figure 17 (b) is a schematic diagram of constructing a joint exploration space by performing spatial pruning on the initial exploration space based on the constraint file.

[0165] After the joint exploration space was constructed and trimmed, the remaining points were all valid. The next goal is to use automated exploration tools to quickly explore the space to obtain the target architecture information.

[0166] In this implementation, the initial exploration space is pruned based on the constraint file to obtain the constructed joint exploration space. This can eliminate unreasonable points in advance, greatly reduce the space size, facilitate the search algorithm to converge quickly, and ensure that the search results are all valid.

[0167] To facilitate understanding, the method for generating architectural information in the embodiments of this application will be further explained below with reference to the accompanying drawings:

[0168] Please see Figure 6 This is a flowchart illustrating the method for generating architecture information based on the DSE framework in this embodiment of the application. The DSE framework consists of two main modules: a mapper and a negotiator.

[0169] First, let's describe the input information for the DSE framework:

[0170] Network Model: This describes the information of each layer of the network model in detail. The model can be in ONNX or PB format.

[0171] Max Architecture specifies the storage hierarchy, maximum capacity, bit width, area, and other details for each storage tier, as well as the dimensions of the MAC array and the total number of MAC units. Some key information of interest to DSE is shown in Table 2 below:

[0172] Table 2:

[0173]

[0174] Constraints: This section describes the constraints that can be set when constructing the Architecture-Mapping joint space to ensure that the solutions represented by the points in the space meet actual requirements. For example, in this embodiment, the number of MAC units allocated in the D2 dimension can be set to no more than 32, and the spatial parallelism of the Kernel's FX and FY can be restricted.

[0175] CostModel: Cost estimation model. It can perform cost analysis for each architecture-mapping solution to obtain the corresponding PPA result under the current architecture configuration and mapping method at each layer.

[0176] Per-Layer-Lookup: This assigns a label to each layer in the model, using the same label for identical layers. For example, if a model has four layers... Figure 6 As shown. If layer 1 and layer 2 are identical Conv layers, both layers can be marked as 1. This way, the entire model actually only has three different layers that need to be analyzed layer by layer. For identical layers, simply copy the analysis results of the layers with the same label. This significantly reduces the time spent exploring the entire DSE framework.

[0177] Then, the functionality of the Mapper is described. The Mapper searches for several suitable configuration results from the joint space of architecture and mapping for each layer of the model using a multi-objective search method.

[0178] Joint Exploration Space (Architecture-Mapping Space): This joint exploration space consists of two parts: an Arch Space and a Map Space.

[0179] Multiobjective exploration: Due to the conflicting nature of the optimization objectives (inference energy consumption, inference latency, and area), the result of the exploration is a set of Pareto solutions, and the designer can choose the most suitable solution for the specific application. Figure 9The advantages of exploring in a multi-objective manner are illustrated graphically. The figure shows the single solution found through single-objective design space exploration (black circle) and the Pareto front (white circle) of the solution found through multi-objective design space exploration, where the size of the circle again represents the area of ​​the system configuration. The existence of multiple Pareto equivalent solutions allows designers to choose the most appropriate trade-offs for a specific application.

[0180] Search-Method: To meet the needs of multi-objective optimization and large-scale spatial search, this application's embodiment modifies the NSGA-II (Non-dominated Sorting Genetic Algorithm II) algorithm to implement multi-objective search. The specific flowchart is as follows... Figure 10 As shown. The genetic algorithm selects from the joint space of Architecture-Mapping through multiple generations and outputs the Pareto front result that meets the requirements. To avoid repetition, it will not be described in detail here.

[0181] The Negotiation phase performs end-to-end system optimization. Its role is to explore the space generated by the Cartesian product of the pareto solution sets provided by the Mapper to derive the pareto set for optimizing system configurations with multiple design objectives optimized in an end-to-end manner.

[0182] Assuming the model is Figure 6 The model in the application has four layers, with the middle two being identical. This embodiment only needs to consider three different layers. The Mapper phase provides a Pareto-efficient population of individuals for each of the three layers, such as... Figure 11 As shown in the illustration. In this embodiment, the Pareto solutions of these three layers are used to generate a space called the Negotiator Space through Cartesian sets. Each individual in the Negotiator Space has a vector for its chromosome, and each position in the vector records the index of a solution from the Pareto solution obtained from the mapper stage of each layer.

[0183] For example, see Figure 6Assuming N is 4, and the four network layers of the network model are denoted as layer0, layer1, layer2, and layer3, if the first architecture information corresponding to layer0 includes four candidate architectures (i.e., the subspace corresponding to layer0 includes four individuals), and these four individuals can be labeled as {0, 1, 2, 3}; if the first architecture information corresponding to layer1 includes three candidate architectures, and the subspace corresponding to layer1 includes three individuals, and the subspace corresponding to layer1 includes three individuals, and the subspace corresponding to layer2 includes three individuals, and the subspace corresponding to layer2 includes three individuals, and the subspace corresponding to layer2 includes three individuals, and the subspace corresponding to layer2 includes five candidate architectures, and the subspace corresponding to layer3 includes five individuals ... Since layer1 and layer2 have the same operator type, they can be represented by the same coordinate system in the Negotiator Space. Based on this, the size of the Negotiator Space is 4*3*5=60.

[0184] The chromosome of an individual in the Negotiator Space can be {2,0,4}, which means that layer0 selected the result with index=2 from its Mapper pareto solution, layer1 selected the result with index=0 from its Mapper pareto solution, and layer4 selected the result with index=4 from its Mapper pareto solution.

[0185] The Negotiation phase also employs a multi-objective genetic algorithm to select suitable individuals from the space. The flowchart of the genetic algorithm is as follows: Figure 12 As shown, to avoid repetition, it will not be elaborated further here.

[0186] The architecture information generation method provided in this application can be executed by an architecture information generation device. This application uses an architecture information generation device executing the architecture information generation method as an example to illustrate the architecture information generation device provided in this application.

[0187] Please see Figure 18 , Figure 18 This application provides a schematic diagram of the structure of an architecture information generation device 1800, which includes:

[0188] The acquisition module 1801 is used to acquire model information of the network model and hardware architecture information of the inference accelerator. The network model includes N network layers, the model information includes N dimension information corresponding to the N network layers, where N is an integer greater than 1, the dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers, and the hardware architecture information includes processing unit information and storage architecture information.

[0189] The construction module 1802 is used to construct a joint exploration space based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, wherein the candidate architecture information includes: a candidate architecture corresponding to a network layer, and the mapping relationship between the network layer and the corresponding candidate architecture;

[0190] The generation module 1803 is used to generate N first architecture information corresponding one-to-one with the N network layers based on the joint exploration space, wherein the first architecture information includes at least two candidate architecture information;

[0191] The generation module 1803 is further configured to generate target architecture information based on the N first architecture information, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

[0192] Optionally, the generation module 1803 is specifically used to perform iterative selection based on the joint exploration space using the multi-objective genetic algorithm NSGA-II to obtain the first architecture information corresponding to the first network layer, wherein the first network layer is any one of the N network layers, and the parent population selected in the first iteration includes multiple candidate architecture information sampled from the joint exploration space, and each candidate architecture information is an individual in the parent population selected in the first iteration;

[0193] The parent population of the i-th selection in the iterative selection is the (i-1)-th offspring population obtained from the (i-1)-th selection, where i is an integer greater than 1;

[0194] The first architecture information corresponding to the first network layer includes: the top k individuals with better PPA results in the offspring population obtained from the last selection in the iterative selection, where k is an integer greater than 1.

[0195] Optionally, the generation module 1803 includes:

[0196] The sampling submodule is used to perform multiple samplings on the joint exploration space using the NSGA-II to obtain multiple sampling architecture information, wherein the sampling architecture information is the architecture information corresponding to the first network layer;

[0197] The cost estimation submodule is used to perform cost estimation on the multiple sampling architecture information based on the cost estimation model, and obtain multiple power, performance and area PPA results that correspond one-to-one with the multiple sampling architecture information.

[0198] The filtering submodule is used to filter the multiple sampling architecture information based on the multiple PPA results to obtain the P candidate architecture information, wherein the P candidate architecture information is the top P sampling architecture information with better PPA results among the multiple sampling architecture information;

[0199] The selection submodule is used to use the P candidate architecture information as the parent population for the first selection, and to perform iterative selection based on NSGA-II to obtain the first architecture information corresponding to the first network layer.

[0200] Optionally, the generation module 1803 includes:

[0201] The identifier submodule is used to set identifier information for the N network layers respectively. Among the N network layers, the identifier information of network layers with the same corresponding operator type is the same, and the identifier information of network layers with different corresponding operator types is different.

[0202] The first determining submodule is configured to, when generating the first architecture information corresponding to the second network layer and the N network layers include other network layers with the same identification information as the second network layer, determine the first architecture information corresponding to the second network layer as the first architecture information corresponding to the other network layer, wherein the second network layer is any one of the N network layers.

[0203] Optionally, the generation module 1803 includes:

[0204] A construction submodule is used to construct a negotiation space based on the N first architecture information, wherein each point in the negotiation space includes one of the candidate architecture information in each of the N first architecture information;

[0205] A generation submodule is used to generate target architecture information based on the negotiation space.

[0206] Optionally, the generation submodule includes:

[0207] The selection unit is used to perform iterative selection based on the negotiation space using NSGA-II to obtain a target individual. The parent population of the first selection in the iterative selection includes multiple points sampled from the negotiation space, each point being an individual from the parent population of the first selection. The parent population of the j-th selection in the iterative selection is the (j-1)-th descendant population obtained from the (j-1)-th selection, where j is an integer greater than 1. The target individual is the individual with the best PPA result among the descendant population obtained from the last selection in the iterative selection.

[0208] A generation unit is used to generate the target architecture information based on the target individual.

[0209] Optionally, the generation unit includes:

[0210] The parsing subunit is used to parse the target individual to obtain N second architecture information that correspond one-to-one with the N network layers. The second architecture information includes: the architecture of the corresponding network layer, and the mapping relationship between the network layer and the corresponding architecture.

[0211] A sub-unit is defined for determining a target architecture based on the N architectures included in the N second architecture information, wherein each architecture parameter in the target architecture is the maximum value among the corresponding N architecture parameters in the N architectures;

[0212] A generation subunit is used to generate mapping relationships between the N network layers and the target architecture based on the N mapping relationships included in the N second architecture information, thereby obtaining N target mapping relationships, wherein the target architecture information includes the target architecture and the N target mapping relationships.

[0213] Optionally, the joint exploration space includes an array size subspace, an indicator factor subspace, a cyclic sorting subspace, and a cyclic allocation subspace. The array size subspace includes multiple processing architecture schemes that can be constructed based on the processing unit information. The indicator factor subspace includes multiple factor decomposition schemes for factoring the dimensional information. The cyclic sorting subspace includes multiple sorting schemes for sorting each factor decomposition scheme. The cyclic allocation subspace includes multiple storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes.

[0214] The candidate architecture information includes: a processing architecture scheme in the array size subspace, a factorization scheme in the index factor subspace, a sorting scheme in the cyclic sorting subspace, and a storage architecture scheme in the cyclic allocation subspace.

[0215] Optionally, the acquisition module 1801 is further configured to acquire a constraint file, wherein the constraint file includes at least one of the following constraint information: first constraint information, second constraint information, third constraint information, and fourth constraint information; the first constraint information is used to constrain the number of processing units in the processing architecture scheme, the second constraint information is used to constrain the number of processing loops in the factorization scheme, the third constraint information is used to constrain the mapping relationship between the processing units and the processing loops, and the fourth constraint information is used to constrain the storage limit of each storage level and the data storage type of each storage level in the storage architecture scheme;

[0216] The construction module 1802 includes:

[0217] A construction submodule is used to construct an initial exploration space based on the model information and the hardware architecture information. The initial exploration space includes an initial array size subspace, an initial indicator factor subspace, an initial cyclic sorting subspace, and an initial cyclic allocation subspace. The initial array size subspace includes all processing architecture schemes that can be constructed based on the processing unit information. The initial indicator factor subspace includes all factor decomposition schemes that factorize the dimensional information. The initial cyclic sorting subspace includes all sorting schemes that sort each factor decomposition scheme. The initial cyclic allocation subspace includes all storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes.

[0218] The pruning submodule is used to perform spatial pruning on the initial exploration space based on the constraint file to obtain the constructed joint exploration space, wherein the candidate architecture information matches the constraint file.

[0219] In this implementation, during the design space exploration process, at least two candidate architectures can be obtained for each network layer. Each candidate architecture can serve as a solution for the corresponding network layer, enabling multi-objective exploration for each network layer. Furthermore, a negotiation space is constructed based on the at least two candidate architectures for each network layer, and target architecture information is generated based on this negotiation space. Because multi-objective exploration can be performed for each network layer, multi-objective optimization is facilitated. This allows for effective balancing of conflicting objectives such as inference energy consumption, inference latency, and area during space exploration, improving the quality of the obtained target architecture information. Therefore, compared to single-objective exploration, this approach enhances the overall performance of the inference accelerator.

[0220] The architecture information generation device 1800 in this application embodiment can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope.

[0221] The architecture information generation device 1800 in this embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment does not specifically limit its use.

[0222] The architecture information generation device 1800 provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0223] This application embodiment also provides a first chip, in which a network model is deployed, wherein the network model is a network model deployed based on target architecture information, and the target architecture information is generated based on the architecture information generation method described in the above embodiments.

[0224] The first chip can serve as various types of inference accelerators. It is understood that the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layers and their corresponding architectures. Therefore, during the deployment of the network model, the architecture corresponding to the network model in the first chip can be determined based on the target architecture information, and the mapping relationship between each network layer and its corresponding architecture can be established based on the target architecture information.

[0225] In this embodiment, since the network model is a network model deployed based on the target architecture information, it is beneficial to improve the overall performance of the first chip.

[0226] This application also provides an accelerator in which a network model is deployed. The network model is a network model deployed based on target architecture information, which is generated based on the architecture information generation method described in the above embodiments.

[0227] It is understood that, since the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layers and their corresponding architectures, the architecture corresponding to the network model in the accelerator can be determined based on the target architecture information, and the mapping relationship between each network layer and its corresponding architecture can be established based on the target architecture information.

[0228] In this embodiment, since the network model is a network model deployed based on the target architecture information, it is beneficial to improve the overall performance of the inference accelerator.

[0229] In some embodiments, such as Figure 19 As shown, this application embodiment also provides an electronic device 1900, including a processor 1901, a memory 1902, and a program or instructions stored in the memory 1902 and executable on the processor 1901. When the program or instructions are executed by the processor 1901, they implement the various processes of the above-described method embodiment for generating architecture information and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0230] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0231] Figure 20 A schematic diagram of the hardware structure of an electronic device according to an embodiment of this application.

[0232] The electronic device 2000 includes, but is not limited to, components such as: radio frequency unit 2001, network module 2002, audio output unit 2003, input unit 2004, sensor 2005, display unit 2006, user input unit 2007, interface unit 2008, memory 2009, and processor 2010.

[0233] The processor 2010 is used to acquire model information of the network model and hardware architecture information of the inference accelerator. The network model includes N network layers, and the model information includes N dimension information corresponding to the N network layers, where N is an integer greater than 1. The dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers. The hardware architecture information includes processing unit information and storage architecture information.

[0234] The processor 2010 is configured to construct a joint exploration space based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, wherein the candidate architecture information includes: a candidate architecture corresponding to a network layer, and a mapping relationship between the network layer and its corresponding candidate architecture;

[0235] The processor 2010 is configured to generate N first architecture information corresponding one-to-one with the N network layers based on the joint exploration space, wherein the first architecture information includes at least two candidate architecture information.

[0236] The processor 2010 is configured to generate target architecture information based on the N first architecture information, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

[0237] Optionally, the processor 2010 is configured to perform iterative selection based on the joint exploration space using the multi-objective genetic algorithm NSGA-II to obtain the first architecture information corresponding to the first network layer, wherein the first network layer is any one of the N network layers, and the parent population selected in the first iteration includes P candidate architecture information sampled from the joint exploration space, each candidate architecture information being an individual in the parent population selected in the first iteration, where P is an integer greater than 2;

[0238] The parent population of the i-th selection in the iterative selection is the (i-1)-th offspring population obtained from the (i-1)-th selection, where i is an integer greater than 1;

[0239] The first architecture information corresponding to the first network layer includes: the top k individuals with better PPA results in the offspring population obtained from the last selection in the iterative selection, where k is an integer greater than 1.

[0240] Optionally, the processor 2010 is configured to use the NSGA-II to sample the joint exploration space multiple times to obtain multiple sampling architecture information, wherein the sampling architecture information is the architecture information corresponding to the first network layer;

[0241] The processor 2010 is used to perform cost estimation on the multiple sampling architecture information based on the cost estimation model, and obtain multiple power, performance and area PPA results that correspond one-to-one with the multiple sampling architecture information.

[0242] The processor 2010 is configured to filter the multiple sampling architecture information based on the multiple PPA results to obtain the P candidate architecture information, wherein the P candidate architecture information are the top P sampling architecture information with better PPA results among the multiple sampling architecture information;

[0243] The processor 2010 is used to perform iterative selection based on the NSGA-II, using the P candidate architecture information as the parent population for the first selection, to obtain the first architecture information corresponding to the first network layer.

[0244] Optionally, the processor 2010 is configured to set identification information for the N network layers respectively, wherein the identification information of network layers with the same corresponding operator type is the same, and the identification information of network layers with different corresponding operator types is different.

[0245] The processor 2010 is configured to, when generating the first architecture information corresponding to the second network layer, and when the N network layers include other network layers with the same identification information as the second network layer, determine the first architecture information corresponding to the second network layer as the first architecture information corresponding to the other network layer, wherein the second network layer is any one of the N network layers.

[0246] Optionally, the processor 2010 is configured to construct a negotiation space based on the N first architecture information, wherein each point in the negotiation space includes one of the candidate architecture information in each of the N first architecture information;

[0247] The processor 2010 is used to generate the target architecture information based on the negotiation space.

[0248] Optionally, the processor 2010 is configured to perform iterative selection using NSGA-II based on the negotiation space to obtain a target individual, wherein the parent population of the first selection in the iterative selection includes multiple points sampled from the negotiation space, and each point is an individual in the parent population of the first selection; the parent population of the j-th selection in the iterative selection is the (j-1)-th offspring population obtained from the (j-1)-th selection, where j is an integer greater than 1; the target individual is the individual with the best PPA result in the offspring population obtained from the last selection in the iterative selection.

[0249] The processor 2010 is used to generate the target architecture information based on the target individual.

[0250] Optionally, the processor 2010 is configured to parse the target individual to obtain N second architecture information corresponding one-to-one with the N network layers, wherein the second architecture information includes: the architecture of the corresponding network layer, and the mapping relationship between the network layer and the corresponding architecture;

[0251] The processor 2010 is configured to determine a target architecture based on the N architectures included in the N second architecture information, wherein each architecture parameter in the target architecture is the maximum value among the corresponding N architecture parameters in the N architectures;

[0252] The processor 2010 is configured to generate mapping relationships between the N network layers and the target architecture based on the N mapping relationships included in the N second architecture information, thereby obtaining N target mapping relationships, wherein the target architecture information includes the target architecture and the N target mapping relationships.

[0253] Optionally, the joint exploration space includes an array size subspace, an indicator factor subspace, a cyclic sorting subspace, and a cyclic allocation subspace. The array size subspace includes multiple processing architecture schemes that can be constructed based on the processing unit information. The indicator factor subspace includes multiple factor decomposition schemes for factoring the dimensional information. The cyclic sorting subspace includes multiple sorting schemes for sorting each factor decomposition scheme. The cyclic allocation subspace includes multiple storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes.

[0254] The candidate architecture information includes: a processing architecture scheme in the array size subspace, a factorization scheme in the index factor subspace, a sorting scheme in the cyclic sorting subspace, and a storage architecture scheme in the cyclic allocation subspace.

[0255] Optionally, the processor 2010 is configured to acquire a constraint file, wherein the constraint file includes at least one of the following constraint information: first constraint information, second constraint information, third constraint information, and fourth constraint information; the first constraint information is used to constrain the number of processing units in the processing architecture scheme, the second constraint information is used to constrain the number of processing loops in the factorization scheme, the third constraint information is used to constrain the mapping relationship between the processing units and the processing loops, and the fourth constraint information is used to constrain the storage limit of each storage level and the data storage type of each storage level in the storage architecture scheme;

[0256] The processor 2010 is configured to construct an initial exploration space based on the model information and the hardware architecture information. The initial exploration space includes an initial array size subspace, an initial index factor subspace, an initial cyclic sorting subspace, and an initial cyclic allocation subspace. The initial array size subspace includes all processing architecture schemes that can be constructed based on the processing unit information. The initial index factor subspace includes all factor decomposition schemes that factorize the dimensional information. The initial cyclic sorting subspace includes all sorting schemes that sort each factor decomposition scheme. The initial cyclic allocation subspace includes all storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes.

[0257] The processor 2010 is configured to perform spatial pruning on the initial exploration space based on the constraint file to obtain the constructed joint exploration space, wherein the candidate architecture information matches the constraint file.

[0258] Those skilled in the art will understand that the electronic device 2000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 2010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 20 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0259] It should be understood that, in this embodiment, the input unit 2004 may include a graphics processing unit (GPU) 20041 and a microphone 20042. The GPU 20041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 2006 may include a display panel 20061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 2007 includes a touch panel 20071 and other input devices 20072. The touch panel 20071 is also called a touch screen. The touch panel 20071 may include two parts: a touch detection device and a touch controller. Other input devices 20072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0260] The memory 2009 can be used to store software programs and various data. The memory 2009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 2009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 2009 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0261] Processor 2010 may include one or more processing units; optionally, processor 2010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 2010.

[0262] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described method for generating architectural information and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0263] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0264] This application embodiment also provides a second chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described method embodiment for generating architecture information, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0265] It should be understood that the second chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0266] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0267] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0268] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for generating architectural information, characterized in that, include: Obtain model information of the network model and hardware architecture information of the inference accelerator. The network model includes N network layers. The model information includes N dimension information corresponding one-to-one with the N network layers. N is an integer greater than 1. The dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers. The hardware architecture information includes processing unit information and storage architecture information. A joint exploration space is constructed based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, and the candidate architecture information includes: a candidate architecture corresponding to a network layer, and the mapping relationship between the network layer and the corresponding candidate architecture; Based on the joint exploration space, N first architecture information corresponding one-to-one with the N network layers are generated, wherein the first architecture information includes at least two candidate architecture information; Based on the N first architecture information, target architecture information is generated, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

2. The method according to claim 1, characterized in that, The generation of N first architecture information points, each corresponding one-to-one with the N network layers, based on the joint exploration space includes: Based on the joint exploration space, the multi-objective genetic algorithm NSGA-II is used for iterative selection to obtain the first architecture information corresponding to the first network layer. The first network layer is any one of the N network layers. The parent population selected in the first iteration includes P candidate architecture information sampled from the joint exploration space. Each candidate architecture information is an individual in the parent population selected in the first iteration. P is an integer greater than 2. The parent population of the i-th selection in the iterative selection is the (i-1)-th offspring population obtained from the (i-1)-th selection, where i is an integer greater than 1; The first architecture information corresponding to the first network layer includes: the top k individuals with better PPA results in the offspring population obtained from the last selection in the iterative selection, where k is an integer greater than 1.

3. The method according to claim 2, characterized in that, The step of iteratively selecting the first architecture information corresponding to the first network layer using the multi-objective genetic algorithm NSGA-II based on the joint exploration space includes: The joint exploration space is sampled multiple times using the NSGA-II to obtain multiple sampled architecture information, wherein the sampled architecture information is the architecture information corresponding to the first network layer; Based on the cost estimation model, the cost of the multiple sampling architecture information is estimated to obtain multiple power, performance and area PPA results that correspond one-to-one with the multiple sampling architecture information. Based on the multiple PPA results, the multiple sampling architecture information is filtered to obtain the P candidate architecture information, wherein the P candidate architecture information are the top P sampling architecture information with better PPA results among the multiple sampling architecture information; Using the P candidate architecture information as the parent population for the first selection, iterative selection is performed based on NSGA-II to obtain the first architecture information corresponding to the first network layer.

4. The method according to claim 1, characterized in that, The generation of N first architecture information points, each corresponding one-to-one with the N network layers, based on the joint exploration space includes: Identification information is set for each of the N network layers. Among the N network layers, the identification information of network layers with the same operator type is the same, and the identification information of network layers with different operator types is different. When generating the first architecture information corresponding to the second network layer, and when the N network layers include other network layers with the same identification information as the second network layer, the first architecture information corresponding to the second network layer is determined to be the first architecture information corresponding to the other network layers, wherein the second network layer is any one of the N network layers.

5. The method according to claim 1, characterized in that, The step of generating target architecture information based on the N first architecture information includes: A negotiation space is constructed based on the N first architecture information, wherein each point in the negotiation space includes one of the candidate architecture information in each of the N first architecture information; The target architecture information is generated based on the negotiation space.

6. The method according to claim 5, characterized in that, The generation of the target architecture information based on the negotiation space includes: Based on the negotiation space, NSGA-II is used for iterative selection to obtain the target individual. The parent population of the first selection in the iterative selection includes multiple points sampled from the negotiation space, each point being an individual from the parent population of the first selection. The parent population of the j-th selection in the iterative selection is the (j-1)-th descendant population obtained from the (j-1)-th selection, where j is an integer greater than 1. The target individual is the individual with the best PPA result among the descendant population obtained from the last selection in the iterative selection. The target architecture information is generated based on the target individual.

7. The method according to claim 6, characterized in that, The generation of the target architecture information based on the target individual includes: The target individual is parsed to obtain N second architecture information that correspond one-to-one with the N network layers. The second architecture information includes: the architecture of the corresponding network layer, and the mapping relationship between the network layer and the corresponding architecture. The target architecture is determined based on the N architectures included in the N second architecture information, wherein each architecture parameter in the target architecture is the maximum value among the corresponding N architecture parameters in the N architectures; Based on the N mapping relationships included in the N second architecture information, a mapping relationship between the N network layers and the target architecture is generated to obtain N target mapping relationships, wherein the target architecture information includes the target architecture and the N target mapping relationships.

8. The method according to any one of claims 1 to 7, characterized in that, The joint exploration space includes an array size subspace, an indicator factor subspace, a cyclic sorting subspace, and a cyclic allocation subspace. The array size subspace includes multiple processing architecture schemes that can be constructed based on the processing unit information. The indicator factor subspace includes multiple factor decomposition schemes for factoring the dimensional information. The cyclic sorting subspace includes multiple sorting schemes for sorting each factor decomposition scheme. The cyclic allocation subspace includes multiple storage architecture schemes corresponding to the processing loops in the factor decomposition schemes. The candidate architecture information includes: a processing architecture scheme in the array size subspace, a factorization scheme in the index factor subspace, a sorting scheme in the cyclic sorting subspace, and a storage architecture scheme in the cyclic allocation subspace.

9. The method according to claim 8, characterized in that, Before constructing the joint exploration space based on the model information and the hardware architecture information, the method further includes: Obtain a constraint file, wherein the constraint file includes at least one of the following constraint information: first constraint information, second constraint information, third constraint information, and fourth constraint information; the first constraint information is used to constrain the number of processing units in the processing architecture scheme, the second constraint information is used to constrain the number of processing loops in the factorization scheme, the third constraint information is used to constrain the mapping relationship between the processing units and the processing loops, and the fourth constraint information is used to constrain the storage limit of each storage level and the data storage type of each storage level in the storage architecture scheme; The construction of the joint exploration space based on the model information and the hardware architecture information includes: An initial exploration space is constructed based on the model information and the hardware architecture information. The initial exploration space includes an initial array size subspace, an initial indicator factor subspace, an initial cyclic sorting subspace, and an initial cyclic allocation subspace. The initial array size subspace includes all processing architecture schemes that can be constructed based on the processing unit information. The initial indicator factor subspace includes all factor decomposition schemes that factorize the dimensional information. The initial cyclic sorting subspace includes all sorting schemes that sort each factor decomposition scheme. The initial cyclic allocation subspace includes all storage architecture schemes corresponding to the processing cycles in the factor decomposition schemes. Based on the constraint file, the initial exploration space is spatially pruned to obtain the constructed joint exploration space, wherein the candidate architecture information matches the constraint file.

10. An apparatus for generating architectural information, characterized in that, include: The acquisition module is used to acquire model information of the network model and hardware architecture information of the inference accelerator. The network model includes N network layers, the model information includes N dimension information corresponding one-to-one with the N network layers, where N is an integer greater than 1, the dimension information includes the dimension parameters and dimension values ​​of the operators in the corresponding network layers, and the hardware architecture information includes processing unit information and storage architecture information. A construction module is used to construct a joint exploration space based on the model information and the hardware architecture information, wherein the joint exploration space includes multiple candidate architecture information, and the candidate architecture information includes: a candidate architecture corresponding to a network layer, and the mapping relationship between the network layer and the corresponding candidate architecture; The generation module is used to generate N first architecture information corresponding one-to-one with the N network layers based on the joint exploration space, wherein the first architecture information includes at least two candidate architecture information; The generation module is further configured to generate target architecture information based on the N first architecture information, wherein the target architecture information includes the architecture corresponding to each of the N network layers in the inference accelerator, and the mapping relationship between the network layer and its corresponding architecture.

11. A first chip, characterized in that, The first chip has a network model deployed thereon, wherein the network model is a network model deployed based on target architecture information, and the target architecture information is generated based on the architecture information generation method according to any one of claims 1 to 9.

12. An accelerator, characterized in that, The accelerator is equipped with a network model, wherein the network model is a network model deployed based on target architecture information, and the target architecture information is generated based on the architecture information generation method described in any one of claims 1 to 9.

13. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and the program or instructions, when executed by the processor, implement the steps of the method for generating architectural information as described in any one of claims 1-9.

14. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method for generating architectural information as described in any one of claims 1-9.

15. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the method for generating architectural information as described in any one of claims 1-9.