Joint search method and device based on network structure, quantization bit width and accelerator architecture

By conducting joint searches based on network structure, quantized bit width and accelerator architecture in the deep learning model, the target neural network, compilation mapping strategy and accelerator parameters are optimized, and the problem of poor deployment and inference speed of deep learning model is solved, and efficient computing and resource use are achieved.

CN120068962APending Publication Date: 2025-05-30TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510141511.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve optimal performance in the deployment and inference speed of deep learning models, especially when optimizing deep learning structures or accelerator architectures alone, it is impossible to effectively combine the specific requirements of deep learning models and the characteristics of the accelerator architecture.

Method used

A joint search method based on network structure, quantized bit width and accelerator architecture is proposed. By building a search space and conducting joint search, the target neural network, compilation mapping strategy and accelerator parameters are optimized, and channel sparse quantization and batch compilation mapping are realized to improve computing efficiency and the use of hardware resources.

Benefits of technology

Through joint search optimization, the computing efficiency and hardware resource use of deep learning models are significantly improved, the consumption of computing resources is reduced, and the performance of the model is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068962A_ABST
    Figure CN120068962A_ABST
Patent Text Reader

Abstract

The invention provides a joint search method and device based on a network structure, a quantization bit width and an accelerator architecture, relates to the technical field of computers, and aims to realize joint search optimization of the network structure, the quantization bit width and the accelerator architecture. The method comprises the following steps: constructing a search space, wherein the search space comprises a first space used for searching operators and quantization bit widths of each layer of network, a second space used for searching accelerator parameters, and a third space used for searching a compilation mapping strategy; performing joint search according to the search space to obtain a target neural network, a target compiling mapping strategy corresponding to each target operator of the target neural network, and a target accelerator parameter; and determining a target loss value according to the target neural network, the target compiling and mapping strategy and the target accelerator parameter, and obtaining a final target neural network, a final target compiling and mapping strategy and a final target accelerator parameter under the condition that the target loss value is smaller than a loss threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a joint search method and apparatus based on a network structure, quantization bit width, and accelerator architecture. Background Art

[0002] Models based on deep learning (DNN) have been widely used in fields such as classification, object detection, and segmentation. However, the parameter scale of these deep models also brings problems, affecting their deployment efficiency and inference speed. To improve the inference speed, hardware-friendly algorithms such as mixed-precision quantization have emerged. These algorithms convert weights and activations of each layer to a lower precision according to their sensitivities. Another method to improve the DNN inference speed involves designing dedicated accelerators, which have many parallel multiply-accumulate (MAC) units, facilitating parallel computing in the output channel dimension of DNNs.

[0003] However, optimizing the DNN structure or accelerator architecture alone will not achieve the best performance. Ideally, the design of the accelerator should consider the specific requirements of the DNN, such as the structure of the operator (including channel depth and convolution kernel size). At the same time, the design of the DNN should consider the characteristics of the accelerator architecture, such as register size or compiler mapping strategy. Therefore, the co-optimization of the DNN structure and accelerator architecture has been widely recognized.

[0004] In related technologies, the focus is on considering the search of a single accelerator or the joint search of a network structure and accelerator architecture. Even when quantization search is added, only the non-challenging 4-bit width precision is explored. Due to parameter coupling and error search problems, using an ultra-low bit width (e.g., 2 bits) in the existing technology will result in unacceptable performance degradation. Summary of the Invention

[0005] In view of the above problems, embodiments of this application provide a joint search method and apparatus based on a network structure, quantization bit width, and accelerator architecture to overcome or at least partially solve the above problems.

[0006] In a first aspect of the embodiments of this application, a joint search method based on a network structure, quantization bit width, and accelerator architecture is disclosed. The method includes: Construct a search space, where the search space includes a first space for searching operators and quantization bit widths of each layer of the network, a second space for searching accelerator parameters, and a third space for searching a compilation mapping strategy, and the compilation mapping strategy is used to map the operator to multiple dimensions, and the multiple dimensions include input channels, output channels, output width, and output height; Perform a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters, where the target operator includes an activation channel quantized according to a target quantization bit width and an unquantized activation channel; Determine a target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameters, and end the joint search when the target loss value is less than a loss threshold to obtain a final target neural network, a final target compilation mapping strategy, and final target accelerator parameters.

[0007] Optionally, the first space includes multiple candidate operators of different network layers and multiple candidate quantization bit widths for each candidate operator, the second space includes multiple candidate accelerator parameters, and the third space includes multiple candidate compilation mapping strategies; Performing a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters includes: Determine the target operator from the multiple candidate operators, and determine the target quantization bit width from the multiple candidate quantization bit widths corresponding to the target operator; Perform channel sparse quantization on the target operator according to the target quantization bit width to obtain a target neural network, where the channel sparse quantization represents quantizing some channels of the target operator; Determine the target accelerator parameters from the multiple candidate accelerator parameters according to the target operator, and determine the target compilation mapping strategy from the multiple candidate compilation mapping strategies.

[0008] Optionally, performing channel sparse quantization on the target operator according to the target quantization bit width to obtain a target neural network includes: Quantize each weight of the target operator according to the target quantization bit width to obtain quantized weights; Quantize each activated target channel of the target operator according to the target quantization bit width to obtain channel-sparse quantized activations, where the target channel is a channel with high importance in the activation; Obtain the target neural network according to the channel-sparse quantized activation and the quantized weight.

[0009] Optionally, the method further includes: Use the scale factor of batch normalization in the target operator as an importance index for activating each channel; According to the importance index of each channel, use the first ratio threshold of channels with high importance as the target channel.

[0010] Optionally, the method further includes: Encoding the multiple candidate operators to obtain multiple operator encoding vectors; Encoding the multiple candidate accelerator parameters to obtain multiple accelerator parameter encoding vectors; Encoding the multiple candidate compilation mapping strategies to obtain multiple compilation mapping strategy encoding vectors; Concatenating each operator encoding vector, each accelerator parameter encoding vector, and each candidate compilation mapping strategy encoding vector to obtain multiple operator-accelerator parameter-compilation mapping pairs, and each operator-accelerator parameter-compilation mapping pair includes: an operator encoding vector, an accelerator parameter encoding vector, and a compilation mapping strategy encoding vector; Determining the target accelerator parameter from the multiple candidate accelerator parameters and determining the target compilation mapping strategy from the multiple candidate compilation mapping strategies according to the target operator, including: Selecting a target operator-accelerator parameter-compilation mapping pair from the multiple operator-accelerator parameter-compilation mapping pairs according to the target operator, where the operator encoding vector in the target operator-accelerator parameter-compilation mapping pair represents the target accelerator parameter, and the accelerator parameter encoding vector in the target operator-accelerator parameter-compilation mapping pair represents the target compilation mapping strategy.

[0011] Optionally, determining a target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameter, including: Determining a network loss value according to the weights of the operators in each layer of the network in the first space, the various target operators of the target neural network, and the target quantization bit width corresponding to each target operator; Determining a hardware metric value according to the accelerator architecture corresponding to the target accelerator parameter, the target operator, and the target quantization bit width corresponding to each target operator, where the hardware metric value includes an energy consumption value, a latency value, and an accelerator area; Obtaining the target loss value according to the network loss value and the hardware metric value.

[0012] Optionally, the final target neural network is obtained in the following manner: When the target loss value is less than a loss threshold, obtaining a final target operator and a final target quantization bit width; Quantizing the weights and activations of the final target operator according to the final target operator to obtain the final target neural network.

[0013] In a second aspect of the embodiments of the present application, an image processing method is disclosed, and the method includes: Obtain an image to be processed; Input the image to be processed into a target neural network, and through an accelerator corresponding to target accelerator parameters, cause the target neural network to process the image to be processed according to a target compilation mapping strategy, so as to obtain a target processing result; Wherein, the target neural network, the target compilation mapping strategy, and the target accelerator parameters are respectively the final target neural network, the final target compilation mapping strategy, and the final target accelerator parameters obtained by the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the first aspect of the embodiments of the present application above.

[0014] In a third aspect of the embodiments of the present application, a joint search device based on a network structure, quantization bit width, and accelerator architecture is disclosed. The device includes: A construction module, configured to construct a search space. The search space includes a first space for searching operators and quantization bit widths of each layer of the network, a second space for searching accelerator parameters, and a third space for searching a compilation mapping strategy. The compilation mapping strategy is used to map the operator to multiple dimensions, and the multiple dimensions include input channels, output channels, output width, and output height; A search module, configured to perform a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters. The target operators include activation channels quantized according to a target quantization bit width and unquantized activation channels; A determination module, configured to determine a target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameters, and end the joint search when the target loss value is less than a loss threshold, so as to obtain a final target neural network, a final target compilation mapping strategy, and a final target accelerator parameters.

[0015] In a fourth aspect of the embodiments of the present application, an electronic device is disclosed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the first aspect of the embodiments of the present application are implemented, or the steps of the image processing method described in the second aspect of the embodiments of the present application are implemented.

[0016] In a fifth aspect of the embodiments of the present application, a computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the first aspect of the embodiments of the present application are implemented, or the steps of the image processing method described in the second aspect of the embodiments of the present application are implemented.

[0017] In a sixth aspect of the embodiments of the present application, a computer program product is disclosed, including a computer program which, when executed by a processor, implements the steps of the joint search method based on network structure, quantization bit width, and accelerator architecture described in the first aspect of the embodiments of the present application, or the steps of the image processing method described in the second aspect of the embodiments of the present application.

[0018] The embodiments of the present application include the following advantages: In the embodiments of the present application, the network structure, ultra-low bit mixed-precision quantization (quantization bit width), and accelerator architecture are innovatively combined for joint search, opening up a new direction in the field of hardware-software joint search. By constructing a search space and performing joint search according to the search space, a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters are obtained; since only partial activation channels of the target operator are quantized during the joint search process, the memory explosion problem during the joint search process is effectively alleviated; and for each joint search, according to the obtained target neural network, target compilation mapping strategy, and target accelerator parameters, a target loss value is determined, so that optimization can be carried out from multiple aspects such as the operation efficiency of the network and the use of hardware resources. While maintaining performance, the consumption of computing resources is significantly reduced. In this way, the joint search optimization based on the network structure, quantization bit width, and accelerator architecture is effectively realized. This method that comprehensively considers quantization accuracy and compilation dimensions not only improves the operation efficiency of the network but also optimizes the use of hardware resources, thereby significantly reducing the consumption of computing resources while maintaining performance. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of the steps of a joint search method based on network structure, quantization bit width, and accelerator architecture provided by the embodiments of the present application; Figure 2 is an overall architecture diagram of a joint search method based on network structure, quantization bit width, and accelerator architecture provided by the embodiments of the present application; Figure 3 is a flowchart of the steps of an image processing method provided by the embodiments of the present application; Figure 4It is a schematic structural diagram of a joint search device based on a network structure, quantization bit width, and accelerator architecture provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0021] To make the above objects, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] Quantization can be divided into two main methods: quantization-aware training (QAT) and post-training quantization (PTQ). Compared with PTQ, QAT usually can obtain higher accuracy. Therefore, only the former method is discussed in the embodiments of the present application. In terms of quantization forms, there are two types: fixed-precision quantization (FPQ) and mixed-precision quantization (MPQ). FPQ assigns a fixed bit width to all layers, while the latter assigns different bit widths to the activations and weights of each layer. More and more hardware supports mixed precision, which further promotes the research of MPQ. Some search methods use reinforcement learning (RL) to train the bit width allocator. To improve the search efficiency, related technologies generally adopt the strategy of weight sharing, integrating all candidate network structures into a shared large super network. The practice of weight sharing enables direct training of the super network and extraction of sub-network structures for evaluation therefrom, which greatly reduces the time required to evaluate the performance of the network.

[0023] To improve the computational performance of CNNs, many dedicated DNN accelerators with fixed bit widths have emerged. These accelerators typically consist of dedicated hardware architectures such as MAC (multiplication) arrays, on-chip caches, and on-chip networks. MPQ has led the development of accelerators with variable bit widths, which support separate width variations for each layer. However, AI accelerators are both laborious and time-consuming, requiring a great deal of hardware expertise. An AI-driven approach is adopted to design accelerators, simplifying the design process by autonomously evaluating multiple design configurations. For example, some methods encode hardware parameters and use an evolution-based framework to design the accelerator architecture; some methods utilize large language models (LLMs) to generate designs for AI accelerators; in addition, some methods adopt reinforcement learning or evolutionary algorithms for hardware-software co-design. These methods require a great deal of training time and have a limited search space. To address this problem, related technologies have adopted differentiable methods for collaborative exploration, effectively exploring a vast search space. However, the joint search of network structure, mixed-precision bit-width allocation, and accelerator architecture remains largely unexplored.

[0024] In summary, most related technologies focus on considering the search for individual accelerators or the joint search of network structure and accelerator architecture. Even when quantization search is added, only the unchallenging 4-bit width precision is explored. Due to parameter coupling and error search problems, the existing work using ultra-low bit widths (e.g., 2 bits) results in unacceptable performance degradation. Therefore, there is currently no solution for the joint optimization of network structure, ultra-low bit mixed-precision quantization, and accelerator architecture.

[0025] To overcome the limitations of related technologies, embodiments of the present application provide a joint search method based on network structure, quantization bit width, and accelerator architecture, taking into account complex ultra-low bit mixed-precision quantization and also adding the search in the compilation dimension. The joint search objective is achieved efficiently and quickly through channel sparse quantization and batch compilation mapping methods.

[0026] According to an embodiment of the present application, there is provided an embodiment of a joint search method based on network structure, quantization bit width, and accelerator architecture. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0027] Refer to Figure 1 shown Figure 1 is a flowchart of the steps of a joint search method based on network structure, quantization bit width, and accelerator architecture provided by an embodiment of the present application. As Figure 1As shown in the figure, the joint search method based on network structure, quantization bit width, and accelerator architecture may include steps S110 to S130: Step S110: Construct a search space, where the search space includes a first space for searching operators and quantization bit widths for each layer of the network, a second space for searching accelerator parameters, and a third space for searching compilation mapping strategies. The compilation mapping strategy is used to map the operators to multiple dimensions, and the multiple dimensions include input channels, output channels, output width, and output height.

[0028] Among them, the first space includes multiple candidate operators for different network layers and multiple candidate quantization bit widths for each candidate operator. For example, if a neural network has 5 layers, the first space includes multiple candidate operators corresponding to each layer of the network, and each operator can be quantized according to the quantization bit width.

[0029] The second space includes multiple candidate accelerator parameters, and the accelerator parameters include the shape and number of processing elements (PEs), the on-chip cache size for storing weights, activations, and outputs, and the interconnection of PEs.

[0030] The third space includes multiple candidate compilation mapping strategies. Compilation mapping is crucial for the inference performance of the neural network model on the accelerator. The accelerator is configured to process one image at a time with a batch size of one. Therefore, each operator of each neural network model needs to be compiled and mapped in four dimensions: input channels, output channels, output width, and output height, that is, the operator is mapped to multiple dimensions through the compilation mapping strategy. To achieve the best performance of model inference on the accelerator, the best compilation mapping strategy for each operator should be determined during the compilation mapping stage.

[0031] Step S120: Perform a joint search according to the search space to obtain a target neural network, the target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters. The target operators include activation channels quantized according to the target quantization bit width and unquantized activation channels.

[0032] Among them, performing a joint search according to the search space means: searching for the operators of the target neural network and the quantization bit width of each operator, the target compilation mapping strategy corresponding to each target operator of the target neural network, and the target accelerator parameters according to the search space.

[0033] Perform multiple joint searches based on the search space. Each joint search obtains a target neural network, the target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters. Only partial activation channels of the target operators are quantized during the joint search process, thus effectively alleviating the memory explosion problem during the joint search process.

[0034] Step S130: Determine a target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameters. When the target loss value is less than a loss threshold, end the joint search to obtain the final target neural network, the final target compilation mapping strategy, and the final target accelerator parameters.

[0035] The loss threshold is determined according to the operation efficiency of the network and the usage of hardware resources. When the target loss value is less than the loss threshold, it indicates that the target neural network with the optimal performance, the optimal compilation mapping strategy, and the optimal accelerator parameters (i.e., the hardware architecture) have been searched.

[0036] Moreover, the target loss value is determined based on the target neural network, the target compilation mapping strategy, and the target accelerator parameters. Therefore, it is optimized from multiple aspects such as the operation efficiency of the network and the usage of hardware resources, and while maintaining the performance, it significantly reduces the consumption of computing resources.

[0037] Adopting the technical solution of the embodiment of the present application, by constructing a search space and performing a joint search according to the search space, a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters are obtained. Since only partial activation channels of the target operator are quantized during the joint search process, the memory explosion problem during the joint search process is effectively alleviated. Moreover, for each joint search, a target loss value is determined according to the obtained target neural network, target compilation mapping strategy, and target accelerator parameters. Therefore, it can be optimized from multiple aspects such as the operation efficiency of the network and the usage of hardware resources, and while maintaining the performance, it significantly reduces the consumption of computing resources. In this way, the joint search optimization based on the network structure, quantization bit width, and accelerator architecture is effectively realized. This method that comprehensively considers the quantization accuracy and compilation dimension not only improves the operation efficiency of the network but also optimizes the usage of hardware resources, thereby significantly reducing the consumption of computing resources while maintaining the performance.

[0038] Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a joint search method based on a network structure, quantization bit width, and accelerator architecture. In this method, the first space includes multiple candidate operators of different network layers and multiple candidate quantization bit widths of each candidate operator, the second space includes multiple candidate accelerator parameters, and the third space includes multiple candidate compilation mapping strategies.

[0039] Further, the step "perform a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters" in the above step S120 specifically includes sub-steps S120-1 to step S120-3: Step S120-1: Determine the target operator from the multiple candidate operators, and determine the target quantization bit width from the multiple candidate quantization bit widths corresponding to the target operator.

[0040] Specifically, based on the differentiable network and quantization bit width joint search framework, search among the multiple candidate operators included in the first space and the multiple candidate quantization bit widths corresponding to the operators to obtain the target operator and the target quantization bit width corresponding to the target operator.

[0041] Step S120-2: Perform channel sparse quantization on the target operator according to the target quantization bit width to obtain a target neural network, where the channel sparse quantization represents quantizing some channels of the target operator.

[0042] Specifically, in order to alleviate the memory explosion problem during the joint search process, perform channel sparse quantization on the target operator according to the target quantization bit width, that is, only quantize some channels with high importance, and the remaining channels remain unchanged, thereby obtaining the target neural network.

[0043] Exemplarily, the target neural network can be expressed as:

[0044] Among them, represents the layer index in the target neural network, that is, represents the (l + 1)-th layer network; represents the number of candidate operators in each layer, represents the number of candidate quantization bit widths for each operator; represents the total sum of the weights quantized at different precisions; represents the total sum of the activations quantized at different precisions; represents the architecture parameter of the weight of the k-th quantization bit width of the i-th operator; represents the architecture parameter of the activation of the k-th quantization bit width of the i-th operator; represents the quantization function.

[0045] In an alternative embodiment, the step of "performing channel sparse quantization on the target operator according to the target quantization bit width to obtain a target neural network" in the above step S120-2 specifically includes the following steps: Step A1: Quantize each weight of the target operator according to the target quantization bit width to obtain quantized weights.

[0046] Step A2: Quantize the target channels of each activation of the target operator according to the target quantization bit width to obtain channel-sparse-quantized activations, where the target channels are the channels with high importance in the activations.

[0047] Step A3: Obtain the target neural network according to the activation of the channel sparse quantization and the quantized weights.

[0048] In the embodiments of the present application, considering that in the joint search process, the memory requirement of weight quantization is negligible, the memory cost bottleneck in the network and quantization bit-width joint search framework can be mainly attributed to the quantization of activations. To alleviate this problem in the joint search process, the embodiments of the present application perform channel sparse quantization on the activations of the target operator, that is, only quantize the most important channels in the activations, and keep the remaining channels unchanged.

[0049] Among them, the activation of the channel sparse quantization includes the channels quantized by the target quantization bit-width and the channels not quantized.

[0050] Exemplarily, the activation of the channel sparse quantization can be expressed as:

[0051] Among them, represents the index of the channels to be quantized by the target operator of the l-th layer network (i.e., the target channels), represents the concatenation operation; represents according to selected all channels of. In this way, only a few channels are quantized in the joint search stage, while other channels are kept unquantized, which significantly reduces the demand for GPU memory.

[0052] In an alternative embodiment, the method for determining the target channels is: taking the scale factor of batch normalization in the target operator as the importance index of each activation channel; according to the importance index of each channel, taking the channels with the top proportion threshold of high importance as the target channels.

[0053] In the embodiments of the present application, in order to obtain better search results, it is necessary to select the most important channels (target channels) from each activation for quantization, and take the scale factor of batch normalization (BN) in the target operator as the importance index of each channel.

[0054] Among them, the scale factor can be expressed as:

[0055] Among them, represents the input of the BN layer, represents the output of the BN layer, represents the average value of the input activations in batch B, represents the standard deviation of the input activations in batch B; the trainable parameters γ and β represent the scaling factor and the offset factor respectively.

[0056] Therefore, in the joint search stage, the embodiments of the present application only quantize the channels of the most important first proportional threshold (e.g., K%) in each activation. Specifically, the target channels can be expressed as:

[0057] wherein, represents the importance index of each channel, is the scaling factor, and N is the output channels of the th layer.

[0058] In the joint search stage, the scaling factor can be trained to dynamically adjust the importance index of each channel. In this way, channel sparse quantization of the target operator is performed by this method, reducing the utilization rate of the GPU.

[0059] Step S120-3: Determine the target accelerator parameter from the multiple candidate accelerator parameters according to the target operator, and determine the target compilation mapping strategy from the multiple candidate compilation mapping strategies.

[0060] In the embodiments of the present application, a search is performed according to the multiple candidate accelerator parameters included in the second space and the multiple candidate compilation mapping strategies included in the third space for the target operator, to obtain the target accelerator parameter and the target compilation mapping strategy.

[0061] Considering that it is very time-consuming to find the best compilation mapping strategy for each target operator and it is not friendly to the end-to-end joint search, in order to efficiently find the best compilation mapping strategy, the embodiments of the present application propose a batch compilation mapping method. By encoding the multiple candidate accelerator parameters included in the second space into different vectors, and encoding the multiple candidate compilation mapping strategies included in the third space into different vectors, the best compilation mapping strategy is determined by the encoded vectors simultaneously, so as to reduce the time overhead.

[0062] In an optional embodiment, determining the target accelerator parameter from the multiple candidate accelerator parameters according to the target operator, and determining the target compilation mapping strategy from the multiple candidate compilation mapping strategies is specifically implemented according to the methods of the following steps B1 to step B5: Step B1: Encode the multiple candidate operators to obtain multiple operator encoding vectors.

[0063] Step B2: Encode the multiple candidate accelerator parameters to obtain multiple accelerator parameter encoding vectors.

[0064] Step B3: Encode the multiple candidate compilation mapping strategies to obtain multiple compilation mapping strategy encoding vectors.

[0065] Among them, encoding the multiple candidate compilation mapping policies means: encoding the compilation mapping policies of each candidate operator in four dimensions, namely the input channels, output channels, output width, and output height.

[0066] Step B4: Concatenate each operator encoding vector, each accelerator parameter encoding vector, and each candidate compilation mapping policy encoding vector to obtain multiple operator-accelerator parameter-compilation mapping pairs. Each operator-accelerator parameter-compilation mapping pair includes: an operator encoding vector, an accelerator parameter encoding vector, and a compilation mapping policy encoding vector.

[0067] Exemplarily, the operator-accelerator parameter-compilation mapping pair can be represented as: <operator encoding vector, accelerator parameter encoding vector, compilation mapping policy encoding vector>.

[0068] Step B5: According to the target operator, select the target operator-accelerator parameter-compilation mapping pair from the multiple operator-accelerator parameter-compilation mapping pairs. The operator encoding vector in the target operator-accelerator parameter-compilation mapping pair represents the target accelerator parameter, and the accelerator parameter encoding vector in the target operator-accelerator parameter-compilation mapping pair represents the target compilation mapping policy.

[0069] Specifically, according to the target operator, selecting the target operator-accelerator parameter-compilation mapping pair from the multiple operator-accelerator parameter-compilation mapping pairs includes: taking the multiple operator-accelerator parameter-compilation mapping pairs as different batches and sending them into the power consumption and latency estimator to simultaneously identify the best compilation mapping policy (i.e., the target compilation mapping policy) of each target operator and the best accelerator architecture parameters (i.e., the target accelerator parameters).

[0070] Adopting the technical solution of the embodiment of the present application to perform joint search on the search space. By combining the channel sparse quantization technology, only partial activation channels of the target operator are quantized during the joint search process, effectively alleviating the memory explosion problem during the joint search process; and, through the batch compilation mapping technology, all compilation mapping policies in the search space are encoded as different vectors, and then the best compilation mapping policy can be determined simultaneously, thus greatly reducing the time overhead. By comprehensively considering the quantization accuracy and compilation dimensions, not only the operation efficiency of the network is improved, but also the use of hardware resources is optimized, thereby significantly reducing the consumption of computing resources while maintaining performance.

[0071] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a joint search method based on network structure, quantization bit width, and accelerator architecture. In this method, the step of "determining the target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameters" in step S130 specifically includes sub-steps S130-1 to step S130-3: Step S130-1: Determine the network loss value according to the weights of the operators of each layer of the network in the first space, each target operator of the target neural network, and the target quantization bit width corresponding to each target operator.

[0072] Step S130-2: Determine the hardware metric values according to the accelerator architecture corresponding to the target accelerator parameters, the target operator, and the target quantization bit width corresponding to each target operator. The hardware metric values include energy consumption value, latency value, and accelerator area.

[0073] Step S130-3: Obtain the target loss value according to the network loss value and the hardware metric values.

[0074] Exemplarily, the target loss value can be expressed as:

[0075] where α represents the operator architecture parameter, i.e., the target operator, N(α) represents the target neural network selected based on α; β represents the quantization bit width architecture parameter, Q(β) represents the quantization bit width selected for each target operator according to β; γ represents the hardware accelerator configuration, i.e., the target accelerator parameter; H(γ) depicts the accelerator architecture based on γ; w represents the weights of the NAS super network, i.e., the weights of the operators of each layer of the network in the first space; represents the network loss value (cross-entropy loss value), represents the hardware metric values (energy consumption, latency, and area); λ is a factor used to balance and .

[0076] In this way, the target loss value is determined based on the target neural network, the target compilation mapping strategy, and the target accelerator parameters, so as to optimize from multiple aspects such as the operation efficiency of the network and the use of hardware resources, and while maintaining performance, significantly reduce the consumption of computing resources.

[0077] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a joint search method based on network structure, quantization bit width, and accelerator architecture. In this method, the final target neural network is obtained according to the following steps C1 to step C2: Step C1: When the target loss value is less than the loss threshold, obtain the final target operator and the final target quantization bit width.

[0078] Step C2: Quantize the weights and activations of the final target operator according to the final target operator to obtain the final target neural network.

[0079] In the embodiments of the present application, during the joint search process, only partial activation channels of the target operator are quantized to alleviate the memory explosion problem during the joint search process. After the joint search is completed, the weights and activations of the final target operator are quantized according to the final target operator, thereby obtaining a quantized target neural network (i.e., the final target neural network).

[0080] Exemplarily, Figure 2 is the overall architecture diagram of a joint search method based on network structure, quantization bit width, and accelerator architecture provided by the embodiments of the present application. This method constructs a search space, which includes a first space for searching operators and quantization bit widths of each layer of the network, a second space for searching accelerator parameters, and a third space for searching compilation mapping strategies; and performs joint search according to the search space to obtain a target neural network, the target compilation mapping strategies corresponding to the respective target operators of the target neural network, and target accelerator parameters. Since only partial activation channels of the target operator are quantized during the joint search process, the memory explosion problem during the joint search process is effectively alleviated; and for each joint search, according to the obtained target neural network, target compilation mapping strategy, and target accelerator parameters, the target loss value is determined, so that optimization can be performed from multiple aspects such as the operation efficiency of the network and the use of hardware resources, and while maintaining performance, the consumption of computing resources is significantly reduced. In this way, the joint search optimization based on network structure, quantization bit width, and accelerator architecture is effectively realized. This method that comprehensively considers quantization accuracy and compilation dimensions not only improves the operation efficiency of the network but also optimizes the use of hardware resources, thereby significantly reducing the consumption of computing resources while maintaining performance.

[0081] The embodiments of the present application also provide an image processing method. Refer to Figure 3 as shown, Figure 3 is the step flowchart of an image processing method provided by the embodiments of the present application. This image processing method may include steps S310 to S320: Step S310: Obtain the image to be processed.

[0082] Step S320: Input the image to be processed into the target neural network, and through the accelerator corresponding to the target accelerator parameters, enable the target neural network to process the image to be processed according to the target compilation mapping strategy to obtain a target processing result; Among them, the target neural network, the target compilation mapping strategy, and the target accelerator parameters are respectively the final target neural network, the final target compilation mapping strategy, and the final target accelerator parameters obtained by the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the embodiments of the present application.

[0083] In the embodiments of the present application, since the target neural network, the target accelerator parameters, and the target compilation mapping strategy are optimized based on joint search, that is, optimized from multiple aspects such as the operation efficiency of the network and the use of hardware resources, the target neural network has good performance and consumes little computing resources.

[0084] The embodiments of the present application also provide a joint search device based on network structure, quantization bit width, and accelerator architecture. Refer to Figure 4 as shown Figure 4 is a schematic structural diagram of a joint search device based on network structure, quantization bit width, and accelerator architecture provided by the embodiments of the present application. The device includes: A construction module 410, configured to construct a search space, where the search space includes a first space for searching operators and quantization bit widths of each layer of the network, a second space for searching accelerator parameters, and a third space for searching a compilation mapping strategy, and the compilation mapping strategy is used to map the operator to multiple dimensions, and the multiple dimensions include input channels, output channels, output width, and output height; A search module 420, configured to perform joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and target accelerator parameters, where the target operator includes an activation channel quantized according to a target quantization bit width and an unquantized activation channel; A determination module 430, configured to determine a target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameters, and end the joint search to obtain a final target neural network, a final target compilation mapping strategy, and a final target accelerator parameter when the target loss value is less than a loss threshold.

[0085] In an alternative embodiment, the first space includes multiple candidate operators of different network layers and multiple candidate quantization bit widths of each candidate operator, the second space includes multiple candidate accelerator parameters, and the third space includes multiple candidate compilation mapping strategies; the search module includes: A first determination module, configured to determine the target operator from the multiple candidate operators and determine the target quantization bit width from the multiple candidate quantization bit widths corresponding to the target operator; A sparse quantization module, configured to perform channel sparse quantization on the target operator according to the target quantization bit width to obtain a target neural network, where the channel sparse quantization represents quantizing some channels of the target operator; A second determination module, configured to determine the target accelerator parameter from the multiple candidate accelerator parameters according to the target operator, and determine the target compilation mapping strategy from the multiple candidate compilation mapping strategies.

[0086] In an optional embodiment, the sparse quantization module is specifically configured to: quantize each weight of the target operator according to the target quantization bit width to obtain quantized weights; quantize each target channel of the activations of the target operator according to the target quantization bit width to obtain channel-sparse-quantized activations, where the target channels are the channels with high importance in the activations; and obtain the target neural network according to the channel-sparse-quantized activations and the quantized weights.

[0087] In an optional embodiment, the apparatus further includes: A target channel determination module, configured to use the scale factor of batch normalization in the target operator as an importance index for activating each channel; and use the channels with the top ratio threshold of high importance as the target channels according to the importance indexes of each channel.

[0088] In an optional embodiment, the apparatus further includes: A first encoding module, configured to encode the multiple candidate operators to obtain multiple operator encoding vectors; A second encoding module, configured to encode the multiple candidate accelerator parameters to obtain multiple accelerator parameter encoding vectors; A third encoding module, configured to encode the multiple candidate compilation mapping strategies to obtain multiple compilation mapping strategy encoding vectors; A vector splicing module, configured to splice each operator encoding vector, each accelerator parameter encoding vector, and each candidate compilation mapping strategy encoding vector to obtain multiple operator-accelerator parameter-compilation mapping pairs, where each operator-accelerator parameter-compilation mapping pair includes: an operator encoding vector, an accelerator parameter encoding vector, and a compilation mapping strategy encoding vector; The second determination module is specifically configured to: select a target operator-accelerator parameter-compilation mapping pair from the multiple operator-accelerator parameter-compilation mapping pairs according to the target operator, where the operator encoding vector in the target operator-accelerator parameter-compilation mapping pair represents the target accelerator parameter, and the accelerator parameter encoding vector in the target operator-accelerator parameter-compilation mapping pair represents the target compilation mapping strategy.

[0089] In an alternative embodiment, the determination module includes: A loss determination module, configured to determine a network loss value according to the weights of the operators of each layer of the network in the first space, each target operator of the target neural network, and the target quantization bit width corresponding to each target operator; determine a hardware metric value according to the accelerator architecture corresponding to the target accelerator parameter, the target operator, and the target quantization bit width corresponding to each target operator, where the hardware metric value includes an energy consumption value, a latency value, and an accelerator area; and obtain the target loss value according to the network loss value and the hardware metric value.

[0090] In an alternative embodiment, the determination module includes: A network determination module, configured to obtain a final target operator and a final target quantization bit width when the target loss value is less than a loss threshold; and quantize the weights and activations of the final target operator according to the final target operator to obtain the final target neural network.

[0091] An embodiment of the present application further provides an electronic device. Refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown, the electronic device 500 includes a memory 510 and a processor 520. The memory 510 is communicatively connected to the processor 520 via a bus. A computer program is stored in the memory 510, and the computer program can run on the processor 520 to implement the steps of the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the foregoing embodiments, or the steps of the image processing method described in the foregoing embodiments.

[0092] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the foregoing embodiments, or the steps of the image processing method described in the foregoing embodiments are implemented.

[0093] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the joint search method based on the network structure, quantization bit width, and accelerator architecture described in the foregoing embodiments, or the steps of the image processing method described in the foregoing embodiments are implemented.

[0094] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0095] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods and apparatuses according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0098] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0099] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising said element.

[0100] The above has introduced in detail a joint search method and device based on a network structure, quantization bit width and accelerator architecture provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A joint search method based on network structure, quantization bit width and accelerator architecture, characterized in that: The method comprises: Constructing a search space, the search space comprising a first space for searching operators and quantization bit widths of each layer of the network, a second space for searching accelerator parameters, and a third space for searching a compilation mapping strategy, the compilation mapping strategy being used to map the operator to multiple dimensions, the multiple dimensions comprising an input channel, an output channel, an output width, and an output height; Performing a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and a target accelerator parameter, wherein the target operator includes an activation channel quantized according to a target quantization bit width and an unquantized activation channel; According to the target neural network, the target compilation mapping strategy and the target accelerator parameters, a target loss value is determined, and when the target loss value is less than a loss threshold, the joint search is terminated to obtain a final target neural network, a final target compilation mapping strategy, and a final target accelerator parameters.

2. The method according to claim 1, characterized in that The first space includes multiple candidate operators of different network layers and multiple candidate quantization bit widths of each candidate operator, the second space includes multiple candidate accelerator parameters, and the third space includes multiple candidate compilation mapping strategies; Performing a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and a target accelerator parameter, including: Determine the target operator from the multiple candidate operators, and determine the target quantization bit width from the multiple candidate quantization bit widths corresponding to the target operator; Performing channel sparse quantization on the target operator according to the target quantization bit width to obtain a target neural network, wherein the channel sparse quantization represents quantization of part of the channels of the target operator; According to the target operator, the target accelerator parameter is determined from the multiple candidate accelerator parameters, and the target compile mapping strategy is determined from the multiple candidate compile mapping strategies.

3. The method according to claim 2, characterized in that The target operator is subjected to channel sparse quantization according to the target quantization bit width to obtain a target neural network, including: quantizing each weight of the target operator according to the target quantization bit width to obtain a quantized weight; quantizing each activated target channel of the target operator according to the target quantization bit width to obtain channel sparse quantized activation, wherein the target channel is a channel with high importance in the activation; The target neural network is obtained according to the channel sparse quantized activation and the quantized weight.

4. The method according to claim 3, characterized in that The method further comprises: Using the scaling factor of batch normalization in the target operator as an importance indicator for activating each channel; According to the importance index of each channel, the channel with a high importance before the ratio threshold is used as the target channel.

5. The method according to claim 2, characterized in that: The method further comprises: Encoding the multiple candidate operators to obtain multiple operator encoding vectors; Encoding the plurality of candidate accelerator parameters to obtain a plurality of accelerator parameter encoding vectors; Encoding the multiple candidate compilation mapping strategies to obtain multiple compilation mapping strategy encoding vectors; Concatenate each operator encoding vector, each accelerator parameter encoding vector, and each candidate compilation mapping strategy encoding vector to obtain multiple operator-accelerator parameter-compilation mapping pairs, each operator-accelerator parameter-compilation mapping pair including: an operator encoding vector, an accelerator parameter encoding vector, and a compilation mapping strategy encoding vector; According to the target operator, determining the target accelerator parameter from the multiple candidate accelerator parameters, and determining the target compile mapping strategy from the multiple candidate compile mapping strategies, including: According to the target operator, a target operator-accelerator parameter-compile mapping pair is selected from the multiple operator-accelerator parameter-compile mapping pairs, the operator encoding vector in the target operator-accelerator parameter-compile mapping pair represents the target accelerator parameter, and the accelerator parameter encoding vector in the target operator-accelerator parameter-compile mapping pair represents the target compile mapping strategy.

6. The method according to any one of claims 1 to 5, characterized in that: Determining a target loss value according to the target neural network, the target compilation mapping strategy, and the target accelerator parameters includes: Determine a network loss value according to the weights of operators of each layer of the network in the first space, each target operator of the target neural network, and a target quantization bit width corresponding to each target operator; Determine a hardware indicator value according to an accelerator architecture corresponding to the target accelerator parameter, the target operator, and a target quantization bit width corresponding to each target operator, wherein the hardware indicator value includes an energy consumption value, a delay value, and an accelerator area; The target loss value is obtained according to the network loss value and the hardware indicator value.

7. The method according to any one of claims 1 to 5, characterized in that: The final target neural network is obtained in the following way: When the target loss value is less than the loss threshold, a final target operator and a final target quantization bit width are obtained; According to the final target operator, the weight and activation of the final target operator are quantized to obtain the final target neural network.

8. An image processing method, characterized in that: The method comprises: Get the image to be processed; Inputting the image to be processed into a target neural network, and using an accelerator corresponding to a target accelerator parameter to enable the target neural network to process the image to be processed according to a target compilation mapping strategy, to obtain a target processing result; Among them, the target neural network, the target compilation mapping strategy, and the target accelerator parameters are respectively the final target neural network, the final target compilation mapping strategy, and the final target accelerator parameters obtained by the joint search method based on the network structure, quantization bit width and accelerator architecture as described in any one of claims 1 to 7 above.

9. A joint search device based on network structure, quantization bit width and accelerator architecture, characterized in that: The device comprises: A construction module is used to construct a search space, wherein the search space includes a first space for searching operators and quantization bit widths of each layer of the network, a second space for searching accelerator parameters, and a third space for searching a compilation mapping strategy, wherein the compilation mapping strategy is used to map the operator to multiple dimensions, wherein the multiple dimensions include input channels, output channels, output widths, and output heights; A search module, configured to perform a joint search according to the search space to obtain a target neural network, a target compilation mapping strategy corresponding to each target operator of the target neural network, and a target accelerator parameter, wherein the target operator includes an activation channel quantized according to a target quantization bit width and an unquantized activation channel; A determination module is used to determine a target loss value based on the target neural network, the target compilation mapping strategy and the target accelerator parameters, and to end the joint search when the target loss value is less than a loss threshold, so as to obtain a final target neural network, a final target compilation mapping strategy, and a final target accelerator parameters.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the joint search method based on network structure, quantization bit width and accelerator architecture described in any one of claims 1-7 are implemented, or the steps of the image processing method described in claim 8 are implemented.