Quantization method and device of neural network model, equipment and medium
Patent Information
- Application Number
- CN202410324596.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-03-20
AI Technical Summary
[0003]复杂的神经网络模型具有良好的性能,但需要占用较多的存储空间和计算资源
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.
Smart Images

Figure CN118114735B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning and chip technology, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for quantizing a neural network model. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] Complex neural network models offer excellent performance but require significant storage and computational resources. Model quantization techniques are used to reduce the storage capacity and computational complexity of neural network models, thereby conserving hardware resources used for model inference and improving inference speed.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for quantizing neural network models.
[0006] According to one aspect of this disclosure, a method for quantizing a neural network model is provided, comprising: determining a plurality of initial operators included in an initial neural network model, wherein the initial neural network model is used to perform at least one of the following tasks: speech recognition, speech synthesis, image processing, natural language processing, object detection, data generation, and trend prediction; for each of the plurality of initial operators, determining a target quantization strategy for the initial operator from a plurality of candidate quantization strategies, wherein a first candidate quantization strategy among the plurality of candidate quantization strategies has a higher quantization rate than other candidate quantization strategies among the plurality of candidate quantization strategies, the quantization rate indicating the number of parameters to be quantized in the initial operator, wherein determining the target quantization strategy for the initial operator from the plurality of candidate quantization strategies comprises: determining the initial operator based on the first... The algorithm evaluates the following: First, the computational accuracy of the initial operator after quantization using a candidate quantization strategy; In response to determining that the first computational accuracy satisfies a first preset condition, the first candidate quantization strategy is determined as the target quantization strategy; In response to determining that the first computational accuracy does not satisfy the first preset condition, a second computational accuracy of the initial operator after quantization based on a second candidate quantization strategy is determined; In response to determining that the second computational accuracy satisfies the first preset condition, the second candidate quantization strategy is determined as the target quantization strategy; In response to determining that the second computational accuracy does not satisfy the first preset condition, the target quantization strategy is determined to be non-quantization; Based on the target quantization strategies of the multiple initial operators, multiple target operators corresponding to the multiple initial operators are determined; Based on the multiple target operators, a target neural network model is determined.
[0007] According to one aspect of this disclosure, a quantization apparatus for a neural network model is provided, comprising: a first determining unit configured to determine a plurality of initial operators included in an initial neural network model, wherein the initial neural network model is used to perform at least one of the following tasks: speech recognition, speech synthesis, image processing, natural language processing, object detection, data generation, and trend prediction; and a second determining unit configured to, for each of the plurality of initial operators, determine a target quantization strategy for that initial operator from a plurality of candidate quantization strategies, wherein a first candidate quantization strategy among the plurality of candidate quantization strategies has a higher quantization rate than other candidate quantization strategies among the plurality of candidate quantization strategies, the quantization rate indicating the number of parameters to be quantized in the initial operator, the second determining unit comprising: a first determining subunit configured to determine a first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy; and a second determining subunit. The system is configured to: 1) determine the first candidate quantization strategy as the target quantization strategy in response to determining that the first computational accuracy satisfies the first preset condition; 2) determine the second computational accuracy of the initial operator after quantization based on the second candidate quantization strategy in response to determining that the first computational accuracy does not satisfy the first preset condition; 3) determine the second candidate quantization strategy as the target quantization strategy in response to determining that the second computational accuracy satisfies the first preset condition; and 4) determine the target neural network model based on the target quantization strategies of the multiple initial operators.
[0008] According to one aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the quantization method of the neural network model described above.
[0009] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to execute the quantization method of the neural network model described above.
[0010] According to one aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, is capable of implementing the quantization method of the aforementioned neural network model.
[0011] According to one or more embodiments of this disclosure, the prediction accuracy of a quantized target neural network model can be improved.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which various methods described herein may be implemented, according to exemplary embodiments of the present disclosure;
[0015] Figure 2 A flowchart illustrating a quantization method for a neural network model according to an exemplary embodiment of the present disclosure is shown;
[0016] Figure 3 A schematic diagram illustrating the quantization process of some parameters of a neural network model according to an exemplary embodiment of the present disclosure is shown.
[0017] Figure 4 A structural block diagram of a quantization device for a neural network model according to an exemplary embodiment of the present disclosure is shown;
[0018] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0019] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0020] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0021] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0022] In related technologies, the target quantization strategy is typically selected by iterating through multiple candidate quantization strategies and considering their respective quantization effects and prediction accuracy. This approach requires significant hardware resources and is inefficient. Based on this, the model can be quantized using the determined target quantization strategy, and the quantized neural network model can then be used to perform various deep learning-based tasks, such as speech recognition, speech synthesis, image processing, natural language processing, object detection, and data generation. However, this method uses a single quantization strategy to quantize the entire model, resulting in coarse-grained quantization control. It fails to consider the compatibility between different operators in the model and different quantization strategies, making it difficult to balance quantization effect and prediction accuracy.
[0023] Based on this, this disclosure provides a model quantization method. When there are multiple candidate strategies with different quantization rates, for each operator in the model, a target quantization strategy that meets the computational accuracy requirements is searched in descending order of quantization rate. The target quantization strategy can indicate whether to perform different levels of quantization or not to perform quantization for the operator. Thus, the target operator can be determined based on the target quantization strategy that fits each operator, thereby obtaining a quantization model with higher accuracy.
[0024] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0025] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0026] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of quantization methods for neural network models.
[0027] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0028] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0029] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to send initial neural network models. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0030] Client devices 101, 102, 103, 104, 105, and / or 106 may include various categories of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various categories and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices can run a variety of different applications, such as various Internet-related applications, communication applications (e.g., email applications), short message service (SMS) applications, and can use various communication protocols.
[0031] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0032] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0033] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0034] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0035] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0036] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different categories. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0037] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be different categories of databases, such as key-value stores, object stores, or regular stores supported by a file system.
[0038] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0039] Figure 2 A flowchart of a quantization method 200 for a neural network model according to an exemplary embodiment of the present disclosure is shown. Figure 2 As shown, method 200 includes:
[0040] Step S210: Determine the multiple initial operators included in the initial neural network model, wherein the initial neural network model is used to perform at least one of the following tasks: speech recognition, speech synthesis, image processing, natural language processing, object detection, data generation, and trend prediction;
[0041] Step S220: For each of the plurality of initial operators, determine the target quantization strategy of the initial operator from a plurality of candidate quantization strategies. The first candidate quantization strategy among the plurality of candidate quantization strategies has a higher quantization rate than the other candidate quantization strategies among the plurality of candidate quantization strategies. The quantization rate indicates the number of parameters to be quantized in the initial operator. Step S220 includes:
[0042] Step S221: Determine the first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy;
[0043] Step S222: In response to determining that the first calculation accuracy meets the first preset condition, determine the first candidate quantization strategy as the target quantization strategy;
[0044] Step S223: In response to determining that the first calculation accuracy does not meet the first preset condition, determine the second calculation accuracy of the initial operator after quantization based on the second candidate quantization strategy;
[0045] Step S224: In response to determining that the second calculation accuracy satisfies the first preset condition, determine the second candidate quantization strategy as the target quantization strategy; and
[0046] Step S225: In response to determining that the second calculation accuracy does not meet the first preset condition, determine that the target quantization strategy is not to perform quantization;
[0047] Step S230: Based on the target quantization strategy of the plurality of initial operators, determine the plurality of target operators corresponding to the plurality of initial operators respectively; and
[0048] Step S240: Determine the target neural network model based on the multiple target operators.
[0049] When model quantization needs to be performed based on multiple candidate quantization strategies, by applying the method 200 described above, the computational accuracy of each operator among the multiple operators included in the model can be checked under the current candidate strategy in the order of trying strategies with high quantization rates first (i.e., trying the first candidate quantization strategy first) and then trying strategies with low quantization rates. When a candidate quantization strategy that meets the requirements is found, it is determined as the target quantization strategy that is suitable for that operator, effectively improving the search efficiency of the target quantization strategy. On this basis, the target operator can be determined based on the target quantization strategy suitable for each operator, thereby obtaining a quantization model with high computational accuracy in a simple and efficient manner.
[0050] Quantization strategy refers to changing the storage or computation of parameters in the initial neural network model from a higher bit-value format to a lower bit-value format. For example, parameters stored as 32-bit floating-point numbers can be changed to 8-bit integers, thereby saving storage space in the target neural network model. Similarly, parameters calculated as 32-bit floating-point numbers can be changed to 8-bit integers, thus saving hardware computing resources used in the inference process of the target neural network model and improving the inference speed. Furthermore, the quantization strategy can include quantization rate configuration information to indicate the number of parameters to be quantized; for example, when the quantization rate is 20%, 20% of the parameters in the initial operator can be quantized accordingly.
[0051] It can be seen that the higher the quantization rate of the target quantization strategy, the better the quantization effect of the model, that is, the target neural network model with less hardware resources (such as storage space or computing resources) can be obtained. Based on this, by selecting the target quantization strategy from multiple candidate quantization strategies in descending order of quantization rate, it is equivalent to selecting the target quantization strategy in descending order of quantization effect while ensuring computational accuracy. This allows for a more efficient selection of a target quantization strategy that balances quantization effect and computational accuracy.
[0052] The above process only describes the steps of determining the computational accuracy for the first and second candidate quantization strategies and checking whether the computational accuracy meets the first preset condition. In some examples, the multiple candidate quantization strategies for the initial neural network model can include a larger number of candidate quantization strategies. In this case, the multiple candidate quantization strategies can be arranged in descending order of quantization rate, so that the computational accuracy of each candidate quantization strategy can be determined sequentially in descending order of quantization rate to check whether it meets the first preset condition, until a target quantization strategy that meets the requirements is found, further improving the flexibility of model quantization.
[0053] In some examples, the quantization strategy may also include configuration information indicating the location of the parameters to be quantized. For example, only one or more specific network layers (e.g., activation layers, weight layers) in the initial neural network model may be quantized. As another example, only one or more specific channels (e.g., channels corresponding to specific feature dimensions) in the initial neural network model may be quantized.
[0054] In some examples, the initial neural network model is a large language model trained with a large corpus. Since the activation values of a large language model have a large dynamic range, quantizing the activation layer parameters can significantly impact the model's inference accuracy. Therefore, multiple candidate quantization strategies can include, for example, quantizing all weight parameters and activation layer parameters, quantizing only some weight parameters and activation layer parameters, quantizing only weight parameters, and not quantizing at all. By flexibly using differentiated quantization strategies for different network layers, the flexibility of model quantization can be improved, balancing the model's inference performance and accuracy. In some examples, the quantized weight parameters or activation layer parameters can be determined based on indicators such as the maximum absolute value of the parameter values or the range of absolute values of the parameter values, to ensure the accuracy of the quantized model.
[0055] In some examples, the quantization strategy may also include configuration information indicating the application of different quantization methods to different parameters to be quantized; that is, indicating the mapping relationship between the distribution characteristics of the original values of the parameters to be quantized and the distribution characteristics of the quantized values of the parameters after quantization. For example, the configuration information can be used to indicate the linear distribution of the quantized values, corresponding to a uniform quantization strategy or a non-uniform quantization strategy. As another example, the configuration information can also be used to indicate the mapping relationship between the zeros of the original values and the zeros of the quantized values, corresponding to a symmetric quantization strategy or an asymmetric quantization strategy. Thus, different quantization strategies can be flexibly used according to the needs of the actual application scenario, further improving the quantization effect of the model.
[0056] In some examples, the initial neural network model is a completed model built or trained based on deep learning methods, and can be a neural network model of various structures, such as convolutional neural network models, feedback neural network models, feedforward neural network models, etc. In some examples, the initial neural network model can be used to perform various deep learning-based tasks, such as speech recognition, speech synthesis, image processing, natural language processing, object detection, data generation, and trend prediction.
[0057] In one example, when the initial neural network model is a speech recognition model, the training and testing data corresponding to the initial neural network model are the speech data to be recognized. Using this speech recognition model, the corresponding speech recognition result can be obtained, which may be, for example, the corresponding text information. By using the method 200 described above to compress the initial speech recognition model, a target speech recognition model that occupies less storage space or less computational resources during inference can be obtained, thereby effectively saving hardware resources required for the speech recognition task.
[0058] In one example, when the initial neural network model is a speech synthesis model, the training and testing data corresponding to the initial neural network model are the text information to be synthesized. Using this speech synthesis model, the synthesized speech data corresponding to the text information to be synthesized can be obtained.
[0059] In one example, when the initial neural network model is an image processing model, the training and testing data corresponding to the initial neural network model are image data. Using this image processing model, the corresponding image processing results can be obtained from the input image data. The image processing results may include, for example, image classification results, image segmentation results (e.g., semantic segmentation results or instance segmentation results), image enhancement results (e.g., image with increased resolution, image sharpening, image denoising), etc.
[0060] In one example, when the initial neural network model is a natural language processing model, the training and testing data corresponding to the initial neural network model are natural language text. Using this initial neural network model, the corresponding text processing results of the input natural language text can be obtained. These results may include, for example, response text generated based on the question information contained in the input text, search results based on the input text, and the results of processing natural language paragraphs contained in the input text (e.g., summary information of the paragraph, entity information extracted from the paragraph, or structured information).
[0061] In one example, when the initial neural network model is an object detection model, the training and testing data for the initial neural network model are image or video data. Using this initial neural network model, the results of detecting or tracking target objects contained in the image or video data can be obtained.
[0062] In one example, when the initial neural network model is a data generation model, the training and test data corresponding to the initial neural network model are data generation prompts. Using this data generation model, the result of data generation based on the prompts can be obtained, and the generated result may include, for example, text or an image.
[0063] In one example, when the initial neural network model is a trend prediction model, the training and testing data corresponding to the initial neural network model are historical data or data that can describe the distribution characteristics of historical data. Historical data may include, for example, historical weather data, historical user behavior data, etc. Using this trend prediction model, trend prediction results can be obtained, which may include, for example, weather forecasts, predictions of whether a user will perform a preset action (such as browsing a page or purchasing a product), etc.
[0064] Similar to the previously described method of using a quantized target neural network model to perform speech recognition tasks, in the above embodiments, method 200 can also be used to quantize at least one of the above speech synthesis model, image processing model, natural language processing model, object detection model, data generation model, and trend prediction model, so that the quantized model can be used to perform the corresponding deep learning-based tasks, thereby saving the hardware resources occupied by the above tasks.
[0065] In some examples, the initial neural network model may include multiple initial operators that can be pre-defined manually, such as convolution operators, normalization operators, etc., so that one or more computational processes contained in each initial operator can be simplified based on the same target quantization strategy, thereby improving the efficiency and effectiveness of model quantization.
[0066] According to some embodiments, determining the first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy in step S221 includes: determining a first inference result of the initial neural network model inference based on test data; determining a first quantization operator corresponding to the initial operator based on the first candidate quantization strategy; determining a first quantization neural network model based on the first quantization operator; determining a second inference result of the first quantization neural network model inference based on the test data; and determining the first computational accuracy based on the first inference result and the second inference result. Therefore, test data can be used to determine the computational accuracy of each operator after quantization, and the impact of the quantization process on the operator accuracy can be more accurately evaluated based on the original inference result and the quantized inference result.
[0067] In some examples, the first computational accuracy can be determined by assessing the similarity between the first and second inference results. For instance, this similarity can be determined by calculating parameters such as Euclidean distance, Manhattan distance, cosine similarity, and Pearson correlation coefficient between the respective vector representations of the first and second inference results. In other examples, the first computational accuracy can also be determined by calculating the difference between the first and second inference results.
[0068] In some examples, the first and second inference results can be the end-to-end output of the entire model to more accurately assess the impact of the quantization of the initial operator on the overall performance of the model.
[0069] According to some embodiments, method 200 further includes: recording first input data when the initial operator performs inference based on the test data, wherein determining the second inference result of the first quantized neural network model performing inference based on the test data includes: determining the second inference result by inputting the first input data into the first quantized operator. By recording the input data of the initial operator when the initial neural network model performs inference based on the test data, calculations can be performed directly from the quantized initial operator as the starting point in determining the first calculation accuracy, without repeating the calculations of operators whose calculation order precedes that of the initial operator. This allows for more accurate indication of the impact of operator quantization on overall inference performance using the model inference result while minimizing computational resources.
[0070] In some examples, only a portion of the input data when the initial operator infers based on the test data can be saved, as long as it meets the requirements for evaluating the accuracy of the first calculation, in order to further save hardware resources.
[0071] According to some embodiments, method 200 further includes: recording first output data when the initial operator performs inference based on the test data, wherein the second inference result includes second output data when the first quantization operator performs inference based on the first input data, and wherein determining the first computational accuracy based on the first inference result and the second inference result includes: determining the first computational accuracy based on the first output data and the second output data. Thus, it is possible to evaluate the change in the output result of a specific operator before and after quantization only for that specific operator, without repeatedly performing calculations for other operators besides that specific operator, further saving computational resources.
[0072] According to some embodiments, the first candidate strategy indicates quantization based on a preset formula including a numerical shrinkage factor. Method 200 further includes: in response to determining that the first computational accuracy does not meet the first preset condition, updating the numerical shrinkage factor in the first candidate quantization strategy using a random search algorithm; and determining the third computational accuracy of the initial operator after quantization based on the updated first candidate quantization strategy. Specifically, in step S223, in response to determining that the first computational accuracy does not meet the first preset condition, determining the second computational accuracy of the initial operator after quantization based on the second candidate quantization strategy includes: in response to determining that both the first computational accuracy and the third computational accuracy do not meet the first preset condition, determining the second computational accuracy of the initial operator after quantization based on the second candidate quantization strategy. Therefore, by utilizing a random search algorithm to determine the numerical shrinkage factor of the quantization strategy, and while keeping the candidate quantization strategy unchanged (i.e., the quantization rate unchanged), it is possible to find numerical shrinkage factors that meet the computational accuracy requirements as much as possible, thereby maximizing the utilization rate of candidate quantization strategies with higher quantization rates and improving the model quantization effect.
[0073] In some examples, certain termination conditions can be pre-configured for the random search algorithm, such as limiting the maximum number of searches. When the termination condition is met and the current numerical shrinkage factor still does not meet the computational accuracy requirements, other candidate quantization strategies with lower quantization rates can be tried to ensure the search efficiency of the target quantization strategy.
[0074] Figure 3 A schematic diagram illustrating the quantization process of some parameters of a neural network model according to an exemplary embodiment of the present disclosure is shown. In this example, the quantization of model parameters can be based on a Smooth-Quant strategy. According to this quantization strategy, both the weight parameters and activation layer parameters of the model need to be quantized. To reduce the impact of the quantization process on the model's inference accuracy, the activation layer parameters need to be transferred to the weight parameters through numerical transformation to achieve a smoothing of the activation layer parameters. In this case, it is necessary to determine an appropriate numerical shrinkage factor to achieve the numerical transformation of the activation parameters.
[0075] See Figure 3 As shown, the original parameter matrices X and W are quantized based on the following preset formula to obtain the quantized parameter matrix. and of:
[0076]
[0077] In the formula, s is the smoothing factor, when the initial neural network model includes C i When the parameters of each channel are given, the smoothing factor s corresponding to channel j is... jIt is determined using the following formula:
[0078]
[0079] In the formula, α j is the numerical shrinkage factor corresponding to channel j.
[0080] In some examples, a random search algorithm can be used to search for α in a continuous interval (e.g., a numerical interval of 0-1). j The value of . In this case, the random search step size can be adjusted based on a strategy of gradually decreasing the step size or an adaptive step size strategy, so as to more efficiently approximate the optimal numerical shrinkage factor that maximizes the computational accuracy in this continuous interval.
[0081] In some examples, the search space of a random search algorithm can be multiple discrete values. In this case, the comprehensiveness of the search can be improved by iterating through the computational accuracy corresponding to each discrete value to determine whether there exists a value of a numerical search factor within the search space that allows the computational accuracy to meet a first preset condition.
[0082] According to some embodiments, method 200 further includes: determining the inference accuracy of the target neural network model; in response to determining that the inference accuracy does not meet a second preset condition, determining at least one quantization operator to be updated from the plurality of target operators; for each of the at least one quantization operator to be updated, adjusting the corresponding target quantization strategy of the quantization operator to be updated, wherein the quantization rate of the adjusted target quantization strategy is less than the quantization rate of the target quantization strategy before adjustment; and determining the updated quantization operator to be updated based on the adjusted target quantization strategy. By applying the above steps, after preliminary quantization is completed for multiple initial operators, the end-to-end inference accuracy of the overall model can be further evaluated. When the end-to-end inference accuracy does not meet the requirements, the quantization strategy of the operator can be flexibly adjusted to ensure that the overall accuracy of the model meets the requirements.
[0083] According to some embodiments, determining at least one quantization operator to be updated from the plurality of target operators includes: determining the computational accuracy of each target operator among the plurality of target operators; and determining the at least one quantization operator to be updated based on the computational accuracy of each target operator. Therefore, when end-to-end inference accuracy is insufficient, operators with lower quantization accuracy (i.e., larger quantization error) can be preferentially selected for updating, resulting in a more efficient quantized target neural network model that meets accuracy requirements.
[0084] According to one aspect of this disclosure, a quantization device for a neural network model is also provided. Figure 4A structural block diagram of a quantization device 400 for a neural network model according to an exemplary embodiment of the present disclosure is shown. Figure 4 As shown, the device 400 includes:
[0085] The first determining unit 410 is configured to determine a plurality of initial operators included in the initial neural network model, wherein the initial neural network model is used to perform at least one of the following tasks: speech recognition, speech synthesis, image processing, natural language processing, object detection, and trend prediction.
[0086] The second determining unit 420 is configured to determine a target quantization strategy for each of the plurality of initial operators from a plurality of candidate quantization strategies, wherein a first candidate quantization strategy among the plurality of candidate quantization strategies has a higher quantization rate than the other candidate quantization strategies among the plurality of candidate quantization strategies, and the quantization rate indicates the number of parameters to be quantized in the initial operator. The second determining unit 420 includes:
[0087] The first determining subunit 421 is configured to determine the first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy; and
[0088] The second determining subunit 422 is configured to determine the first candidate quantization strategy as the target quantization strategy in response to determining that the first calculation accuracy meets a first preset condition.
[0089] The first determining subunit 421 is further configured to, in response to determining that the first calculation accuracy does not meet the first preset condition, determine the second calculation accuracy of the initial operator after quantization based on the second candidate quantization strategy.
[0090] The second determining subunit 422 is further configured to determine the second candidate quantization strategy as the target quantization strategy in response to determining that the second calculation accuracy satisfies the first preset condition.
[0091] The second determining subunit 422 is further configured to determine, in response to determining that the target quantization strategy is not to perform quantization, in response to determining that the second calculation accuracy does not meet the first preset condition;
[0092] The third determining unit 430 is configured to determine, based on the target quantization strategy of the plurality of initial operators, a plurality of target operators corresponding to the plurality of initial operators; and
[0093] The fourth determining unit 440 is configured to determine the target neural network model based on the plurality of target operators.
[0094] According to some embodiments, the first determining subunit 421 includes: a first determining module configured to determine a first inference result of the initial neural network model inference based on test data; a second determining module configured to determine a first quantization operator corresponding to the initial operator based on the first candidate quantization strategy; a third determining module configured to determine a first quantized neural network model based on the first quantization operator; a fourth determining module configured to determine a second inference result of the first quantized neural network model inference based on the test data; and a fifth determining module configured to determine the first calculation accuracy based on the first inference result and the second inference result.
[0095] According to some embodiments, the apparatus 400 further includes: a recording unit configured to record first input data when the initial operator performs inference based on the test data, wherein the fourth determining module is configured to: determine the second inference result by inputting the first input data into the first quantization operator.
[0096] According to some embodiments, the recording unit is further configured to record first output data when the initial operator performs inference based on the test data, wherein the second inference result includes second output data when the first quantization operator performs inference based on the first input data, and wherein the fifth determining module is configured to determine the first calculation accuracy based on the first output data and the second output data.
[0097] According to some embodiments, the first candidate strategy indicates quantization based on a preset formula including a numerical shrinkage factor. The device 400 further includes an update unit configured to update the numerical shrinkage factor in the first candidate quantization strategy using a random search algorithm in response to determining that the first calculation accuracy does not meet the first preset condition. The first determining subunit 421 is further configured to determine the third calculation accuracy of the initial operator after quantization based on the updated first candidate quantization strategy, and the first determining subunit 421 is further configured to determine the second calculation accuracy of the initial operator after quantization based on the second candidate quantization strategy in response to determining that both the first calculation accuracy and the third calculation accuracy do not meet the first preset condition.
[0098] According to some embodiments, the apparatus 400 further includes: a fifth determining unit configured to determine the inference accuracy of the target neural network model; a sixth determining unit configured to determine at least one quantization operator to be updated from the plurality of target operators in response to determining that the inference accuracy does not meet a second preset condition; and an adjusting unit configured to adjust the target quantization strategy corresponding to each of the at least one quantization operator to be updated, wherein the quantization rate of the adjusted target quantization strategy is less than the quantization rate of the target quantization strategy before adjustment, wherein the third determining unit 430 is further configured to determine the updated quantization operator to be updated based on the adjusted target quantization strategy.
[0099] According to some embodiments, the sixth determining unit includes: a third determining subunit configured to determine the computational accuracy of each of the plurality of target operators; and a fourth determining subunit configured to determine the at least one quantization operator to be updated based on the computational accuracy of each target operator.
[0100] It should be understood that Figure 4 The operation of each unit of the quantization device 400 in the neural network model shown can be synchronized with... Figure 2 The steps in the quantization method 200 for the described neural network model correspond to each other. Therefore, the operations, features, and advantages described above for method 200 also apply to device 400 and its constituent units. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0101] According to one aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the quantization method of the neural network model described above.
[0102] According to one aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to execute the quantization method of the neural network model described above.
[0103] According to one aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the quantization method of the neural network model described above.
[0104] refer to Figure 5The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0105] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0106] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to device 500. Input unit 506 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, a hard disk and an optical disk. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0107] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the quantization method of a neural network model. For example, in some embodiments, the quantization method of a neural network model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the quantization method of the neural network model described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the quantization method of a neural network model by any other suitable means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0113] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0114] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0115] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited to these embodiments or examples. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A quantization method for a neural network model, comprising: The initial neural network model is determined to include multiple initial operators, wherein the initial neural network model is used to perform at least one of the following tasks: speech recognition, speech synthesis, and image processing; For each of the plurality of initial operators, a target quantization strategy is determined from a plurality of candidate quantization strategies. A first candidate quantization strategy among the plurality of candidate quantization strategies has a higher quantization rate than the other candidate quantization strategies. The quantization rate indicates the number of parameters to be quantized in the initial operator. Determining the target quantization strategy from the plurality of candidate quantization strategies includes: Determine the first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy; In response to determining that the first calculation accuracy meets the first preset condition, the first candidate quantization strategy is determined to be the target quantization strategy; In response to determining that the first calculation accuracy does not meet the first preset condition, the second calculation accuracy of the initial operator after quantization based on the second candidate quantization strategy is determined. In response to determining that the second calculation accuracy satisfies the first preset condition, the second candidate quantization strategy is determined as the target quantization strategy; and In response to determining that the second calculation accuracy does not meet the first preset condition, the target quantization strategy is determined to be not to perform quantization; Based on the target quantization strategy of the plurality of initial operators, determine the plurality of target operators corresponding to the plurality of initial operators respectively; and Based on the aforementioned multiple target operators, the target neural network model is determined.
2. The method as described in claim 1, wherein, Determining the first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy includes: Determine the first inference result of the initial neural network model based on the test data; Based on the first candidate quantization strategy, determine the first quantization operator corresponding to the initial operator; Based on the first quantization operator, a first quantization neural network model is determined; Determine the second inference result of the first quantized neural network model inferring based on the test data; and Based on the first reasoning result and the second reasoning result, the first calculation accuracy is determined.
3. The method of claim 2, further comprising: Record the first input data when the initial operator performs inference based on the test data. Wherein, determining the second inference result of the first quantization neural network model inference based on the test data includes: The second inference result is determined by inputting the first input data into the first quantization operator.
4. The method of claim 3, further comprising: Record the first output data of the initial operator when it performs inference based on the test data. The second inference result includes the second output data of the first quantization operator when performing inference based on the first input data. Furthermore, determining the first calculation accuracy based on the first reasoning result and the second reasoning result includes: The accuracy of the first calculation is determined based on the first output data and the second output data.
5. The method according to any one of claims 1-4, wherein, The first candidate quantization strategy indicates quantization based on a preset formula including a numerical shrinkage factor, and the method further includes: In response to determining that the first calculation accuracy does not meet the first preset condition, the numerical shrinkage factor in the first candidate quantization strategy is updated using a random search algorithm; and Determine the third computational accuracy of the initial operator after quantization based on the updated first candidate quantization strategy. Wherein, the step of determining the second computational accuracy of the initial operator after quantization based on the second candidate quantization strategy in response to determining that the first computational accuracy does not meet the first preset condition includes: In response to determining that neither the first calculation accuracy nor the third calculation accuracy satisfies the first preset condition, the second calculation accuracy of the initial operator after quantization based on the second candidate quantization strategy is determined.
6. The method according to any one of claims 1-4, further comprising: Determine the inference accuracy of the target neural network model; In response to determining that the inference accuracy does not meet the second preset condition, at least one quantization operator to be updated is determined from the plurality of target operators; For each of the at least one quantization operators to be updated, Adjust the target quantization strategy corresponding to the quantization operator to be updated, wherein the quantization rate of the adjusted target quantization strategy is less than the quantization rate of the target quantization strategy before adjustment; as well as Based on the adjusted target quantization strategy, the updated quantization operator to be updated is determined.
7. The method of claim 6, wherein, The step of determining at least one quantization operator to be updated from the plurality of target operators includes: Determine the computational accuracy of each of the plurality of target operators; and Based on the computational accuracy of each target operator, at least one quantization operator to be updated is determined.
8. A quantization device for a neural network model, comprising: The first determining unit is configured to determine a plurality of initial operators included in the initial neural network model, wherein the initial neural network model is used to perform at least one of the following tasks: speech recognition, speech synthesis, and image processing; The second determining unit is configured to, for each of the plurality of initial operators, determine a target quantization strategy for that initial operator from a plurality of candidate quantization strategies, wherein a first candidate quantization strategy among the plurality of candidate quantization strategies has a higher quantization rate than the other candidate quantization strategies among the plurality of candidate quantization strategies, and the quantization rate indicates the number of parameters to be quantized in the initial operator. The second determining unit includes: The first determining subunit is configured to determine the first computational accuracy of the initial operator after quantization based on the first candidate quantization strategy; and The second determining subunit is configured to determine the first candidate quantization strategy as the target quantization strategy in response to determining that the first calculation accuracy meets a first preset condition. The first determining subunit is further configured to, in response to determining that the first computational accuracy does not meet the first preset condition, determine the second computational accuracy of the initial operator after quantization based on the second candidate quantization strategy. The second determining subunit is further configured to determine the second candidate quantization strategy as the target quantization strategy in response to determining that the second calculation accuracy satisfies the first preset condition. The second determining subunit is further configured to determine the target quantization strategy as not performing quantization in response to determining that the second calculation accuracy does not meet the first preset condition; The third determining unit is configured to determine, based on the target quantization strategy of the plurality of initial operators, a plurality of target operators corresponding to the plurality of initial operators; and The fourth determining unit is configured to determine the target neural network model based on the plurality of target operators.
9. The apparatus of claim 8, wherein, The first determining subunit includes: The first determining module is configured to determine a first inference result of the initial neural network model inference based on test data; The second determining module is configured to determine the first quantization operator corresponding to the initial operator based on the first candidate quantization strategy; The third determining module is configured to determine the first quantization neural network model based on the first quantization operator; The fourth determining module is configured to determine a second inference result of the first quantized neural network model inference based on the test data; and The fifth determining module is configured to determine the first calculation accuracy based on the first reasoning result and the second reasoning result.
10. The apparatus of claim 9, further comprising: The recording unit is configured to record the first input data when the initial operator performs inference based on the test data. The fourth determining module is configured as follows: The second inference result is determined by inputting the first input data into the first quantization operator.
11. The apparatus of claim 10, wherein, The recording unit is also configured to record the first output data when the initial operator performs inference based on the test data. The second inference result includes the second output data of the first quantization operator when performing inference based on the first input data. Furthermore, the fifth determining module is configured as follows: The accuracy of the first calculation is determined based on the first output data and the second output data.
12. The apparatus according to any one of claims 8-11, wherein, The first candidate quantization strategy indicates quantization based on a preset formula including a numerical shrinkage factor, and the device further includes: The update unit is configured to update the numerical shrinkage factor in the first candidate quantization strategy using a random search algorithm in response to determining that the first calculation accuracy does not meet the first preset condition. The first determining subunit is further configured to determine the third computational accuracy of the initial operator after quantization based on the updated first candidate quantization strategy. Furthermore, the first determining subunit is also configured as follows: In response to determining that neither the first calculation accuracy nor the third calculation accuracy satisfies the first preset condition, the second calculation accuracy of the initial operator after quantization based on the second candidate quantization strategy is determined.
13. The apparatus of any one of claims 8-11, further comprising: The fifth determining unit is configured to determine the inference accuracy of the target neural network model; The sixth determining unit is configured to determine at least one quantization operator to be updated from the plurality of target operators in response to determining that the inference accuracy does not meet the second preset condition; as well as The adjustment unit is configured to adjust the target quantization strategy corresponding to each of the at least one quantization operators to be updated, wherein the quantization rate of the adjusted target quantization strategy is less than the quantization rate of the target quantization strategy before adjustment. The third determining unit is further configured to determine the updated quantization operator to be updated based on the adjusted target quantization strategy.
14. The apparatus of claim 13, wherein, The sixth determining unit includes: The third determining subunit is configured to determine the computational accuracy of each of the plurality of target operators; and The fourth determining subunit is configured to determine the at least one quantization operator to be updated based on the computational accuracy of each target operator.
15. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
17. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Neural-network-model compression method, system and device and readable storage medium
CN108229681A
Quantization method and device of neural network model and computer storage medium
CN111814955A