Artificial intelligence model hybrid quantification method and device based on multi-objective optimization
By combining transfer learning, evolutionary algorithms and Bayesian neural networks with multi-objective optimization methods, the optimization problems of precision and size during model quantization are solved, the optimal balance of model accuracy and size is achieved, and the multi-objective global optimal solution set is provided.
Patent Information
- Application Number
- CN202510249719.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
AI Technical Summary
In application scenarios with insufficient data, it is difficult for the prior art to optimize the accuracy and size of the artificial intelligence model at the same time, especially the accuracy loss that needs to be considered when quantizing the model, and the prior art cannot effectively solve the problem of multi-objective optimization.
Using a hybrid quantization method of artificial intelligence model based on multi-objective optimization, combining transfer learning, evolutionary algorithms and Bayesian neural networks, the optimal mixed quantization samples that achieve the optimal balance of model accuracy and size are screened through parallel quantization evaluation and the combination of Bayesian neural network alternative model and evolutionary algorithms.
It realizes the optimal balance between model accuracy and size in scenarios with insufficient data, provides a globally optimal solution set for multi-objectives, and improves the overall performance of the model.
Smart Images

Figure CN120146155A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an artificial intelligence model hybrid quantization method and device based on multi-objective optimization. Background Art
[0002] Model pruning and model quantization are common means for sparsifying models. Model pruning can remove unimportant redundant parameters in the model to obtain a compact model structure. In contrast, model quantization uses low-precision data types (such as INT8, FP16, and the video memory space occupied by FP16 data is half of that of FP32, and INT8 is even lower) to represent the activations and weights of the model, thereby reducing the computational complexity of the model and the storage cost.
[0003] However, when performing model quantization, it is necessary to consider the accuracy loss caused by sparsification, especially in application scenarios with insufficient data. When the data in the application scenario is insufficient, such as in fields like special tumor detection, rare plant and animal detection, etc., it is very difficult to meet the model accuracy when the data is severely insufficient, and it is even more difficult to simultaneously optimize the model accuracy and the model size. The model accuracy loss becomes a factor that must be considered during quantization. Currently, the mainstream quantization is to convert the data types of all weights and activation functions into low-precision types, which is difficult to meet the high-precision requirements, especially when the data is insufficient. To prevent the loss after quantization, although the mixed-precision model is a good idea, which weights should be low-precision and which weights should be high-precision is a non-convex optimization problem, and the prior art is difficult to give the optimal result. In addition, the existing model hybrid quantization problems are all single-objective optimizations, not multi-objective optimizations (for example, simultaneously optimizing the model accuracy and size, the higher the accuracy, the better, and the smaller the size, the better), and the prior art cannot optimize the optimal solution set of multiple objectives for users to choose. Summary of the Invention
[0004] In view of the above defects or deficiencies in the prior art, the present invention provides an artificial intelligence model hybrid quantization method and device based on multi-objective optimization. The present invention is applied to the field of artificial intelligence model quantization with insufficient data, and combines multi-objective optimization algorithms assisted by transfer learning, evolutionary algorithms, and Bayesian neural networks to improve the model accuracy while taking into account the model size, and can achieve the global optimum of multiple objectives.
[0005] In one aspect of the present invention, an artificial intelligence model hybrid quantization method based on multi-objective optimization is provided, including the following steps: Randomly generate a number of hybrid quantization samples that are super-average for the target model, and the hybrid quantization samples are mixed vector representations of different precision data types of model weights or activation functions; Performing parallel hybrid quantization on the target model using the hybrid quantization samples to obtain a hybrid quantization result representing the accuracy and size of the hybrid quantization model, and forming a data pair with each hybrid quantization sample and the corresponding hybrid quantization result, and all the data pairs form a hybrid quantization dataset; Using an evolutionary algorithm and a Bayesian neural network model to iteratively screen out a set of optimal hybrid quantization samples from the hybrid quantization dataset; Screening out the best hybrid quantization sample that achieves the best balance between model accuracy and size from the set of optimal hybrid quantization samples, and quantizing the target model with the best hybrid quantization sample to obtain the best quantization model, and fine-tuning the best quantization model using the target domain dataset to obtain the final quantization model.
[0006] On the other hand, the present invention also provides a hybrid quantization device for an artificial intelligence model based on multi-objective optimization, including: A first module for randomly generating a number of hybrid quantization samples that are super-average for the target model, where the hybrid quantization samples are hybrid vector representations of different precision data types of model weights or activation functions; A second module for performing parallel hybrid quantization on the target model using the hybrid quantization samples to obtain a hybrid quantization result representing the accuracy and size of the hybrid quantization model, and forming a data pair with each hybrid quantization sample and the corresponding hybrid quantization result, and all the data pairs form a hybrid quantization dataset; A third module for using an evolutionary algorithm and a Bayesian neural network model to iteratively screen out a set of optimal hybrid quantization samples from the hybrid quantization dataset; A fourth module for screening out the best hybrid quantization sample that achieves the best balance between model accuracy and size from the set of optimal hybrid quantization samples, and quantizing the target model with the best hybrid quantization sample to obtain the best quantization model, and fine-tuning the best quantization model using the target domain dataset to obtain the final quantization model.
[0007] The hybrid quantization method and device for an artificial intelligence model based on multi-objective optimization provided by the present invention have the following beneficial effects: To address the field of artificial intelligence model quantization with insufficient data, the present invention uses a Bayesian neural network to replace the model-assisted multi-objective optimization algorithm for selection during model quantization, thus ensuring global optimality. Specifically, the present invention accelerates the convergence of the design space through the combination of parallel quantization evaluation, Bayesian neural network surrogate model, and evolutionary algorithm. The uncertainty of the Bayesian neural network can balance the exploration and development of the design space. Finally, by weighing the model accuracy and model size, the optimal combination solution that takes into account both model size and accuracy is selected for the user to choose. In addition, a method based on transfer learning is also used to ensure the improvement of the target model accuracy, and the model after fine-tuning multi-objective quantization is used to further improve the model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings: Figure 1 is a schematic flowchart of a method for hybrid quantization of an artificial intelligence model based on multi-objective optimization provided by an embodiment of the present application; Figure 2 is a schematic flowchart of a transfer learning algorithm provided by an embodiment of the present application; Figure 3 is a schematic flowchart of a multi-objective optimization method assisted by a Bayesian neural network surrogate model and an evolutionary algorithm provided by an embodiment of the present application; Figure 4 is a schematic diagram of an evolutionary algorithm provided by an embodiment of the present application; Figure 5 is a schematic structural diagram of a device for hybrid quantization of an artificial intelligence model based on multi-objective optimization provided by an embodiment of the present application; Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0010] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention are also intended to include the plural forms unless the context clearly indicates otherwise.
[0011] It should be understood that although terms such as first, second, and third may be used to describe the acquisition modules in the embodiments of the present invention, these acquisition modules should not be limited to these terms. These terms are only used to distinguish the acquisition modules from each other.
[0012] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0013] It should be noted that the orientation terms such as "upper", "lower", "left", and "right" described in the embodiments of the present invention are described from the angles shown in the drawings and should not be construed as limitations on the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that an element is formed "on" or "under" another element, it can not only be directly formed "on" or "under" another element, but also be indirectly formed "on" or "under" another element through an intermediate element.
[0014] In view of the requirements of high energy efficiency and fast response of the edge AI chip in the present invention, in order to cope with the AI scenarios with insufficient data and better quantize artificial intelligence models such as machine learning models and deep learning models, a method and device for hybrid quantization of artificial intelligence models based on multi-objective optimization are proposed.
[0015] See Figure 1 , an embodiment of the present invention provides a method for hybrid quantization of an artificial intelligence model based on multi-objective optimization. Taking the YOLO series object detection model as an example, the method steps of this embodiment will be described in detail below.
[0016] Step S101, randomly generate a number of hybrid quantization samples that are super-average for the target model, and the hybrid quantization samples are mixed vector representations of different precision data types of model weights or activation functions.
[0017] Take the YOLOv5n model as an example. Given the source domain dataset Ds = {xs, ys}, with the data size Ns, where Ns is generally greater than 1000, the loss function Loss of the YOLOv5n model is composed of the weighted sum of three parts: localization loss, confidence loss, and classification loss. The source domain model is trained based on the Adam optimizer. After epochs of training, the source domain model fs is obtained. Copy the weights Ws and activation function As of the source domain model fs to the initial target model. Copying the weights Ws can avoid training the target model from scratch and reduce the dependence on the target data volume; copying the activation function As can maintain the stability of the network structure and the consistency of feature extraction, and avoid feature distribution deviation caused by changing the activation function. Based on the target domain dataset Dt = {xt, yt}, with the data size Nt, where Nt is much smaller than Ns, for example, only 100, perform local fine-tuning on the initial target model.
[0018] More preferably, an early stopping technique is adopted during fine-tuning to prevent overfitting. Since the present invention is for the scenario of insufficient training data, the present invention considers freezing a part of the non-critical network parameters and fine-tuning other critical parameters. For example, in this embodiment, the Backbone layer of the backbone can be frozen, and only the subsequent Neck and Head layers are trained, and a small number of critical parameters are fine-tuned to obtain the best transfer learning effect. The training of the initial target model is based on the model fine-tuning method. After epocht times of training, the final target model ft is obtained. See the target detection transfer learning process in Figure 2 .
[0019] Furthermore, after obtaining the target model, the Latin hypercube algorithm is used to initialize the sample population of mixed quantization. Specifically, several mixed quantization samples that are super-averaged are randomly generated for the target model. The mixed quantization samples are the mixed vector representations of different precision data types of the model weights or activation functions.
[0020] Step S102, perform parallel mixed quantization on the target model using the mixed quantization samples to obtain a mixed quantization result representing the precision and size of the mixed quantization model. Form data pairs with each mixed quantization sample and the corresponding mixed quantization result, and all the data pairs form a mixed quantization dataset.
[0021] Exemplarily, given that the data state range of each weight and activation function is [0, 1], x is regarded as an individual in the sample population, the population number is N, and super-average initial samples are randomly generated based on the Latin hypercube algorithm. Then, the hybrid quantization of the target model is performed in parallel to form N real quantization individuals. Each quantization individual can be a data pair composed of a hybrid quantization sample and the corresponding hybrid quantization result, and the data pair is added to the dataset D. Among them, the optional value of N is 60 - 200, and N in this embodiment is given as 200. For each individual x, if the value of the i-th dimension is 0, the hybrid quantization operation needs to be performed, and the data is IN8; conversely, if the value of the i-th dimension is 1, the Float32 type is retained. Each x corresponds to a hybrid quantization model, and the hybrid quantization models corresponding to each individual have different precisions and sizes.
[0022] Step S103, iteratively screen out the set of optimal hybrid quantization samples from the hybrid quantization dataset by using an evolutionary algorithm and a Bayesian neural network model.
[0023] Specifically, in this embodiment, a brand-new algorithm with fast iteration speed and strong global optimization ability is adopted to screen out the Pareto optimal solution set for the user to select appropriate hybrid-precision quantization samples from it.
[0024] The idea of this brand-new algorithm is as follows: Through non-dominated sorting and crowding degree sorting, a set of hybrid quantization samples is iteratively selected from the dataset D as the parent population of the evolutionary algorithm. Through crossover operation and mutation operation, the offspring population is obtained, and then the offspring population and the quantization results of the offspring population are put into the dataset D again, and the next generation of population is iteratively optimized until the highest number of iterations is reached, and the Pareto optimal solution set is screened out. It should be noted that in this embodiment, the hybrid quantization results of the offspring population are not directly used to calculate the target model, but the trained Bayesian neural network is used to predict the hybrid quantization results of the offspring population. Although constructing a Bayesian neural network requires a certain cost, the prediction based on the Bayesian neural network is very fast. By replacing the target model to screen out the optimal elite individuals, the unnecessary hybrid quantization process of other individuals can be reduced, and the performance evaluation efficiency of each generation of population can be accelerated.
[0025] See Figure 3 , the above-mentioned brand-new algorithm uses an evolutionary algorithm and a Bayesian neural network model to perform multi-objective optimization and iteratively screen out the optimal hybrid quantization samples, including the following steps (the step numbers are only used to distinguish steps and do not limit the step order): Step S1031, based on non-dominated sorting and crowding degree sorting, select several optimal data pairs from the hybrid quantization dataset, train to obtain a Bayesian neural network model, and select several optimal hybrid quantization samples from the hybrid quantization dataset as the parent population.
[0026] See Figure 4 Figure 4 , several optimal individuals are selected from the data set D as the parental population P, and non-dominated sorting and crowding degree sorting are used as the selection criteria. Based on the mutation operation and crossover operation of the evolutionary algorithm, the offspring population Q is generated. The size of the parental population P is the same as that of the offspring population Q. Preferably, in this embodiment, the simulated binary crossover (SBX) method is used. This method can explore unknown quantization schemes. When two parental individuals cross, they will exchange the data state variable values of a certain dimension with a certain probability, thereby generating two new offspring individuals. Then, the mutation operation is used to introduce randomness to prevent the algorithm from falling into a local optimal solution. This way is beneficial to achieving the global optimum.
[0027] For the generated offspring population, based on non-dominated sorting and crowding degree sorting, several data pairs composed of mixed quantization samples and mixed quantization results are selected from the mixed quantization data set. For example, the optimal b number of individuals b less than or equal to N are selected from the data set D. In this embodiment, the value of b is 150. These selected data pairs are used to train the Bayesian neural network model. In this way, the trained Bayesian neural network model can accurately predict the size and accuracy of the target model corresponding to the individuals in each generation of the population.
[0028] In this embodiment, the selection rule for selecting the optimal several mixed quantization samples from the mixed quantization data set based on non-dominated sorting and crowding degree sorting is as follows: if the non-dominated degrees of the individuals in the population composed of the mixed quantization data set are different, then select the individuals with the higher non-dominated sorting; if the non-dominated degrees of the individuals in the population composed of the mixed quantization data set are the same, then select the individuals with the higher crowding degree sorting; wherein, the individuals include the data pairs or mixed quantization samples of the mixed quantization data set. That is to say, this embodiment will select the individuals with a higher sorting level as the preferred data pairs of mixed quantization samples and mixed quantization results. Generally speaking, they will have an advantage in terms of model accuracy or model size. If the non-dominated sorting levels of two individuals are the same, then select the individual with a larger crowding degree as the preferred data pair of mixed quantization samples and mixed quantization results.
[0029] Step S1032, generate the offspring population of the parental population through the crossover operation and the mutation operation; Step S1033, predict the individuals in the offspring population through the trained Bayesian neural network model to obtain the size and accuracy of the mixed-quantized target model corresponding to the individuals in the offspring population; Step S1034, through non-dominated sorting and crowding degree sorting, select several optimal new individuals in the offspring population corresponding to the prediction result of the Bayesian neural network; wherein, the selection rule is the same as the selection rule in step S1031.
[0030] Step S1035: Use the new individual to perform hybrid quantization on the target model, that is: for each new individual, if the value of the i-th dimension is 0, then the hybrid quantization operation needs to be performed, and the quantized data is IN8; conversely, if the value of the i-th dimension is 1, then the hybrid quantization operation is not performed and the Float32 type is retained; and calculate the accuracy and size of the quantized model, and add the new individual and its quantized model accuracy and size as a new data pair to the hybrid quantization dataset. Step S1036: Perform iterative operations on steps S1031 to S1035 above until the maximum number of iterations is met, and use several optimal new individuals in the offspring population generated in the last iteration as a set of optimal hybrid quantization samples.
[0031] The novel algorithm for performing multi-objective optimization using an evolutionary algorithm and a Bayesian neural network model proposed in this embodiment has the advantages of fast iteration speed and strong global optimization ability. Due to the use of the elitist retention strategy of the evolutionary algorithm, the optimal solution can be searched more quickly. At the same time, due to the use of the Bayesian neural network model to perform quantization evaluation instead of the target model, the global optimization ability becomes stronger.
[0032] Step S104: Screen out the best hybrid quantization sample that achieves the best balance between model accuracy and size from the set of optimal hybrid quantization samples, and use this best hybrid quantization sample to quantize the target model to obtain the best quantized model, and use the target domain dataset to fine-tune the best quantized model to obtain the final quantized model.
[0033] Specifically, according to the special requirements for model size and model accuracy, the user screens out a best hybrid quantization sample from the set of optimal hybrid quantization samples, uses it to quantize the target model to obtain the best quantized model. Use the Dt dataset in the target domain to perform fine-tuning operations on the model. Try to avoid large changes in the model during fine-tuning and set a small initial learning rate. Further, in order to determine the initial learning rate, preferably use the Bayesian optimization algorithm to optimize this parameter. In order to avoid overfitting during the fine-tuning of the target model, the early stopping method is used to prevent the validation set error from rising. In addition, all data types are fine-tuned during fine-tuning. Finally, deploy the fine-tuned model to the hardware system, and this hardware system supports mixed-precision calculation.
[0034] See Figure 5 , Another embodiment of the present invention also provides an artificial intelligence model hybrid quantization device 200 based on multi-objective optimization, including a first module 201, a second module 202, a third module 203, and a fourth module 204. The device 200 can execute the artificial intelligence model hybrid quantization method based on multi-objective optimization in the method embodiment.
[0035] Specifically, the artificial intelligence model hybrid quantization device 200 based on multi-objective optimization includes: The first module 201 is used to randomly generate several hybrid quantization samples that are super-average for the target model. The hybrid quantization samples are the hybrid vector representations of different precision data types of model weights or activation functions; The second module 202 is used to perform parallel hybrid quantization on the target model using the hybrid quantization samples, obtain hybrid quantization results representing the precision and size of the hybrid quantization model, form data pairs with each hybrid quantization sample and the corresponding hybrid quantization result, and all the data pairs constitute a hybrid quantization data set; The third module 203 is used to iteratively screen out a set of optimal hybrid quantization samples from the hybrid quantization data set using an evolutionary algorithm and a Bayesian neural network model; The fourth module 204 is used to screen out the best hybrid quantization sample from the set of optimal hybrid quantization samples that achieves the best balance between model precision and size, and use this best hybrid quantization sample to quantize the target model to obtain the best quantization model, and use the target domain data set to fine-tune the best quantization model to obtain the final quantization model.
[0036] Furthermore, it further includes: a transfer learning module, which is used to obtain a source domain model trained according to a source domain data set, and perform transfer learning on the source domain model based on the target domain data set to obtain the target model.
[0037] Furthermore, the transfer learning module is also used to copy the weights and activation functions of the source domain model to the initial target model, fine-tune the initial target model using the target domain data set, freeze non-critical network parameters during fine-tuning, and adjust critical network parameters to obtain the final target model.
[0038] Furthermore, if the non-dominance degrees of the individuals in the population constituted by the hybrid quantization data set are different, then select the individuals with a higher non-dominance ranking; if the non-dominance degrees of the individuals in the population constituted by the hybrid quantization data set are the same, then select the individuals with a higher crowding degree ranking; wherein, the individuals include the data pairs or hybrid quantization samples of the hybrid quantization data set.
[0039] Further, the third module 203 is further configured to: based on non-dominated sorting and crowding degree sorting, select a number of optimal data pairs from the mixed quantization dataset, train a Bayesian neural network model, and select a number of optimal mixed quantization samples from the mixed quantization dataset as the parent population; generate an offspring population of the parent population through crossover operation and mutation operation; predict individuals in the offspring population through the Bayesian neural network model to obtain the size and accuracy of the target model after mixed quantization corresponding to the individuals in the offspring population; through non-dominated sorting and crowding degree sorting, select a number of optimal new individuals in the offspring population corresponding to the prediction result of the Bayesian neural network; perform mixed quantization on the target model with the new individuals, and calculate the quantization model accuracy and model size, and add the new individuals and their quantization model accuracy and model size as new data pairs to the mixed quantization dataset; perform iterative operations on the above steps until the maximum number of iterations is satisfied, and use a number of optimal new individuals in the offspring population generated in the last iteration as the set of optimal mixed quantization samples.
[0040] It should be noted that the technical solution of the artificial intelligence model mixed quantization device 200 based on multi-objective optimization provided in this embodiment can be used to execute the technical solutions of the method embodiments. Its implementation principle and technical effects are similar to those of the method, and will not be elaborated here.
[0041] Figure 6 FIG. 7 is a schematic structural diagram of an electronic device 300 provided in another embodiment of the present invention. The electronic device 300 is used to implement the method for mixed quantization of an artificial intelligence model based on multi-objective optimization in the method embodiments. The electronic device 300 in the embodiments of the present invention may include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a PC, a server, etc. Figure 6 The shown electronic device 300 is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0042] As Figure 6 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can execute various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303 to implement the method of the embodiments as described in the present invention. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0043] Typically, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0044] The above description is only a preferred embodiment of the present invention. Those skilled in the art should understand that the disclosed scope in the present invention is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solution formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.
Claims
1. A hybrid quantization method for artificial intelligence models based on multi-objective optimization, characterized in that: The steps include: Randomly generate a plurality of super-average mixed quantized samples for the target model, where the mixed quantized samples are mixed vector representations of data types with different precisions of the model weights or activation functions; Performing parallel hybrid quantization on the target model using the hybrid quantization samples to obtain a hybrid quantization result representing the accuracy and size of the hybrid quantization model, and forming a data pair with each hybrid quantization sample and the corresponding hybrid quantization result, and all the data pairs forming a hybrid quantization data set; Iteratively screening out a set of optimal mixed quantization samples from the mixed quantization data set using an evolutionary algorithm and a Bayesian neural network model; The best hybrid quantization sample that achieves the best balance between the accuracy and size of the target model is selected from the set of the best hybrid quantization samples, and the target model is quantized with the best hybrid quantization sample to obtain the best quantization model. The best quantization model is fine-tuned using the target domain data set to obtain the final quantization model.
2. The hybrid quantization method of artificial intelligence model based on multi-objective optimization according to claim 1 is characterized in that: Also includes: A source domain model is obtained by training according to a source domain data set, and transfer learning is performed on the source domain model based on a target domain data set to obtain the target model.
3. The hybrid quantization method of artificial intelligence model based on multi-objective optimization according to claim 2 is characterized in that: The step of performing transfer learning on the source domain model based on the target domain dataset comprises: The weights and activation functions of the source domain model are copied to the initial target model, and the initial target model is fine-tuned using the target domain dataset. During the fine-tuning, non-critical network parameters are frozen, and critical network parameters are adjusted to obtain the final target model.
4. The hybrid quantization method of artificial intelligence model based on multi-objective optimization according to claim 1 is characterized in that: If the non-dominance degrees of individuals in the population formed by the mixed quantized data set are different, then selecting the individual with the highest non-dominance degree ranking; If the non-dominated degrees of individuals in the population formed by the mixed quantitative data set are the same, then the individual with the highest crowding degree ranking is selected; The individuals include data pairs or mixed quantized samples of the mixed quantized data set.
5. The hybrid quantization method of artificial intelligence model based on multi-objective optimization according to claim 1 is characterized in that: The step of iteratively screening out a set of optimal mixed quantization samples from the mixed quantization data set by using an evolutionary algorithm and a Bayesian neural network model comprises: Based on non-dominated sorting and crowding sorting, the best several data pairs are selected from the mixed quantization data set, a Bayesian neural network model is obtained by training, and the best several mixed quantization samples are selected from the mixed quantization data set as the parent population; Generate a child population of the parent population through crossover and mutation operations; Predicting the individuals in the offspring population by using the Bayesian neural network model to obtain the size and accuracy of the mixed quantized target model corresponding to the individuals in the offspring population; Selecting the best new individuals in the offspring population corresponding to the prediction results of the Bayesian neural network through non-dominated sorting and crowding sorting; Using the new individual to perform hybrid quantization on the target model, and calculating the model accuracy and model size after quantization, and adding the new individual and its model accuracy and model size after quantization as a new data pair into the hybrid quantization data set; The above steps are iterated until the maximum number of iterations is met, and the best new individuals in the offspring population generated by the last iteration are used as the set of optimal mixed quantization samples.
6. A hybrid quantization device for artificial intelligence model based on multi-objective optimization, characterized in that: include: The first module is used to randomly generate a plurality of super-average mixed quantized samples for the target model, where the mixed quantized samples are mixed vector representations of data types with different precisions of the model weight or activation function; The second module is used to perform parallel hybrid quantization on the target model using the hybrid quantization samples to obtain a hybrid quantization result representing the accuracy and size of the hybrid quantization model, and each hybrid quantization sample and the corresponding hybrid quantization result constitute a data pair, and all data pairs constitute a hybrid quantization data set; The third module is used to iteratively select a set of optimal mixed quantization samples from the mixed quantization data set by using an evolutionary algorithm and a Bayesian neural network model; The fourth module is used to screen out the best hybrid quantization sample that achieves the best balance between model accuracy and size from the set of the best hybrid quantization samples, and use the best hybrid quantization sample to quantize the target model to obtain the best quantization model, and use the target domain data set to fine-tune the best quantization model to obtain the final quantization model.
7. The hybrid quantization device of artificial intelligence model based on multi-objective optimization according to claim 6 is characterized in that: Also includes: The transfer learning module is used to obtain a source domain model trained according to a source domain data set, and perform transfer learning on the source domain model based on a target domain data set to obtain the target model.
8. The hybrid quantization device of artificial intelligence model based on multi-objective optimization according to claim 7 is characterized in that: The transfer learning module is also used to copy the weights and activation functions of the source domain model to the initial target model, fine-tune the initial target model using the target domain dataset, freeze non-critical network parameters during fine-tuning, adjust critical network parameters, and obtain the final target model.
9. The hybrid quantization device of artificial intelligence model based on multi-objective optimization according to claim 6, characterized in that: If the non-dominance degrees of individuals in the population formed by the mixed quantized data set are different, then selecting the individual with the highest non-dominance degree ranking; If the non-dominated degrees of individuals in the population formed by the mixed quantitative data set are the same, then the individual with the highest crowding degree ranking is selected; The individuals include data pairs or mixed quantized samples of the mixed quantized data set.
10. The hybrid quantization device of artificial intelligence model based on multi-objective optimization according to claim 6, characterized in that: The third module is further used for: Based on non-dominated sorting and crowding sorting, the best several data pairs are selected from the mixed quantization data set, a Bayesian neural network model is obtained by training, and the best several mixed quantization samples are selected from the mixed quantization data set as the parent population; Generate a child population of the parent population through crossover and mutation operations; Predicting the individuals in the offspring population by using the Bayesian neural network model to obtain the size and accuracy of the mixed quantized target model corresponding to the individuals in the offspring population; Selecting the best new individuals in the offspring population corresponding to the prediction results of the Bayesian neural network through non-dominated sorting and crowding sorting; Using the new individual to perform hybrid quantization on the target model, and calculating the model accuracy and model size after quantization, and adding the new individual and its model accuracy and model size after quantization as a new data pair into the hybrid quantization data set; The above steps are iterated until the maximum number of iterations is met, and the best new individuals in the offspring population generated by the last iteration are used as the set of optimal mixed quantization samples.
Citation Information
Cited By
Automatic design method for NPU chip
CN122471952A