Model Compression Method, Model Compression System, Server, and Storage Medium
By combining the closed-loop operation of neural network architecture search and computing power compression, a high-accuracy compression model is generated, which solves the problem of time-consuming and labor-consuming artificially designed efficient neural network structures in the existing technology, and realizes efficient model compression.
Patent Information
- Application Number
- CN202111266014.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-10-28
AI Technical Summary
The existing model compression method cannot take into account accuracy and save labor costs. The accuracy of parameter quantization and network pruning methods is limited. Manually designing efficient neural network structures requires a lot of manpower investment and time.
Combining neural network architecture search and computing power compression, multiple transformation models are generated by continuous polling, and the transformation model obtained by neural network architecture search is used to compress computing power, forming a closed-loop operation to improve the accuracy of the compression model and reduce manual participation.
It realizes improving the accuracy of the compression model during the model compression process, reducing manual participation, reducing labor costs, and accelerating the model compression speed.
Smart Images

Figure CN114004334B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network architecture search, and in particular to a model compression method, a model compression system, a server and a storage medium. Background Art
[0002] Neural networks have achieved outstanding results in the field of machine learning, such as object classification and target detection. Intelligent algorithms based on neural networks have changed the way people produce and live. However, due to the high computational complexity of neural networks and the large number of model parameters, model computing power compression technology has become a hot topic in academic and industrial research in recent years.
[0003] At present, the commonly used methods of model compression include: parameter quantization, network pruning, and manually designed efficient neural network structures; however, parameter quantization and network pruning have great limitations in accuracy; and the design of manually designed efficient neural network structures requires a lot of manpower investment and research, and requires a lot of time for manual experiments and parameter adjustment; it can be seen that the current model compression methods cannot take into account both accuracy and saving labor costs. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose a model compression method, a model compression system, a server and a storage medium, which take into account both obtaining the accuracy of the compressed model and reducing labor costs.
[0005] To achieve the above-mentioned purpose, an embodiment of the present application provides a model compression method, including: receiving a candidate model; performing a neural network architecture search on the candidate model to obtain multiple transformation models; performing computing power compression on the candidate model to obtain a compressed model; using the transformation model and the compressed model as the candidate model, and re-performing the neural network architecture search and the computing power compression.
[0006] An embodiment of the present application also provides a model compression system, including: a receiving module, a model automatic search module, and a computing power compression module; the receiving module is used to receive a candidate model; the model automatic search module is used to perform a neural network architecture search on the candidate model, obtain multiple transformation models and input them into the receiving module as the candidate model; the computing power compression module is used to perform computing power compression on the candidate model to obtain a compressed model, and the compressed model is input into the receiving module as the candidate model.
[0007] An embodiment of the present application also provides a server, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model compression method as described above.
[0008] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned model compression method is implemented.
[0009] The model compression method of the present application can obtain as many neural network models as possible through neural network architecture search, thereby increasing the probability of obtaining a better model. And the present application can compress multiple transformation models obtained from the previous polling through computing power compression. It can be seen that the present application combines model compression with neural network architecture search, uses the transformation models obtained from neural network architecture search as candidate models for computing power compression to obtain compressed models in the next polling, and then uses these compressed models as candidate models again. The compressed models obtained through computing power compression can also be used for neural network architecture search to obtain transformation models in the next polling. Through continuous polling, after polling to a certain extent, a compressed model with a relatively high accuracy can be obtained, thus taking into account the accuracy of obtaining the compressed model, and the entire model compression process reduces manual participation and lowers labor costs. In addition, compared with the related art, the present application can automatically obtain compressed models and also speeds up the model compression process. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation to the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.
[0011] Figure 1 is a flowchart of a model compression method according to an embodiment of the present application;
[0012] Figure 2 is a flowchart of the sub-steps of step 103 of the model compression method according to an embodiment of the present application;
[0013] Figure 3 is a flowchart of a model compression method according to an embodiment of the present application;
[0014] Figure 4 is a flowchart of a model compression method according to an embodiment of the present application;
[0015] Figure 5 is a schematic structural diagram of a model compression system according to an embodiment of the present application;
[0016] Figure 6 is a schematic structural diagram of a model compression system according to an embodiment of the present application;
[0017] Figure 7It is a schematic structural diagram of a server according to an embodiment of the present application. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are presented to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.
[0019] An embodiment of the present application relates to a model compression method, as Figure 1 shown, which is a schematic flowchart of the model compression method of this embodiment, and specifically includes the following steps:
[0020] Step 101, receive a candidate model.
[0021] Specifically, the execution subject of this embodiment is a model compression system. During the process of model compression, the model compression system needs to first receive a candidate model.
[0022] Step 102, perform neural network architecture search on the candidate model to obtain multiple transformed models.
[0023] Specifically, neural network architecture search (NAS) can effectively liberate human input, and NAS can explore a larger space of network structures, and has the opportunity to generate high-performance structures beyond human design, and has been widely used in the field of neural networks. Therefore, the model compression method of this embodiment combines NAS to solve the problem of human input, and can also obtain a more accurate compressed model. In one embodiment, performing neural network architecture search on the candidate model to obtain multiple transformed models includes: performing neural network architecture search based on a preset search space, and randomly transforming the candidate model to obtain multiple transformed models.
[0024] Specifically, the NAS search process traverses a preset search space based on a preset search strategy, and randomly transforms the candidate model during each NAS search to obtain multiple transformed models. This randomness ensures the diversity of the generated model structures and is conducive to searching for high-performance models. It should be noted that the search space of this embodiment can be set by the user or be the system default, and the search strategy can also be set by the user or be the system default; among them, the NAS search strategy includes those based on reinforcement learning, evolutionary algorithms, gradient descent, etc.
[0025] It should be noted that the search strategy involved in this embodiment can be any search strategy, which has strong robustness and can be set to the corresponding search strategy according to the needs of the user or the system. This embodiment does not make specific limitations.
[0026] Step 103, perform computing power compression on the candidate model to obtain a compressed model.
[0027] The order of the above steps 102 and 103 is not specifically limited in this embodiment. Step 102 can be before step 103, step 102 can also be after step 103, and steps 102 and 103 can also be performed simultaneously.
[0028] In one embodiment, performing computing power compression on the candidate model to obtain a compressed model, that is, step 104, includes the following sub-steps. The specific process schematic diagram is as Figure 2 shown, and includes:
[0029] Step 1031, calculate the computing power of each layer in the candidate model.
[0030] Specifically, computing power compression is responsible for directionally compressing the computing power of the candidate model. For example, it will detect the computing power intensive area of the model and perform directional compression processing on the layer with the largest computing power consumption in the model.
[0031] Specifically, according to the input of the candidate model and the parameters of each layer, the computing power consumption of each layer in the candidate model can be calculated. The computing power consumption of each layer is mainly determined by the size of the feature map input to this layer and the size of the weight parameter quantity that needs to be trained for the operation of this layer, and can be approximately expressed by the following formula:
[0032] FLOPs(l) = H × W × C × Params(l);
[0033] Among them, l represents an operation layer (such as a convolutional layer) in the model, H and W represent the size of the feature map input to this layer, C represents the number of channels of the input feature map, and Params represents the parameter quantity of the current layer.
[0034] Step 1032, compare the computing power of each layer to obtain the layer with the largest computing power.
[0035] Specifically, by comparing the computing power of each layer obtained in step 1041, the layer with the largest computing power in the model, namely the "computing power intensive area", can be located. The computing power compression step mainly performs computing power compression on the "computing power intensive area" of the model.
[0036] Step 1033, compress the computing power of the layer with the largest computing power to obtain a compressed model.
[0037] In one embodiment, the computing power of the layer with the largest computing power is compressed, including: reducing the parameter amount of the layer with the largest computing power, or removing the layer with the largest computing power, or reducing the size of the input feature map of the layer with the largest computing power.
[0038] Specifically, from the above formula for computing power consumption, we can see that the factors affecting computing power mainly include the feature map of the input layer and the number of parameters of the current layer. Therefore, there are three possible directions for compressing the computing power of a layer in the model:
[0039] (1) Reduce the number of parameters in the current layer. For example, for a convolution operation, the number of parameters is determined by the size and number of its convolution kernels. Therefore, the size of the convolution kernel can be compressed or the number of filters proposed after the search task is started can be reduced. Reducing the number of parameters in the current layer is equivalent to changing the hyperparameters of the current layer, but the direction of change is in the direction of decreasing the number of parameters.
[0040] (2) Remove the current layer, which is the extreme case of case (1), that is, reduce the number of parameters of this layer to 0.
[0041] (3) Reduce the size of the feature map input to this layer. Reducing the size of the input feature map means that the operation layer in front of the "computational power intensive area" needs to be changed, because only by changing the previous operation layer can the input of the current layer be changed. Specific operations include: ① Add a pooling layer in front of the "computational power intensive area" to compress the features; ② Tracing back to the layers in front of the "computational power intensive area" that consume computing power and changing their hyperparameters, such as reducing the number of channels of the output feature or increasing the step size, to achieve the effect of downsampling the feature map.
[0042] Step 104: train the transformation model and the compression model as candidate models.
[0043] Specifically, the models generated by the neural network architecture search and computing power compression are jointly used as candidate models, and the process returns to step 101 to re-perform the neural network architecture search and computing power compression; that is, the above steps 101 to 104 are a polling process, and after step 104 is completed, step 101 is executed again, thereby forming a closed-loop operation of the system.
[0044] Specifically, during the first polling process, the candidate model received by the model compression system is the baseline model input by the user.
[0045] It should be noted that each candidate model in this embodiment will be set with an identity (ID). When the user inputs the baseline model into the system, as well as the models generated through neural network architecture search and computing power compression, an ID will be assigned to each of them. This ID is unique, which can ensure that there will be no problem of duplicate training for the trained models.
[0046] The model compression method of this embodiment can obtain as many neural network models as possible through neural network architecture search, thereby increasing the probability of obtaining a better model. And in this application, computing power compression can be used to compress multiple transformed models obtained from the previous polling. It can be seen that this application combines model compression with neural network architecture search, uses the transformed models obtained from neural network architecture search as candidate models to perform computing power compression to obtain compressed models during the next polling, and then uses these compressed models as candidate models again. The compressed models obtained by computing power compression can also be used to perform neural network architecture search to obtain transformed models during the next polling. Through continuous polling, after polling to a certain extent, a compressed model with a relatively high accuracy can be obtained, thus taking into account the accuracy of obtaining the compressed model, and the entire model compression process reduces manual participation and lowers labor costs. In addition, compared with the related technologies, this application can automatically obtain compressed models and also speeds up the model compression speed.
[0047] An embodiment of this application relates to a model compression method. This embodiment is substantially the same as the previous embodiment. The main difference is that after receiving the candidate model, it further includes: training the candidate model to obtain a trained model; obtaining the accuracy and computing power of the trained model; and obtaining the evaluation parameters of the candidate model corresponding to the trained model based on the accuracy and computing power. For the sake of simplicity of description, the same or corresponding parts of this embodiment and the previous embodiment will not be described again.
[0048] The flow diagram of the model compression method of this embodiment is as Figure 3 shown, and specifically includes the following steps:
[0049] Step 201, receive the candidate model.
[0050] Step 202, train the candidate model to obtain a trained model.
[0051] Step 203, obtain the accuracy and computing power of the trained model.
[0052] Step 204, obtain the evaluation parameters of the candidate model corresponding to the trained model based on the accuracy and computing power.
[0053] Step 205: Conduct neural network architecture search on the candidate model to obtain multiple transformed models.
[0054] Step 206: Compress the computing power of the candidate model to obtain a compressed model.
[0055] Step 207: Use the transformed model and the compressed model as candidate models.
[0056] The above Steps 201, 205 to 207 are the same as Steps 101 to 104 in the previous embodiment, and will not be elaborated here.
[0057] It should be noted that the above Steps 201 to 207 are a process of one round of polling. After the model compression system completes Step 207, it will enter Step 201 again, thus forming a closed-loop operation of the system.
[0058] Specifically, after receiving the candidate model, train the candidate model to obtain a trained model, and conduct performance evaluation on the trained model, so that the candidate model required by the user can be screened according to the user's needs. Specifically, in the process of the first round of polling, the benchmark model input by the user is used as the candidate model, and only the benchmark model input by the user is subjected to performance evaluation. In the subsequent polling process, the number of candidate models is relatively large, and each candidate model needs to be trained to obtain a relatively large number of trained models. In the process of performance evaluation, all the trained models need to be traversed, and performance evaluation is conducted on each trained model.
[0059] Specifically, the performance evaluation of this embodiment needs to rely on two parameters, namely the accuracy and computing power of the candidate model. Therefore, after obtaining the candidate model in this embodiment, obtain the accuracy and computing power of the candidate model. Among them, the accuracy refers to the accuracy rate of the candidate model during use, and the computing power refers to the intensity of the amount of operations of the candidate model during use.
[0060] Specifically, since the accuracy and computing power are different parameters, when conducting evaluation, the values of each parameter need to be converted to the same dimension, that is, based on the accuracy and computing power, obtain the evaluation parameter of the candidate model corresponding to the trained model, and finally give the comprehensive performance score of the candidate model corresponding to the trained model, that is, the evaluation parameter. The evaluation parameter can characterize the performance of the candidate model after weighing multiple indicators.
[0061] In one embodiment, the evaluation parameter includes an evaluation score; based on accuracy and computing power, the evaluation parameters of the candidate model corresponding to the trained model are obtained, including: calculating the accuracy and computing power according to a preset weight ratio to obtain the evaluation score of the candidate model. Specifically, in this embodiment, different weights are set for accuracy and computing power respectively, so as to calculate the evaluation score of the candidate model. Different weights can be set according to the user's needs. For example, if the user has a higher requirement for accuracy, a higher weight can be set for accuracy; or if the user has a higher requirement for computing power, a higher weight can be set for computing power. Specifically, for the evaluation score, it can be set that the higher the evaluation score, the better the performance of the obtained candidate model.
[0062] In one embodiment, the evaluation parameter includes an evaluation score; after obtaining the evaluation parameters of the candidate model corresponding to the trained model based on accuracy and computing power, it further includes: retaining the candidate models whose evaluation scores are greater than or equal to a preset score, or retaining the N candidate models with the highest evaluation scores. Specifically, except for the first round of polling, there will be multiple candidate models in the subsequent polling process. If each candidate model enters the subsequent neural network architecture search and computing power compression, it will lead to a large computational amount of the system and also cause waste of resources. Therefore, in this embodiment, a preset rule will be set, that is, retaining the candidate models whose evaluation scores are greater than or equal to a preset score, or retaining the N candidate models with the highest evaluation scores, so as to retain the candidate models that meet the preset rule and delete the candidate models that do not meet the preset rule.
[0063] Specifically, in the process of the first round of polling, only the benchmark model input by the user is trained. In the subsequent polling process, the number of candidate models increases and screening is required. The candidate models can be sorted according to the evaluation scores. A preset score is set in advance, and the candidate models whose evaluation scores are greater than or equal to the preset score are retained, or a value N is set in advance, and the N candidate models with the highest evaluation scores are retained, so as to enter the subsequent steps of neural network architecture search and computing power compression, and delete the candidate models that do not meet the preset rule, so that the models with higher performance can be fully utilized, thereby reducing the computational amount of the system and avoiding waste of resources.
[0064] It should be noted that if the user has no special requirements, the default rule is to set a value N and retain the N candidate models with the highest evaluation scores.
[0065] Specifically, in the subsequent polling process, other methods can also be used to retain some candidate models to enter the subsequent neural network architecture search and computing power compression, such as roulette selection method, tournament selection method, etc.
[0066] One embodiment of the present application relates to a model compression method. This embodiment is substantially the same as the previous embodiment. The main difference is that after obtaining the evaluation scores of the candidate models corresponding to the trained model based on accuracy and computing power, this embodiment further includes: determining whether the maximum evaluation score obtained in this polling is greater than the maximum evaluation score obtained in the previous polling; for the sake of convenience of description, the same or corresponding parts of this embodiment and the previous embodiment will not be described again.
[0067] The flowchart of the model compression method of this embodiment is as Figure 4 shown, and specifically includes the following steps:
[0068] Step 301, receive a candidate model.
[0069] Step 302, train the candidate model to obtain a trained model.
[0070] Step 303, obtain the accuracy and computing power of the trained model.
[0071] Step 304, obtain the evaluation parameters of the candidate model corresponding to the trained model based on accuracy and computing power.
[0072] Step 305, determine whether the maximum evaluation score obtained in this polling is greater than the maximum evaluation score obtained in the previous polling. If so, go to Step 307 and Step 308; if not, go to Step 306.
[0073] Step 306, increment the counter value by 1, and determine whether the counter value is less than the preset threshold. If so, go to Step 307 and Step 308; if not, the polling ends.
[0074] Step 307, perform neural network architecture search on the candidate model to obtain multiple transformed models.
[0075] Step 308, perform computing power compression on the candidate model to obtain a compressed model.
[0076] Step 309, use the transformed model and the compressed model as candidate models.
[0077] The above steps 301 to 304, steps 307 to 309 are the same as steps 201 to 207 of the previous embodiment, and will not be described again here.
[0078] Specifically, after each polling to evaluate the performance of the candidate models, the evaluation scores of each candidate model are calculated, and among them, there will be a model with the highest evaluation score, that is, the best-performing model; then, compare the maximum evaluation score of this polling with the maximum evaluation score obtained in the previous polling. If the maximum evaluation score of this polling is greater than the maximum evaluation score obtained in the previous polling, it means that the candidate models are continuously being optimized, and a model with better performance may be obtained in the next polling. At this time, continue to step 307 and step 308; if the maximum evaluation score of this polling is less than or equal to the maximum evaluation score obtained in the previous polling, it means that it is very likely that the best-performing model has been obtained at present. At this time, enter step 306, increment the counter value by 1, and determine whether the counter value is less than the preset threshold. If the counter value is less than the preset threshold, the number of polling times has not reached the set standard, and a model with a higher evaluation score may still appear, and it is necessary to enter step 307 and step 308 again; if the counter value is greater than or equal to the preset threshold, the number of polling times has reached the set standard, and the polling ends.
[0079] Specifically, the model compression system is provided with a counter. If the maximum evaluation score of this polling is less than or equal to the maximum evaluation score obtained in the previous polling, the counter will be incremented by 1, and the initial value of the counter value is 0.
[0080] Specifically, the user can also actively issue a stop instruction to indicate the system to end the polling. When the system has obtained a model with relatively high performance in steps 303 and 304, and the user believes that there is no need to conduct further polling, the user can actively end the entire polling process.
[0081] In one embodiment, in the case of the end of polling, a target candidate model is obtained among multiple candidate models; the accuracy of the target candidate model is greater than the preset accuracy and the computing power of the target candidate model is less than the preset computing power. Specifically, after the polling ends, that is, after the search stops, the system will organize the log information recorded during the search process, including the structure information, accuracy rate, computing power consumption, and comprehensive performance information of the candidate models in each round of iteration. In this embodiment, the candidate model that meets the conditions of having an accuracy rate greater than the preset accuracy and a computing power less than the preset computing power will be used as the target candidate model, and the target candidate model will be organized into a file and output to the user. After the system completes the result output, it releases the computing resources occupied by the search task and ends the entire process. Among them, the system pre-sets a benchmark model, the preset accuracy is the accuracy of the benchmark model, and the preset computing power is the computing power of the benchmark model.
[0082] Specifically, in the case where the polling ends, the user can also sample other methods to obtain the target candidate model. For example, after calculating the evaluation scores of each candidate model, select the candidate model with the highest evaluation score as the target candidate model, organize the target candidate model into a file and output it to the user to ensure that the user uses the optimal candidate model.
[0083] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this patent.
[0084] An embodiment of the present invention relates to a model compression system, and the specific structural schematic diagram is as Figure 5 shown, including: a receiving module 401, a model automatic search module 402, and a computing power compression module 403.
[0085] Specifically, the receiving module 401 is used to receive candidate models; the model automatic search module 402 is used to perform neural network architecture search on the candidate models to obtain multiple transformed models and input them as candidate models into the receiving module 401; the computing power compression module 403 is used to perform computing power compression on the candidate models to obtain compressed models, and the compressed models are input as candidate models into the receiving module 401.
[0086] In one embodiment, the model compression system further includes: a model training module and a performance evaluation module.
[0087] As Figure 6 shown, it is the structural schematic diagram of this embodiment, including: a receiving module 401, a model automatic search module 402, a computing power compression module 403, a model training module 404, and a performance evaluation module 405.
[0088] Specifically, the receiving module 401 is connected to the model training module 404, the model training module 404 is connected to the performance evaluation module 405, the performance evaluation module 405 is respectively connected to the model automatic search module 402 and the computing power compression module 403, and the model automatic search module 402 and the computing power compression module 403 are connected to the receiving module 401; the model training module 404 is used to receive candidate models and train the candidate models to obtain trained models; the performance evaluation module 405 is used to perform performance evaluation on the trained models.
[0089] Specifically, the model training module 404 mainly has two functions: ① training candidate models to obtain trained models; ② validating each candidate model to obtain the accuracy and computing power of each candidate model. The model training module will input these two model metrics into the performance evaluation module. To efficiently utilize system resources, the model training module adopts a distributed parallel training method.
[0090] Specifically, the performance evaluation module 405 will receive the accuracy and computing power of each candidate model, evaluate the comprehensive performance of each candidate model, obtain the evaluation score of each candidate model, and retain the candidate models that meet the preset rules, while deleting the candidate models that do not meet the preset rules.
[0091] Specifically, the model automatic search module 402 receives candidate models and conducts neural network architecture search based on the candidate models. Network architecture search is performed according to the search strategy for the search space set by the user or the default full search space.
[0092] Specifically, the computing power compression module 403 receives candidate models. This module is the core module of the system and is responsible for directionally compressing the computing power of candidate models. It will detect the computing power intensive areas of the models and perform directional compression processing on the layers that consume the most computing power of the models.
[0093] It is not difficult to find that this embodiment is a system embodiment corresponding to the previous embodiment, and this embodiment can be implemented in cooperation with the previous embodiment. The relevant technical details mentioned in the previous embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the previous embodiment.
[0094] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0095] An embodiment of the present invention relates to a server, as Figure 7 shown, including at least one processor 501; and a memory 502 communicatively connected to the at least one processor 501; wherein, the memory 502 stores instructions executable by the at least one processor 501, and the instructions are executed by the at least one processor 501 to enable the at least one processor 501 to execute the communication control method as described above.
[0096] Among them, the memory 502 and the processor 501 are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 501 and the memory 502 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 501 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 501.
[0097] The processor 501 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 502 can be used to store data used by the processor 501 when executing operations.
[0098] An embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiment is implemented.
[0099] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by a program instructing relevant hardware. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0100] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.
Claims
1. A model compression method, characterized in that, Comprising: Receiving a candidate model; Training the candidate model to obtain a trained model; Obtaining the accuracy and computing power of the trained model; Based on the accuracy and the computing power, obtaining an evaluation parameter of the candidate model corresponding to the trained model, the evaluation parameter including an evaluation score; Retaining the candidate models whose evaluation scores are greater than or equal to a preset score, or retaining the N candidate models with the largest evaluation scores; Performing neural network architecture search on the candidate model to obtain multiple transformed models; Calculating the computing power of each layer in the candidate model, wherein the consumption formula of the computing power of each layer is, FLOPs(l) = H × W × C × Params(l); where l represents an operation layer in the model, H and W represent the sizes of the feature maps input to the operation layer, C represents the number of channels of the input feature map, and Params represents the number of parameters of the current layer; Comparing the computing power of each layer to obtain the layer with the largest computing power; Reducing the number of parameters of the layer with the largest computing power, or removing the layer with the largest computing power, or reducing the size of the input feature map of the layer with the largest computing power to obtain the compressed model; Taking the transformed model and the compressed model as the candidate models, and re-performing the neural network architecture search and the computing power compression.
2. The model compression method according to claim 1, wherein The performing neural network architecture search on the candidate model to obtain multiple transformed models includes: Performing the neural network architecture search based on a preset search space, and randomly transforming the candidate model to obtain multiple transformed models.
3. The model compression method according to claim 1, wherein The steps from receiving the candidate model to taking the transformed model and the compressed model as the candidate models are one round of polling; the evaluation parameter includes an evaluation score; after obtaining the evaluation parameter of the candidate model corresponding to the trained model based on the accuracy and the computing power, further including: Judging whether the maximum evaluation score obtained in this round of polling is greater than the maximum evaluation score obtained in the previous round of polling; If so, entering the steps of neural network architecture search and computing power compression; If not, adding 1 to the counter value, and judging whether the counter value is less than a preset threshold; the initial value of the counter value is 0; If the counter value is less than the preset threshold, entering the steps of neural network architecture search and computing power compression; if the counter value is greater than or equal to the preset threshold, the polling ends.
4. The model compression method according to claim 3, wherein In the case where the polling ends, obtaining the target candidate model among the multiple candidate models; the accuracy of the target candidate model is greater than a preset accuracy and the computing power of the target candidate model is less than a preset computing power.
5. The model compression method according to claim 1, wherein The candidate model received in the first round of polling is the benchmark model input by the user.
6. A model compression system, characterized in that, Comprising: A receiving module, a model automatic search module, a computing power compression module, and a selection module; The receiving module is used to receive a candidate model; The model automatic search module is used to perform neural network architecture search on the candidate model, obtain multiple transformed models and input them into the receiving module as the candidate models; The computing power compression module is used to calculate the computing power of each layer in the candidate model. Among them, the consumption formula of the computing power of each layer is FLOPs(l) = H × W × C × Params(l); where l represents an operation layer in the model, H and W represent the size of the feature map input to the operation layer, C represents the number of channels of the input feature map, and Params represents the number of parameters of the current layer; compare the computing power of each layer to obtain the layer with the maximum computing power; reduce the number of parameters of the layer with the maximum computing power, or remove the layer with the maximum computing power, or reduce the size of the input feature map of the layer with the maximum computing power to obtain the compressed model, and use the compressed model as the candidate model to input to the receiving module; It further includes: a model training module and a performance evaluation module; The model training module is connected to the receiving module, the model training module is further connected to the performance evaluation module, the performance evaluation module is respectively connected to the model automatic search module and the computing power compression module, and the model automatic search module and the computing power compression module are respectively connected to the receiving module; The model training module is used to train the candidate model to obtain a trained model; The performance evaluation module is used to obtain the accuracy and computing power of the trained model, and based on the accuracy and the computing power, obtain the evaluation parameters of the candidate model corresponding to the trained model, where the evaluation parameters include an evaluation score; retain the candidate models whose evaluation scores are greater than or equal to the preset score, or retain the N candidate models with the largest evaluation scores.
7. A server, characterized in that, It includes: At least one processor; And, A memory communicatively connected to the at least one processor; where, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model compression method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the model compression method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Model selection method and device
CN118302772A
Neural network compression method, apparatus and device, and storage medium
US20230297846A1
Target recognition method based on image, and neural network model processing method
WO2024156255A1