Model compression method, apparatus and system

By tuning the candidate population on the hardware platform, the optimal compression method for each quantization layer is determined, and the problem of unreasonable selection of model compression method in the prior art is solved, and the efficient operation and performance optimization of the model on the hardware is achieved.

WO2025156909A1PCT designated stage expired Publication Date: 2025-07-31YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/142024
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-23
Filing Date
2024-12-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

When choosing the appropriate compression method, existing model compression technology cannot guarantee the model accuracy and performance, resulting in the model being unable to fully utilize its performance in hardware.

Method used

By performing tuning processing on the candidate population on the hardware platform, the optimal compression method for each quantization layer is determined, and the target network model is generated to ensure the performance and efficiency of the model when running on the hardware.

Benefits of technology

On the premise of ensuring model accuracy, the model compression strategy is optimized, the model performance is fully utilized, and the model's adaptability and prediction accuracy are improved in hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024142024_31072025_PF_FP_ABST
    Figure CN2024142024_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applied to the field of artificial intelligence. Provided are a model compression method, apparatus and system. The method comprises: acquiring a first network model; on the basis of the probability that each quantization layer in the first network model uses different compression modes, generating a first candidate population, wherein the first candidate population comprises m first candidate network models, m being a positive integer; and on the basis of a hardware platform, executing tuning processing on the first candidate population, so as to determine a target network model, wherein the target network model meets a preset condition. The present application can reduce overheads, optimize model compression, ensure model precision, and give full play to model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Model compression method, device and system

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 23, 2024, with application number 202410101427.4 and application name “Model Compression Method, Device and System”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to a model compression method, device, and system. Background Art

[0003] Currently, the models used in deep learning are becoming increasingly larger, requiring more storage and computing resources. Furthermore, larger models also increase the hardware cost and speed of model inference. Therefore, model compression has become a key technology in model deployment.

[0004] Common model compression methods include sparsification and quantization. Specifically, sparsification reduces the number of parameters in the model by setting some to zero, thereby reducing the computational complexity and storage space. Quantization reduces the computational complexity and storage space by converting floating-point parameters to low-precision fixed-point numbers or integers, thereby reducing the number of parameter bits.

[0005] However, when compressing a model, it is crucial to choose the appropriate compression method for each compressible layer (also known as a quantization layer) in the model. If the model compression method is not chosen properly, the model accuracy may not be guaranteed, and the model performance may not be fully utilized. Summary of the Invention

[0006] The present application provides a model compression method, device and system, which can optimize model compression, ensure model accuracy and give full play to model performance.

[0007] In order to achieve the above objectives, this application provides the following technical solutions:

[0008] In a first aspect, the present application provides a model compression method, which includes: obtaining a first network model; generating a first candidate population based on the probability of each quantization layer in the first network model using different compression methods, the first candidate population including m first candidate network models, where m is a positive integer; based on a hardware platform, performing tuning processing on the first candidate population to determine a target network model, and the target network model meets preset conditions.

[0009] In this application, during the model compression process, the target network model is determined based on hardware platform tuning, and the model compression strategy is optimized to ensure that the optimal compression method is selected for each quantization layer, and the model performance is fully utilized while ensuring model accuracy.

[0010] According to the first aspect, based on the hardware platform, the first candidate population is tuned to determine the target network model, including: in a first round of training, the m first candidate network models are respectively run based on the hardware platform to obtain a first intermediate network model; wherein, the first intermediate network model is the first candidate network model that meets the preset conditions and has the highest score among the m first candidate network models; in the i-th round of training, the i-th candidate population obtained after processing the first candidate population for the i-1th time is run based on the hardware platform to obtain the i-th intermediate network model; wherein, the i-th candidate population of the i-th round of training includes m i-th candidate network models, 2≤i≤j, i and j are positive integers; and the target network model is determined, which is the intermediate network model that meets the preset conditions and has the highest score among the k intermediate network models obtained after k rounds of training, k≤j, k is a positive integer.

[0011] In some examples, the preset condition is a preset model accuracy that the compressed model should meet.

[0012] According to the first aspect, or any implementation of the first aspect above, during the first round of training, the above-mentioned m first candidate network models are respectively run based on the hardware platform to obtain a first intermediate network model, including: during the first round of training, the above-mentioned m first candidate network models are respectively run based on the hardware platform to determine the scores of the above-mentioned m first candidate network models; determining the accuracy of the n first candidate network models with the highest scores among the above-mentioned m first candidate network models, 2<n≤m, n is a positive integer; and determining the first candidate network model with the highest score among the n first candidate network models whose accuracy meets the preset conditions as the first intermediate network model.

[0013] According to the first aspect, or any implementation of the first aspect above, the m first candidate network models are respectively run based on the hardware platform to determine the scores of the m first candidate network models, including: running the m first candidate network models respectively based on the hardware platform, and collecting performance parameters of the m first candidate network models; and determining the scores of the m first candidate network models respectively according to the performance parameters of the m first candidate network models.

[0014] In some examples, the performance parameters include performance data such as inference latency, consumed memory, and power consumption when the first candidate network model runs on a hardware platform.

[0015] In some examples, each candidate network model is run on a hardware platform, and performance data of each candidate network model during its run is collected using a performance analysis tool of the hardware platform.

[0016] In some examples, based on the performance parameters of the first candidate network model, a score corresponding to each performance parameter of the first candidate network model is determined, and the score of the first candidate network model is determined based on the score of each performance parameter of the first candidate network model.

[0017] In this application, the performance parameters of the model running on the hardware platform can more comprehensively and accurately evaluate the model's performance in real-world environments, verify the model's reliability, stability, and adaptability, and enable corresponding adjustments and optimizations. This ensures that the model runs well on a variety of hardware devices, improves the model's adaptability and predictive accuracy, and provides strong support for the model's practical application.

[0018] In some instances, the method of determining the accuracy of the n first candidate network models with the highest scores among the above-mentioned m first candidate network models includes: sorting the above-mentioned m first candidate network models according to the scores, and selecting the n first candidate network models with the highest scores; distilling the accuracy of the above-mentioned n first candidate network models according to a preset distillation model; and determining the accuracy of the above-mentioned n first candidate network models after distillation.

[0019] In this application, the first candidate population is loaded onto a hardware platform, and the performance data of each candidate network model in the first candidate population is collected during runtime using the hardware platform's performance analysis tools. Each candidate network is scored based on its real-time performance data, and a first intermediate network model is determined based on its accuracy and score. This ensures the performance and efficiency of the first intermediate network model when running on the hardware platform, and optimizes the model compression strategy. The first intermediate network model is a neural network model with the highest accuracy and the highest model score, while ensuring model accuracy. Selecting the neural network model with the best performance when running on the hardware is beneficial to fully utilizing the model's performance.

[0020] According to the first aspect, or any implementation of the first aspect above, after collecting the performance parameters of the above-mentioned m first candidate network models, the method also includes: analyzing the performance parameters of the above-mentioned m first candidate network models to form an optimization strategy, which is used to select the i-th candidate population.

[0021] In this application, the performance parameters of each quantization layer in each candidate network model can be obtained based on the hardware platform. The performance parameters can then be used to evaluate the compression method used by each quantization layer. An optimization strategy is then formed based on the performance parameters. When a candidate population is generated during subsequent training, the candidate population is pruned according to the optimization strategy, with poorly performing compression methods removed and high-performing ones prioritized. This optimization strategy ensures that the quantization layer in the generated i-th candidate population uses a compression method with better performance.

[0022] According to the first aspect, or any implementation of the first aspect above, the method further includes: obtaining an i-th candidate population for the i-th round of training by performing an i-1-th crossover mutation on the first candidate population.

[0023] In this application, except for the first round of training, the candidate populations used in all subsequent rounds of training are generated based on the candidate population used in the previous round of training. The candidate population in each round of training has better genes than the candidate population in the previous round of training. Through crossover and mutation, the individuals in the candidate population are continuously evolved and improved to determine the optimal candidate network model and give full play to the model performance.

[0024] According to the first aspect, or any implementation of the first aspect above, during the i-th round of training, the i-th candidate population obtained after processing the above-mentioned first candidate population for the i-1th time is run based on the hardware platform to obtain the i-th intermediate network model, including: during the i-th round of training, the i-th candidate population is run based on the hardware platform, and the scores of the m i-th candidate network models in the i-th candidate population are determined respectively; the accuracy of the n i-th candidate network models with the highest scores among the above-mentioned m i-th candidate network models is determined, 2<n≤m, and n is a positive integer; and the i-th candidate network model whose accuracy meets the preset conditions and has the highest score among the above-mentioned n i-th candidate network models is determined as the i-th target network model.

[0025] According to the first aspect, or any implementation of the first aspect above, the training termination condition is that the number of training times reaches a preset number of times. When k equals j, the number of training times reaches the maximum number of training times, then the training is terminated, and from the j intermediate network models obtained after j rounds of training, the intermediate network model that meets the preset conditions and has the highest score is selected as the target network model

[0026] According to the first aspect, or any implementation of the first aspect above, the training termination condition is that the difference between the scores of the intermediate network models obtained after two consecutive training rounds meets a preset range. When k is less than j, the difference between the score of the intermediate network model determined at the end of the kth round of training and the score of the intermediate network model determined at the end of the k-1th round of training meets the preset range. The training is terminated, and from the k intermediate network models obtained after the k rounds of training, the intermediate network model that meets the preset conditions and has the highest score is selected as the target network model.

[0027] In this application, the training end condition can ensure that the model fully learns and adapts to the training data while avoiding overfitting or premature stopping.

[0028] According to the first aspect, or any implementation of the first aspect above, a first candidate population is generated based on the probability that each quantization layer in the first network model uses a different compression method, including: determining the probability that each quantization layer in the above-mentioned first network model uses a different compression method; generating multiple first candidate network models based on the probability that each quantization layer in the above-mentioned first network model uses a different compression method; and selecting the top m first candidate network models from the above-mentioned multiple first candidate network models to form a first candidate population.

[0029] In some examples, after obtaining the first network model, the role, number of parameters, parameter distribution, importance, and network structure of each network layer in the first network model are analyzed to determine the quantization layer in the first network model.

[0030] According to the first aspect, or any implementation of the first aspect above, determining the probability that each quantization layer in the above-mentioned first network model uses different compression methods, including: calculating the output data distribution of each quantization layer in the above-mentioned first network model using different compression methods; for each of the above-mentioned quantization layers, respectively calculating the similarity between the output data distribution of the quantization layer using different compression methods and the output data distribution of the quantization layer not using compression methods, and based on the similarity, calculating the probability that the quantization layer uses different compression methods.

[0031] In some examples, for each quantization layer, all other quantization layers are fixed as non-quantization layers, and the output data distribution of the quantization layer using different compression methods and the quantization layer without compression is inferred. For the same quantization layer, the similarity between the output data distribution of the quantization layer using different compression methods and the output data distribution of the quantization layer without compression is calculated, and the probability corresponding to each compression method that can be used by the quantization layer is calculated based on the similarity.

[0032] In this application, for each quantization layer, only the quantization method variable is changed during inference, which can improve the accuracy of the output data distribution and probability corresponding to different compression methods used in each quantization layer, so that a better quantization method can be selected for each quantization layer subsequently.

[0033] In some examples, multiple first candidate network models are generated based on the probability of each quantization layer in the above-mentioned first network model using different compression methods; and the top m first candidate network models are selected from the above-mentioned multiple first candidate network models to form a first candidate population.

[0034] In this application, the first candidate network model with the highest ranking is selected to form the first candidate population. The compression method with higher probability is used in the quantization layer of each candidate network model in the first candidate population, so that a better compression method for each quantization layer can be quickly determined later to give full play to the model performance.

[0035] In the second aspect, the present application provides a model compression device, which includes: a processor and a memory, the memory being coupled to the processor, the memory being used to store computer-readable instructions, and when the processor reads the computer-readable instructions from the memory, the model compression device executes the method of the first aspect and any one of the embodiments of the first aspect.

[0036] In the third aspect, the present application provides a model compression system, which includes a hardware platform and a model compression device as described in the second aspect. The model compression device compresses the first network model to obtain a target network model, and sends the target network model to the hardware platform, which runs the target network model.

[0037] In a fourth aspect, the present application provides a vehicle comprising a body and a processor, wherein the processor is used to run the target network model output by the model compression device as described in the second aspect.

[0038] Exemplary vehicles include new energy vehicles, electric vehicles, cars, trucks, motorcycles, buses, lawn mowers, recreational vehicles, amusement park vehicles, construction equipment, trams, golf carts, trains, and the like, without particular limitation in this application. These vehicles may be powered by gasoline, diesel, electricity, solar energy, or hydrogen energy.

[0039] In a fifth aspect, the present application provides a chip system comprising at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions, and the at least one processor is used to execute the method of the first aspect and any one of the embodiments of the first aspect.

[0040] In a sixth aspect, the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on a computer, the computer executes the method of the first aspect and any one of the embodiments of the first aspect.

[0041] The technical effects corresponding to the second to sixth aspects and any implementation method of each aspect can be referred to the technical effects corresponding to the above-mentioned first aspect and any implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is a schematic diagram of a processing flow of model compression provided by an embodiment of the present application;

[0043] FIG2 is a schematic diagram of a model compression system architecture according to an embodiment of the present application;

[0044] FIG3 is a second schematic diagram of the model compression system architecture provided in an embodiment of the present application;

[0045] FIG4 is a schematic diagram of the hardware structure of a model compression device provided in an embodiment of the present application;

[0046] FIG5 is a schematic diagram of the vehicle structure provided in an embodiment of the present application;

[0047] FIG6 is a flow chart of a model compression method according to an embodiment of the present application;

[0048] FIG7 is a second flow chart of the model compression method provided in an embodiment of the present application;

[0049] FIG8 is a schematic diagram of the architecture of the distillation model accuracy provided in an embodiment of the present application;

[0050] FIG9 is a schematic diagram of a model compression scenario provided in an embodiment of the present application;

[0051] FIG10 is a second schematic diagram of a model compression scenario provided in an embodiment of the present application;

[0052] FIG11 is a schematic structural diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in this article is merely a way to describe the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0054] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the quantity of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features.

[0055] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more. In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.

[0056] In some examples, mixed-precision quantization is performed on convolutional neural networks based on differentiable neural architecture search. Specifically, architecture search and quantization techniques are combined, using differentiable search algorithms (such as neural architecture search (NAS)) to automatically discover the optimal neural network structure, and using mixed-precision quantization techniques to reduce the computational complexity of the neural network.

[0057] As shown in Figure 1, there is a quantization layer between each data node (V1, V2...Vn-1, Vn), and each quantization layer can use different compression methods. For example, there is a quantization layer between data node V1 and data node V2, and this quantization layer has k compression methods, each of which is abstracted into an edge. For example, the compression method between V1 and V2 is: e1 1,2 , e2 1,2 ,……ek 1,2 , the compression method between Vn-1 and Vn is: e1 n-1,n , e2 n-1,n ,……ek n-1,n . And set a mask for each compression method. Each time a search is performed, only one compression method's mask will be 1, and the rest will be 0. When performing a structural search, each quantization layer selects a compression method for training to obtain the loss value (loss, L) of this training. This loss value is related to the training weight (ω) and the training parameter (θ). According to the loss value of this training, the training parameter (θ) is adjusted so that a new round of training can be performed based on the adjusted training parameters. This process continues until the loss value of the nth training meets the convergence condition and the training is terminated. During each training, the probability of each compression method used in this training is calculated according to Formula 1 and the training parameter (θ). After the training is completed, for each quantization layer, the compression method with the highest probability is selected to form a quantization model.

[0058] Wherein, i and j are data node numbers, i and j are positive integers, ij represents the quantization layer between data node Vi and data node Vj; k is the compression method number, k is a positive integer; Represents the probability that the quantization layer between data node Vi and data node Vj adopts quantization method k; GumbelSoftmax represents the probability distribution based on Gumbel distribution and Softmax activation function; Indicates the training parameters of the quantization layer between data node Vi and data node Vj using quantization method k, θ is the training parameter; τ is the number of training iterations; is a random value that follows the Gumbel distribution.

[0059] In other examples, training is performed based on a neural network model. For each quantization layer, a variable representing the bit length (also described as the number of bits) is introduced, so that the neural network model automatically learns and trains the quantization accuracy of each quantization layer based on the variable. The variable can be an integer or a decimal, and is used to represent the bit length of each weight and activation function. In order to maintain continuity, the interpolation method of Formula 2 is used to represent non-integer bit lengths. Q r (V, b + α) = (1 - α) Q i (V, b) + αQ i (V, b+1) Formula 2

[0060] Among them, V is a floating point number, b and represents the bit length / bit number used for quantization, b is an integer, α is a decimal, 0<α<1, r is a real number, i is an integer, Q represents the quantized data, Q r Represents the value quantized with real bit length, Q i Indicates the value quantized with integer bit length, Q r (V, b+α) is the value after the floating point number V is quantized using the real bit length b+α, Q i (V, b) is the value after the floating point number V is quantized using the integer bit length b, Q i (V, b+1) is the value after the floating point number V is quantized using an integer bit length of b+1.

[0061] In the above examples, searching for compression methods or quantization accuracy based on a neural network results in long training times and high training costs. Furthermore, this approach only generates an applicable quantization model based on neural network training, without considering the performance of this quantization model in actual hardware applications. The quantization model generated based on network training cannot ensure that the appropriate quantization method is used at each quantization layer in the quantization model, failing to fully utilize the quantization model's performance. Instead, it can lead to problems such as reduced accuracy, performance loss, or quantization regression.

[0062] In order to solve the technical problems described above, an embodiment of the present application provides a model compression method. The method includes: obtaining a first network model; generating a first candidate population based on the probability of each quantization layer in the first network model using different compression methods, the first candidate population including m first candidate network models, where m is a positive integer; based on the hardware platform, performing tuning processing on the first candidate population to determine the target network model, and the target network model meets the preset conditions. The method provided in the embodiment of the present application, in the process of compressing the model, determines the target network model based on hardware platform tuning, selects a suitable compression method for each quantization layer, and gives full play to the model performance while ensuring the accuracy of the model.

[0063] The model compression method in the embodiment of the present application can be applied to application scenarios in different fields where the neural network model needs to be compressed, such as application scenarios where the neural network model is quantized or application scenarios where the neural network model is structured and sparse. For example, edge computing, privacy protection, cloud reasoning, network transmission, autonomous driving, application development and other scenarios. For example, in the fields of autonomous driving or artificial intelligence, when running a neural network model on a device or system, it is compressed to reduce computing overhead, so that the neural network model can run efficiently on a mobile device. The embodiment of the present application does not impose any special restrictions on the application scenarios of model compression.

[0064] The model compression method in the embodiment of the present application can be applied to various devices and / or various systems using neural network models, quantify the neural network models in various devices and / or various systems, reduce the resource requirements of the neural network model, improve the execution efficiency of the neural network model, and operate efficiently. Various devices may include various means of transportation such as new energy vehicles, electric vehicles, buses, and cars. It may also include various electronic devices such as mobile phones, tablet computers, personal computers (PCs), wearable devices, etc. Various systems may include electronic equipment systems, smart home systems, edge computing systems, cloud service systems, autonomous driving systems, Internet of Things systems, and other systems. The embodiments of the present application do not impose any special restrictions on the specific forms of devices and systems.

[0065] For example, referring to FIG2 , FIG2 shows a schematic diagram of the architecture of a model compression system 20 provided in an embodiment of the present application. As shown in FIG2 , the model compression system 20 includes a model compression device 21 and a hardware platform 22. The model compression device 21 and the hardware platform 22 are connected and communicate with each other.

[0066] In an embodiment of the present application, the model compression device 21 shown in Figure 2 is used to execute the model compression method provided by the present application to send the compressed model to the hardware platform 22 and run it on the hardware platform 22.

[0067] It is understood that the model compression device 21 is a server. The server can be a Linux server, a Windows server, or other server device that can provide simultaneous access to multiple devices. It can also be a server cluster consisting of multiple regions, multiple computer rooms, and multiple servers. The hardware platform 22 can be a means of transportation such as a new energy vehicle, electric vehicle, bus, or car. The hardware platform 22 can also be an electronic device such as a mobile phone, tablet computer, personal computer, wearable device, handheld terminal, smart speaker, etc.

[0068] Exemplarily, the model compression device 21 may be connected to the hardware platform 22 via a wired network or a wireless network, such as a local area network, a cellular network, or wireless fidelity (Wi-Fi).

[0069] Specifically, the model compression device 21 is configured to obtain a first network model; generate a first candidate population based on the probability of using different compression methods for each quantization layer in the first network model; and perform tuning processing on the first candidate population based on the hardware platform to determine a target network model. The first candidate population includes m first candidate network models, where m is a positive integer; and the target network model satisfies a preset condition.

[0070] In the embodiment of the present application, the model compression device 21 sends the target network model to the hardware platform 22, and the hardware platform 22 runs the target network model.

[0071] It should be understood that the hardware platform 22 deployed with the above-mentioned compressed model can use the compressed model to perform tasks such as image recognition, target detection, and behavior prediction. Compared with the model before compression, the hardware platform 22 uses the compressed model to perform the above-mentioned tasks, and its model inference time can be reduced, thereby improving the model operation efficiency, giving full play to the model performance, saving resources and costs, and also improving the decision-making speed and response time of the hardware platform 22.

[0072] For example, taking the autonomous driving system as an example, in the field of autonomous driving, the autonomous driving system can quickly identify obstacles through compressed models, etc., which can enhance the safety and reliability of the autonomous driving system.

[0073] Optionally, the model compression device 21 may be a server. As an example, the model compression device 21 may be a server of an intelligent transportation system, such as a physical server or a cloud server, which is not limited in this embodiment of the present application.

[0074] Optionally, hardware platform 22 may be an intelligent driving computing platform. This platform implements intelligent driving, decision-making, planning, and control functions and is a core component of the entire vehicle. This platform interacts with various components in the vehicle, acquiring real-time data from each component and controlling its operation. This is not a limitation in the present embodiment.

[0075] For example, the cloud server obtains the neural network model (i.e., the first network model) used in the vehicle to perform image recognition tasks, processes the first network model according to the model compression method provided in the embodiment of the present application, obtains the target network model, and sends the target network model to the intelligent driving computing platform. The intelligent driving computing platform runs the target network model to speed up the vehicle's image recognition speed. If an obstacle is identified, the vehicle can quickly avoid the obstacle and avoid danger, which can increase the safety of vehicle driving and improve the user experience.

[0076] For example, referring to Figure 3, Figure 3 shows a schematic diagram of the architecture of another model compression system 30 provided in an embodiment of the present application. As shown in Figure 3, model compression system 30 may include an acquisition module 31, a first processing module 32, and a second processing module 33. The acquisition module 31, the first processing module 32, and the second processing module 33 are connected and communicate with each other.

[0077] In the embodiment of the present application, the acquisition module 31 is used to acquire the first network model.

[0078] In the embodiment of the present application, the first processing module 32 is configured to generate a first candidate population based on the probability of each quantization layer in the first network model using different compression methods, wherein the first candidate population includes m first candidate network models, where m is a positive integer.

[0079] In the embodiment of the present application, the second processing module 33 is used to perform an optimization process on the first candidate population based on the hardware platform to determine a target network model, wherein the target network model meets a preset condition.

[0080] The modules in the above-mentioned model compression system are divided according to functional logic, and may actually be divided in other ways. In addition, the above-mentioned modules can be named by other names. In addition, each module can be implemented by hardware, or by software, or by a combination of hardware and software. Whether a specific module is implemented in hardware, software, or a combination of hardware and software depends on the specific application and design constraints of the technical solution. Different modules can be implemented by different hardware, and multiple modules can also be implemented by the same hardware. The embodiments of the present application do not specifically limit this.

[0081] It is understandable that the model compression system 20 and the model compression system 30 in the above examples are possible system architectures of the model compression system. The embodiments of the present application do not limit the specific implementation of the system architecture of the model compression system.

[0082] It can be understood that the system architecture and business scenarios described in this application are intended to more clearly illustrate the technical solutions of this application, and do not constitute the sole limitation on the technical solutions provided by this application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by this application are also applicable to similar technical problems.

[0083] Refer to FIG4 , which shows a hardware structure of a model compression device 21 provided in an embodiment of the present application.

[0084] As shown in Fig. 4 , the model compression device 21 includes a processor 41, a memory 42, a communication interface 43, and a bus 44. The processor 41, the memory 42, and the communication interface 43 may be connected via the bus 44.

[0085] The processor 41 is used to manage and control the model compression device 21 and / or to execute the model compression method described below. The memory 42 is used to store program code and data of the model compression device 21. The communication interface 43 is used to support communication between the model compression device 21 and other network entities.

[0086] Among them, the above-mentioned processor 41 (or described as a controller) is the control center of the model compression device 21, and can implement or execute various exemplary logic blocks, unit modules and circuits described in combination with the contents disclosed in this application. The processor or controller can be a general-purpose central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc. The processor 41 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc.

[0087] As an example, the processor 41 may include one or more CPUs, such as CPU 0 and CPU 1 shown in FIG. 4 .

[0088] The memory 42 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or a flash memory, a hard disk or a solid-state drive; it can also be an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 42 can also include a combination of the above-mentioned types of memories. In one possible implementation, the memory 42 can exist independently of the processor 41. The memory 42 can be connected to the processor 41 via a bus 44 for storing data, instructions or program codes. When the processor 41 calls and executes the instructions or program codes stored in the memory 42, the model compression method provided in the embodiment of the present application can be implemented.

[0089] In another possible implementation, the memory 42 may also be integrated with the processor 41 .

[0090] Communication interface 43 is used to connect model compression device 21 to other devices via a communication network. The communication network can be a transceiver circuit, Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. Communication interface 43 can include a receiving unit for receiving data and a sending unit for sending data.

[0091] Bus 44 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG4 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.

[0092] It should be pointed out that the structure shown in FIG4 does not constitute a limitation on the model compression device 21. In addition to the components shown in FIG4, the model compression device 21 may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0093] FIG5 is a schematic diagram of the structure of a vehicle 500 provided in an embodiment of the present application. Referring to FIG5 , vehicle 500 may include various subsystems, such as a travel system 510, a sensor system 520, a control system 530, one or more peripheral devices 540, a power supply 550, a computer system 560, and a user interface 570. Optionally, vehicle 500 may include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 500 may be interconnected via wired or wireless connections.

[0094] Propulsion system 510 may include components that provide powered motion for vehicle 500. Engine 511 may be an electric motor or other types of engine combinations. Engine 511 converts energy source 512 into mechanical energy. Examples of energy source 512 include solar panels, batteries, and other sources of electricity. Transmission 513 may transmit mechanical power from engine 511 to wheels 514.

[0095] The sensor system 520 may include a number of sensors that sense information about the environment surrounding the vehicle 500. For example, the sensor system 520 may include a positioning system 521, such as a global positioning system (GPS), a BeiDou system, or other positioning systems, an inertial measurement unit (IMU) 522, a radar 523, a laser rangefinder 524, and a camera 525.

[0096] Control system 530 controls the operation of vehicle 500 and its components. Control system 530 may include various components, including a steering system 531, a throttle 532, a brake unit 533, a computer vision system 534, a path control system 535, and an obstacle avoidance system 536, which may also be referred to as an obstacle avoidance system.

[0097] Vehicle 500 interacts with external sensors, other vehicles, other computer systems, or users via peripheral devices 540. Peripheral devices 540 may include a wireless communication system 541, an onboard computer 542, a microphone 543, and / or a speaker 544.

[0098] Power source 550 may provide power to various components of vehicle 500 .

[0099] Some or all functions of vehicle 500 are controlled by computer system 560. Computer system 560 may include at least one processor 561 that executes instructions 5621 stored in a non-transitory computer-readable medium such as memory 562. Computer system 560 may also be a plurality of computing devices that control individual components or subsystems of vehicle 500 in a distributed manner.

[0100] The processor 561 may be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, the processor may be a dedicated device such as an application-specific integrated circuit (ASIC) or other hardware-based processor.

[0101] In some embodiments, memory 562 may include instructions 5621 (e.g., program logic) that are executable by processor 561 to perform various functions of vehicle 500. Memory 562 may also include additional instructions, including instructions for sending data to, receiving data from, interacting with, and / or controlling one or more of travel system 510, sensor system 520, control system 530, and peripherals 540.

[0102] In addition to instructions 5621, memory 562 may also store data such as road maps, route information, the vehicle's location, direction, speed, and other vehicle data, and other information. This information may be used by vehicle 500 and computer system 560 during operation of vehicle 500 in autonomous, semi-autonomous, and / or manual modes.

[0103] The user interface 570 is used to provide information to or receive information from a user of the vehicle 500 .

[0104] Computer system 560 may control functions of vehicle 500 based on input received from various subsystems (eg, travel system 510 , sensor system 520 , and control system 530 ) and from user interface 570 .

[0105] In some embodiments, vehicle 500 may also include a vehicle controller (not shown in FIG. 5 ), which can also be described as a powertrain controller or intelligent driving computing platform. It is the core control component of the entire vehicle. It collects input information from various systems and components, makes decisions based on this input, and controls the operation of various components in vehicle 500, driving vehicle 500.

[0106] Specifically, as the command and management center for vehicle 500, the vehicle controller (VCU) performs key functions, including driving torque control, optimized braking energy control, vehicle energy management, maintenance and management of the controller area network (CAN), fault diagnosis and troubleshooting, and vehicle status monitoring. It controls vehicle operation. Therefore, the quality of the VCU directly determines the stability and safety of the vehicle.

[0107] Alternatively, one or more of the above components may be installed or associated separately from the vehicle 500. For example, the memory 562 may be partially or completely separate from the vehicle 500. The above components may be communicatively coupled together in a wired and / or wireless manner.

[0108] Optionally, the above components are only an example. In actual applications, the components in the above modules may be added or deleted according to actual needs. Figure 5 should not be understood as a limitation on the embodiments of the present application.

[0109] The vehicle 500 may be a new energy vehicle, an electric vehicle, a car, a truck, a motorcycle, a bus, a boat, an airplane, a helicopter, a lawn mower, an amusement vehicle, an amusement park vehicle, construction equipment, a tram, a golf cart, or a train, etc., and is not particularly limited in this embodiment of the present application. The vehicle may be powered by gasoline, diesel, electricity, solar energy, hydrogen energy, etc.

[0110] In other embodiments of the present application, the vehicle may further include hardware structures and / or software modules to implement the aforementioned functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular one of the aforementioned functions is implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.

[0111] The method provided in the embodiments of the present application is described below with reference to the accompanying drawings.

[0112] To optimize the model compression strategy, an appropriate compression method is selected for each quantization layer of the model to maximize model performance. This application proposes a model compression method. The execution subject of this method can be a vehicle or other devices outside the vehicle, such as mobile phones, computers, and other electronic devices. It can also be a processor on the vehicle or other devices outside the vehicle, such as processor 41 or processor 561 mentioned above.

[0113] The present application embodiment is introduced using an autonomous driving scenario as an example. Referring to FIG6 , FIG6 shows a flow chart of a model compression method provided by the present application embodiment. The model compression method includes the following steps S601-S603:

[0114] S601: The server obtains a first network model.

[0115] In an embodiment of the present application, the server is used to train, optimize, and compress a first network model to be run on a vehicle. The server sends the trained compressed model to the vehicle, which runs the compressed model to implement the corresponding function.

[0116] In an embodiment of the present application, the first network model is a neural network model that can be run on the vehicle to implement vehicle perception, decision-making, and control functions. The neural network model can be a convolutional neural network (CNN), a recurrent neural network (RNN), a feedforward neural network (FNN), etc. The neural network model can be a pre-configured neural network model in the vehicle or a neural network model obtained from an open source platform or cloud service.

[0117] For example, the first network model can be a neural network model for object detection and recognition. It assists the driver or vehicle in perceiving the surrounding environment by identifying environmental information such as pedestrians, vehicles, and traffic signs. The first network model can also be a neural network model for behavior prediction. By analyzing the vehicle's surroundings (such as obstacles) and the behavior of other vehicles, it predicts their movement trajectory and intentions, allowing the vehicle to make rational decisions and avoid danger. The first network model can also be a neural network model for route planning and navigation. By analyzing information such as maps, traffic, and user needs, it selects the optimal driving route to improve driving efficiency and comfort. The first network model can also be a neural network model for detecting driver status. By analyzing the driver's physiological and behavioral characteristics, it monitors the driver's status, prevents fatigue, and provides timely driver alerts to improve driving safety. The first network model can also be a neural network model for vehicle diagnostics and predictive maintenance. By analyzing vehicle sensor data and driving records, it diagnoses the vehicle's operating status and potential faults to reduce the risk of vehicle failure and optimize vehicle performance. The first network model can also be a neural network model for voice recognition and human-computer interaction. By acquiring voice commands, it assists the driver in interacting with the vehicle system platform or other user interactions, thereby enhancing vehicle intelligence.

[0118] It is understandable that the embodiments of the present application do not specifically limit the type of the first network model and the functions that can be implemented by the first network model.

[0119] In one possible implementation, the server obtains a first network model to be run in the vehicle. The first network model may be a neural network model to be configured in the vehicle.

[0120] In some examples, the server obtains a neural network model that matches the function to be implemented by the vehicle.

[0121] S602: The server generates a first candidate population based on the probability of each quantization layer in the first network model using different compression methods. The first candidate population includes m first candidate network models, where m is a positive integer.

[0122] In the embodiment of the present application, the quantization layer is a network layer in the first network model that requires a large amount of calculation, such as a convolutional layer, a fully connected layer, etc. The embodiment of the present application does not limit the quantization layer.

[0123] In the embodiments of the present application, compression methods include quantization and sparsification. For example, quantization includes converting 32-bit floating point numbers into 16-bit floating point numbers (FP16), converting 32-bit floating point numbers into 8-bit integer format (INT8), and converting 32-bit floating point numbers into 4-bit integer format (INT4), among other quantization methods. Sparsification includes 4:2 structured sparsification, etc. The embodiments of the present application do not limit the compression method.

[0124] In an embodiment of the present application, the first candidate network model is a network quantization model in which each quantization layer in the first network model uses a compression method. Optionally, the compression methods used by each quantization layer in the first candidate network model are the same or different.

[0125] In the embodiment of the present application, the first candidate population is a set of different network quantization models composed of different quantization methods used in each quantization layer of the first network model. That is, the first candidate population includes multiple first candidate network models.

[0126] Optionally, as shown in FIG7 , the probability of the server using different compression methods based on each quantization layer in the first network model in step S602 may be specifically implemented as follows: S6021-S6022:

[0127] S6021. The server determines a quantization layer in the first network model.

[0128] In an embodiment of the present application, after obtaining the first network model, the server analyzes factors such as the role, number of parameters, parameter distribution, importance and network structure of each network layer in the first network model to determine the quantization layer in the first network model.

[0129] For example, taking the first network model including multiple network layers as an example, if the number of parameters of a certain network layer (such as a convolutional layer) is large and the data distribution is concentrated, then the network layer is determined to be a quantization layer.

[0130] It is understandable that the embodiments of the present application do not limit the specific implementation method of the server determining the quantization layer.

[0131] S6022. The server determines the probability of each quantization layer in the first network model using a different compression method.

[0132] In an embodiment of the present application, the server calculates the output data distribution of each quantization layer in the above-mentioned first network model using different compression methods; for each of the above-mentioned quantization layers, the server respectively calculates the similarity between the output data distribution of the quantization layer using different compression methods and the output data distribution of the quantization layer not using compression methods, and based on the similarity, calculates the probability of the quantization layer using different compression methods.

[0133] In one possible implementation, for each quantization layer, the server sets all other quantization layers as non-quantization layers and infers the output data distribution of the quantization layer using different compression methods and the output data distribution of the quantization layer without compression. For the same quantization layer, the server calculates the similarity between the output data distribution of the quantization layer using different compression methods and the output data distribution of the quantization layer without compression, and based on the similarity, calculates the probability corresponding to each compression method that can be used for the quantization layer.

[0134] In some examples, the similarity is calculated using relative entropy (which can also be described as KL divergence), and the probability of each compression method is calculated based on the similarity.

[0135] For example, the similarity is calculated using the following formula 3, and the probability of each compression method is calculated using formula 4.

[0136] Among them, D KL is the KL divergence, D KL (p||q) represents the KL divergence of probability distribution p and probability distribution q, x is the output data, p(x) is the probability distribution of the output data using the compression method, and q(x) is the probability distribution of the output data without using the compression method.

[0137] In the embodiment of the present application, the smaller the KL divergence, the more similar the data distribution, indicating that the subsequent impact of the compression method on the model accuracy is smaller, and the probability of selecting the quantization method is higher.

[0138] Among them, K represents the quantization method, p i represents the probability of the i-th quantization method, D i Represents the KL divergence of the i-th quantization method, where i is an integer.

[0139] In this application, for each quantization layer in the first network model, when analyzing and inferring the output data distribution of the quantization layer using different compression methods, only the quantization layer uses the compression method, and the other quantization layers are not compressed, and the output data distribution of the quantization layer using the different compression methods is inferred. The similarity of the output data distribution of each quantization layer using the different compression methods and the output data distribution of the quantization layer not using the compression method is calculated, and the probability of each quantization layer using the different quantization methods is determined based on the similarity.

[0140] It should be understood that by using the above method for each quantization layer, only the quantization method variable is changed during inference, which can improve the accuracy of the output data distribution and probability corresponding to different compression methods used in each quantization layer, so that a better quantization method can be selected for each quantization layer subsequently.

[0141] It is understandable that the embodiments of the present application do not limit the specific implementation method of the server determining the probability of using different compression methods for each quantization layer.

[0142] In an embodiment of the present application, after determining the probability that each quantization layer in the first network model uses a different compression method, the server generates a first candidate population based on the probability that each quantization layer in the first network model uses a different compression method.

[0143] In one possible implementation, the server generates multiple first candidate network models based on the probability of each quantization layer in the above-mentioned first network model using different compression methods; the server selects the top m first candidate network models from the above-mentioned multiple first candidate network models to form a first candidate population.

[0144] Specifically, the server integrates the different quantization methods used by different quantization layers based on the probability of each quantization layer in the above-mentioned first network model using different compression methods to obtain multiple first candidate network models, and sorts the above-mentioned multiple first candidate network models according to probability, and selects the top m first candidate network models to form a first candidate population.

[0145] Exemplarily, the first network model includes two quantization layers, namely quantization layer 1 and quantization layer 2. Each quantization layer can use two compression methods, namely compression method 1 and compression method 2. The probability that quantization layer 1 uses compression method 1 is p1, the probability that quantization layer 1 uses compression method 2 is p2, the probability that quantization layer 2 uses compression method 1 is p3, and the probability that quantization layer 2 uses compression method 2 is p4. Arbitrary combination of different compression methods for each quantization layer to obtain multiple first candidate network models, and calculate the probability of each first candidate network model. And sort the multiple first candidate network models according to the probability, and select the top two to form the first candidate population. For example, the first candidate population includes the first candidate network model 1 and the first candidate network model 3.

[0146] In this application, the first candidate network model with the highest ranking is selected to form the first candidate population. The compression method used in the quantization layer of each candidate network model in the first candidate population is a compression method with a higher probability, so that the appropriate compression method for each quantization layer can be quickly determined subsequently to give full play to the model performance.

[0147] It is understandable that the embodiment of the present application does not limit the specific implementation method of the server determining the first candidate population.

[0148] S603: The server performs an optimization process on the first candidate population based on the hardware platform to determine a target network model, where the target network model meets a preset condition.

[0149] In the embodiments of the present application, the hardware platform is a platform for running the target network model. For example, the hardware platform is a vehicle's intelligent driving computing platform. A server interacts with the intelligent driving computing platform. The server determines a first candidate population, and the intelligent driving computing platform runs each first candidate network model in the first candidate population. The server obtains the run results and performs optimization processing based on the run results to determine the target network model.

[0150] In an embodiment of the present application, the target network model is a neural network model after compression and optimization of the first network model.

[0151] In the embodiment of the present application, the preset conditions are pre-set constraints that the compressed neural network model should satisfy, such as the accuracy constraints of the compressed neural network model, or the performance constraints of the compressed neural network model.

[0152] In an embodiment of the present application, the server performs multiple rounds of training based on a hardware platform, and performs tuning processing on the above-mentioned first candidate population to determine the target network model.

[0153] In one possible implementation, during the first round of training, the server runs the above-mentioned m first candidate network models based on the hardware platform to obtain a first intermediate network model; wherein, the first intermediate network model is the first candidate network model that meets the preset conditions and has the highest score among the above-mentioned m first candidate network models; during the i-th round of training, the server runs the i-th candidate population obtained after processing the above-mentioned first candidate population for the i-1th time based on the hardware platform to obtain the i-th intermediate network model; wherein, the i-th candidate population of the i-th round of training includes m i-th candidate network models, 2≤i≤j, i and j are positive integers; the server determines the target network model, which is the intermediate network model that meets the preset conditions and has the highest score among the k intermediate network models obtained after k rounds of training, k≤j, k is a positive integer.

[0154] The preset condition is a preset model accuracy that the compressed model should meet. For example, the first intermediate network model is determined from the candidate network models whose model accuracy is higher than the preset condition in the first candidate network models.

[0155] Among them, the first intermediate network model is the neural network model determined in the first round of training, which has the highest score and meets the preset accuracy conditions among the first candidate network models.

[0156] The i-th candidate population is the candidate population used in the i-th round of training. The i-th candidate population is generated by processing the first candidate population for the i-1th time. The number of i-th candidate network models included in the i-th candidate population is the same as the number of first candidate network models included in the first candidate population. The i-th intermediate network model is the neural network model determined in the i-th round of training that meets the preset accuracy conditions and has the highest score among the i-th candidate network models.

[0157] Specifically, during the first round of training, the server runs the above-mentioned m first candidate network models based on the hardware platform to obtain the first intermediate network model. The method includes: during the first round of training, running the above-mentioned m first candidate network models based on the hardware platform to determine the scores of the above-mentioned m first candidate network models; the server determines the precision of the n first candidate network models with the highest scores among the above-mentioned m first candidate network models, 2<n≤m, n is a positive integer; the server determines the first candidate network model with the highest score among the n first candidate network models whose precision meets the preset conditions as the first intermediate network model.

[0158] Specifically, the server runs the above-mentioned m first candidate network models respectively based on the hardware platform to determine the scores of the above-mentioned m first candidate network models. The method includes: running the above-mentioned m first candidate network models respectively based on the hardware platform and collecting performance parameters of the above-mentioned m first candidate network models; the server determines the scores of the above-mentioned m first candidate network models respectively according to the performance parameters of the above-mentioned m first candidate network models.

[0159] Among them, the performance parameters include performance data such as inference latency, memory consumption, and power consumption when the first candidate network model runs on the hardware platform.

[0160] In a possible implementation, each candidate network model is run on a hardware platform, and performance data of each candidate network model during its run is collected using a performance analysis tool of the hardware platform.

[0161] In some examples, after the server determines m first candidate network models, it sends the above m first candidate network models to the hardware platform so that the hardware platform runs the above m first candidate network models respectively to determine the performance parameters of the above m first candidate network models; the server obtains the performance parameters of the above m first candidate network models and determines the scores of the above m first candidate network models.

[0162] It should be understood that in this application, during the process of determining the target network model, candidate network models are directly run on the hardware platform, and the candidate network models are evaluated based on the performance parameters of each candidate network model during operation. Model quantization is optimized based on the model quantization effect and the performance parameters of actual inference to adjust the neural network model to be more suitable for the hardware platform, ensure the model runs well, and improve the model's prediction accuracy and application effect.

[0163] In an embodiment of the present application, after collecting the performance parameters of the m first candidate network models, the server may further analyze the performance parameters of the m first candidate network models to form an optimization strategy, which is used by the server to select the i-th candidate population.

[0164] For example, if it is determined that the delay of using 4:2 structured sparse for the same quantization layer is greater than the delay of using FP16, or the delay of the low-bit quantization method is greater than the delay of the high-bit quantization method, then the 4:2 structured sparse or low-bit quantization method will not be selected for the quantization layer.

[0165] It should be understood that the hardware platform can be used to obtain performance parameters for each quantization layer in each candidate network model. Based on these performance parameters, the quality of the compression method used by each quantization layer can be evaluated. An optimization strategy is then formed based on these performance parameters. When generating candidate populations during subsequent training, the candidate population is pruned according to the optimization strategy, with poorly performing compression methods removed and high-performing ones prioritized. This optimization strategy ensures that the quantization layer in the generated i-th candidate population uses the best-performing compression method.

[0166] It is understandable that the embodiments of the present application do not limit the optimization strategy.

[0167] In a possible implementation, the server determines a score corresponding to the performance parameter of the first candidate network model according to the performance parameter of the first candidate network model, and determines the score of the first candidate network model according to the score of the performance parameter of the first candidate network model.

[0168] It should be understood that in the embodiments of the present application, the smaller the inference delay, consumed memory and power consumption, the higher the corresponding score, indicating that the quantization model is better and the model performance is better.

[0169] In some examples, the score of each first candidate network model is calculated using a model scoring formula (such as the following formula 5). S = ∑ (w1A + w2B + w3C) Formula 5

[0170] Among them, S is the model score, w1, w2, and w3 are the weights corresponding to inference latency, consumed memory, and power consumption, respectively. A is the inference latency score, B is the consumed memory score, and C is the power consumption score.

[0171] It should be understood that in the embodiments of the present application, a higher score of the quantization model indicates better performance of the quantization model and a greater probability of selecting the quantization model.

[0172] It is understandable that the embodiment of the present application does not limit the specific implementation method of the server determining the score of the first candidate network model in the first candidate population.

[0173] In this application, the performance parameters of the model when running on the hardware platform can more comprehensively and accurately evaluate the performance of the model in the real environment, verify the reliability, stability and adaptability of the model, so as to make corresponding adjustments and optimizations, ensure that the model can run well on various hardware devices, improve the adaptability and prediction accuracy of the model, and provide strong support for the practical application of the model.

[0174] In an embodiment of the present application, the server determines the accuracy of the n first candidate network models with the highest scores among the above-mentioned m first candidate network models in the following manner: the server sorts the above-mentioned m first candidate network models according to the scores, and selects the n first candidate network models with the highest scores; the server distills the accuracy of the above-mentioned n first candidate network models according to a preset distillation model; and the server determines the accuracy of the above-mentioned n first candidate network models after distillation.

[0175] Exemplarily, as shown in FIG8 , the first candidate network model includes an input layer, a convolutional layer, and an output layer. The preset distillation model includes an input layer, a convolutional layer, and an output layer. The preset distillation model is an FP16 neural network model using full precision. The accuracy of the output result of the quantization layer (i.e., the convolutional layer) in the preset distillation model is standard accuracy, and the accuracy of the first candidate network model is supervised and guided based on the standard accuracy, so that the updated / distilled first candidate network model meets the standard accuracy.

[0176] It should be understood that when distilling the accuracy of the first candidate network model, there may be a situation where, even if the performance data score of the compression method used in the quantization layer of the first candidate network model is very high, the compression method loses a lot of accuracy, resulting in the accuracy of the first candidate network model after distillation failing to meet the standard accuracy. Therefore, the first candidate network model whose accuracy after distillation does not meet the standard accuracy should be deleted, and the first intermediate network model should be selected from the first candidate network models that meet the standard accuracy.

[0177] In an embodiment of the present application, the server determines the first candidate network model whose accuracy meets the preset conditions and has the highest score among the n first candidate network models as the first intermediate network model.

[0178] In this application, the first intermediate network model is a neural network model whose accuracy meets the preset conditions and has the highest model score. While ensuring the accuracy of the model, selecting the neural network model with the best performance when running on the hardware is conducive to giving full play to the model performance.

[0179] Exemplarily, based on the example of S602 above, the first candidate population includes a first candidate network model 1 and a first candidate network model 3. The two first candidate network models are run separately on the hardware platform, and the performance parameters of each first candidate network model during operation are collected. The score of each first candidate network model is calculated based on the performance parameters. For example, the score of the first candidate network model 1 is S1, and the score of the second candidate network model 3 is S2, where S1 is greater than S2. The accuracies of the first candidate network model 1 and the first candidate network model 3 are then distilled separately according to the preset distillation model. If the accuracies of the first candidate network model 1 and the first candidate network model 3 after distillation both meet the preset conditions, the first intermediate network model is determined to be the first candidate network model 1.

[0180] In this application, the server loads the first candidate population onto the hardware platform and, using the hardware platform's performance analysis tools, collects runtime performance data for each candidate network model in the first candidate population. The server scores each candidate network based on its performance data and determines a first intermediate network model based on its accuracy and score. This ensures the performance and efficiency of the first intermediate network model when running on the hardware platform, optimizing the model compression strategy.

[0181] It is understood that the above example is a specific implementation method for the first round of server training. The training method for subsequent training rounds is the same, but the candidate population for each training round is different. The following details the method for determining the candidate population for each training round.

[0182] In some examples, after each round of training, the candidate population used in the current round is subjected to crossover mutation to obtain the candidate population used in the next round of training.

[0183] Exemplarily, during the first round of training, the server runs the first candidate population based on the hardware platform to determine the first intermediate network model corresponding to the first round; before the second round of training, the server performs cross-mutation on the first candidate population to obtain a second candidate population, and during the second round of training, the server runs the second candidate population based on the hardware platform so that the server determines the second intermediate network model corresponding to the second round of training; before the i-th round of training, the server cross-edits the i-1th candidate population used in the i-1th round of training to obtain the i-th candidate population, and during the i-th round of training, the server runs the i-th candidate population based on the hardware platform so that the server determines the i-th intermediate network model corresponding to the i-th round of training.

[0184] It is understandable that the above description uses the number of training rounds to illustrate how to determine the candidate population used in each round of training. The number of crossover mutations may also be used to indicate the candidate population used in each round of training.

[0185] In some other examples, the i-th candidate population for the i-th round of training is obtained by performing the i-1-th crossover mutation on the first candidate population.

[0186] For example, the first candidate population used in the first round of training has not undergone crossover mutation, the second candidate population used in the second round of training is the candidate population after the first candidate population has undergone one crossover mutation, the third candidate population used in the third round of training is the candidate population after the first candidate population has undergone two crossover mutations, and the i-th candidate population used in the i-th round of training is the candidate population after the first candidate population has undergone i-1 crossover mutations.

[0187] Exemplarily, based on the above example, the first candidate population includes the first candidate network model 1 and the first candidate network model 3, and the first candidate population is cross-mutated to obtain the second candidate network model 1 and the second candidate network model 2 to form the second candidate population.

[0188] In this application, except for the first round of training, the candidate populations used in all subsequent rounds of training are generated based on the candidate population used in the previous round of training. The candidate population in each round of training has better genes than the candidate population in the previous round of training. Through crossover and mutation, the individuals in the candidate population are continuously evolved and improved to determine the optimal candidate network model and give full play to the model performance.

[0189] In an embodiment of the present application, during the i-th round of training, the server runs the i-th candidate population obtained after processing the above-mentioned first candidate population for the i-1th time based on the hardware platform to obtain the i-th intermediate network model, including: during the i-th round of training, running the i-th candidate population based on the hardware platform, and determining the scores of m i-th candidate network models in the i-th candidate population respectively; the server determines the precision of the n i-th candidate network models with the highest scores among the above-mentioned m i-th candidate network models, 2<n≤m, n is a positive integer; the server determines the i-th candidate network model with the highest score among the above-mentioned n i-th candidate network models whose precision meets the preset conditions as the i-th target network model.

[0190] It can be understood that the specific implementation method of the server determining the i-th intermediate network model is described above and will not be repeated here.

[0191] It is understandable that in order to ensure that the model fully learns and adapts to the training data and avoid overfitting or premature stopping, a training end condition is set for the model compression training.

[0192] In the embodiment of the present application, j is the preset maximum number of training times, and k is the number of training times.

[0193] In a possible implementation, the training end condition is that the number of training times reaches a preset number.

[0194] Specifically, when k is equal to j and the number of training times reaches the maximum number of training times, the training is terminated, and the intermediate network model that meets the preset conditions and has the highest score is selected as the target network model from the j intermediate network models obtained after j rounds of training.

[0195] In another possible implementation, the training termination condition is that the difference between the scores of the intermediate network models obtained after two consecutive training sessions meets a preset range, where the preset range is pre-set, such as one thousandth.

[0196] Specifically, when k is less than j, the difference between the score of the intermediate network model determined after the kth round of training and the score of the intermediate network model determined after the k-1th round of training meets the preset range. That is, the difference between the score of the kth intermediate network model obtained in the kth round of training and the score of the k-1th intermediate network model obtained in the k-1th round of training is very small, and the model score will not change even if retrained. Therefore, the training is ended, and from the k intermediate network models obtained after the kth round of training, the intermediate network model that meets the preset conditions and has the highest score is selected as the target network model.

[0197] Exemplarily, as shown in FIG9 , it is a schematic diagram of iterative training to determine the target network model. During the first round of training, the first intermediate network model is determined based on the m first candidate network models. During the second round of training, m second candidate network models are obtained based on the cross-mutation of the m first candidate network models, and then the second intermediate network model is determined based on the m second candidate network models. During the i-th round of training, m i-th candidate network models are obtained based on the cross-mutation of the m i-1th candidate network models used in the i-1th round of training, and then the i-th intermediate network model is determined based on the m i-1th candidate network models. It can be understood that during the i-th round of training, the i-1th intermediate network model is determined based on the m i-1th candidate network models. The difference between the score of the intermediate network model determined at the end of the i-th round of training and the score of the intermediate network model determined at the end of the i-1th round of training meets the preset range, and the training ends. The target network model is determined from the first intermediate network model, the second intermediate network model, and the i-th intermediate network model.

[0198] For example, as shown in FIG10 , which is a specific flow chart for the first round of training, the first candidate population includes m first candidate network models, namely, first candidate network model 1, first candidate network model 2, ..., first candidate network model m. Each first candidate network model is run based on the hardware platform to obtain the corresponding performance parameters of each first candidate network model, namely, performance parameter 1, performance parameter 2, ..., performance parameter m. Then, based on the performance parameters, a score of each first candidate network model is determined, namely, score 1, score 2, ..., score m. Based on the scores, the accuracy of the n first candidate network models with the highest scores is determined, namely, accuracy 1 ..., accuracy n. From the n first candidate network models, the first candidate network model with the highest score and the accuracy that meets the preset conditions is selected as the first intermediate network model.

[0199] In this application, the training is completed in the above manner, and the intermediate network model obtained from each training that meets the preset conditions and has the highest score is selected as the target network model. The target network model is the neural network model with the best performance when running on the hardware platform when the model accuracy meets the preset conditions. Each quantization layer uses the optimal compression method to fully utilize the model performance.

[0200] In this application, the above-mentioned method is used to determine the probability of different quantization methods for each quantization layer through reasoning, thereby optimizing the construction of the quantization model; based on the selection of compression methods for hardware performance optimization, the quantization model is intuitively evaluated through performance data, and the compression methods of the quantization layers in the quantization model are optimized; and the quantization model is distilled layer by layer using a distillation model, which can achieve higher model accuracy using less data and reduce training time. The model compression method provided in this application optimizes the model compression strategy, can ensure that the optimal compression method is selected for each quantization layer, and give full play to the model performance.

[0201] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0202] An embodiment of the present application also provides a chip system, as shown in Figure 11, the chip system 1100 includes at least one processor 1101 and at least one interface circuit 1102. As an example, when the chip system 1100 includes one processor and one interface circuit, the one processor may be the processor 1101 shown in the solid box in Figure 11 (or the processor 1101 shown in the dotted box), and the one interface circuit may be the interface circuit 1102 shown in the solid box in Figure 11 (or the interface circuit 1102 shown in the dotted box). When the chip system 1100 includes two processors and two interface circuits, the two processors include the processor 1101 shown in the solid box in Figure 11 and the processor 1101 shown in the dotted box, and the two interface circuits include the interface circuit 1102 shown in the solid box in Figure 11 and the interface circuit 1102 shown in the dotted box. This is not limited.

[0203] The processor 1101 and the interface circuit 1102 can be interconnected via a line. For example, the interface circuit 1102 can be used to receive signals. For another example, the interface circuit 1102 can be used to send signals to other devices (such as the processor 1101). For example, the interface circuit 1102 can read instructions stored in a memory and send the instructions to the processor 1101. When the instructions are executed by the processor 1101, the firewall device can execute the various steps in the above-mentioned embodiment. Of course, the chip system can also include other discrete components, which are not specifically limited in the embodiments of the present application.

[0204] Exemplarily, the chip system can be a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system on a chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD) or other integrated chip.

[0205] It should be understood that each step in the above method embodiment can be completed by hardware integrated logic circuits in a processor or by software instructions. The method steps disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in a processor.

[0206] An embodiment of the present application further provides a computer-readable storage medium storing one or more computer programs, wherein the one or more computer programs include instructions that, when executed by a computer, enable the computer to execute the corresponding process of the model compression method in the above embodiment.

[0207] In some embodiments, the disclosed methods may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of manufacture.

[0208] An embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement the model compression method in the above-mentioned embodiment.

[0209] The apparatus, computer-readable storage medium, computer program product, or chip provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0210] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A model compression method, characterized in that, The method includes: Obtain a first network model; Based on the probabilities of using different compression methods for each quantization layer in the first network model, generate a first candidate population, where the first candidate population includes m first candidate network models, and m is a positive integer; Based on the hardware platform, perform tuning processing on the first candidate population to determine a target network model that meets the preset conditions.

2. The method according to claim 1, wherein The performing tuning processing on the first candidate population based on the hardware platform to determine a target network model includes: In the first round of training, run the m first candidate network models respectively based on the hardware platform to obtain a first intermediate network model, where the first intermediate network model is the first candidate network model that meets the preset conditions and has the highest score among the m first candidate network models; In the i-th round of training, run the i-th candidate population obtained after processing the first candidate population for the (i - 1)-th time based on the hardware platform to obtain an i-th intermediate network model, where the i-th candidate population in the i-th round of training includes m i-th candidate network models, 2 ≤ i ≤ j, and i and j are positive integers; Determine the target network model, where the target network model is the intermediate network model that meets the preset conditions and has the highest score among the k intermediate network models obtained after k rounds of training, k ≤ j, and k is a positive integer.

3. The method according to claim 2, wherein The running the m first candidate network models respectively based on the hardware platform in the first round of training to obtain a first intermediate network model includes: In the first round of training, run the m first candidate network models respectively based on the hardware platform to determine the scores of the m first candidate network models; Determine the accuracy of the n first candidate network models with the highest scores among the m first candidate network models, 2 < n ≤ m, and n is a positive integer; Determine the first candidate network model with the highest score among the n first candidate network models whose accuracy meets the preset conditions as the first intermediate network model.

4. The method according to claim 2 or 3, characterized in that, The method further includes: Obtain the i-th candidate population for the i-th round of training by performing (i - 1)-th cross mutation on the first candidate population.

5. The method according to claim 3, wherein The running the m first candidate network models respectively based on the hardware platform to determine the scores of the m first candidate network models includes: Run the m first candidate network models respectively based on the hardware platform and collect the performance parameters of the m first candidate network models; Based on the performance parameters of the m first candidate network models, determine the scores of the m first candidate network models respectively.

6. The method according to claim 5, wherein After collecting the performance parameters of the m first candidate network models, the method further includes: Analyze the performance parameters of the m first candidate network models to form an optimization strategy, where the optimization strategy is used to select the i-th candidate population.

7. The method according to any one of claims 1 to 6, characterized in that The generating a first candidate population based on the probabilities of using different compression methods for each quantization layer in the first network model includes: Determine the probabilities of using different compression methods for each quantization layer in the first network model; Generate a plurality of first candidate network models based on the probabilities of using different compression methods for each quantization layer in the first network model; Select the top m first candidate network models from the plurality of first candidate network models to form the first candidate population.

8. The method according to claim 7, wherein The determining the probabilities of using different compression methods for each quantization layer in the first network model includes: Calculate the output data distributions of using different compression methods for each quantization layer in the first network model; For each quantization layer, calculate the similarity between the output data distribution of using different compression methods for the quantization layer and the output data distribution of not using the compression method for the quantization layer, and calculate the probability of using different compression methods for the quantization layer according to the similarity.

9. The method according to any one of claims 2 to 8, characterized in that, When k is less than j, the difference between the score of the intermediate network model determined at the end of the k-th round of training and the score of the intermediate network model determined after the (k - 1)-th round of training satisfies a preset range.

10. A model compression device, characterized in that, Includes: A processor and a memory, the memory is coupled to the processor, the memory is used to store computer-readable instructions, and when the processor reads the computer-readable instructions from the memory, the model compression device executes the method according to any one of claims 1-9.

11. A model compression system, characterized in that, Includes a hardware platform and the model compression device according to claim 10; The model compression device compresses the first network model to obtain a target network model, and sends the target network model to the hardware platform, and the hardware platform runs the target network model.

12. A vehicle, characterized in that, The vehicle includes a vehicle body and a processor, and the processor is used to run the target network model output by the model compression device according to claim 10.

13. A chip system, characterized in that, Includes at least one processor and at least one interface circuit, the at least one interface circuit is used to perform the transceiver function and send instructions to the at least one processor, the at least one processor executes the instructions, and the at least one processor executes the method according to any one of claims 1-9.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program, and when the computer program runs on a computer, the computer executes the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Quantization strategy determination method of neural network and image identification method and device

    CN110348562A

  • Network structure adjustment method, device, storage medium and electronic equipment

    CN113569886A

  • Model compression method and device, storage medium and electronic equipment

    CN115543945A

  • Model compression method, device and system

    CN118153653A

Cited By

  • Model parameter determination method and device, chip, electronic equipment, storage medium and computer program product

    CN120873877A