Global sparsification method and device for convolutional neural network model
By using Latin hypercube sampling algorithm and differential evolution algorithm to optimize the threshold of the convolution kernel channel in the convolution neural network model, the problem of difficulty in achieving global optimization and perceived importance of convolutional layers in the existing technology is solved, and efficient sparseness and precision optimization of the model are achieved.
Patent Information
- Application Number
- CN202510106212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The sparseness scheme of existing convolutional neural network models is difficult to ensure global optimization, and the same comparison threshold is generally set for the convolution kernel, so the importance of different convolutional layers cannot be perceived, making it difficult to achieve the optimal balance of the sparse model accuracy and size.
The Latin hypercube sampling algorithm is used to set different random thresholds for each convolution kernel channel, and the threshold of each convolution kernel channel is iteratively optimized through the differential evolution algorithm until the preset number of iterations or the model size and accuracy meet the constraints, achieving global sparseness.
The global sparseness of the convolutional neural network model is realized, taking into account both the model accuracy and model size, and improving the model's inference ability and energy efficiency.
Smart Images

Figure CN119940428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a global sparsification method and device for a convolutional neural network model. Background Art
[0002] At present, with the rapid development of emerging applications such as AI, the automotive industry, and the Internet of Things, the development of embedded AI MCUs has been stimulated. In order to pursue low costs and cope with complex and fast AI tasks, especially image recognition tasks, edge AI MCU applications have emerged. From the perspective of SOC chip design, the key storage and computing units of the system play a decisive role in chip cost, response speed, energy efficiency and other performance. For example, the target detection convolutional neural network currently deployed in the edge AI MCU is mainly the YOLO series model, which is a popular target detection algorithm known for its fast speed and good performance. The model is small but can simultaneously process target detection and give detection contours. The small YOLO model can reduce the computing and storage costs of the MCU chip. The compressed YOLO model can greatly improve the energy efficiency of the MCU chip, reduce the chip size, and increase the speed of model reasoning.
[0003] In order to achieve compression or sparsification of convolutional neural network models such as the YOLO target detection model, many solutions are adopted in the prior art, such as model pruning, model sparse regularization, etc. However, the main disadvantages of the prior art are: (1) Existing model sparsification schemes mainly focus on how to set thresholds or use strategies to reduce model sparsity. However, the optimization of thresholds basically uses the Bayesian optimization algorithm to find the optimal solution. Although the Bayesian optimization algorithm is fast, it is difficult to ensure global optimality due to its limited modeling capabilities. Therefore, a robust global optimization algorithm is needed to compress the model while taking into account model accuracy. (2) Existing model sparsification schemes generally set the same comparison threshold for the convolution kernels, which makes it difficult to perceive the importance of different convolutional layers of the convolutional neural network model, resulting in unfair convolution kernel thresholds and making it difficult for the sparsified model to achieve global optimization. Summary of the invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a global sparsification method and device for a convolutional neural network model, which can perfectly solve the technical problems existing in the background technology.
[0005] In one aspect of the present invention, a global sparsification method for a convolutional neural network model is provided, comprising: a constraint setting step: setting a constraint condition for model compression for a trained model, wherein the constraint condition includes a target model accuracy and a target model size; a model compression step: determining the number of convolution kernels to be optimized according to the target model size, setting different random thresholds for each convolution kernel channel to be optimized through a Latin hypercube sampling algorithm, performing a sparsification operation on the model based on the random threshold, and calculating the current model accuracy and the current model size. If the current model accuracy and the current model size do not meet the constraint condition, a new threshold for each convolution kernel channel is iteratively evolved through a differential evolution algorithm, and a model sparsification operation is performed based on the new threshold value evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraint condition; wherein the sparsification operation includes deleting the convolution kernel when the number of weights and bias items of the convolution kernel that are zero is greater than the threshold; a loop execution step: if the current model accuracy and the current model size do not meet the constraint condition, the model compression step is repeatedly executed until the current model accuracy and the current model size meet the constraint condition.
[0006] Another aspect of the present invention further provides a global sparseness device for a convolutional neural network model, comprising: a constraint setting module, used to set model compression constraints for a trained model, wherein the constraints include a target model accuracy and a target model size; a model compression module, used to determine the number of convolution kernels to be optimized according to the target model size, set different random thresholds for each convolution kernel channel to be optimized through a Latin hypercube sampling algorithm, perform a sparseness operation on the model based on the random threshold, and calculate the current model accuracy and the current model size. If the current model accuracy and the current model size do not meet the constraints, If a constraint condition is met, a new threshold value of each convolution kernel channel is iteratively evolved through a differential evolution algorithm, and a sparse operation of the model is performed based on the new threshold value evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraint condition; wherein the sparse operation includes deleting the convolution kernel when the number of weights and bias items of the convolution kernel that are zero is greater than the threshold value; and a loop execution module is used to repeatedly execute the steps in the model compression module if the current model accuracy and the current model size do not meet the constraint condition until the current model accuracy and the current model size meet the constraint condition.
[0007] The global sparsification method and device of the convolutional neural network model provided by the present invention first use the Latin hypercube sampling algorithm to initialize multiple convolution kernel thresholds, use different random thresholds to perceive the importance of different layers of the convolutional neural network model, and form relatively fair thresholds. Then, the differential evolution algorithm is used to optimize the threshold in each channel of the convolution kernel to obtain the optimal model size, while ensuring the model accuracy to ensure reasoning ability. The global sparsification method and device of the convolutional neural network model of the present invention have strong global search capabilities and good robustness, and can take into account both model accuracy and model size. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 This is a schematic diagram of a global sparsification method for a convolutional neural network model provided by an embodiment of the present application. Figure 1 ; Figure 2 This is a schematic diagram of a global sparsification method for a convolutional neural network model provided by an embodiment of the present application. Figure 2 ; Figure 3 It is a structural schematic diagram of a global sparsification device for a convolutional neural network model provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0009] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0010] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.
[0011] It should be understood that although the terms first, second, third, etc. may be used to describe the acquisition modules in the embodiments of the present invention, the acquisition modules should not be limited to these terms. These terms are only used to distinguish the acquisition modules from each other.
[0012] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0013] It should be noted that the directional words such as "upper", "lower", "left", and "right" described in the embodiments of the present invention are described at the angles shown in the drawings and should not be understood as limiting the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that an element is formed "on" or "under" another element, it can not only be formed directly "on" or "under" another element, but also be formed "on" or "under" another element indirectly through an intermediate element.
[0014] In order to meet the requirements of high energy efficiency and fast response of the end-side AI MCU, it is necessary to compress the size of convolutional neural network models such as target detection networks, and the most common way to compress models is to perform model sparsification. An embodiment of the present application takes the sparsification of the YOLO v5n target detection network model as an example, and proposes a global sparsification method and device for a convolutional neural network model.
[0015] See also Figure 1 , 2 The global sparsification method of the convolutional neural network model proposed in this embodiment includes the following steps: Step S101, constraint setting step: setting constraints for model compression for the trained model, wherein the constraints include target model accuracy and target model size.
[0016] First, prepare an image dataset based on the application scenario of the YOLO model. For example, collect 200 video data of road driving. Sort the image data in the video into 8,000 data, annotate the learning targets in the data, and form a dataset. . Divide the dataset into training, validation, and test data in a ratio of 8:1:1.
[0017] Next, download the v5n model provided by YOLO. Although the model has been trained, it can be retrained on new tasks. For example, GPU is used for training. The entire environment is based on Pytorch and Python environments. The model evaluation function is set to BCE and CIoU functions. The initial model learning rate is set to 0.0001, the epoch is 10000, and the cosine learning rate reduction strategy is used. In order to prevent overfitting, it is preferred to introduce an early stopping strategy to avoid overfitting. During training, the model is verified on the validation data set and tested on the test data set. The model accuracy is obtained after the weighted loss function. . Save the model as YOLOv5n.pt, the model size is .
[0018] Then, the target model accuracy and target model size are determined, and the optimization task is to minimize the model size while meeting the target model accuracy. For example, set the target F(x) as the model size and the constraint , that is, the maximum decrease in model accuracy is 5%.
[0019] Step S102, model compression step: determine the number of convolution kernels to be optimized according to the target model size, set different random thresholds for each convolution kernel channel to be optimized through the Latin hypercube sampling algorithm, perform a sparse operation on the model based on the random threshold, and calculate the current model accuracy and the current model size. If the current model accuracy and the current model size do not meet the constraints, iteratively evolve new thresholds for each convolution kernel channel through the differential evolution algorithm, and perform a sparse operation on the model based on the new thresholds evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraints; wherein the sparse operation includes deleting the convolution kernel when the number of weights and bias items of the convolution kernel that are zero is greater than the threshold.
[0020] Specifically, a different random threshold is set for each convolution kernel channel to be optimized through the Latin hypercube sampling algorithm, including: selecting the number of convolution kernels to be optimized according to the target model size and experience, or directly setting all convolution kernels as the target to be optimized. Then, a different random threshold is set for each convolution kernel channel to be optimized through the Latin hypercube sampling algorithm. The threshold here represents the comparison threshold of the total number of 0s in the convolution kernel weight and the bias term. Specifically, the threshold is set through the following steps:
[0021] Step 1: Given convolution kernel objects to be optimized, and set the threshold range of the i-th dimension of each convolution kernel to be optimized to [ ]; Step 2: Divide the i-th dimension of each convolution kernel to be optimized into N intervals, and the threshold length of each interval is recorded as: ; Step 3: Set the random sampling probability for each interval of the i-th dimension of each convolution kernel to be optimized ,in, represents the random sampling probability of the Nth interval of the i-th dimension; Step 4: Calculate the sample threshold of the jth interval on the i-th dimension of each convolution kernel to be optimized according to the following formula: ,in, represents the random sampling probability of the jth interval of dimension i, [0,N]; Step 5: The N sample thresholds on each dimension of each convolution kernel to be optimized are respectively used as the thresholds.
[0022] The Latin hypercube sampling algorithm is used to perform super-average initialization on the convolution kernel threshold. Different convolution layers use different random thresholds to perceive the importance of different layers of the model and form a more fair threshold. After setting different thresholds for the convolution kernel to be optimized, the model is trained and the convolution kernel weight is determined. Is the total number of 0 or very close to 0 (0 is an invalid calculation) in the sum bias terms greater than the set threshold? If so, it means that the convolution kernel does not participate in subsequent operations and is deleted (i.e., sparse processing). After the sparse operation, save the current model and calculate the current model accuracy and current model size. If the current model accuracy and current model size meet the constraints, output the optimized model; if the current model accuracy and current model size do not meet the constraints, iteratively optimize the threshold of each convolution kernel channel through the differential evolution algorithm.
[0023] Specifically, the differential evolution algorithm is used to iteratively optimize the threshold of each convolution kernel channel, including the following steps: Step 1: Set the number of generations to be iterated , population size , ,set up For model size, problem orientation is a minimization problem, set constraints Model accuracy, considering convolution kernel objects to be optimized, all convolution kernels constitute the design sample , set the threshold range of each convolution kernel to [ ]; Step 2: Use Latin hypercube sampling algorithm to randomly initialize the population, starting from [ ]Super League average generation Sample threshold, evaluating the size and accuracy of the sparsely-calculated model when each sample threshold is calculated as the threshold; Step three, generate the next generation population through mutation, crossover and selection operations.
[0024] Specifically, the mutation operation is performed on the parent population based on the following formula: in, is a scaling factor between 0 and 2; and are two random sample thresholds; It represents the sample threshold that minimizes the model when the model accuracy constraint is met, or the sample threshold that makes the model closest to the constraint when the model accuracy constraint is not met; Specifically, the crossover operation is performed in the following ways: The dimension of the threshold for the i-th sample , using binomial crossover, using each dimension Generate random numbers to perform mutation: in, Indicates Generation population, CR represents the crossover probability, Represents a uniform random number between 0 and 1. represents a random dimension, Indicates the first Daidi No. The sample threshold of the dimension, Indicates the mutation Daidi No. The sample threshold of the dimension, Indicates the crossover Daidi No. Sample threshold of dimension; Specifically, the selection operation is performed in the following manner: Perform the sparse operation and model evaluation, and select the next generation of population to be evaluated according to the following formula : in, represents the objective function of the model, Indicates the crossover Daidi Sample threshold, Indicates the number before mutation Daidi Sample threshold.
[0025] Step 4: Repeat the population evolution process until the preset number of iterations of the population is reached, or the current model size and the current model accuracy meet the constraints.
[0026] Until the preset number of iterations is reached, or the current model size and the current model accuracy meet the constraints. Set to 60, the scaling factor F is set to 0.7, the crossover probability CR is set to 0.5, and the population generation number G is set to 100. When the model accuracy meets the requirements and the model size reaches When the size is greater than 0.01 (B is the model size), stop the optimization process. Output the pruned model and save the model as YOLOv5n_new.pt.
[0027] Step S103, loop execution step: if the current model accuracy and the current model size do not satisfy the constraint condition, then repeatedly execute the model compression step until the current model accuracy and the current model size satisfy the constraint condition.
[0028] Specifically, the threshold is iteratively optimized by the differential evolution algorithm, and the model may still not meet the constraint condition after reaching the preset number of iterations of the population. In this case, step S102 needs to be re-executed, that is, the initial random threshold is reset using the Latin hypercube sampling algorithm, and then the differential evolution algorithm is used to iteratively optimize the threshold of each channel. In the evolution process of each generation of thresholds, sparse processing will also be performed based on the new threshold until the current model accuracy and the current model size meet the constraint condition.
[0029] Furthermore, the global sparsification method of the convolutional neural network model of this embodiment also includes: a model fine-tuning step for fine-tuning the architecture parameters of the model after the sparsification processing to improve the model accuracy.
[0030] Specifically, fine-tuning is performed based on the model after sparse processing. Since the accuracy of the model will inevitably decrease after sparse processing, fine-tuning the model architecture parameters can further improve the model accuracy. During fine-tuning, the learning rate should be controlled to ensure that the model does not undergo significant changes to further improve the model accuracy. Preferably, an early stopping strategy is used during the model fine-tuning process to prevent fine-tuning overfitting. For example, the model accuracy is finally determined to be , the model size is , save the model size as YOLOv5n_small.pt.
[0031] The model is deployed in an embedded AI MCU for target detection tasks. Although the sparse compressed model is small, it has high accuracy and can reduce the computing requirements of the MCU, thereby alleviating the pressure on MCU chip design, reducing chip size and improving chip energy efficiency.
[0032] The global sparsification method and device of the convolutional neural network model provided by the present invention first use the Latin hypercube sampling algorithm to initialize multiple convolution kernel thresholds, use different random thresholds to perceive the importance of different layers of the convolutional neural network model, and form relatively fair thresholds. Then, the differential evolution algorithm is used to optimize the threshold in each channel of the convolution kernel to obtain the optimal model size, while ensuring the model accuracy to ensure reasoning ability. The global sparsification method and device of the convolutional neural network model of the present invention have strong global search capabilities and good robustness, and can take into account both model accuracy and model size.
[0033] See also Figure 3 Another embodiment of the present invention further provides a global sparsification device 200 for a convolutional neural network model, comprising a constraint setting module 201, a model compression module 202, and a loop execution module 203. The global sparsification device 200 for a convolutional neural network model can execute the global sparsification method for a convolutional neural network model in the method embodiment.
[0034] Specifically, the global sparseness device 200 of the convolutional neural network model includes: A constraint setting module 201 is used to set the constraint conditions of model compression for the trained model, wherein the constraint conditions include target model accuracy and target model size; The model compression module 202 is used to determine the number of convolution kernels to be optimized according to the target model size, set different random thresholds for each convolution kernel channel to be optimized through the Latin hypercube sampling algorithm, perform a sparse operation on the model based on the random threshold, and calculate the current model accuracy and the current model size. If the current model accuracy and the current model size do not meet the constraints, iteratively evolve a new threshold for each convolution kernel channel through the differential evolution algorithm, and perform a sparse operation on the model based on the new threshold evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraints; wherein the sparse operation includes deleting the convolution kernel when the number of weights and bias items of the convolution kernel that are zero is greater than the threshold; The loop execution module 203 is used to repeatedly execute the steps in the model compression module if the current model accuracy and the current model size do not meet the constraint conditions until the current model accuracy and the current model size meet the constraint conditions.
[0035] Furthermore, a model fine-tuning module is included for fine-tuning the architecture parameters of the model after the sparse processing to improve the model accuracy. The model fine-tuning module adopts an early stopping strategy to prevent fine-tuning overfitting.
[0036] It should be noted that the global sparsification device 200 for the convolutional neural network model provided in this embodiment corresponds to a technical solution that can be used to execute each method embodiment. Its implementation principle and technical effect are similar to the method and will not be repeated here.
[0037] See also Figure 4 Another embodiment of the present invention provides a schematic diagram of the structure of an electronic device 300, which is used to implement the global sparsification method of the convolutional neural network model in the method embodiment. The electronic device 300 in the embodiment of the present invention may include but is not limited to a smartphone, a tablet computer, a PC, a laptop computer, and a server. Figure 4 The electronic device 300 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0038] like Figure 4 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303 to implement the method of the embodiment of the present invention. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 305. An input / output (I / O) interface 304 is also connected to the bus 305.
[0039] Typically, the following devices may be connected to the I / O interface 304: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0040] The above description is only a preferred embodiment of the present invention. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but also should cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present invention (but not limited to) to form a technical solution.
Claims
1. A global sparsification method for a convolutional neural network model, characterized in that include: Constraint setting step: setting model compression constraints for the trained model, wherein the constraints include target model accuracy and target model size; Model compression step: determine the number of convolution kernels to be optimized according to the target model size, set different random thresholds for each convolution kernel channel to be optimized through the Latin hypercube sampling algorithm, perform a sparse operation on the model based on the random threshold, and calculate the current model accuracy and the current model size. If the current model accuracy and the current model size do not meet the constraints, iteratively evolve new thresholds for each convolution kernel channel through the differential evolution algorithm, and perform a sparse operation on the model based on the new thresholds evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraints; wherein the sparse operation includes deleting the convolution kernel when the number of weights and bias items of the convolution kernel that are zero is greater than the threshold; Loop execution step: if the current model accuracy and the current model size do not satisfy the constraint condition, then repeatedly execute the model compression step until the current model accuracy and the current model size satisfy the constraint condition.
2. The global sparsification method of a convolutional neural network model according to claim 1, characterized in that: The Latin hypercube sampling algorithm is used to set different random thresholds for each convolution kernel channel to be optimized, including: Set the threshold range of the i-th dimension of each convolution kernel to be optimized to [ ]; The i-th dimension of each convolution kernel to be optimized is divided into N intervals, and the threshold length of each interval is recorded as: ; Set the random sampling probability for each interval of the i-th dimension of each convolution kernel to be optimized : ; in, represents the random sampling probability of the Nth interval of the i-th dimension; The sample threshold of the jth interval on the i-th dimension of each convolution kernel to be optimized is calculated according to the following formula: : ; in, represents the random sampling probability of the jth interval of dimension i, [0,N]; The N sample thresholds on each dimension of each convolution kernel to be optimized are respectively used as the random thresholds.
3. The global sparsification method of a convolutional neural network model according to claim 2, characterized in that: The step of iteratively evolving a new threshold for each convolution kernel channel through a differential evolution algorithm, and performing a sparse operation on the model based on the new threshold evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraint conditions, includes: Generate multiple sample thresholds as parent populations from the threshold range of each convolution kernel through the Latin hypercube sampling algorithm; Perform mutation operation on the parent population based on the following formula: in, is a scaling factor between 0 and 2; and are two random sample thresholds; It represents the sample threshold that minimizes the model when the model accuracy constraint is met, or the sample threshold that makes the model closest to the constraint when the model accuracy constraint is not met; The dimension of the threshold for the i-th sample , using binomial crossover, using each dimension Generate random numbers to perform mutation: in, Indicates Generation population, CR represents the crossover probability, Represents a uniform random number between 0 and 1. represents a random dimension, Indicates the first Daidi No. The sample threshold of the dimension, Indicates the mutation Daidi No. The sample threshold of the dimension, Indicates the crossover Daidi No. Sample threshold of dimension; Perform the sparse operation and model evaluation, and select the next generation of population to be evaluated according to the following formula : in, represents the objective function of the model, Indicates the crossover Daidi Sample threshold, Indicates the number before mutation Daidi Sample threshold; The population evolution process is repeated until the preset number of iterations of the population is reached, or the current model size and the current model accuracy meet the constraints.
4. The global sparsification method of a convolutional neural network model according to claim 3, characterized in that: It also includes a model fine-tuning step, which is used to fine-tune the architecture parameters of the model after sparsification processing to improve the model accuracy.
5. The global sparsification method of a convolutional neural network model according to claim 4, characterized in that: An early stopping strategy is used during model fine-tuning to prevent overfitting.
6. A global sparseness device for a convolutional neural network model, characterized in that include: A constraint setting module, used to set the constraint conditions of model compression for the trained model, wherein the constraint conditions include target model accuracy and target model size; A model compression module, used to determine the number of convolution kernels to be optimized according to the target model size, set different random thresholds for each convolution kernel channel to be optimized through the Latin hypercube sampling algorithm, perform a sparse operation on the model based on the random threshold, and calculate the current model accuracy and the current model size. If the current model accuracy and the current model size do not meet the constraints, iteratively evolve a new threshold for each convolution kernel channel through the differential evolution algorithm, and perform a sparse operation on the model based on the new threshold evolved in each generation until a preset number of iterations is reached, or the current model size and the current model accuracy meet the constraints; wherein the sparse operation includes deleting the convolution kernel when the number of weights and bias items of the convolution kernel that are zero is greater than the threshold; A loop execution module is used to repeatedly execute the steps in the model compression module if the current model accuracy and the current model size do not meet the constraint conditions until the current model accuracy and the current model size meet the constraint conditions.
7. The global sparsification device for a convolutional neural network model according to claim 6, characterized in that: The model compression module is also used for: Set the threshold range of the i-th dimension of each convolution kernel to be optimized to [ ]; The i-th dimension of each convolution kernel to be optimized is divided into N intervals, and the threshold length of each interval is recorded as: ; Set the random sampling probability for each interval of the i-th dimension of each convolution kernel to be optimized : ; in, represents the random sampling probability of the Nth interval of the i-th dimension; The sample threshold of the jth interval on the i-th dimension of each convolution kernel to be optimized is calculated according to the following formula: : ; in, represents the random sampling probability of the jth interval of dimension i, [0,N]; The N sample thresholds on each dimension of each convolution kernel to be optimized are respectively used as the random thresholds.
8. The global sparsification device for a convolutional neural network model according to claim 7, characterized in that: The model compression module is also used for: Generate multiple sample thresholds as parent populations from the threshold range of each convolution kernel through the Latin hypercube sampling algorithm; Perform mutation operation on the parent population based on the following formula: in, is a scaling factor between 0 and 2; and are two random sample thresholds; It represents the sample threshold that minimizes the model when the model accuracy constraint is met, or the sample threshold that makes the model closest to the constraint when the model accuracy constraint is not met; The dimension of the threshold for the i-th sample , using binomial crossover, using each dimension Generate random numbers to perform mutation: in, Indicates Generation population, CR represents the crossover probability, Represents a uniform random number between 0 and 1. represents a random dimension, Indicates the first Daidi No. The sample threshold of the dimension, Indicates the mutation Daidi No. The sample threshold of the dimension, Indicates the crossover Daidi No. Sample threshold of dimension; Perform the sparse operation and model evaluation, and select the next generation of population to be evaluated according to the following formula : in, represents the objective function of the model, Indicates the crossover Daidi Sample threshold, Indicates the number before mutation Daidi Sample threshold; The population evolution process is repeated until the preset number of iterations of the population is reached, or the current model size and the current model accuracy meet the constraints.
9. The global sparsification device for a convolutional neural network model according to claim 8, characterized in that: It also includes a model fine-tuning module, which is used to fine-tune the architecture parameters of the model after sparsification processing to improve the model accuracy.
10. The global sparsification device for a convolutional neural network model according to claim 9, characterized in that: The model fine-tuning module adopts an early stopping strategy to prevent fine-tuning overfitting.
Citation Information
Patent Citations
Method for compensating positioning errors of robot based on deep neural network
CN110385720A
Deep learning network model compression method based on network layer pruning
CN111401523A
Full convolutional neural network density peak pruning method for image segmentation
CN114742997A
High-precision fiber bragg grating blast furnace temperature detection method and system
CN116839756A
Network pruning optimization method based on network activation and sparsification
WO2021129570A1