Image recognition model compression method and device, equipment and storage medium

By constructing and using the order code index of the image recognition model, the storage and computing resource requirements of the model are reduced, the deployment problem of the image recognition model in resource-constrained environments is solved, and the energy consumption in the inference process is reduced.

CN119940452APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510091121.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Due to the huge number of parameters, the image recognition model has extremely high demand for storage and computing resources, making it difficult to deploy in resource-constrained environments, and the energy consumption during the inference process is also very large.

Method used

By constructing the order code index corresponding to the weighted order code and storing the order code index into the order code table, the footprint of the image recognition model is reduced. The specific steps include obtaining the initial image recognition model, constructing the target order code table of each neural network layer, determining the first and second target order code indexes, and reducing the target weight and bit reduction based on these indexes.

Benefits of technology

By reducing the space required for storage order codes and the space occupied by the model, the overall spatial demand of the image recognition model is reduced, the model's deployment capability is improved in resource-constrained environments, and the energy consumption in the inference process is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940452A_ABST
    Figure CN119940452A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition model compression method and device, equipment and a storage medium, and relates to the technical field of neural networks, and the method comprises the steps: obtaining an initial image recognition model, and constructing each target order code table according to the order code corresponding to the target weight of each neural network layer in the initial image recognition model; constructing a corresponding order code index according to each order code, and storing each order code index to a target order code table; determining a first target order code index and a second target order code index from each order code by using the number of each order code and a corresponding target order code table; and reducing the number of the target weights based on the first target order code index and the second target order code index, and performing bit reduction on mantissas corresponding to the target weights to obtain a target image recognition model corresponding to the initial image recognition model constructed based on the neural network. By constructing the order code index corresponding to the order code of the weight and storing the order code index in the order code table, the occupied space of the image recognition model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network technology, and in particular to an image recognition model compression method, device, equipment and storage medium. Background Art

[0002] With the development of artificial intelligence, the scale of image recognition models continues to increase, with the number of parameters reaching millions or even billions, which places extremely high demands on storage and computing resources. In pursuit of higher accuracy, the complexity of image recognition models continues to increase, making it difficult to deploy models in resource-constrained environments such as IoT devices and mobile platforms. At the same time, as the scale of image recognition models continues to grow, the energy consumption during the inference process is also becoming larger and larger. Therefore, how to compress image recognition models has become a problem that needs to be solved. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide an image recognition model compression method, device, equipment and storage medium, which can reduce the space occupied by the image recognition model by constructing the order code corresponding to the order code of the weight and storing the order code index in the order code table. The specific scheme is as follows:

[0004] In a first aspect, the present application provides an image recognition model compression method, comprising:

[0005] Acquire an initial image recognition model constructed based on a neural network, and construct target order tables corresponding to each neural network layer according to the order corresponding to the target weight of each neural network layer in the initial image recognition model;

[0006] Constructing a corresponding order code index according to each of the order codes, and storing each of the order code indexes in a corresponding target order code table;

[0007] Determine a first target order index and a second target order index from each target order table according to the number of each order; wherein the first target order index is an order index corresponding to orders whose number is not less than a preset number threshold, and the second target order index is an order index corresponding to orders whose number is less than the preset number threshold;

[0008] The number of the target weights is reduced based on the first target order index and the second target order index, and the number of digits of the decimals corresponding to each target weight is reduced to obtain a target image recognition model corresponding to the initial image recognition model.

[0009] Optionally, constructing target order tables corresponding to the neural network layers includes:

[0010] Counting the number of different order codes in each of the neural network layers;

[0011] The index bit corresponding to the order code index is calculated according to the number of the order codes, and the target order code tables corresponding to the neural network layers are constructed according to the index bits.

[0012] Optionally, determining the first target order index and the second target order index from each target order table according to the number of each order includes:

[0013] The order code indexes in each target order code table are sorted in descending order according to the order from large to small in terms of the number of each order code, and the first target order code indexes and the second target order code indexes are determined from the sorted target order code tables based on the order corresponding to each order code index and the number of the orders.

[0014] Optionally, reducing the number of the target weights based on the first target order index and the second target order index includes:

[0015] Obtaining a significant weight corresponding to the first target order index, and obtaining a non-significant weight corresponding to the second target order index;

[0016] Convert each of the non-significant weights into a significant weight with the closest value in the same neural network layer, so as to reduce the number of the target weights;

[0017] The significant weights are target weights whose quantity is not less than the preset quantity threshold, and the non-significant weights are target weights whose quantity is less than the preset quantity threshold.

[0018] Optionally, the step of reducing the number of digits of the mantissa corresponding to each of the target weights to obtain a target image recognition model corresponding to the initial image recognition model includes:

[0019] The mantissa corresponding to each of the target weights is reduced in number to obtain a compressed image recognition model corresponding to the initial image recognition model;

[0020] The compressed image recognition model is trained using a target neural network training algorithm to obtain a target image recognition model corresponding to the initial image recognition model.

[0021] Optionally, after obtaining the target image recognition model corresponding to the initial image recognition model, the method further includes:

[0022] Obtaining the orders corresponding to the target weights using the order indexes in the target order tables in the target neural network;

[0023] The target weight is converted into a standard floating-point format according to the sign bit and mantissa corresponding to each of the exponents and the target weight, and a model inference operation corresponding to the target image recognition model is performed according to the target weight in the standard floating-point format.

[0024] Optionally, converting the target weight into a standard floating point format includes:

[0025] The FPGA chip is used to combine the sign bit and mantissa corresponding to each of the exponents and the target weight to convert the target weight into a standard floating point format.

[0026] In a second aspect, the present application provides an image recognition model compression device, comprising:

[0027] An order code table construction module is used to obtain an initial image recognition model constructed based on a neural network, and to construct target order code tables corresponding to each neural network layer according to the order codes corresponding to the target weights of each neural network layer in the initial image recognition model;

[0028] An order code index storage module, used for constructing a corresponding order code index according to each order code, and storing each order code index in a corresponding target order code table;

[0029] An order index determination module, configured to determine a first target order index and a second target order index from each target order table according to the number of each order; wherein the first target order index is an order index corresponding to an order not less than a preset number threshold, and the second target order index is an order index corresponding to an order less than the preset number threshold;

[0030] A model compression module is used to reduce the number of the target weights based on the first target order code index and the second target order code index, and to reduce the number of digits of the decimals corresponding to each of the target weights to obtain a target image recognition model corresponding to the initial image recognition model.

[0031] In a third aspect, the present application provides an electronic device, including:

[0032] Memory, used to store computer programs;

[0033] A processor is used to execute the computer program to implement the aforementioned image recognition model compression method.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which implements the aforementioned image recognition model compression method when executed by a processor.

[0035] The present application first obtains an initial image recognition model constructed based on a neural network, and constructs target order code tables corresponding to each neural network layer according to the order codes corresponding to the target weights of each neural network layer in the initial image recognition model, then constructs corresponding order code indexes according to each order code, and stores each order code index in the corresponding target order code table, and then determines a first target order code index and a second target order code index from each target order code table according to the number of each order code; wherein the first target order code index is an order code index corresponding to an order code whose number is not less than a preset number threshold, and the second target order code index is an order code index corresponding to an order code whose number is less than the preset number threshold, and finally, based on the first target order code index and the second target order index, the number of target weights is reduced, and the number of digits of the decimals corresponding to each target weight is reduced to obtain a target image recognition model corresponding to the initial image recognition model. It can be seen that the present application reduces the space required to store the order code by constructing the order code index corresponding to the order code of the weight and storing the order code index in the order code table, thereby reducing the space occupied by the image recognition model; by reducing the number of weights in the model and reducing the decimals corresponding to the weights, the space occupied by the weight data in the image recognition model is reduced, thereby reducing the space occupied by the image recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0037] Figure 1 A flow chart of an image recognition model compression method disclosed in this application;

[0038] Figure 2 This is a schematic diagram of the structure of an image recognition model compression device disclosed in this application;

[0039] Figure 3 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] At present, the scale of image recognition models continues to increase, and the number of parameters reaches millions or even billions, which places extremely high demands on storage and computing resources. To this end, the present application provides an image recognition model compression method, which reduces the space occupied by the image recognition model by constructing an order code index corresponding to the order code of the weight and storing the order code index in an order code table.

[0042] See also Figure 1 As shown, an embodiment of the present invention discloses an image recognition model compression method, comprising:

[0043] Step S11, obtaining an initial image recognition model constructed based on a neural network, and constructing target order code tables corresponding to each neural network layer according to the order codes corresponding to the target weights of each neural network layer in the initial image recognition model.

[0044] The image recognition model in this embodiment is an image recognition model built based on a neural network, that is, the image recognition model in this embodiment is a neural network model for image recognition. The above process of constructing target order code tables corresponding to each neural network layer according to the order codes corresponding to the target weights of each neural network layer in the initial image recognition model can specifically include: counting the number of different order codes in each neural network layer; calculating the index bits corresponding to the order code index according to the number of order codes, and constructing target order code tables corresponding to each neural network layer according to the index bits; the above process is also a process of realizing order code sharing in units of layers by constructing multiple order code tables. For example, in the Float32 format, the order code is 8 bits and can represent 256 values, but the order code distribution of weights of different layers in the actual model is often concentrated in a few values. By constructing an order code table, different order codes in each layer are stored and accessed through indexes. Replace the exponent field in the floating point format with the index field to reduce the space required to store the exponent. For example, in the BFloat32 format, the exponent occupies 8 bits, which can save up to 43.75% of the storage space. Specifically, for a given pre-trained model, that is, the initial image recognition model mentioned above, analyze the distribution of the exponents of the weights by layer. Taking the BFloat16 format as an example, count the number of different exponents in each layer, and calculate the required index bits log2 (e). Construct an exponent table of index bit length, replace the exponents in the weights with the corresponding index, and store them in a separate table. It can be understood that the neural network model includes multiple neural network layers, wherein each neural network layer includes a number of weights, and the weight value of each weight in the neural network layer exists in the form of a floating point number, and each weight value has its own corresponding order code; and the order code table in this embodiment is constructed by counting the number of order codes, calculating the index bit according to the number of order codes, and constructing it according to the index bit, so the order code table constructed in this embodiment will not have extra free storage space, that is, each target order code table in this embodiment is an order code table that can store the order code index corresponding to all the order codes in each layer and has no redundant storage space. By constructing the order code table, the order code field in floating point format can be replaced by the index field, reducing the storage space required for the model; by counting the number of order codes and constructing each order code table according to the index bit, it is ensured that the storage resources occupied by the order code table are minimized, further reducing the storage space required for the neural network model.

[0045] Step S12: construct a corresponding order index according to each order, and store each order index in a corresponding target order table.

[0046] In this embodiment, the process of constructing the corresponding order code index according to the order code and storing the order code index in the target order code table is performed simultaneously in units of layers. It can be understood that the number and value of weights in each neural network layer in the model are different, so the order code indexes stored in different order code tables are also different. It should be noted that the same order code may exist in different neural network layers. When the order code index corresponding to a certain order code is generated by analysis and calculation in a certain neural network layer, there is no need to calculate it again in other neural network layers, and the same order code index can be directly stored in the order code table of this layer. By replacing the order code field in floating point format with the index field, the storage space required for the model is reduced; by reusing the order code index corresponding to the same order code in different neural network layers, the calculation time and computing resources required to generate the order code index are reduced, and the efficiency of model compression is improved.

[0047] Step S13, determining a first target order index and a second target order index from each target order table according to the number of each order; wherein the first target order index is an order index corresponding to an order whose number is not less than a preset number threshold, and the second target order index is an order index corresponding to an order whose number is less than the preset number threshold.

[0048] In this embodiment, the process of determining the first target order code index and the second target order code index from each target order code table according to the number of each order code may specifically include: sorting the order code indexes in each target order code table in descending order according to the order of the number of each order code from large to small, and determining each first target order code index and each second target order code index from each sorted target order code table based on the order corresponding to each order code index and the number of orders; wherein the first target order code index is the order code index corresponding to the order code whose number is not less than a preset number threshold, and the second target order code index is the order code index corresponding to the order code whose number is less than the preset number threshold; for example, in a specific implementation, the order code indexes in the order code table are sorted in descending order according to the number of each order code appearing in the corresponding neural network layer, the number threshold is set to 200, and the order code indexes are screened from the sorted order code table: the order code index corresponding to the order code whose number of orders is greater than or equal to 200 is determined as the first target order code index, and correspondingly, the order code index corresponding to the order code whose number of orders is less than 200 is determined as the second target order code index. By sorting the order code indexes in each target order code table in descending order according to the number of each order code from large to small, the process of determining the first order code index and the second order code index is faster, thereby improving the compression efficiency of the neural network model.

[0049] Step S14, reducing the number of the target weights based on the first target order index and the second target order index, and reducing the number of digits of the mantissa corresponding to each target weight to obtain a target image recognition model corresponding to the initial image recognition model.

[0050] In this embodiment, the process of reducing the number of target weights based on the first target order code index and the second target order code index may specifically include: obtaining the significant weight corresponding to the first target order code index, and obtaining the non-significant weight corresponding to the second target order code; converting each non-significant weight into a significant weight with the closest value in the same neural network layer to reduce the number of target weights; wherein the significant weight is a target weight whose number is not less than a preset number threshold, and the non-significant weight is a target weight whose number is less than the preset number threshold; specifically, after sorting the order code table, it is necessary to iterate the initial image recognition model, and reduce the length of the order code table in each iteration: that is, each iteration selects the first n significant weights. The order codes, that is, the order codes greater than the preset number threshold, the rest are non-significant order codes, and the corresponding weights are non-significant weights, and the non-significant weights are adjusted to the significant weights with the closest values ​​in the same layer; in this embodiment, the root mean square of the non-significant weights and the significant weights can be calculated in the same neural network layer to determine the significant weights with the closest values; corresponding to the above-mentioned weight approximation process, after determining the significant order codes and the non-significant order codes, according to the above-mentioned iterative approximation method, the order table of each layer is processed in each iteration, that is, the order index corresponding to the non-significant weight is replaced by the order index corresponding to the significant weight with the closest value, and the replaced order indexes are merged in the order table to reduce the length of the order table. It should be noted that the model accuracy is tested after each iteration. If the accuracy drops by more than the user-set threshold or the model memory reaches the user-set saving threshold, the iteration is stopped. In a specific implementation, if after several iterations, the accuracy of the model is less than the user-set accuracy threshold, for example, the model accuracy is 80% of the original, the iterative operation on the model is stopped; in another specific implementation, if after several iterations, the compression degree of the model has reached the user-set saving threshold, for example, the model volume has been compressed to 70% of the original, the model iteration is stopped.

[0051] In this embodiment, the process of reducing the number of digits of the mantissa corresponding to each target weight to obtain the target image recognition model corresponding to the initial image recognition model can specifically include: reducing the number of digits of the mantissa corresponding to each target weight to obtain the compressed image recognition model corresponding to the initial image recognition model; training the compressed image recognition model using the target neural network training algorithm to obtain the target image recognition model corresponding to the initial image recognition model; specifically, first calculating the number of digits k of the mantissa in the weight value in floating point format. In the iterative process, the precision of the weight mantissa is reduced by 1 bit each iteration, that is, the mantissa approximation method is used in the model iteration process, and the precision of the weight mantissa part is reduced by 1 bit each iteration: first calculate the number of digits k of the mantissa, and then adjust the weight mantissa to k-1 bits after the order code is shared in each layer, and then use the conventional training algorithm in the field, that is, the target neural network training algorithm, to train the model to minimize the error under the current precision, and test the precision impact, and repeat this process until the user-defined memory saving ratio or precision threshold is met.

[0052] In this embodiment, after obtaining the target image recognition model corresponding to the initial image recognition model, it also includes: using the index of each order code in the target order code table in the target neural network to obtain the order codes corresponding to each target weight; converting the target weight into a standard floating point format according to the sign bit and mantissa corresponding to each order code and the target weight, and performing the model reasoning operation corresponding to the target image recognition model according to the target weight in the standard floating point format; that is, during model reasoning, the order code is obtained from the order code table through the index, and the order code is converted into a standard floating point format in combination with the sign bit and mantissa for image recognition. In addition, in this embodiment, the above-mentioned process of converting the target weight into a standard floating point format can specifically include: using an FPGA chip to combine the sign bit and mantissa corresponding to each order code and the target weight to convert the target weight into a standard floating point format; that is, the operation of combining the sign bit, the order code and the mantissa can be implemented in hardware through FPGA and other solutions. By trimming the mantissa of the weight, the storage space occupied by the weight value is reduced, thereby reducing the storage space of the model; by stopping the compression of the model according to the model accuracy threshold and the model saving threshold, the compressed model achieves a better balance between resource usage and accuracy, thereby reducing the deployment difficulty and energy consumption of the model within an acceptable accuracy; by using FPGA to convert the weight into a standard floating-point format, the target image recognition model can be used for model inference operations, thereby improving the efficiency of model inference.

[0053] It can be seen that the present application reduces the space required to store the order code by constructing the order code index corresponding to the order code of the weight and storing the order code index in the order code table, thereby reducing the space occupied by the image recognition model; by reducing the number of weights in the model and reducing the decimals corresponding to the weights, the space occupied by the weight data in the image recognition model is reduced, thereby reducing the space occupied by the image recognition model.

[0054] See also Figure 2 As shown, an embodiment of the present invention discloses an image recognition model compression device, comprising:

[0055] The order code table construction module 11 is used to obtain an initial image recognition model constructed based on a neural network, and construct each target order code table corresponding to each neural network layer according to the order code corresponding to the target weight of each neural network layer in the initial image recognition model;

[0056] An order code index storage module 12, used for constructing a corresponding order code index according to each order code, and storing each order code index in a corresponding target order code table;

[0057] The order index determination module 13 is used to determine a first target order index and a second target order index from each target order table according to the number of each order; wherein the first target order index is an order index corresponding to an order not less than a preset number threshold, and the second target order index is an order index corresponding to an order less than the preset number threshold;

[0058] The model compression module 14 is used to reduce the number of the target weights based on the first target order code index and the second target order code index, and reduce the number of digits of the decimals corresponding to each target weight to obtain a target image recognition model corresponding to the initial image recognition model.

[0059] It can be seen that the present application reduces the space required to store the order code by constructing the order code index corresponding to the order code of the weight and storing the order code index in the order code table, thereby reducing the space occupied by the image recognition model; by reducing the number of weights in the model and reducing the decimals corresponding to the weights, the space occupied by the weight data in the image recognition model is reduced, thereby reducing the space occupied by the image recognition model.

[0060] In some specific embodiments, the order table construction module 11 may specifically include:

[0061] An order code number counting unit, used to count the number of different order codes in each of the neural network layers;

[0062] An order code table construction unit is used to calculate the index bit corresponding to the order code index according to the number of the order codes, and to construct each target order code table corresponding to each neural network layer according to the index bit.

[0063] In some specific embodiments, the rank index determination module 13 may specifically include:

[0064] The order code index determination unit is used to sort the order code indexes in each target order code table in descending order according to the order from large to small of the number of each order code, and determine each first target order code index and each second target order code index from each sorted target order code table based on the order corresponding to each order code index and the number of the order codes.

[0065] In some specific embodiments, the model compression module 14 may specifically include:

[0066] A weight acquisition unit, used to acquire a significant weight corresponding to the first target order index, and acquire a non-significant weight corresponding to the second target order index;

[0067] A weight quantity reduction unit, used for converting each of the non-significant weights into a significant weight with the closest value in the same neural network layer, so as to reduce the quantity of the target weights;

[0068] The significant weights are target weights whose quantity is not less than the preset quantity threshold, and the non-significant weights are target weights whose quantity is less than the preset quantity threshold.

[0069] In some specific embodiments, the model compression module 14 may specifically include:

[0070] A mantissa reduction unit, used for reducing the number of mantissas corresponding to each of the target weights to obtain a compressed image recognition model corresponding to the initial image recognition model;

[0071] The model training unit is used to train the compressed image recognition model using a target neural network training algorithm to obtain a target image recognition model corresponding to the initial image recognition model.

[0072] In some specific embodiments, the model compression module 14 further includes:

[0073] An order code acquisition unit, used for acquiring each order code corresponding to each target weight by using each order code index in each target order code table in the target neural network;

[0074] The format conversion submodule is used to convert the target weight into a standard floating-point format according to the sign bit and mantissa corresponding to each of the exponents and the target weight, and perform a model inference operation corresponding to the target image recognition model according to the target weight in the standard floating-point format.

[0075] In some specific embodiments, the format conversion submodule may specifically include:

[0076] The format conversion unit is used to combine the sign bit and mantissa corresponding to each of the exponents and the target weight using an FPGA chip to convert the target weight into a standard floating point format.

[0077] Furthermore, the present application also discloses an electronic device. Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0078] Figure 3 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the image recognition model compression method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0079] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0080] In addition, the memory 22 as a carrier for storing resources may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0081] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the image recognition model compression method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0082] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the image recognition model compression method disclosed above is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the above embodiments, and no further description will be given here.

[0083] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0084] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0085] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0086] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0087] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for compressing an image recognition model, characterized in that: include: Acquire an initial image recognition model constructed based on a neural network, and construct target order tables corresponding to each neural network layer according to the order corresponding to the target weight of each neural network layer in the initial image recognition model; Constructing a corresponding order code index according to each of the order codes, and storing each of the order code indexes in a corresponding target order code table; Determine a first target order index and a second target order index from each target order table according to the number of each order; wherein the first target order index is an order index corresponding to orders whose number is not less than a preset number threshold, and the second target order index is an order index corresponding to orders whose number is less than the preset number threshold; The number of the target weights is reduced based on the first target order index and the second target order index, and the number of digits of the decimals corresponding to each target weight is reduced to obtain a target image recognition model corresponding to the initial image recognition model.

2. The image recognition model compression method according to claim 1, characterized in that: The constructing of each target order code table corresponding to each neural network layer includes: Counting the number of different order codes in each of the neural network layers; The index bit corresponding to the order code index is calculated according to the number of the order codes, and the target order code tables corresponding to the neural network layers are constructed according to the index bits.

3. The image recognition model compression method according to claim 1, characterized in that: The determining of the first target order index and the second target order index from each target order table according to the number of each order comprises: The order code indexes in each target order code table are sorted in descending order according to the order from large to small in terms of the number of each order code, and the first target order code indexes and the second target order code indexes are determined from the sorted target order code tables based on the order corresponding to each order code index and the number of the orders.

4. The image recognition model compression method according to claim 3, characterized in that: The reducing the number of the target weights based on the first target order index and the second target order index includes: Obtaining a significant weight corresponding to the first target order index, and obtaining a non-significant weight corresponding to the second target order index; Convert each of the non-significant weights into a significant weight with the closest value in the same neural network layer, so as to reduce the number of the target weights; The significant weights are target weights whose quantity is not less than the preset quantity threshold, and the non-significant weights are target weights whose quantity is less than the preset quantity threshold.

5. The image recognition model compression method according to any one of claims 1 to 4, characterized in that: The step of reducing the number of digits of the mantissa corresponding to each of the target weights to obtain the target image recognition model corresponding to the initial image recognition model includes: The mantissa corresponding to each of the target weights is reduced in number to obtain a compressed image recognition model corresponding to the initial image recognition model; The compressed image recognition model is trained using a target neural network training algorithm to obtain a target image recognition model corresponding to the initial image recognition model.

6. The image recognition model compression method according to claim 1, characterized in that: After obtaining the target image recognition model corresponding to the initial image recognition model, the method further includes: Obtaining the orders corresponding to the target weights using the order indexes in the target order tables in the target neural network; The target weight is converted into a standard floating-point format according to the sign bit and mantissa corresponding to each of the exponents and the target weight, and a model inference operation corresponding to the target image recognition model is performed according to the target weight in the standard floating-point format.

7. The image recognition model compression method according to claim 6, characterized in that: The converting the target weight into a standard floating point format comprises: The FPGA chip is used to combine the sign bit and mantissa corresponding to each of the exponents and the target weight to convert the target weight into a standard floating point format.

8. An image recognition model compression device, characterized in that: include: An order code table construction module is used to obtain an initial image recognition model constructed based on a neural network, and to construct target order code tables corresponding to each neural network layer according to the order codes corresponding to the target weights of each neural network layer in the initial image recognition model; An order code index storage module, used for constructing a corresponding order code index according to each order code, and storing each order code index in a corresponding target order code table; An order index determination module, configured to determine a first target order index and a second target order index from each target order table according to the number of each order; wherein the first target order index is an order index corresponding to an order not less than a preset number threshold, and the second target order index is an order index corresponding to an order less than the preset number threshold; A model compression module is used to reduce the number of the target weights based on the first target order code index and the second target order code index, and to reduce the number of digits of the decimals corresponding to each of the target weights to obtain a target image recognition model corresponding to the initial image recognition model.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the image recognition model compression method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the image recognition model compression method as described in any one of claims 1 to 7.