Model quantification method and device, electronic equipment and storage medium
By determining the target network layer in the floating-point model and quantizing its input features according to the channel dimension, the problem of reducing quantization accuracy caused by large numerical differences between different semantic inputs is solved, and a fixed-point model of higher quantization accuracy and integer type is achieved.
Patent Information
- Application Number
- CN202311523242.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
When existing model quantization methods deal with models with non-image inputs, it is difficult to effectively deal with the problem of large numerical differences between different semantic inputs, resulting in a reduced quantization accuracy.
By determining the target network layer whose input features are splicing types in the floating-point model, and quantizing the input features of this layer according to the channel dimension, the quantization loss of input data in different channels is reduced and the quantization accuracy is improved.
By quantifying the input data on each channel separately, the quantization loss is reduced, the model quantization accuracy is improved, and a fixed-point model of integer type is obtained.
Smart Images

Figure CN120011702A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a model quantization method, device, electronic device and storage medium. Background Art
[0002] Model quantization is a process of converting full-precision (FP32, 32-bit single precision) or half-precision (FP16) parameters into integers or fixed-point numbers (i.e. INT8 or INT16) to reduce the number of model parameters and speed up the operation. Commonly used quantization methods are PTQ (Post-Training Quantization) and QAT (Quantization Aware Training). Both methods require adding a scaling factor to the model, but the scaling factor in the PTQ method is calculated based on statistical methods, while the scaling factor added in the QAT method needs to be learned through the model training process.
[0003] However, both methods only add scaling factors according to the feature dimension. For models with non-image inputs, that is, input features concatenated from multiple different semantic inputs, since the numerical differences between different semantic inputs are relatively large, if the same scaling factor is used for unified quantization, the input information with larger or smaller values will be lost, which will ultimately lead to reduced model quantization accuracy. Summary of the invention
[0004] The purpose of this application is to propose a model quantization method, device, electronic device and storage medium to address the deficiencies of the above-mentioned prior art, and this purpose is achieved through the following technical solutions.
[0005] The first aspect of the present application proposes a model quantization method, the method comprising:
[0006] Get the trained floating point model;
[0007] Determine a target network layer in the floating point model whose input feature is a concatenated feature;
[0008] The input features of the target network layer are quantized according to the channel dimension, and the model parameters in the floating-point model are quantized to obtain a fixed-point model.
[0009] A second aspect of the present application provides a model quantization device, the device comprising:
[0010] Model acquisition module, used to obtain the trained floating-point model;
[0011] A determination module, used to determine a target network layer in the floating point model where the input feature is a concatenated feature;
[0012] The quantization module is used to quantize the input features of the target network layer according to the channel dimension, and quantize the model parameters in the floating-point model to obtain a fixed-point model.
[0013] The third aspect of the present application proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the first aspect above.
[0014] The fourth aspect of the present application proposes a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the steps of the method described in the first aspect above.
[0015] Based on the model quantization method and device described in the first and second aspects above, the present application has at least the following beneficial effects or advantages:
[0016] By determining that the input features in the floating-point model are of the target network layer of the concatenation type, that is, the input features include multiple inputs with different physical semantics, and by quantizing the input features of the target network layer according to the channel dimension, the input data representing certain physical semantics on each channel is quantized separately, so as to reduce the quantization loss of input data in different channels and improve the quantization accuracy. After quantizing the model parameters in the floating-point model, a fixed-point model of integer type is obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 A schematic diagram of an existing quantitative method shown in this application;
[0019] Figure 2 This is a flow chart of an embodiment of a model quantization method according to an exemplary embodiment of the present application;
[0020] Figure 3 This is a schematic diagram of a quantization network structure according to an exemplary embodiment of the present application;
[0021] Figure 4 This is a schematic diagram of a specific structure of a quantization network according to an exemplary embodiment of the present application;
[0022] Figure 5 This is a schematic diagram of the structure of a model quantization device according to an exemplary embodiment of the present application;
[0023] Figure 6This is a schematic diagram of the hardware structure of an electronic device according to an exemplary embodiment of the present application;
[0024] Figure 7 The figure is a schematic diagram of the structure of a storage medium according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0025] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0026] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0027] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0028] At present, the model quantization process adds scaling factors to the model, which are only added and calculated in the feature dimension, that is, one scaling factor is added for each feature. For models with different semantic inputs, because the numerical differences between different semantic inputs are large, if the same scaling factor is used for quantization, the input information with large or small values will be lost, resulting in reduced model quantization accuracy. Figure 1 As shown in FIG. 1 , the input feature of a network layer of the model is composed of four features representing different physical semantics. These four features represent different physical meanings. When adding a scaling factor for quantization, a scaling factor is set for the input feature for quantization.
[0029] To solve the above technical problems, the present application proposes a model quantization method. After obtaining a trained floating-point model, the target network layer whose input features in the floating-point model are concatenated type is determined, that is, the input features include multiple different physical semantic input features, and the input features of the target network layer are quantized according to the channel dimension, so that the input data representing certain semantics on each channel is quantized separately, so as to reduce the quantization loss of input data in different channels and improve the quantization accuracy. After the model parameters in the floating-point model are quantized, a fixed-point model of integer type is obtained.
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0031] Figure 2 This is a flow chart of an embodiment of a model quantization method according to an exemplary embodiment of the present application, the method comprising the following steps:
[0032] Step 201: Obtain the trained floating-point model.
[0033] In this step, the floating-point model refers to a trained model that has been trained to achieve a certain prediction function. The data type of the model parameters in the model is a floating-point type, such as FP32 or FP16. If the model has data input, a series of operations on the data in the model are all floating-point operations.
[0034] The embodiments of the present application are applicable to models with multiple different semantic inputs. For example, for a trajectory prediction model, its input includes multiple different semantic physical quantities, including speed, heading, position, acceleration, etc. This is different from the image recognition model. Although the input image also involves multi-dimensional features, these multi-dimensional features are all image features and have the same physical semantics.
[0035] Step 202: Determine the target network layer whose input feature in the floating point model is the concatenated feature.
[0036] In this step, the target network layer may be one or more, and may be located in a shallow network of a floating-point model or in a deep network of a floating-point model.
[0037] For example, assuming that the input data of the model has multiple physical quantities, these multiple physical quantities need to be spliced into one feature and then subjected to a series of subsequent operations. Therefore, the target network layer is the first operation processing layer, that is, it is located in the shallow network of the model.
[0038] It is also assumed that in the intermediate process of a series of computational processing of the model, in addition to receiving a physical semantic feature transmitted from an adjacent layer, it is also necessary to introduce other physical semantic features or new input data transmitted from other layers, so the target network layer is located in the deep network.
[0039] In a feasible implementation, a concatenation layer in a floating-point model may be determined, and then the next network layer of the concatenation layer may be determined as a target network layer whose input features are concatenation features.
[0040] Among them, in the model, the splicing layer is usually used to splice different features into a splicing feature by channel and then input it into the next network layer for calculation processing. Therefore, the next network layer of the splicing layer can be determined as the target network layer.
[0041] In another feasible implementation, it is also possible to pre-analyze which network layers in the model need to be quantized by channel in terms of input features, and write these network layer identifiers into the configuration file. Based on this, when determining the target network layer, the configuration file can be directly read, and the network layer indicated by the configuration file in the floating-point model can be determined as the target network layer.
[0042] Step 203: quantize the input features of the target network layer according to the channel dimension, quantize the model parameters in the floating-point model, and obtain a fixed-point model.
[0043] In this step, the fixed-point model refers to a trained integer model used to implement a certain prediction function. The model parameters of the model are stored in integer type. If the model has data input, a series of operations on the data in the model are all integer operations.
[0044] It can be understood that the quantization method of the model parameters in the floating-point model is implemented using relevant technologies, and the embodiments of the present application do not limit this, that is, the QAT quantization method or the PTQ quantization method can be used.
[0045] It should be noted that due to hardware limitations, except for depthwise separable convolution, other convolution layers do not support quantization by setting scaling factors according to the channel dimension of the input features. In other words, most conventional convolution layers do not support this setting operation.
[0046] In the embodiment of the present application, the scaling factor represents the quantization scale of the floating-point value. The quantization principle is usually to multiply the floating-point value by the scaling factor and then round it to the nearest integer.
[0047] Based on this, for the process of quantizing the input features of the target network layer according to the channel dimension, a quantization network is inserted between the input features and the target network layer, and a scaling factor is set for each channel of the input features in the quantization network, so that the input features are quantized according to the channel dimension by using the scaling factor set by channel through the quantization network.
[0048] The quantization network is an operation network for realizing channel-based quantization of input features. Therefore, a scaling factor can be set for each channel of the input feature in the quantization network to quantize the input feature according to the channel.
[0049] Based on the above description, the target network layer in the embodiment of the present application can be any conventional convolutional layer or other computing layer whose input features are splicing features. By inserting a quantization network that supports setting the scaling factor according to the input feature channel, the conventional convolutional layer or other computing layer can support setting the scaling factor by channel.
[0050] It is worth noting that these scaling factors set according to the channel can be applied to the quantization of the corresponding channel features respectively, so that the information of the large or small values in the channel will not be lost, thereby reducing the quantization loss and achieving the purpose of improving the quantization accuracy of the model.
[0051] In one possible implementation, if Figure 3 As shown in the figure, the quantization network includes a quantization layer and a denoising layer. In the quantization layer, a scaling factor is set for each channel of the input feature. Assuming that the input feature is 5*4*4 concatenated from four inputs, a scaling factor is set for each 5*4*1 feature. Since the quantization layer introduces noise after performing operator operations and quantization on the input features, a denoising layer is used after the quantization layer to eliminate the influence of noise on the input features.
[0052] Based on this, the input features are quantized according to the channel dimension using the scaling factor through the quantization network, including:
[0053] The input features are calculated through the quantization layer, and the features obtained by the calculation are quantized according to the channel dimension using the set scaling factor. The quantized features output by the quantization layer are denoised through the denoising layer to obtain the quantized input features.
[0054] Among them, the operation operator of the quantization layer can be any operator such as a multiplication operator (mul), an addition operator (add), a depthwise separable convolution operator (depthwise_conv), etc.
[0055] The specific processing of the quantization layer and the denoising layer is explained by taking the multiplication operator (mul) as an example.
[0056] like Figure 4 As shown in the figure, the multiplication operator needs to have two multiplier inputs for multiplication operations. Therefore, the input of the quantization layer needs to set two multipliers. One multiplier is the input feature to be quantized, and the other multiplier is the initialization feature (the same size as the input feature, and the feature values all use the default value). In the quantization layer, after the input feature and the initialization feature are calculated by the multiplication operator, the set scaling factor is used to quantize the obtained features according to the channel.
[0057] Since the multiplication operator in the quantization layer introduces an initialization feature for operation, the information of the initialization feature is introduced into the quantization feature output by the quantization layer. The quantization feature needs to be passed into the denoising layer to remove the influence of the initialization feature on the quantization feature. Since the initialization feature is introduced by multiplication, the second input of the denoising layer is set to the inverse of the feature value of the initialization feature, and the numerical multiplication operator mul_scalar is used in the denoising layer to remove the information of the initialization feature. For example, if the feature values of the initialization features are all 2, then the second input of the denoising layer is set to 1 / 2. Therefore, in the denoising layer, the numerical multiplication operator is used to multiply the quantization feature output by the quantization layer with the inverse of the initialization feature to obtain the final quantized input feature.
[0058] It should be noted here that in order to improve the operational efficiency of the quantization layer and the denoising layer, the initialization features of the input quantization layer can use integer values, and the input features can first be converted into integer features using a scaling factor.
[0059] The process of setting a scaling factor for each channel of the input feature in the quantization network may include the following two implementation methods.
[0060] The first implementation method is to use the model parameters of the target network layer to determine the scaling factor, and then set the scaling factor to the quantization network according to the channels of the input features. Among them, each channel of the model parameter corresponds to a scaling factor, and the number of channels of the model parameter is the same as the number of channels of the input feature. Specifically, the corresponding scaling factor can be determined based on the statistical value of each channel model parameter (such as maximum value, minimum value, average value, etc.).
[0061] Based on the above method, the scaling factor is calculated by statistical method and can be used to quantize the input features after being set in the quantization network.
[0062] The second implementation method is: setting a preset value as a scaling factor for each channel of the input feature in the quantization network, and then obtaining the adjusted scaling factor by training the floating-point model with the set scaling factor.
[0063] That is, a preset value is set as a scaling factor for each channel of the input feature in the quantization layer of the quantization network. The scaling factor is a learnable scaling factor, that is, the scaling factor is learned during model training to ensure that the quantization error is minimized. Therefore, during the secondary training of the floating-point model, the scaling factor in the quantization layer of the quantization network is adjusted to minimize the quantization loss.
[0064] In a specific implementation, the adjusted scaling factor is obtained by training the floating-point model for setting the scaling factor, including:
[0065] The training data is obtained, the numerical range of the training data is normalized, and then the floating-point model with the scaling factor set is trained using the normalized training data. During the training process, the scaling factor is adjusted according to the quantization loss calculation of the floating-point model.
[0066] Based on the above method, by normalizing the values of the training data to a uniform value range, the value range of the entire model parameters can be controlled to be relatively stable, making it easier to adjust the scaling factor.
[0067] The normalization process may include: calculating the statistical distribution of the training data, such as the maximum value, minimum value, average value, median and other statistics, and then normalizing each training data to [-1, 1] according to the statistical distribution. That is, for positive values, first normalize to [0, 1], then subtract 0.5 and multiply by 2, and normalize to [-1, 1]. For negative values, directly normalize to [-1, 1].
[0068] It should be noted that after obtaining the fixed-point model, the accuracy of the fixed-point model can be checked to determine whether the prediction effect of the fixed-point model meets the requirements.
[0069] Based on this, by obtaining test data, using the test data to determine the precision index of the floating-point model, and using the test data to determine the precision index of the fixed-point model, if the difference between the precision index of the floating-point model and the precision index of the fixed-point model is within a preset range, it means that the prediction effect of the fixed-point model meets the requirements. Among them, the precision index can include any one or more of the model recall rate, accuracy, average error, etc.
[0070] At this point, the above Figure 2 The model quantization process shown in the figure determines that the input features in the floating-point model are of the target network layer of the splicing type, that is, the input features include multiple inputs with different physical semantics, and quantizes the input features of the target network layer according to the channel dimension, so that the input data representing certain physical semantics on each channel is quantized separately, so as to reduce the quantization loss of input data of different channels and improve the quantization accuracy. After the model parameters in the floating-point model are quantized, a fixed-point model of integer type is obtained.
[0071] Corresponding to the embodiments of the aforementioned model quantization method, the present application also provides embodiments of a model quantization device.
[0072] Figure 5 FIG. 1 is a schematic diagram of a structure of a model quantization device according to an exemplary embodiment of the present application. The device is used to execute the model quantization method provided in any of the above embodiments, such as Figure 5 As shown, the model quantization device includes:
[0073] A model acquisition module 510 is used to acquire a trained floating-point model;
[0074] A determination module 520, configured to determine a target network layer in the floating point model whose input feature is a concatenated feature;
[0075] The quantization module 530 is used to quantize the input features of the target network layer according to the channel dimension, and quantize the model parameters in the floating-point model to obtain a fixed-point model.
[0076] In some embodiments of the present application, the quantization module 530 is specifically used to insert a quantization network between the input features and the target network layer during the process of quantizing the input features of the target network layer according to the channel dimension; set a scaling factor for each channel of the input features in the quantization network; and quantize the input features according to the channel dimension using the scaling factor through the quantization network.
[0077] In some embodiments of the present application, the quantization network includes a quantization layer and a denoising layer, and a scaling factor is set for each channel of the input feature in the quantization layer;
[0078] The quantization module 530 is specifically used to, during the process of quantizing the input features according to the channel dimension using the scaling factor through the quantization network, operate on the input features through the quantization layer, and quantize the features obtained by the operation according to the channel dimension using the set scaling factor; perform denoising operation on the quantized features output by the quantization layer through the denoising layer to obtain the quantized input features.
[0079] In some embodiments of the present application, the quantization module 530 is specifically used to determine the scaling factor using the model parameters of the target network layer in the process of setting a scaling factor for each channel of the input feature in the quantization network; a scaling factor is determined for each channel of the model parameter, and the number of channels of the model parameter is the same as the number of channels of the input feature; and the scaling factor is set in the quantization network according to the channel of the input feature.
[0080] In some embodiments of the present application, the quantization module 530 is specifically used to set a preset value as a scaling factor for each channel of the input feature in the quantization network during the process of setting a scaling factor for each channel of the input feature in the quantization network; and obtain the adjusted scaling factor by training a floating-point model for setting the scaling factor.
[0081] In some embodiments of the present application, the quantization module 530 is specifically used to obtain training data in the process of training a floating-point model for setting a scaling factor to obtain an adjusted scaling factor; normalize the numerical range of the training data; use the normalized training data to train the floating-point model for setting the scaling factor, and during the training process, adjust the scaling factor according to the calculated quantization loss.
[0082] In some embodiments of the present application, the determination module 520 is specifically used to determine the splicing layer in the floating-point model, and determine the next network layer of the splicing layer as the target network layer whose input feature is the splicing feature.
[0083] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0084] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0085] The embodiment of the present application also provides an electronic device corresponding to the model quantization method provided in the above embodiment to execute the above model quantization method.
[0086] Figure 6 This is a hardware structure diagram of an electronic device according to an exemplary embodiment of the present application, and the electronic device includes: a communication interface 601, a processor 602, a memory 603 and a bus 604; wherein the communication interface 601, the processor 602 and the memory 603 communicate with each other through the bus 604. The processor 602 can execute the model quantization method described above by reading and executing the machine executable instructions corresponding to the control logic of the model quantization method in the memory 603. The specific content of the method is referred to the above embodiment, and will not be repeated here.
[0087] The memory 603 mentioned in this application can be any electronic, magnetic, optical or other physical storage device, and can contain storage information, such as executable instructions, data, etc. Specifically, the memory 603 can be RAM (Random Access Memory), flash memory, storage drive (such as hard disk drive), any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 601 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0088] The bus 604 may be an ISA bus, a PCI bus or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 603 is used to store programs, and the processor 602 executes the programs after receiving execution instructions.
[0089] The processor 602 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 602. The above processor 602 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be combined and executed.
[0090] The electronic device provided in the embodiment of the present application and the model quantization method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.
[0091] The present application also provides a computer-readable storage medium corresponding to the model quantization method provided in the above embodiment. Figure 7 As shown, the computer-readable storage medium shown is a CD 30 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the model quantization method provided by any of the aforementioned embodiments will be executed.
[0092] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0093] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the model quantization method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0094] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0095] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0096] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A model quantization method, characterized in that: The method comprises: Get the trained floating point model; Determine a target network layer in the floating point model whose input feature is a concatenated feature; The input features of the target network layer are quantized according to the channel dimension, and the model parameters in the floating-point model are quantized to obtain a fixed-point model.
2. The method according to claim 1, characterized in that The step of quantizing the input features of the target network layer according to the channel dimension includes: Inserting a quantization network between the input feature and the target network layer; Setting a scaling factor for each channel of the input feature in the quantization network; The input features are quantized according to the channel dimension using the scaling factor through the quantization network.
3. The method according to claim 2, characterized in that The quantization network comprises a quantization layer and a denoising layer, wherein a scaling factor is set for each channel of the input feature in the quantization layer; Quantizing the input feature according to the channel dimension by using the scaling factor through the quantization network includes: The input features are operated through the quantization layer, and the features obtained by the operation are quantized according to the channel dimension using the set scaling factor; The denoising layer performs denoising operations on the quantized features output by the quantization layer to obtain quantized input features.
4. The method according to claim 2 or 3, characterized in that: Setting a scaling factor for each channel of the input feature in the quantization network includes: Determine a scaling factor using a model parameter of the target network layer; a scaling factor is determined for each channel of the model parameter, and the number of channels of the model parameter is the same as the number of channels of the input feature; The scaling factor is set in the quantization network according to the channel of the input feature.
5. The method according to claim 2 or 3, characterized in that: Setting a scaling factor for each channel of the input feature in the quantization network includes: Setting a preset value as a scaling factor for each channel of the input feature in the quantization network; The adjusted scaling factor is obtained by training the floating-point model with the scaling factor set.
6. The method according to claim 5, characterized in that By training the floating-point model with the scaling factor set, the adjusted scaling factor is obtained, including: Get training data; Normalizing the numerical range of the training data; The normalized training data is used to train a floating-point model with a set scaling factor. During training, the scaling factor is adjusted based on the calculated quantization loss.
7. The method according to claim 1, characterized in that The step of determining that the input feature in the floating point model is a target network layer of a splicing feature includes: A concatenation layer in the floating-point model is determined, and a next network layer of the concatenation layer is determined as a target network layer whose input feature is the concatenation feature.
8. A model quantization device, characterized in that: The device comprises: Model acquisition module, used to obtain the trained floating-point model; A determination module, used to determine a target network layer in the floating point model where the input feature is a concatenated feature; The quantization module is used to quantize the input features of the target network layer according to the channel dimension, and quantize the model parameters in the floating-point model to obtain a fixed-point model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the steps of the method according to any one of claims 1 to 7.