Integral quantization model storage method, task processing method, device and equipment
Patent Information
- Application Number
- CN202211118919.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-09-13
AI Technical Summary
[0004]数据反复在CPU和整数量化加速装置之间交换拷贝,间接导致模型的执行速度下降,使得执行效率降低
[0014]本公开的上述各个实施例具有如下有益效果:通过本公开的一些实施例的整数量化模型存储方法提高了模型的执行速度以及执行效率。具体来说,造成模型执行速度下降,以及执行效率降低的原因在于:数据反复在CPU和整数量化加速装置之间交换拷贝,间接导致模型的执行速度下降,使得执行效率降低。基于此,本公开的一些实施例的整数量化模型存储方法,首先,确定目标神经网络的网络结构中的、模型量化噪声满足预设条件的非线性算子集,以筛选出对应模型量化效果不好、对应取值分布与均匀分布差异较大的非线性算子。在这里,非线性算子集中的非线性算子为待替换算子。然后,对于上述非线性算子集中的每个非线性算子,执行算子替换步骤:第一步,利用范围求取算子,确定上述非线性算子的算子范围。通过确定非线性算子的算子范围,便于后续确定可替换的、在算子范围内等价于非线性算子的至少一个多项式替换公式。第二步,利用替换多项式求取算子,根据上述算子范围,可以准确地生成至少一个多项式替换公式的多项式系数,得到至少一个多项式替换公式,作为替换算子。其中,多项式替换公式包括:多个乘法加法算子。多项式系数为多个乘法加法算子的系数。第三步,将上述非线性算子替换为上述替换算子。在这里,通过非线性算子替换为对应多项式替换公式的形式,可以有效解决模型量化噪声较大、且取值分布与均匀分布差异较大的问题,使得后续的整数量化处理更为精准。接着,对算子替换后的目标神经网络进行整数量化处理,得到量化模型。通过对算子替换后的目标神经网络进行模型量化,可以减少内存带宽和存储空间的占用。最后,对上述量化模型进行模型存储,以便于量化模型的后续使用。由此,对模型量化噪声大的算子进行算子替换处理,使得非线性算子无需复制至cpu以进行处理,仅需要在整数量化加速装置内进行处理,提高了执行速度以及执行效率。同时,通过此种方式,减少了数据计算量,降低了计算量,使得减少了针对内存的占用。
Smart Images

Figure CN115470898B_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, specifically to integer quantization model storage methods, task processing methods, apparatus, and devices. Background Technology
[0002] With the development of deep learning, neural networks have been widely applied in various fields. However, as model performance improves, it also introduces a huge number of parameters and computational load. Model quantization is a technique that converts floating-point calculations into low-ratio, specific-point calculations, which can effectively reduce the computational intensity, number of parameters, and memory consumption of the model. The typical approach to model quantization is as follows: in response to the determination that there are nonlinear operators with significant quantization noise in the model, operator operations targeting the nonlinear operators are run in the Central Processing Unit (CPU). After the operation is completed, integer quantization acceleration devices are used to accelerate the execution of the operator logic for the remaining operators in the model.
[0003] However, the inventors discovered that when using the above method for model quantization, the following technical problems often arise:
[0004] Data is repeatedly copied and exchanged between the CPU and the integer quantization acceleration device, which indirectly leads to a decrease in the execution speed of the model and reduces execution efficiency.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of this disclosure propose integer quantization model storage methods, task processing methods, apparatuses, and devices to address the technical problems mentioned in the background section above.
[0008] In a first aspect, some embodiments of this disclosure provide an integer quantization model storage method, comprising: determining a set of nonlinear operators in the network structure of a target neural network whose model quantization noise satisfies preset conditions; for each nonlinear operator in the set of nonlinear operators, performing an operator replacement step: using range to obtain the operator, determining the operator range of the nonlinear operator; using a replacement polynomial to obtain the operator, generating polynomial coefficients of at least one polynomial replacement formula based on the operator range, obtaining at least one polynomial replacement formula as a replacement operator, wherein the polynomial replacement formula includes: multiple multiplication and addition operators, and the polynomial coefficients are the coefficients of the multiple multiplication and addition operators; replacing the nonlinear operator with the replacement operator; performing integer quantization processing on the target neural network after operator replacement to obtain a quantized model; and storing the quantized model.
[0009] Secondly, some embodiments of this disclosure provide an integer quantization model storage device, comprising: a determining unit configured to determine a set of nonlinear operators in the network structure of a target neural network whose model quantization noise satisfies preset conditions; an execution unit configured to perform an operator replacement step for each nonlinear operator in the set of nonlinear operators: determining the operator range of the nonlinear operator using a range calculation; determining the operator range using a replacement polynomial; generating polynomial coefficients of at least one polynomial replacement formula based on the operator range, thereby obtaining at least one polynomial replacement formula as a replacement operator, wherein the polynomial replacement formula includes: multiple multiplication and addition operators, and the polynomial coefficients are the coefficients of the multiple multiplication and addition operators; replacing the nonlinear operator with the replacement operator; a quantization processing unit configured to perform integer quantization processing on the target neural network after operator replacement to obtain a quantized model; and a storage unit configured to store the quantized model.
[0010] Thirdly, some embodiments of this disclosure provide a task processing method, including: acquiring task data for a target task, wherein the target task includes: an image processing task and a point cloud processing task, and the task data includes: an image corresponding to the image processing task and point cloud data corresponding to the point cloud processing task; in response to determining that a pre-stored quantization model is a quantization model for processing the image processing task, inputting the image corresponding to the image processing task into the quantization model to output an image processing result, wherein the quantization model is generated by a method described in any implementation of the first aspect, and multiple substitution operators corresponding to the target neural network corresponding to the quantization model are determined based on range-finding operators and substitution polynomial-finding operators; in response to determining that the pre-stored quantization model is a quantization model for processing the point cloud processing task, inputting the point cloud data corresponding to the point cloud processing task into the quantization model to output a point cloud processing result.
[0011] Fourthly, some embodiments of this disclosure provide a task processing apparatus, including: an acquisition unit configured to acquire task data for a target task, wherein the target task includes: an image processing task and a point cloud processing task, and the task data includes: an image corresponding to the image processing task and point cloud data corresponding to the point cloud processing task; a first input unit configured to, in response to determining that a pre-stored quantization model is a quantization model for processing the image processing task, input the image corresponding to the image processing task into the quantization model to output an image processing result, wherein the quantization model is generated by a method described in any implementation of the first aspect, and a plurality of substitution operators corresponding to the target neural network corresponding to the quantization model are determined based on a range-finding operator and a substitution polynomial-finding operator; and a second input unit configured to, in response to determining that a pre-stored quantization model is a quantization model for processing the point cloud processing task, input the point cloud data corresponding to the point cloud processing task into the quantization model to output a point cloud processing result.
[0012] Fifthly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any of the implementations of the first and third aspects.
[0013] Sixthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any of the implementations of the first and third aspects.
[0014] The above embodiments of this disclosure have the following beneficial effects: the integer quantization model storage method of some embodiments of this disclosure improves the execution speed and efficiency of the model. Specifically, the reason for the decrease in model execution speed and efficiency is that data is repeatedly exchanged and copied between the CPU and the integer quantization acceleration device, which indirectly leads to a decrease in model execution speed and thus a decrease in execution efficiency. Based on this, the integer quantization model storage method of some embodiments of this disclosure first determines the set of nonlinear operators in the network structure of the target neural network whose model quantization noise meets preset conditions, so as to screen out nonlinear operators with poor corresponding model quantization effects and corresponding value distributions that differ greatly from uniform distributions. Here, the nonlinear operators in the set of nonlinear operators are the operators to be replaced. Then, for each nonlinear operator in the set of nonlinear operators, an operator replacement step is performed: First, the operator is obtained by using range calculation to determine the operator range of the nonlinear operator. By determining the operator range of the nonlinear operator, it is convenient to subsequently determine at least one polynomial replacement formula that is equivalent to the nonlinear operator within the operator range. The second step involves using a substitution polynomial to obtain the operator. Based on the range of operators mentioned above, at least one polynomial substitution formula can be accurately generated, resulting in at least one polynomial substitution formula as the substitution operator. This polynomial substitution formula includes multiple multiplication and addition operators. The polynomial coefficients are the coefficients of these multiple multiplication and addition operators. The third step involves replacing the aforementioned nonlinear operators with the substitution operators. Here, replacing nonlinear operators with corresponding polynomial substitution formulas effectively solves the problems of high model quantization noise and significant differences between the value distribution and the uniform distribution, making subsequent integer quantization processing more accurate. Next, integer quantization processing is performed on the target neural network after operator replacement to obtain the quantized model. Quantizing the target neural network after operator replacement reduces memory bandwidth and storage space usage. Finally, the quantized model is stored for subsequent use. Therefore, by replacing operators with high model quantization noise, nonlinear operators do not need to be copied to the CPU for processing; they only need to be processed within the integer quantization acceleration device, improving execution speed and efficiency. At the same time, this method reduces the amount of data computation, thereby reducing memory usage. Attached Figure Description
[0015] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0016] Figure 1This is a flowchart of some embodiments of the integer quantization model storage method according to the present disclosure;
[0017] Figure 2 These are flowcharts of some embodiments of the task processing method according to this disclosure;
[0018] Figure 3 These are schematic diagrams illustrating the structure of some embodiments of the integer quantization model storage device according to the present disclosure;
[0019] Figure 4 These are schematic diagrams of the structure of some embodiments of the task processing apparatus according to the present disclosure;
[0020] Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0022] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0023] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0026] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0027] refer to Figure 1The diagram illustrates a flow 100 of some embodiments of an integer quantization model storage method according to the present disclosure. This integer quantization model storage method includes the following steps:
[0028] Step 101: Determine the set of nonlinear operators in the network structure of the target neural network that satisfy the preset conditions for model quantization noise.
[0029] In some embodiments, the execution entity (e.g., an electronic device) of the above-described integer quantization model storage method can determine a set of nonlinear operators in the network structure of the target neural network whose model quantization noise satisfies preset conditions. The target neural network can be a neural network that processes various tasks. For example, various tasks can include, but are not limited to, at least one of the following: computer vision tasks and natural language processing tasks. Computer vision tasks can be tasks related to computer vision (CV). For example, computer vision tasks can be, but are not limited to, one of the following: object detection tasks, image segmentation tasks, and image generation tasks. Natural language processing tasks can be tasks related to natural language processing (NLP). For example, natural language processing tasks can be, but are not limited to, one of the following: text part-of-speech tagging tasks, sentiment analysis tasks, text generation tasks, speech recognition tasks, and speech-to-text conversion tasks. For object detection tasks, the target neural network can be an object detection neural network. For image segmentation tasks, the target neural network can be an image segmentation neural network. For speech recognition tasks, the target neural network can be a speech recognition neural network. For sentiment analysis tasks, the target neural network can be a sentiment analysis neural network. Model quantization noise can be error information resulting from the quantization processing of nonlinear operators. The preset condition can be that the nonlinear operator is an operator in the network structure whose model quantization noise is less than a predetermined quantization noise.
[0030] Here, the value range of each nonlinear operator in the set of nonlinear operators differs significantly from the value range corresponding to a uniform distribution, making it impossible to support nonlinear operators on integer quantization acceleration devices due to excessive model quantization noise. For example, an integer quantization acceleration device could be an artificial intelligence chip, such as a TPU (Tensor Processing Unit).
[0031] Step 102: For each nonlinear operator in the above set of nonlinear operators, perform the operator replacement step:
[0032] Step 1021: Use the range to find the operator and determine the operator range of the above nonlinear operator.
[0033] In some embodiments, the execution entity may utilize a range-finding operator to determine the operator range of the nonlinear operator. The range-finding operator can be an operator used to determine the operator range of the nonlinear operator. The operator range can be the range of values that the nonlinear operator can take. The operator range of the nonlinear operator may include: a domain range and a value range.
[0034] Optionally, each nonlinear operator in the set of nonlinear operators has a corresponding range-finding operator. This range-finding operator can be a pre-defined operator by those skilled in the art to determine the range of the corresponding nonlinear operator.
[0035] As an example, the aforementioned execution entity can obtain the operator by acquiring a pre-constructed range that has a one-to-one correspondence with the aforementioned nonlinear operator. Then, the aforementioned execution entity can determine the operator range of the aforementioned nonlinear operator by executing the range to obtain the operator.
[0036] In some optional implementations of certain embodiments, determining the operator range of the nonlinear operator by utilizing the range-finding operator may include the following steps:
[0037] The first step is to use the above range to obtain the operator and determine the actual physical range of the above nonlinear operator.
[0038] The range determination operators include: the actual physical range determination operator and the domain range determination operator. The actual physical range determination operator can characterize the mapping relationship between the actual physical range of the nonlinear operator and the information range of the target neural network output information. The domain range determination operator can characterize the mapping relationship between the actual physical range of the nonlinear operator and the domain range.
[0039] For example, the target neural network is an image segmentation neural network. The nonlinear operator is a power function operator. The operator for determining the actual physical range can characterize the mapping relationship between the actual physical range of the power function operator and the range of segmentation results output by the image segmentation neural network. The aforementioned operators for determining the actual physical range and the operator for determining the domain range can be pre-defined operators. The operator for determining the actual physical range can be based on prediction-related information in each dimension of the network output vector matrix of each sub-network of the target neural network to determine the actual physical range of the nonlinear operator. The operator for determining the domain range can be based on the operator characteristics of the nonlinear operator and the range of actual physical range values to determine the domain range of the nonlinear operator.
[0040] As an example, the aforementioned execution entity can execute the actual physical range range determination operator to determine the actual physical range range corresponding to the aforementioned nonlinear operator.
[0041] The second step is to use the above range to obtain the operator, and determine the domain range of the above nonlinear operator based on the above actual physical value range.
[0042] As an example, the aforementioned execution entity can perform a domain range determination operator to determine the domain range of the aforementioned nonlinear operator.
[0043] Optionally, the nonlinear operators in the above set of nonlinear operators are one of the following: exponential operators, logarithmic operators, non-integer power function operators, and composite operators based on at least one of exponential, logarithmic, and power function operators.
[0044] In practice, composite operators can be composite functions based on exponential operators. Composite operators can also be composite functions based on logarithmic operators. Composite functions can also be composite functions based on power function operators. For example, a composite operator could be 1 / (1+e^(-1 / 2)). x ).
[0045] Optionally, determining the actual physical range of the nonlinear operator may include the following steps:
[0046] The first step is to determine the bounding box information of the prediction rectangle corresponding to the above nonlinear operator in response to the determination that the above nonlinear operator is the target exponential operator.
[0047] The target index operator described above represents the operational relationship between the prediction vector and the target ratio. That is, assuming the target ratio is a and the prediction vector is b, then a = e. b The aforementioned target ratio refers to the size ratio between the predicted bounding box and the prior anchor box. The prior anchor box can be a pre-set candidate box used for subsequent object prediction. The number of prior anchor boxes can be at least one. The bounding box information corresponding to the predicted bounding box of the target exponential operator can be the matrix information of the vector matrix output by the target exponential operator. The matrix information (i.e., bounding box information) can include: the matrix size information of the vector matrix (i.e., the bounding box size information of the predicted bounding box), and the matrix angle information of the vector matrix (i.e., the bounding box angle information of the predicted bounding box). The predicted bounding box can be a multi-dimensional bounding box, and the corresponding bounding box angle information is multi-dimensional angle information. For example, the predicted bounding box can be a three-dimensional bounding box, and the corresponding bounding box angle information is three-dimensional angle information. The aforementioned prediction vector is the output vector of the internal sub-networks included in the aforementioned target neural network. For example, the aforementioned internal sub-network can be the backbone network in the target neural network.
[0048] It should be noted that the target exponent operator is an operator that characterizes the operational relationship between the predicted vector and the target ratio, based on the target neural network being an image processing-related network. The aforementioned image processing-related network can be, but is not limited to, at least one of the following: an image segmentation neural network, or a target object detection neural network for an image. Based on the target neural network being an image processing-related network, the target neural network needs to output the corresponding prediction matrix information according to the anchor boxes.
[0049] The second step is to determine the actual physical range of the nonlinear operator based on the above-mentioned border information.
[0050] As an example, the aforementioned execution entity can determine the actual physical range of the aforementioned nonlinear operator based on the aforementioned border information, including:
[0051] The first step, in response to determining that the aforementioned bounding box information is the frame angle information of the predicted rectangle, is to determine the first angle range corresponding to the predicted rectangle. The first angle range includes at least one first sub-angle range. For example, the at least one first sub-angle range could be three first sub-angle ranges for each dimension of three-dimensional space.
[0052] The second step is to normalize the above-mentioned at least one first sub-angle range to obtain at least one first sub-angle normalized range.
[0053] The third step is to determine the second angle range corresponding to the prior anchor frame, wherein the second angle range includes at least one second sub-angle range. Furthermore, the second sub-angle range within the at least one second sub-angle range has a dimensional correspondence with the first sub-angle range within the at least one first sub-angle range.
[0054] For example, at least one second sub-angle range includes: a second sub-angle range for the first dimension, a second sub-angle range for the second dimension, and a second sub-angle range for the third dimension. At least one first sub-angle range includes: a first sub-angle range for the first dimension, a first sub-angle range for the second dimension, and a first sub-angle range for the third dimension.
[0055] The fourth step is to normalize the above-mentioned at least one second sub-angle range to obtain at least one normalized range of the second sub-angle.
[0056] Fifth step: Based on the above-mentioned at least one first sub-angle normalization range and the above-mentioned at least one second sub-angle normalization range, generate at least one angle ratio range.
[0057] As an example, for each of the first sub-angle normalization ranges within at least one first sub-angle normalization range, the aforementioned execution entity can divide the maximum value of the first sub-angle normalization range by the minimum value of the second sub-angle normalization range in the same dimension to obtain the maximum value of the angle ratio range. The aforementioned execution entity can also divide the minimum value of the first sub-angle normalization range by the maximum value of the second sub-angle normalization range in the same dimension to obtain the minimum value of the angle ratio range. Based on the minimum and maximum values, the angle ratio range for the aforementioned first sub-angle normalization range is determined.
[0058] The sixth step is to determine at least one of the above multiple angle ratio ranges as the actual physical value range of the above target exponential operator.
[0059] Optionally, determining the actual physical range of the nonlinear operator based on the bounding box information may include the following steps:
[0060] In response to determining that the border information is at least one of a width range, a height range, and a depth range, the width range, the height range, and the depth range are determined as a first width range, a first height range, and a first depth range.
[0061] The width range refers to the width range of the predicted bounding box. The height range refers to the height range of the predicted bounding box. The depth range refers to the depth range of the predicted bounding box.
[0062] For example, the target neural network is an object detection network. The input image of the target neural network (i.e., the image to be detected) has a width of W pixels, a height of H pixels, and a depth of D pixels. The first width range is [2 pixels, W pixels]. The first height range is [2 pixels, H pixels]. The first depth range is [2 pixels, D pixels].
[0063] The second step is to normalize the first width range, the first height range, and the first depth range to obtain the first normalized width range, the first normalized height range, and the first normalized depth range.
[0064] For example, the first width range is [2 pixels, W pixels]. The first height range is [2 pixels, H pixels]. The first depth range is [2 pixels, D pixels]. The corresponding first normalized width range is [2 / W, 1]. The corresponding first normalized height range is [2 / H, 1]. The corresponding first normalized depth range is [2 / D, 1].
[0065] The third step is to determine the width, height, and depth ranges of the prior anchor frame, which will serve as the second width range, second height range, and second depth range.
[0066] As an example, the aforementioned execution entity can determine the width, height, and depth ranges of the prior anchor boxes based on the image information of the input image of the target neural network, serving as the second width range, second height range, and second depth range. The aforementioned image information includes: image size and image depth.
[0067] For example, the target neural network is an object detection network. The input image of the target neural network (i.e., the image to be detected) has a width of W pixels, a height of H pixels, and a depth of D pixels. The second width range is [2 pixels, W pixels]. The second height range is [2 pixels, H pixels]. The second depth range is [2 pixels, D pixels].
[0068] The fourth step is to normalize the second width range, the second height range, and the second depth range to obtain the second normalized width range, the second normalized height range, and the second normalized depth range.
[0069] For example, the second width range is [2 pixels, W pixels]. The second height range is [2 pixels, H pixels]. The second depth range is [2 pixels, D pixels]. The corresponding second normalized width range is [2 / W, 1]. The corresponding second normalized height range is [2 / H, 1]. The corresponding second normalized depth range is [2 / D, 1].
[0070] Fifth step: Based on the first normalized width range, the first normalized height range, the first normalized depth range, the second normalized width range, the second normalized height range, and the second normalized depth range, generate the height ratio range, width ratio range, and depth ratio range.
[0071] As an example, the aforementioned execution entity can divide the maximum value of the first normalized width range by the minimum value of the second normalized width range to obtain the maximum value of the width ratio range. The aforementioned execution entity can divide the minimum value of the first normalized width range by the maximum value of the second normalized width range to obtain the minimum value of the width ratio range. Based on the minimum and maximum values of the aforementioned width ratio ranges, the width ratio range is determined. The aforementioned execution entity can divide the maximum value of the first normalized height range by the minimum value of the second normalized height range to obtain the maximum value of the height ratio range. The aforementioned execution entity can divide the minimum value of the first normalized height range by the maximum value of the second normalized height range to obtain the minimum value of the height ratio range. Based on the minimum and maximum values of the aforementioned height ratio ranges, the height ratio range is determined. The aforementioned execution entity can divide the maximum value of the first normalized depth range by the minimum value of the second normalized depth range to obtain the maximum value of the depth ratio range. The aforementioned execution entity can divide the minimum value of the first normalized depth range by the maximum value of the second normalized depth range to obtain the minimum value of the depth ratio range. The depth ratio range is determined based on the minimum and maximum values of the aforementioned depth ratio range.
[0072] For example, the first normalized width range is [2 / W, 1]. The first normalized height range is [2 / H, 1]. The first normalized depth range is [2 / D, 1]. The second normalized width range is [2 / W, 1]. The second normalized height range is [2 / H, 1]. The second normalized depth range is [2 / D, 1]. The width ratio range can be [2 / W, W / 2]. The height ratio range can be [2 / H, H / 2]. The depth ratio range can be [2 / D, D / 2].
[0073] The sixth step is to determine at least one of the above-mentioned height ratio range, the above-mentioned width ratio range, and the above-mentioned depth ratio range as the actual physical value range of the above-mentioned target index operator.
[0074] Optionally, determining the domain range of the nonlinear operator based on the actual physical range may include the following steps:
[0075] The first step is to determine the height range, width range, and depth range of the predicted vector based on the height ratio range, width ratio range, and depth ratio range mentioned above, and use these as the third height range, third width range, and third depth range.
[0076] There is an exponential correspondence between the height ratio range and the third height range. There is also an exponential correspondence between the width ratio range and the third width range. Finally, there is an exponential correspondence between the depth ratio range and the third depth range.
[0077] As an example, the aforementioned execution entity can determine the third height range corresponding to the height ratio range based on the exponential correspondence. The aforementioned execution entity can determine the third width range corresponding to the width ratio range based on the exponential correspondence. The aforementioned execution entity can determine the third depth range corresponding to the depth ratio range based on the exponential correspondence.
[0078] For example, the width ratio range can be [2 / W, W / 2]. The height ratio range can be [2 / H, H / 2]. The depth ratio range can be [2 / D, D / 2]. Then the third width range is [ln(2 / W), ln(W / 2)]. The third height range is [ln(2 / H), ln(H / 2)]. The third depth range is [ln(2 / D), ln(D / 2)].
[0079] The second step is to determine at least one of the above-mentioned third height range, the above-mentioned third width range, and the above-mentioned third depth range as the domain range of the above-mentioned target index operator.
[0080] Step 1022: Obtain the operator using the substitution polynomial. Based on the range of the operators mentioned above, generate the polynomial coefficients of at least one polynomial substitution formula to obtain at least one polynomial substitution formula as the substitution operator.
[0081] In some embodiments, the execution entity may utilize a substitution polynomial to obtain operators, and based on the range of operators, generate polynomial coefficients for at least one polynomial substitution formula, thereby obtaining at least one polynomial substitution formula as a substitution operator. The polynomial substitution formula includes multiple multiplication and addition operators. The polynomial coefficients are the coefficients of the multiple multiplication and addition operators.
[0082] For example, the polynomial substitution formula is ax + bx 2 +cx 3 +d. The polynomial substitution formula corresponds to three multiplication and addition operators. These three operators are: ax corresponds to the multiplication and addition operator, bx... 2 Corresponding multiplication and addition operators and cx 3 Corresponding to multiplication and addition operators.
[0083] For example, when the polynomial substitution formula is a composite function, it can be one of the following: f2 = f1 * b * x + b1, f3 = f2 * a * x + a1. Here, f1 can be a first-order multiplication-addition operator, i.e., f1 = c * x + c1. f2 = c * b * x 2 +c1*b*x+b1. That is, f2 can be a first-order multiplication and addition operator corresponding to the cascaded f1. f3=c*b*a*x 3 +c1*b*a*x 2+b1*x+a1. That is, f3 corresponds to the first-order multiplication and addition operators cascaded with f2. Through multiple cascaded first-order multiplication and addition operators, a higher-order polynomial of x is realized. Therefore, the higher the power, the more times the composition occurs, and so on.
[0084] Optionally, the above-mentioned at least one polynomial substitution formula includes: a first polynomial substitution formula for the width.
[0085] Based on the aforementioned operator range, generating at least one polynomial coefficient for a polynomial substitution formula to obtain at least one polynomial substitution formula as a substitution operator may include the following steps:
[0086] For the third width range within the corresponding domain, perform the first polynomial substitution formula generation step:
[0087] Sub-step 1: Select the target number of values from the third width range mentioned above as the first target values to obtain the first target value set.
[0088] The target number can be N times the number of quantization intervals, or it can be determined through multiple trials. N is an integer. For example, N is 10.
[0089] Sub-step 2: Determine the first exponential function value corresponding to each first target value in the aforementioned first target value set, and obtain the first exponential function value set.
[0090] As an example, the aforementioned execution entity can substitute each first target value in the first target value set into the target exponential operator to generate the first exponential function value, thereby obtaining the first exponential function value set.
[0091] Sub-step 3: Based on the first target value set and the first exponential function value set, a polynomial substitution formula for the third width range can be generated in various ways as the first polynomial substitution formula.
[0092] Optionally, the aforementioned at least one polynomial substitution formula includes: a second polynomial substitution formula for height. And the process of generating at least one polynomial substitution formula based on the aforementioned operator range to obtain at least one polynomial substitution formula as a substitution operator may include the following steps:
[0093] For the third height range within the corresponding domain, perform the second polynomial substitution formula generation step:
[0094] Sub-step 1: Select a number of target values from the above third height range as the second target values to obtain the second target value set.
[0095] Sub-step 2: Determine the second exponential function value corresponding to each second target value in the above second target value set, and obtain the second exponential function value set.
[0096] As an example, the aforementioned execution entity can substitute each second target value in the second target value set into the target exponential operator to generate the second exponential function value, thus obtaining the second exponential function value set.
[0097] Sub-step 3: Based on the second target value set and the second exponential function value set, a polynomial substitution formula for the third height range can be generated in various ways as the second polynomial substitution formula.
[0098] Optionally, the above-mentioned at least one polynomial substitution formula includes: a third polynomial substitution formula for depth.
[0099] Based on the aforementioned operator range, generating at least one polynomial coefficient for a polynomial substitution formula to obtain at least one polynomial substitution formula as a substitution operator may include the following steps:
[0100] For the third depth range within the corresponding domain, perform the third polynomial substitution formula generation step:
[0101] Sub-step 1: Select the target number of values from the above third depth range as the third target values to obtain the third target value set.
[0102] Sub-step 2: Determine the third exponential function value corresponding to each third target value in the above third target value set, and obtain the third exponential function value set.
[0103] As an example, the aforementioned execution entity can substitute each third target value in the third target value set into the target exponential operator to generate a third exponential function value, thus obtaining a set of third exponential function values.
[0104] Sub-step 3: Based on the aforementioned third target value set and the aforementioned third exponential function value set, a polynomial substitution formula for the aforementioned third depth range can be generated in various ways, serving as the third polynomial substitution formula.
[0105] Optionally, generating a polynomial substitution formula for the third width range based on the first target numerical set and the first exponential function value set, as the first polynomial substitution formula, may include the following steps:
[0106] The first step is to determine the order of the initial formula.
[0107] For example, the initial order of the formula is 3. The formula corresponding to the initial order is: ax + bx 2 +cx 3+d.
[0108] For example, the initial formula has an order of 4. The initial order corresponds to the formula ax + bx. 2 +cx 3 +dx 4 +e.
[0109] The second step is to generate the error value based on the initial formula order:
[0110] Sub-step 1: Based on the first target value set and the first exponential function value set, generate polynomial coefficients for the order of the initial formula.
[0111] As an example, the aforementioned execution entity can input the first target numerical set and the first exponential function value set into the target coefficient solution formula to generate polynomial coefficients for the order of the aforementioned initial formula.
[0112] The formula for solving polynomial coefficients can be:
[0113] A = Y(X) T X) -1 X T ,
[0114] Where A can be a row vector of polynomial coefficients. Y can be the vector corresponding to the first set of exponential function values. X can be the vector corresponding to the first target numerical set. T It can be the transpose of X.
[0115] Here, the formula for solving the polynomial coefficients can be obtained through the following steps:
[0116] The first sub-step is to determine the initial formula order corresponding to the formula.
[0117] Here, assuming the initial formula order is K, the formula corresponding to the initial formula order can be:
[0118] f(x) = a0 + a1x 1 +a2x 2 +···+a k-1 x k-1 +a k x k .
[0119] Where, a0, a1, a2, a k-1 a k These are the polynomial coefficients.
[0120] The second sub-step involves determining the error value for N sample points. These N sample points include N first target values and N first exponential function values. The error value can be of various types depending on the error calculation formula. For example, if the error calculation formula is the mean square error formula, the error value is the mean square error.
[0121] As an example, the aforementioned execution entity can determine the mean squared error value for N sample points using the following formula:
[0122]
[0123] Where MSE is the mean squared error value.
[0124] The third sub-step involves differentiating the polynomial coefficients based on the error, making the derivative equal to zero, which minimizes the error. Setting the derivative to zero and simplifying, we obtain the following equation:
[0125]
[0126] Among them, [a0 a1 a2…a k-1 a k [y0 y1 y2…y] can be a vector of polynomial coefficients. N-1 y N ] can be a vector of N values of the first exponential function.
[0127] This is the Vandermonde matrix.
[0128] The fourth sub-step involves using the Vandermonde matrix X. T The determinant of X is not zero, the inverse matrix exists, and the formula for solving the coefficients of the generator polynomial.
[0129] Sub-step 2: Generate a polynomial formula for the above polynomial coefficients, as the first initial polynomial formula.
[0130] As an example, the aforementioned executing entity can substitute the polynomial coefficients into the initial polynomial formula to obtain the first initial polynomial formula. Here, the initial polynomial formula is a formula whose coefficients are not yet determined.
[0131] Sub-step 3 involves determining the error value for the first initial polynomial formula and the target exponential operator based on the aforementioned first target numerical set and first exponential function value set. The specific implementation of sub-step 3 can be found in the process of generating the polynomial coefficient solution formula, and will not be elaborated upon here.
[0132] Sub-step 4: In response to determining that the error value is less than a predetermined threshold, the first initial polynomial formula is determined as the first polynomial substitution formula. For example, the predetermined threshold is 0.8.
[0133] Third, in response to determining that the error value is greater than or equal to the predetermined threshold, the initial formula order is added to a preset value to generate the summed formula order, which is then used as the initial formula order. The error value generation step continues. For example, the preset value is 1.
[0134] Optionally, generating a polynomial substitution formula for the third width range based on the first target numerical set and the first exponential function value set, as the first polynomial substitution formula, may include the following steps:
[0135] The first step is to obtain the correspondence between the preset error value and the order of the formula.
[0136] It should be noted that each formula order has a corresponding preset error value. The correspondence between the formula order and the corresponding preset error value is based on multiple tests conducted by relevant technical personnel. The preset error value can also be a range of error values.
[0137] For example, a preset error value of 0.4 corresponds to a formula of order 2. A preset error value of 0.3 corresponds to a formula of order 3. A preset error value of 0.2 corresponds to a formula of order 4. A preset error value of 0.1 corresponds to a formula of order 7.
[0138] For example, a preset error value of [0.4, 0.5] corresponds to a formula of order 2. A preset error value of [0.3, 0.4] corresponds to a formula of order 3. A preset error value of [0.2, 0.3] corresponds to a formula of order 4. A preset error value of [0.1, 0.2] corresponds to a formula of order 7.
[0139] The second step is to determine the order of the formula corresponding to the target preset error value based on the above correspondence, and use it as the order of the polynomial formula.
[0140] The third step is to determine the coefficients of the polynomial substitution formula for the third width range based on the order of the polynomial formula.
[0141] As an example, the aforementioned execution entity can determine the coefficients of the polynomial substitution formula for the third width range based on the order of the polynomial formula using a polynomial fitting method. Specific polynomial fitting methods will not be elaborated upon here.
[0142] The fourth step is to generate the first polynomial substitution formula based on the order and coefficients of the polynomial formulas mentioned above.
[0143] As an example, the aforementioned execution entity can substitute the formula coefficients into the order of the aforementioned polynomial formula to generate the first polynomial substitution formula.
[0144] Step 1023: Replace the above nonlinear operator with the above replacement operator.
[0145] In some embodiments, the execution entity may replace the nonlinear operator with the replacement operator.
[0146] Step 103: Perform integer quantization on the target neural network after operator replacement to obtain the quantized model.
[0147] In some embodiments, the aforementioned execution entity may perform integer quantization on the target neural network after operator replacement to obtain a quantized model.
[0148] As an example, the aforementioned execution entity can convert the 32-bit floating-point numbers (e.g., model weights) in the target neural network after operator replacement into 8-bit fixed-point numbers to obtain the converted model, which serves as the quantized model.
[0149] Step 104: Store the above quantization model.
[0150] In some embodiments, the execution entity may store the quantization model.
[0151] As an example, the aforementioned execution entity can store the model parameters of the quantization model in a predetermined database in the form of files.
[0152] The above embodiments of this disclosure have the following beneficial effects: the integer quantization model storage method of some embodiments of this disclosure improves the execution speed and efficiency of the model. Specifically, the reason for the decrease in model execution speed and efficiency is that data is repeatedly exchanged and copied between the CPU and the integer quantization acceleration device, which indirectly leads to a decrease in model execution speed and thus a decrease in execution efficiency. Based on this, the integer quantization model storage method of some embodiments of this disclosure first determines the set of nonlinear operators in the network structure of the target neural network whose model quantization noise meets preset conditions, so as to screen out nonlinear operators with poor corresponding model quantization effects and corresponding value distributions that differ greatly from uniform distributions. Here, the nonlinear operators in the set of nonlinear operators are the operators to be replaced. Then, for each nonlinear operator in the set of nonlinear operators, an operator replacement step is performed: First, the operator is obtained by using range calculation to determine the operator range of the nonlinear operator. By determining the operator range of the nonlinear operator, it is convenient to subsequently determine at least one polynomial replacement formula that is equivalent to the nonlinear operator within the operator range. The second step involves using a substitution polynomial to obtain the operator. Based on the range of operators mentioned above, at least one polynomial substitution formula can be accurately generated, resulting in at least one polynomial substitution formula as the substitution operator. This polynomial substitution formula includes multiple multiplication and addition operators. The polynomial coefficients are the coefficients of these multiple multiplication and addition operators. The third step involves replacing the aforementioned nonlinear operators with the substitution operators. Here, replacing nonlinear operators with corresponding polynomial substitution formulas effectively solves the problems of high model quantization noise and significant differences between the value distribution and the uniform distribution, making subsequent integer quantization processing more accurate. Next, integer quantization processing is performed on the target neural network after operator replacement to obtain the quantized model. Quantizing the target neural network after operator replacement reduces memory bandwidth and storage space usage. Finally, the quantized model is stored for subsequent use. Therefore, by replacing operators with high model quantization noise, nonlinear operators do not need to be copied to the CPU for processing; they only need to be processed within the integer quantization acceleration device, improving execution speed and efficiency. At the same time, this method reduces the amount of data computation, thereby reducing memory usage.
[0153] refer to Figure 2 The flowchart 200 illustrates some embodiments of a task processing method according to the present disclosure. The task processing method includes the following steps:
[0154] Step 201: Obtain task data for the target task.
[0155] In some embodiments, the execution entity of the above-described task processing method (e.g., an electronic device) can acquire task data for a target task via wired or wireless means. The target task includes an image processing task and a point cloud processing task, and the task data includes an image corresponding to the image processing task and point cloud data corresponding to the point cloud processing task. The image processing task can be an image-based processing task (i.e., a task associated with computer vision). For example, an image processing task can be, but is not limited to, at least one of the following: an image segmentation task, or a target object recognition task for an image. The point cloud processing task can be a point cloud processing task. For example, a point cloud processing task can be, but is not limited to, at least one of the following: a ground point cloud segmentation task, or a target object recognition task for point cloud data.
[0156] Step 202: In response to determining that the pre-stored quantization model is the quantization model for processing the above image processing task, the image corresponding to the above image processing task is input into the above quantization model to output the image processing result.
[0157] In some embodiments, in response to determining that a pre-stored quantization model is a quantization model for processing the image processing task, the executing entity can input the image corresponding to the image processing task into the quantization model to output the image processing result. The quantization model is generated using an integer quantization model storage method according to some embodiments of this disclosure, and the multiple substitution operators corresponding to the target neural network of the quantization model are determined based on range-finding operators and substitution polynomial-finding operators.
[0158] For example, the image processing task is image segmentation, and the image processing result is the image segmentation result.
[0159] Step 203: In response to determining that the pre-stored quantization model is the quantization model for processing the point cloud processing task, the point cloud data corresponding to the point cloud processing task is input into the quantization model to output the point cloud processing result.
[0160] In some embodiments, in response to determining that a pre-stored quantization model is a quantization model for processing the point cloud processing task, the execution entity may input the point cloud data corresponding to the point cloud processing task into the quantization model to output the point cloud processing result.
[0161] For example, the point cloud processing task is a ground point cloud segmentation task, and the point cloud processing result is ground point cloud data.
[0162] The above embodiments of this disclosure have the following beneficial effects: by using the task processing methods of some embodiments of this disclosure and utilizing the stored quantization model, the target task can be executed accurately and efficiently.
[0163] Further reference Figure 3 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an integer quantization model storage device, which are similar to... Figure 1 Corresponding to the method embodiments shown, this integer quantization model storage device can be specifically applied to various electronic devices.
[0164] like Figure 3 As shown, an integer quantization model storage device 300 includes: a determining unit 301, an execution unit 302, a quantization processing unit 303, and a storage unit 304. The determining unit 301 is configured to determine a set of nonlinear operators in the network structure of a target neural network whose model quantization noise satisfies preset conditions. The execution unit 302 is configured to perform an operator replacement step for each nonlinear operator in the set of nonlinear operators: determining the operator range of the nonlinear operator using range calculation; determining the operator range using a replacement polynomial; generating polynomial coefficients of at least one polynomial replacement formula based on the operator range, obtaining at least one polynomial replacement formula as a replacement operator, wherein the polynomial replacement formula includes multiple multiplication and addition operators, and the polynomial coefficients are the coefficients of the multiple multiplication and addition operators; replacing the nonlinear operator with the replacement operator; the quantization processing unit 303 is configured to perform integer quantization processing on the target neural network after operator replacement to obtain a quantized model; and the storage unit 304 is configured to store the quantized model.
[0165] It is understandable that the units recorded in the integer quantization model storage device 300 are related to the reference. Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the integer quantization model storage device 300 and the units contained therein, and will not be repeated here.
[0166] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a task processing apparatus, which are similar to... Figure 2 Corresponding to the method embodiments shown, this task processing device can be specifically applied to various electronic devices.
[0167] like Figure 4As shown, a task processing device 400 includes: an acquisition unit 401, a first input unit 402, and a second input unit 403. The acquisition unit 401 is configured to acquire task data for a target task, wherein the target task includes an image processing task and a point cloud processing task, and the task data includes an image corresponding to the image processing task and point cloud data corresponding to the point cloud processing task. The first input unit 402 is configured to, in response to determining that a pre-stored quantization model is a quantization model for processing the image processing task, input the image corresponding to the image processing task into the quantization model to output an image processing result. The quantization model is generated using an integer quantization model storage method according to some embodiments of this disclosure, and the multiple substitution operators corresponding to the target neural network of the quantization model are determined based on range-finding operators and substitution polynomial-finding operators. The second input unit 403 is configured to, in response to determining that a pre-stored quantization model is a quantization model for processing the point cloud processing task, input the point cloud data corresponding to the point cloud processing task into the quantization model to output a point cloud processing result.
[0168] It is understandable that the units described in the task processing device 400 and the reference Figure 2 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the task processing device 400 and the units contained therein, and will not be repeated here.
[0169] The following is for reference. Figure 5 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 500 suitable for implementing some embodiments of the present disclosure. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0170] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0171] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 5 Each box shown can represent a device or multiple devices as needed.
[0172] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from ROM 502. When the computer program is executed by processing device 501, it performs the functions defined in the methods of some embodiments of this disclosure.
[0173] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0174] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0175] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: determine a set of nonlinear operators in the network structure of the target neural network whose model quantization noise satisfies preset conditions; for each nonlinear operator in the aforementioned set of nonlinear operators, perform an operator replacement step: using range calculation to determine the operator range of the aforementioned nonlinear operator; using a replacement polynomial to calculate the operator, and based on the aforementioned operator range, generate polynomial coefficients of at least one polynomial replacement formula to obtain at least one polynomial replacement formula as a replacement operator, wherein the polynomial replacement formula includes: multiple multiplication and addition operators, and the polynomial coefficients are the coefficients of the multiple multiplication and addition operators; replace the aforementioned nonlinear operators with the aforementioned replacement operators; perform integer quantization processing on the target neural network after operator replacement to obtain a quantized model; and store the aforementioned quantized model. The process involves acquiring task data for a target task, where the target task includes an image processing task and a point cloud processing task. The task data includes an image corresponding to the image processing task and point cloud data corresponding to the point cloud processing task. In response to determining that a pre-stored quantization model is the quantization model for processing the image processing task, the image corresponding to the image processing task is input into the quantization model to output an image processing result. The quantization model is generated using an integer quantization model storage method according to some embodiments of this disclosure. Multiple substitution operators corresponding to the target neural network of the quantization model are determined based on range-finding operators and substitution polynomial-finding operators. In response to determining that a pre-stored quantization model is the quantization model for processing the point cloud processing task, the point cloud data corresponding to the point cloud processing task is input into the quantization model to output a point cloud processing result.
[0176] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0177] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0178] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a determination unit, an execution unit, a quantization processing unit, and a storage unit. The names of these units do not necessarily limit the specific unit; for example, a storage unit may also be described as "a unit that stores the aforementioned quantization model."
[0179] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0180] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for storing an integer quantization model, comprising: Obtain task data for a target task, wherein the target task includes: an image processing task and a point cloud processing task, and the task data includes: the image corresponding to the image processing task and the point cloud data corresponding to the point cloud processing task; In response to determining that a pre-stored quantization model is the quantization model for processing the image processing task, the image corresponding to the image processing task is input into the quantization model to output the image processing result. The multiple substitution operators corresponding to the target neural network of the quantization model are determined based on range-finding operators and substitution polynomial-finding operators. The quantization model is generated through the following steps: A set of nonlinear operators in the network structure of a target neural network that satisfy a preset condition for model quantization noise is determined. The model quantization noise is the error information resulting from quantization processing of the nonlinear operators. The nonlinear operators in the set of nonlinear operators are one of the following: exponential operators, logarithmic operators, non-integer power function operators, and composite operators based on at least one of exponential, logarithmic, and power function operators. For each nonlinear operator in the set of nonlinear operators, perform the operator replacement step: The range-finding operator is used to determine the operator range of the nonlinear operator. The range-finding operator is used to determine the operator range of the nonlinear operator, which is the range of values for the nonlinear operator. The operator range of the nonlinear operator includes: the domain range and the value range. The range-finding operator includes: the actual physical value range range operator and the domain range operator. Using the range to derive the operator, the actual physical range corresponding to the nonlinear operator is determined, including: In response to determining that the nonlinear operator is a target exponential operator, the bounding box information of the prediction rectangle corresponding to the target exponential operator is determined, wherein the target exponential operator represents the operational relationship between the prediction vector and the target ratio, the target ratio is the box size ratio between the prediction rectangle and the prior anchor box, and the prediction vector is the output vector of the internal sub-networks included in the target neural network; Based on the border information, the actual physical value range corresponding to the nonlinear operator is determined; The operator is obtained using the range, and the domain range of the nonlinear operator is determined based on the actual physical value range. Operators are obtained by using substitution polynomials. Based on the range of operators, polynomial coefficients of at least one polynomial substitution formula are generated to obtain at least one polynomial substitution formula as a substitution operator. The polynomial substitution formula includes multiple multiplication and addition operators, and the polynomial coefficients are the coefficients of multiple multiplication and addition operators. Replace the nonlinear operator with the replacement operator; The target neural network after operator replacement is subjected to integer quantization to obtain a quantized model; The quantization model is stored. In response to determining that the pre-stored quantization model is the quantization model for processing the point cloud processing task, the point cloud data corresponding to the point cloud processing task is input into the quantization model to output the point cloud processing result.
2. The method according to claim 1, wherein, The step of determining the actual physical value range corresponding to the nonlinear operator based on the border information includes: In response to determining that the border information is at least one of a width range, a height range, and a depth range, the width range, the height range, and the depth range are determined as a first width range, a first height range, and a first depth range; The first width range, the first height range, and the first depth range are normalized to obtain a first normalized width range, a first normalized height range, and a first normalized depth range. Determine the width, height, and depth ranges of the prior anchor frame, which will serve as the second width, second height, and second depth ranges. The second width range, the second height range, and the second depth range are normalized to obtain the second normalized width range, the second normalized height range, and the second normalized depth range; Based on the first normalized width range, the first normalized height range, the first normalized depth range, the second normalized width range, the second normalized height range, and the second normalized depth range, a height ratio range, a width ratio range, and a depth ratio range are generated; At least one of the height ratio range, the width ratio range, and the depth ratio range is determined as the actual physical value range of the target exponential operator.
3. The method according to claim 2, wherein, Determining the domain range of the nonlinear operator based on the actual physical value range includes: Based on the height ratio range, the width ratio range, and the depth ratio range, the height range, width range, and depth range of the prediction vector are determined as the third height range, the third width range, and the third depth range; At least one of the third height range, the third width range, and the third depth range is determined as the domain range of the target exponential operator.
4. The method according to claim 3, wherein, The at least one polynomial substitution formula includes: a first polynomial substitution formula for the width; and The step of generating at least one polynomial coefficient for a polynomial substitution formula based on the operator range, thereby obtaining at least one polynomial substitution formula as a substitution operator, includes: For the third width range within the corresponding domain, perform the first polynomial substitution formula generation step: Select a number of target values from the third width range to obtain the first target value set; Determine the first exponential function value corresponding to each first target value in the first target value set to obtain the first exponential function value set; Based on the first target value set and the first exponential function value set, a polynomial substitution formula for the third width range is generated as the first polynomial substitution formula.
5. The method according to claim 3, wherein, The at least one polynomial substitution formula includes: a second polynomial substitution formula for height; and The step of generating at least one polynomial coefficient for a polynomial substitution formula based on the operator range, thereby obtaining at least one polynomial substitution formula as a substitution operator, includes: For the third height range within the corresponding domain, perform the second polynomial substitution formula generation step: Select a number of target values from the third height range to obtain the second target value set; Determine the second exponential function value corresponding to each second target value in the second target value set to obtain the second exponential function value set; Based on the second target value set and the second exponential function value set, a polynomial substitution formula for the third height range is generated as the second polynomial substitution formula.
6. The method according to claim 3, wherein, The at least one polynomial substitution formula includes: a third polynomial substitution formula for depth; and The step of generating at least one polynomial coefficient for a polynomial substitution formula based on the operator range, thereby obtaining at least one polynomial substitution formula as a substitution operator, includes: For the third depth range within the corresponding domain, perform the third polynomial substitution formula generation step: Select a number of target values from the third depth range to obtain the third target value set; Determine the third exponential function value corresponding to each third target value in the third target value set to obtain the third exponential function value set; Based on the third target value set and the third exponential function value set, a polynomial substitution formula for the third depth range is generated as the third polynomial substitution formula.
7. The method according to claim 4, wherein, The step of generating a polynomial substitution formula for the third width range based on the first target value set and the first exponential function value set, as the first polynomial substitution formula, includes: Determine the order of the initial formula; Based on the initial formula order, perform the error value generation step: Based on the first target numerical set and the first exponential function value set, generate polynomial coefficients for the initial formula order; Generate a polynomial formula for the polynomial coefficients as the first initial polynomial formula. Based on the first target value set and the first exponential function value set, determine the error value for the first initial polynomial formula and the target exponential operator; In response to determining that the error value is less than a predetermined threshold, the first initial polynomial formula is determined as the first polynomial substitution formula; In response to determining that the error value is greater than or equal to the predetermined threshold, the initial formula order is added to a preset value to generate an added formula order, which is used as the initial formula order, and the error value generation step is continued.
8. The method according to claim 4, wherein, The step of generating a polynomial substitution formula for the third width range based on the first target value set and the first exponential function value set, as the first polynomial substitution formula, includes: Obtain the correspondence between preset error values and formula order; Based on the aforementioned correspondence, the order of the formula corresponding to the target preset error value is determined as the order of the polynomial formula; Based on the order of the polynomial formula, determine the formula coefficients of the polynomial substitution formula for the third width range; Based on the order and coefficients of the polynomial formula, a first polynomial substitution formula is generated.
9. An integer quantization model storage device, comprising: The acquisition unit is configured to acquire task data for a target task, wherein the target task includes an image processing task and a point cloud processing task, and the task data includes an image corresponding to the image processing task and point cloud data corresponding to the point cloud processing task. The first input unit is configured to, in response to determining that a pre-stored quantization model is the quantization model for processing the image processing task, input the image corresponding to the image processing task into the quantization model to output the image processing result. The multiple substitution operators corresponding to the target neural network of the quantization model are determined based on range-finding operators and substitution polynomial-finding operators. The quantization model is generated through the following steps: determining a set of nonlinear operators in the network structure of the target neural network whose model quantization noise satisfies preset conditions, wherein the model quantization noise is the error information resulting from the quantization processing of the nonlinear operators. The nonlinear operators in the set of nonlinear operators are one of the following: exponential operators, logarithmic operators, non-integer power function operators, and composite operators based on at least one of exponential, logarithmic, and power function operators; for each nonlinear operator in the set of nonlinear operators, an operator substitution step is performed: using a range-finding operator to determine the operator range of the nonlinear operator, wherein the range-finding operator is an operator used to determine the operator range of the nonlinear operator, the operator range is the range of values of the nonlinear operator, the operator range of the nonlinear operator includes: the domain range and the value range, and the range-finding operator includes: the actual physical value range. The method for determining the range of operators and the domain range includes: using the range-determining operator to determine the actual physical value range corresponding to the nonlinear operator, including: in response to determining that the nonlinear operator is a target exponential operator, determining the bounding box information of the prediction rectangle corresponding to the target exponential operator, wherein the target exponential operator represents the operational relationship between the prediction vector and the target ratio, the target ratio is the ratio of the box size between the prediction rectangle and the prior anchor box, and the prediction vector is the output vector of the internal subnetworks included in the target neural network; and determining the actual physical value range corresponding to the nonlinear operator based on the bounding box information. The physical value range is defined; the operator is obtained using the range, and the domain range of the nonlinear operator is determined based on the actual physical value range; the operator is obtained using a substitution polynomial, and the polynomial coefficients of at least one polynomial substitution formula are generated based on the operator range, resulting in at least one polynomial substitution formula as the substitution operator, wherein the polynomial substitution formula includes: multiple multiplication and addition operators, and the polynomial coefficients are the coefficients of multiple multiplication and addition operators; the nonlinear operator is replaced with the substitution operator; the target neural network after operator substitution is subjected to integer quantization processing to obtain a quantized model; the quantized model is stored. The second input unit is configured to, in response to determining that a pre-stored quantization model is the quantization model for processing the point cloud processing task, input the point cloud data corresponding to the point cloud processing task into the quantization model to output the point cloud processing result.
10. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.
11. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Model quantification method, server, electronic equipment and medium
CN114065622A
Model quantification method and device, equipment and storage medium
CN114936619A