A method, system and readable storage medium for deployment of a machine learning model
By performing distributed compression and transformation on the weight data of machine learning models, the problem of low deployment efficiency on lightweight embedded devices is solved, enabling the efficient application of machine learning models on embedded devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2022-12-08
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, machine learning models are inefficient to deploy on lightweight embedded devices and cannot be deployed efficiently.
The weight data is compressed based on the weight data distribution of the machine learning model using a compression device, generating a compressed model with reduced bit count. This compressed model is then converted to floating-point numbers for computation on an embedded device, enabling lightweight deployment.
This effectively reduces the storage space requirements for weight data, improves the deployment efficiency of the model on lightweight embedded devices, and ensures the application of the model in embedded devices.
Smart Images

Figure CN115829056B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of model deployment, in particular to a machine learning model deployment method, system and readable storage medium. BACKGROUND
[0002] At present, a mature machine learning model is deployed into an embedded device to make the machine learning model realize application in the embedded device. However, the memory volume of the machine learning model is generally large, and the efficiency of deploying the machine learning model into a lightweight (i.e. small memory) embedded device is low, and the machine learning model may even be unable to be deployed into the lightweight embedded device. Therefore, how to efficiently deploy the machine learning model in the lightweight embedded device has become a problem to be solved. SUMMARY
[0003] The present application provides a machine learning model deployment method, system and readable storage medium, which can improve the lightweight deployment efficiency of the machine learning model.
[0004] To solve the above technical problems, the technical scheme adopted by the present application is to provide a machine learning model deployment method applied to a compression device in a machine learning model deployment system, the machine learning model deployment method comprising: obtaining a machine learning model from an artificial intelligence open platform, compressing weight data of the machine learning model based on the distribution of the weight data of the machine learning model to obtain a compressed machine learning model; wherein the sum of the bit numbers of all weight data of the compressed machine learning model is less than the sum of the bit numbers of all weight data of the machine learning model before compression; and sending the compressed machine learning model to an embedded device to be deployed, so that the embedded device receives the compressed machine learning model, converts the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performs operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0005] To solve the above technical problems, another technical solution adopted by the present application is to provide a machine learning model deployment method applied to an embedded device in a machine learning model deployment system, the machine learning model deployment method comprising: receiving a compressed machine learning model sent by a compression device from the compression device; wherein the compression device is configured to obtain a machine learning model from an artificial intelligence open platform, compress weight data of the machine learning model based on a distribution of the weight data of the machine learning model, and obtain a compressed machine learning model; wherein the sum of the bit numbers of all weight data of the compressed machine learning model is less than the sum of the bit numbers of all weight data of the machine learning model before compression; converting the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performing operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0006] To solve the above technical problems, another technical solution adopted by the present application is to provide a machine learning model deployment system, the deployment system comprising an artificial intelligence open platform, a compression device, and an embedded device, the artificial intelligence open platform being connected to the compression device, the artificial intelligence open platform being connected to the compression device, the artificial intelligence open platform being configured to train a machine learning model; the compression device being configured to obtain the machine learning model from the artificial intelligence open platform, compress weight data of the machine learning model based on a distribution of the weight data of the machine learning model, and obtain a compressed machine learning model; wherein the sum of the bit numbers of all weight data of the compressed machine learning model is less than the sum of the bit numbers of all weight data of the machine learning model before compression; the compressed machine learning model being sent to an embedded device to be deployed; the embedded device being connected to the compression device, the embedded device being configured to receive the compressed machine learning model, convert the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and perform operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0007] To solve the above technical problems, another technical solution adopted by the present application is to provide a computer readable storage medium for storing a computer program, the computer program being configured to implement the machine learning model deployment method in the above technical solution when executed by a processor.
[0008] By the above scheme, the beneficial effects of the present application are: the compression device obtains the machine learning model from the artificial intelligence open platform, compresses the trained machine learning model based on the distribution of the weight data of the machine learning model, so that the sum of the bit numbers of all weight data of the compressed machine learning model is less than the sum of the bit numbers of all weight data of the machine learning model before compression; then the weight data of the received compressed machine learning model is operated through the embedded device, so as to realize the lightweight deployment of the machine learning model on the embedded device; by quantizing the large amount of weight data of the machine learning model through the compression device, the storage space occupied by the weight data can be reduced before deployment to the embedded device, the machine learning model is compressed, thereby solving the problem that the weight data of the machine learning model occupies a large storage space and cannot be deployed on a lightweight embedded device, improving the efficiency of model deployment, and enabling a larger machine learning model to be applied in a lightweight embedded device. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:
[0010] Figure 1 is a flowchart of an embodiment of the machine learning model deployment method provided by the present application;
[0011] Figure 2 is a flowchart of an embodiment of step 12 provided by the present application;
[0012] Figure 3 is a flowchart of another embodiment of step 12 provided by the present application;
[0013] Figure 4 is a flowchart of step 37 provided by the present application;
[0014] Figure 5 is a flowchart of another embodiment of the machine learning model deployment method provided by the present application;
[0015] Figure 6 is a flowchart of still another embodiment of the machine learning model deployment method provided by the present application;
[0016] Figure 7 is a structural diagram of an embodiment of the machine learning model deployment system provided by the present application;
[0017] Figure 8is a structural schematic diagram of an embodiment of the computer readable storage medium provided in the present application. DETAILED DESCRIPTION
[0018] The present application will be further described below in conjunction with the accompanying drawings and embodiments. It is particularly pointed out that the following embodiments are only used to illustrate the present application, but do not limit the scope of the present application. Similarly, the following embodiments are only part of the embodiments of the present application, not all embodiments, and all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of the present application.
[0019] In the present application, the term "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0020] It should be noted that the terms "first", "second", "third" in the present application are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second", "third" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically limited. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0021] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the machine learning model deployment method provided in the present application. In the following, the machine learning model deployment method in the present embodiment is introduced in the interaction mode of artificial intelligence platform, compression device and embedded device. The method comprises:
[0022] Step 11: training the machine learning model by the artificial intelligence open platform.
[0023] Training the machine learning model by the artificial intelligence open platform.
[0024] The compression device obtains the machine learning model from the artificial intelligence open platform, compresses the weight data of the machine learning model based on the distribution of the weight data of the machine learning model, and obtains the compressed machine learning model.
[0025] The compression device can be a computer that obtains the machine learning model from the artificial intelligence open platform, compresses the weight data of the machine learning model based on the distribution of the weight data of the machine learning model, and then obtains the compressed machine learning model, so that the sum of the bit numbers of all the weight data of the compressed machine learning model is less than the sum of the bit numbers of all the weight data of the machine learning model before compression, i.e., the sum of the bit numbers of all the quantized weight data is less than the sum of the bit numbers of all the weight data before quantization.
[0026] Specifically, in an embodiment, referring to Figure 2 , Figure 2 is a flowchart of an embodiment of step 12 provided by the present application, and the method comprises:
[0027] Step 21: Obtain a plurality of weight data of a machine learning model.
[0028] The compression process of the machine learning model is also the quantization process of the weight data in the machine learning model, which can be obtained by the artificial intelligence open platform and is not limited herein.
[0029] Step 22: Equidistantly divide the plurality of weight data to obtain at least two weight division intervals.
[0030] The weight data is equidistantly divided to obtain at least two weight division intervals; specifically, the interval ranges of each weight division interval are the same, the minimum value and the maximum value of the weight data of the machine learning model can be counted first to obtain the numerical interval [w min , w max ] of the weight data, wherein w min is the minimum value of the weight data, and w max is the maximum value of the weight data, and then the numerical interval [w min , w max ] is equidistantly divided into at least two weight division intervals.
[0031] It can be understood that the number of weight division intervals is generally not more than half of the total number of weight data, so as to ensure that there is an appropriate amount of weight data in each weight division interval, avoid the number of weight data falling in each weight division interval being too small, and reduce the poor weight quantization effect caused by the small number of weight data. The specific number of weight division intervals can be self-defined according to actual conditions, for example, 4096, which is not limited herein.
[0032] Step 23: Count the number of weight data in each weight division interval.
[0033] The weight division intervals can be counted by histogram, so as to obtain the number of weight data in each weight division interval.
[0034] Step 24: Adjust the weight number data to obtain calibrated number data.
[0035] The number of weight data corresponding to each weight division interval can be adjusted by limiting the number of weights contained in each weight division interval within a preset number interval, so as to obtain calibrated number data, thereby avoiding the situation that the number of weights in the weight division interval is too concentrated, resulting in the need to occupy a large number of shared weights after weight quantization, and also avoiding the situation that the number of weights in the weight division interval is too sparse, resulting in too large error of weight quantization. Understandably, the specific value of the preset number interval can be set according to actual conditions, which is not limited here.
[0036] Step 25: Based on the calibrated number data corresponding to each weight division interval, the plurality of weight data is re-divided to obtain at least two weight quantization intervals, so that the number of weight data in each weight quantization interval is balanced.
[0037] The value interval [w min , w max ] of the plurality of weight data can be re-divided based on the calibrated number data corresponding to each weight division interval to obtain at least two weight quantization intervals, so that each weight quantization interval obtained after division will contain balanced weight data, so that the weight data of the compressed machine learning model is scattered and distributed at equal distances compared with the weight data of the machine learning model before compression, thereby realizing non-uniform quantization of weight data, so that the weight data in each weight quantization interval can share the same weight value, thereby reducing the quantization error of weight, improving the weight quantization precision, and further improving the compression effect of the machine learning model, improving the precision of the compressed machine learning model, thereby improving the deployment efficiency of the machine learning model in the embedded device, and ensuring the deployment quality and effect of the compressed machine learning model in the embedded device; taking division into K weight quantization intervals as an example, the number of weight data contained in each weight quantization interval can be 1 / K of the total amount of weight data.
[0038] It can be understood that the number of weight quantization intervals is the same as the number of target storage characters corresponding to all quantized weight data, and the number of target storage characters can be set according to actual application conditions, for example: the weight data is quantized to obtain eight-bit data quantization results, that is, the number of target storage characters corresponding to all quantized weight data is 256, and then 256 weight quantization intervals can be set to divide the weight data.
[0039] Step 26: quantizing the weight data in the weight quantization interval to make the sum of the bit numbers of all quantized weight data less than the sum of the bit numbers of all weight data before quantization.
[0040] The sum of the bit numbers of all quantized weight data is less than the sum of the bit numbers of all weight data before quantization. Specifically, quantizing the weight data in the weight quantization interval can quantize all weight data in each weight quantization interval into the same shared weight value, and the shared weight value is used as the weight data of the compressed machine learning model. Further, using the shared weight value to represent all compressed weight data in the weight quantization interval can reduce the storage space of the weight data and achieve compression of the machine learning model.
[0041] In a specific embodiment, the average value of all weight data in the weight quantization interval can be calculated, and the average value is used as the quantized weight data corresponding to each weight data in the weight quantization interval, that is, the shared weight value corresponding to the weight quantization interval, or the average value of the maximum value of all weight data in the weight quantization interval and the minimum value of all weight data in the weight quantization interval is calculated, and the average value is used as the quantized weight data corresponding to each weight data in the weight quantization interval, or the median value of all weight data in the weight quantization interval is calculated, and the median value is used as the quantized weight data corresponding to each weight data in the weight quantization interval; in other embodiments, other calculation methods can be used to obtain the shared weight value, which is not limited here.
[0042] Step 13: The compression device sends the compressed machine learning model to the embedded device to be deployed.
[0043] Step 14: The embedded device receives the compressed machine learning model, converts the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performs operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0044] The compression device can send the compressed machine learning model to the embedded device to be deployed. During the deployment process, the embedded device can convert the shared weight values (i.e., the weight data of the compressed machine learning model) into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and then perform operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device. In a specific implementation, the step of operating the weight data of the machine learning model before compression can include: performing model inference on the machine learning model based on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0045] After the machine learning model is trained, the machine learning model can be deployed to the embedded device to realize the application of the machine learning model. By quantizing the large amount of weight data of the machine learning model through the compression device, the storage space occupied by the weight data can be reduced before deployment to the embedded device, the machine learning model can be compressed, the problem that the weight data of the machine learning model occupies a large storage space and cannot be deployed on a lightweight embedded device can be solved, and the efficiency of model deployment can be improved, so that a larger machine learning model can be applied in a lightweight embedded device.
[0046] The deployment method of the machine learning model in the present embodiment will be described below taking the machine learning model as a license plate recognition model and the embedded device as a camera. The artificial intelligence open platform can first train the license plate recognition model to obtain a mature license plate recognition model with license plate recognition capability. Then, the compression device obtains the license plate recognition model from the artificial intelligence open platform, compresses the license plate recognition model based on the distribution of the weight data of the license plate recognition model, obtains a compressed license plate recognition model, and sends the compressed license plate recognition model to the camera, so that the camera can operate the weight data of the compressed license plate recognition model, thereby deploying a larger license plate recognition model on a lightweight camera device.
[0047] The embodiment trains a machine learning model through an artificial intelligence open platform, compresses the trained machine learning model based on the distribution of weight data of the machine learning model, and obtains a compressed machine learning model with equidistantly scattered weight data. Then, the weight data of the compressed machine learning model received is operated through an embedded device, so as to realize lightweight deployment of the machine learning model on the embedded device, reduce transmission cost, and improve deployment efficiency. Moreover, the embodiment can first divide the plurality of weight data at equal intervals to obtain at least two weight division intervals, then count the number of weight data in each weight division interval to obtain weight quantity data, and then adjust, divide, and quantize the number of weight data corresponding to each weight division interval, so as to realize quantization of the weight data. The weight distribution can be used as a basis to control the quantization result, effectively avoiding the situation that the weight data is too dense or too sparse, thereby improving the precision of weight quantization, improving the compression effect of the machine learning model, and further improving the precision of the compressed machine learning model, thereby improving the deployment efficiency of the machine learning model in the embedded device, ensuring the deployment quality of the compressed machine learning model in the embedded device, and ensuring the deployment effect of the compressed machine learning model in the embedded device. Moreover, the weight quantization can be realized without using a clustering algorithm, which can improve the efficiency of weight quantization, shorten the compression time of the machine learning model by the compression device, and further improve the overall deployment efficiency of the machine learning model.
[0048] Please refer to Figure 3 , Figure 3 is a flowchart of another embodiment of step 12 provided by the present application, and the method comprises the following steps:
[0049] Step 31: Obtain a plurality of weight data of a machine learning model.
[0050] Step 31 is the same as step 21 in the above embodiment, and will not be described here.
[0051] Step 32: Divide the plurality of weight data at equal intervals to obtain at least two weight division intervals.
[0052] Step 32 is the same as step 22 in the above embodiment, and will not be described here.
[0053] Step 33: Count the number of weight data in each weight division interval.
[0054] Step 33 is the same as step 23 in the above embodiment, and will not be described here.
[0055] Step 34: Numerical limit processing is performed on the number of weight data corresponding to each weight division interval to obtain first weight quantity data.
[0056] The number of weight data corresponding to each weight division interval is numerically limited to obtain the first weight quantity data. Specifically, the quantity of one weight data can be taken from the quantity of weight data corresponding to all weight division intervals in turn to obtain the current quantity. It is then determined whether the current quantity meets the preset conditions. If the current quantity meets the preset conditions, the current quantity is not adjusted. If the current quantity does not meet the preset conditions, the current quantity is adjusted to the preset minimum value when it falls within the first quantity range, the current quantity is not adjusted when it falls within the second quantity range, and the current quantity is adjusted to the preset maximum value when it falls within the third quantity range.
[0057] In one specific implementation, when the current quantity is a preset value, it can be determined that the current quantity meets the preset conditions, and the preset value can be 0; the preset value is less than the minimum value of the first quantity range, the maximum value of the first quantity range is less than the minimum value of the second quantity range, the maximum value of the second quantity range is less than the minimum value of the third quantity range; the minimum value of the second quantity range is a preset minimum value, and the maximum value of the second quantity range is a preset maximum value.
[0058] Furthermore, the maximum value of the first quantity range can be a preset minimum value, and the maximum value of the second quantity range can be a preset maximum value, using thresh. min This indicates the minimum preset value, thresh max This indicates the preset maximum value; that is, the first quantity range can be (0, thresh) min The second quantity range can be [thresh] min thresh max The third quantity range can be (thresh) max The preset minimum and maximum values can be set based on experience or actual conditions, and are not limited here.
[0059] When the current quantity is 0, it means there is no weight data within the weight division interval, and no weight quantization operation is needed, so no adjustment to the weight data is required; when the current quantity falls within the first quantity range (0, thresh...), it means there is no weight data within the weight division interval, and no weight data needs to be adjusted. min If the number of weights within the specified range is too small, then the current number will be increased to the preset minimum value (thresh). min ; the current quantity falls within the second quantity range [thresh] min thresh maxWhen the current quantity falls in the third quantity range (threshmax, ∞), it indicates that the weight quantity corresponding to the weight division interval is too large, and the current quantity is adjusted to a preset maximum value thresh max , and is shown in the following formula (1) :
[0060]
[0061] In the above formula (1), thresh min represents a preset minimum value, thresh max represents a preset maximum value, h j represents the quantity of weight data corresponding to the weight division interval, h j represents the first weight quantity data, wherein j∈{0, 1, …, M-1}, which represents the label of the weight division interval, and M is the quantity of the weight division interval.
[0062] By setting the first quantity range, the second quantity range, and the third quantity range, the maximum value and the minimum value of the quantity of weight data corresponding to all weight division intervals can be limited, the problems of sparse quantity and too dense quantity of weight data in the weight division interval are solved, thereby reducing the quantization error and ensuring the weight quantization precision. Understandably, the above adjustment of the current quantity is only the processing of the quantity of weight data corresponding to each weight division interval, and the original weight data content is not adjusted.
[0063] Step 35: Transforming and normalizing the first weight quantity data to obtain calibrated quantity data.
[0064] The first weight quantity data is transformed and normalized to obtain calibrated quantity data. Specifically, a current operation function can be selected from a preset function library, and the first weight quantity data is input into the current operation function to obtain operation statistical data. Then, the operation statistical data is normalized to obtain calibrated quantity data.
[0065] Further, in the process of weight quantization, quantization error is prone to occur, which affects the weight quantization effect. Generally, Lp norm is used to measure the quantization error. The smaller the Lp norm, the smaller the quantization error, and the better the weight quantization effect. A suitable current operation function can be selected according to the current actual situation to reduce the corresponding Lp norm, thereby realizing the minimum control of different Lp norms to reduce the quantization error.
[0066] For example, a constant function f(x) = c (c > 0) can be selected as the current operation function, so that the subsequent weight quantization intervals are uniformly divided, so that the maximum value of the quantization error is minimized, that is, the L ∞ norm is minimized; or f(x) = x can be selected as the current operation function, so that each weight division interval will contain an equal number of weight data, so that the L 1 norm of the quantization error is reduced; in particular, f(x) = sqrt(x) can also be selected as the current operation function, so that the number of weight data in each weight division interval is proportional to the square root of the distribution density of the weight data, so that the L 2 norm of the quantization error is reduced; it can be understood that the above only takes a few operation functions as examples for illustration, and the current operation function is not limited, and can be selected according to actual conditions.
[0067] In a specific embodiment, the scheme for normalizing the operation statistical data to obtain the calibrated quantity data can include: adding all the operation statistical data to obtain a first value; then dividing the operation statistical data by the first value to obtain the corresponding calibrated weight quantity, which is specifically shown in the following formula (2):
[0068]
[0069] In the above formula (2), h j represents the calibrated quantity data, h j represents the operation statistical data, j ∈ {0, 1, …, M-1}.
[0070] It can be understood that in other embodiments, the number of weight data corresponding to each weight division interval can also be transformed to obtain second weight quantity data; then the second weight quantity data is subjected to numerical limit processing and normalization processing to obtain the calibrated quantity data; the order of numerical limit processing and data transformation is not limited; in other embodiments, the number of weight data corresponding to each weight division interval can also be directly subjected to numerical limit processing and normalization processing, that is, the number of weight data corresponding to each weight division interval is first subjected to numerical limit processing, and then normalization processing is performed to obtain the calibrated quantity data; or the number of weight data corresponding to each weight division interval is directly subjected to transformation and normalization processing, that is, the number of weight data corresponding to each weight division interval is first subjected to data transformation, and then normalization processing is performed to obtain the calibrated quantity data, which is not limited.
[0071] Step 36: Accumulating the calibrated quantity data to obtain an accumulated array.
[0072] The calibration quantity data is accumulated to obtain an accumulated array. Specifically, the accumulated array can include accumulated values, and the calibration quantity data can include at least two calibration values. A preset value (such as 0) can be determined as the first accumulated value in the accumulated array. The sum of the (n-1)th accumulated value and the (n-1)th calibration value in the calibration quantity data is calculated to obtain the nth accumulated value in the accumulated array, where n is an integer greater than or equal to 1 and less than or equal to the number of weight division intervals. The specific expression of the accumulated array is shown in the following formula (3):
[0073]
[0074] In the above formula (3), H n is the accumulated array. It can be understood that the accumulated array is an increasing array, that is, H n ≥H n-1 , and H0=1 and H M =1.
[0075] Step 37: Based on the accumulated value, the minimum weight value, the maximum weight value, and the number of weight division intervals, the plurality of weight data is re-divided to obtain at least two weight quantization intervals.
[0076] The minimum weight value is the minimum value of all weights in the weight data, and the maximum weight value is the maximum value of all weights in the weight data. Based on the accumulated value, the minimum weight value, the maximum weight value, and the number of weight division intervals, the plurality of weight data can be re-divided to obtain at least two weight quantization intervals, as shown in the following formula (5): Figure 4 The specific scheme for obtaining the weight quantization interval includes steps 41-44.
[0077] Step 41: Based on the number of at least two weight quantization intervals, a preset array is generated.
[0078] The preset array includes a plurality of preset values, and the preset array can be an increasing array. Specifically, the first preset value of the preset array is the reciprocal of the number of at least two weight quantization intervals, the tolerance of the preset array is the reciprocal of the number of at least two weight quantization intervals, the number of preset values in the preset array is the number of weight quantization intervals minus one, and the specific expression of the preset array is shown in the following formula (4):
[0079]
[0080] In the above formula (4), T k represents the preset array, and K represents the number of at least two weight quantization intervals.
[0081] Step 42: Based on the preset array and the accumulated array, accumulated values in the accumulated array that meet a preset segmentation condition are screened out to obtain candidate accumulated values.
[0082] Based on the preset array and the accumulated array, the accumulated values in the accumulated array satisfying the preset segmentation condition are filtered out to obtain candidate accumulated values. Specifically, a preset value in the preset array can be selected as a current preset value in sequence. Then, it is determined whether the current preset value falls in the comparison interval. If the current preset value falls in the comparison interval, it is determined that the preset segmentation condition is satisfied, and the two adjacent accumulated values are determined as the candidate accumulated values.
[0083] Further, the comparison interval can be composed of two adjacent accumulated values in the accumulated array. By traversing all the preset values in the preset array, the preset values are matched with the intervals composed of the accumulated array in sequence until the condition H n ≤T k ≤H n+1 is satisfied, the comparison interval [H n , H n+1 ] is obtained, and H n and H n+1 are determined as the candidate accumulated values.
[0084] Step 43: generating interval segmentation points based on the candidate accumulated values, the minimum weight value and the maximum weight value.
[0085] Based on the candidate accumulated values, the minimum weight value and the maximum weight value, interval segmentation points can be generated. The two adjacent accumulated values contained in the candidate accumulated values are referred to as a first accumulated value (i.e., H n ) and a second accumulated value (i.e., H n+1 ), respectively. Specifically, the current preset value can be subtracted from the first accumulated value to obtain a second value. Then, the second accumulated value is subtracted from the first accumulated value to obtain a third value. The second value is divided by the third value to obtain a fourth value. The fourth value is added to the number of items corresponding to the first accumulated value (i.e., the value of n) to obtain a fifth value. The maximum weight value is subtracted from the minimum weight value to obtain a sixth value. The sixth value is divided by the number of weight division intervals to obtain a seventh value. The fifth value is multiplied by the seventh value to obtain an eighth value. The eighth value is added to the minimum weight value to obtain the interval segmentation point. Specifically, as shown in the following formula (5):
[0086]
[0087] In the above formula (5), x k represents the interval segmentation point, w max represents the maximum weight value, and w min represents the minimum weight value.
[0088] Step 44: re-dividing the plurality of weight data based on all the interval segmentation points to obtain at least two weight quantization intervals.
[0089] The number of interval division points calculated by using the above formula (5) is the same as the number of preset values in the preset array, that is, the number of the plurality of weight quantization intervals minus one, so that the plurality of weight data is re-divided by using all the generated interval division points, and the plurality of weight data can be divided into the target number of weight quantization intervals, while ensuring that the number of weight data in each weight quantization interval is the same.
[0090] Step 38: quantizing the weight data in the weight quantization interval, so that the sum of the bit numbers of all the quantized weight data is less than the sum of the bit numbers of all the weight data before quantization.
[0091] Step 38 is the same as step 26 in the above embodiment, and will not be described here.
[0092] By limiting the number of weight data corresponding to each weight division interval, the embodiment can effectively avoid the situation that the quantized weight data is too dense or too sparse, and ensure the quantization accuracy. By performing function transformation and normalization processing on the number of weight data corresponding to each weight division interval, the corresponding operation function can be selected according to the actual demand to reduce the quantization error, so as to achieve a better quantization effect, improve the compression effect of the machine learning model, improve the deployment efficiency of the machine learning model in the embedded device, improve the accuracy of the compressed machine learning model, and ensure the deployment quality and effect of the compressed machine learning model in the embedded device. In addition, weight quantization can be realized without mean clustering algorithm, which can further improve the quantization efficiency, shorten the compression time of the compression device for the machine learning model, and thus improve the overall deployment efficiency of the machine learning model.
[0093] Please refer to Figure 5 , Figure 5 is a flowchart of another embodiment of the deployment method of the machine learning model provided by the present application, which is applied to the compression device in the deployment system of the machine learning model, that is, the present embodiment takes the compression device as the execution subject to introduce the deployment method of the machine learning model, and the method comprises:
[0094] Step 51: obtaining the machine learning model from the artificial intelligence open platform.
[0095] Step 52: compressing the weight data of the machine learning model based on the distribution of the weight data of the machine learning model, to obtain a compressed machine learning model.
[0096] Steps 51-52 are the same as steps 12 in the above embodiment, and will not be described here.
[0097] Step 53: sending the compressed machine learning model to the embedded device to be deployed, so that the embedded device receives the compressed machine learning model, converts the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performs operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0098] Step 53 is the same as steps 13-14 in the above embodiment, and will not be repeated here.
[0099] The embodiment compresses the trained machine learning model based on the distribution of the weight data of the machine learning model by using the compression device, so that the compressed machine learning model with equidistantly scattered weight data is obtained. The embedded device performs operations on the received weight data of the compressed machine learning model to realize the lightweight deployment of the machine learning model on the embedded device, reduce the transmission cost, and improve the deployment efficiency. The weight distribution is used to control the quantization result, effectively avoiding the situation of too dense and too sparse weight data, thereby improving the accuracy of weight quantization, enhancing the compression effect of the machine learning model, and further improving the accuracy of the compressed machine learning model, thereby improving the deployment efficiency of the machine learning model in the embedded device, ensuring the deployment quality of the compressed machine learning model in the embedded device, and ensuring the deployment effect of the compressed machine learning model in the embedded device.
[0100] Please refer to Figure 6 , Figure 6 is a flowchart of another embodiment of the machine learning model deployment method provided by the present application, which is applied to an embedded device in a machine learning model deployment system, i.e., the present embodiment introduces the machine learning model deployment method with the embedded device as the execution subject. The method comprises:
[0101] Step 61: receiving the compressed machine learning model sent by the compression device.
[0102] The compression device is configured to obtain the machine learning model from an artificial intelligence open platform, compress the weight data of the machine learning model based on the distribution of the weight data of the machine learning model, and obtain the compressed machine learning model. The sum of the bit numbers of all weight data of the compressed machine learning model is less than the sum of the bit numbers of all weight data of the machine learning model before compression. Specifically, the compression process of the compression device on the machine learning model is the same as step 12 in the above embodiment, and will not be repeated here.
[0103] Step 62: converting the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performing operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device.
[0104] Step 62 is the same as step 14 in the above embodiment, and will not be repeated here.
[0105] The embodiment converts the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performs operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device, thereby reducing transmission costs and improving deployment efficiency.
[0106] Please refer to Figure 7 , Figure 7 is a structural schematic diagram of an embodiment of the machine learning model deployment system provided by the present application. The machine learning model deployment system 70 includes an artificial intelligence open platform 71, a compression device 72, and an embedded device 73. The artificial intelligence open platform 71 is connected to the compression device 72, and the artificial intelligence open platform 71 is used to train a machine learning model.
[0107] The compression device 72 is used to obtain the machine learning model from the artificial intelligence open platform 71, compress the weight data of the machine learning model based on the distribution of the weight data of the machine learning model, obtain a compressed machine learning model, and send the compressed machine learning model to the embedded device 73 to be deployed. The sum of the bit numbers of all weight data of the compressed machine learning model is less than the sum of the bit numbers of all weight data of the machine learning model before compression.
[0108] The embedded device 73 is connected to the compression device 72, and the embedded device 73 is used to receive the compressed machine learning model, convert the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and perform operations on the floating-point numbers to complete the lightweight deployment of the machine learning model on the embedded device 73.
[0109] The embodiment can realize the lightweight deployment of the machine learning model in the embedded device through the cooperation between the artificial intelligence open platform, the compression device, and the embedded device, improve the deployment efficiency, and ensure the deployment quality at the same time.
[0110] Please refer to Figure 8 , Figure 8is a structural schematic diagram of an embodiment of the computer readable storage medium provided in the present application, the computer readable storage medium 80 is used to store a computer program 81, the computer program 81 is used to implement the deployment method of the machine learning model in the above embodiment when executed by a processor.
[0111] The computer readable storage medium 80 can be a server, a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0112] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the above-described device embodiment is only schematic, for example, the division of the module or unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0113] The unit described as a separate component can be or can not be physically separated, and the component displayed as a unit can be or can not be a physical unit, that is, it can be located in one place, or it can be distributed to a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.
[0114] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0115] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or, on the device for processing personal information, the personal information processing rules are informed by using obvious marks / information, and the personal authorization is obtained by means of pop-up information or asking the individual to upload his / her personal information. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.
[0116] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A method for deploying a machine learning model, characterized in that, The application discloses a compression device applied to a deployment system of a machine learning model, the deployment system further comprising an artificial intelligence open platform and an embedded device, the artificial intelligence open platform being connected with the compression device, and the compression device being connected with the embedded device, and the compression device comprises the following steps: obtaining a machine learning model from the artificial intelligence open platform, compressing weight data of the machine learning model based on a distribution of the weight data to obtain a compressed machine learning model; sending the compressed machine learning model to the embedded device to be deployed, so that the embedded device receives the compressed machine learning model, converts weight data of the compressed machine learning model into floating-point numbers with the same value range as weight data of the machine learning model before compression, and performs operation on the floating-point numbers to complete lightweight deployment of the machine learning model on the embedded device; wherein the step of compressing the weight data of the machine learning model based on the distribution of the weight data comprises the following steps: obtaining a plurality of weight data of a machine learning model, and equally dividing the plurality of weight data to obtain at least two weight division intervals; counting the number of weight data in each weight division interval, and performing numerical limit processing and normalization processing on the number of weight data corresponding to each weight division interval to obtain calibration quantity data corresponding to the weight division interval; based on the calibration quantity data corresponding to each weight division interval, re-dividing the plurality of weight data to obtain at least two weight quantization intervals, so that the number of weight data in each weight quantization interval is balanced; calculating a shared weight value of all weight data in each weight quantization interval, and taking the shared weight value as weight data of the compressed machine learning model to realize quantization processing of the weight data in the weight quantization interval, wherein the sum of bit numbers of all quantized weight data is less than the sum of bit numbers of all weight data before quantization.
2. The method of claim 1, wherein, The step of obtaining the calibration quantity data corresponding to the weight division interval further comprises: counting the number of weight data in each weight division interval, and performing transformation and normalization processing on the number of weight data corresponding to each weight division interval to obtain the calibration quantity data; or counting the number of weight data in each weight division interval, and performing numerical limit processing on the number of weight data corresponding to each weight division interval to obtain first weight quantity data; performing transformation and normalization processing on the first weight quantity data to obtain the calibration quantity data; or counting the number of weight data in each weight division interval, and performing transformation on the number of weight data corresponding to each weight division interval to obtain second weight quantity data; performing numerical limit processing and normalization processing on the second weight quantity data to obtain the calibration quantity data.
3. The method of claim 2, wherein, The step of performing numerical limit processing on the number of weight data corresponding to each weight division interval comprises: sequentially taking one weight data quantity from all the weight division interval corresponding weight data quantities, to obtain a current quantity; judging whether the current quantity meets a preset condition; if yes, not adjusting the current quantity; if no, adjusting the current quantity to a preset minimum value when the current quantity falls in a first quantity range, not adjusting the current quantity when the current quantity falls in a second quantity range, and adjusting the current quantity to a preset maximum value when the current quantity falls in a third quantity range.
4. The method of claim 2, wherein, The step of transforming and normalizing the first weight quantity data to obtain the calibration quantity data comprises: selecting a current operation function from a preset function library, and inputting the first weight quantity data into the current operation function to obtain operation statistical data; normalizing the operation statistical data to obtain the calibration quantity data.
5. The method of claim 1, wherein, The step of re-dividing the plurality of weight data based on the calibration quantity data corresponding to each weight division interval to obtain at least two weight quantization intervals comprises: accumulating the calibration quantity data to obtain an accumulation array, the accumulation array comprising accumulation values; re-dividing the plurality of weight data based on the accumulation values, a weight minimum value, a weight maximum value, and the number of weight division intervals to obtain the at least two weight quantization intervals, the weight minimum value being the minimum value of all the weight data, and the weight maximum value being the maximum value of all the weight data.
6. The method of claim 5, wherein, The calibration quantity data comprises at least two calibration values, and the step of accumulating the calibration quantity data to obtain an accumulation array comprises: determining a preset value as the first accumulation value in the accumulation array; calculating the sum of an n-1th accumulation value and an n-1th calibration value in the calibration quantity data to obtain an nth accumulation value in the accumulation array, n being an integer greater than or equal to 1 and less than or equal to the number of weight division intervals.
7. The method of claim 5, wherein, The step of re-dividing the plurality of weight data based on the accumulation values, a weight minimum value, a weight maximum value, and the number of weight division intervals to obtain the at least two weight quantization intervals comprises: generating a preset array based on the number of the at least two weight quantization intervals; selecting accumulation values in the accumulation array that meet a preset segmentation condition based on the preset array and the accumulation array to obtain candidate accumulation values; generating interval segmentation points based on the candidate accumulation values, the weight minimum value, and the weight maximum value; re-dividing the plurality of weight data based on all the interval segmentation points to obtain the at least two weight quantization intervals.
8. The method of claim 7, wherein, The step of selecting accumulation values in the accumulation array that meet a preset segmentation condition based on the preset array and the accumulation array to obtain candidate accumulation values comprises: sequentially selecting a preset value from the preset array as a current preset value; determining whether the current preset value falls in a comparison interval, the comparison interval being composed of two adjacent accumulated values in the accumulated array; if yes, determining that the preset segmentation condition is met, and determining the two adjacent accumulated values as the candidate accumulated values. 9.A method for deploying a machine learning model, the method comprising: An embedded device applied to a deployment system of a machine learning model, comprising: receiving a compressed machine learning model sent by a compression device; wherein the compression device is configured to obtain a machine learning model from an artificial intelligence open platform, compress weight data of the machine learning model based on a distribution of the weight data of the machine learning model, and obtain a compressed machine learning model; wherein the sum of bit numbers of all weight data of the compressed machine learning model is less than the sum of bit numbers of all weight data of the machine learning model before compression; converting the weight data of the compressed machine learning model into floating-point numbers with the same value range as the weight data of the machine learning model before compression, and performing operations on the floating-point numbers to complete lightweight deployment of the machine learning model on the embedded device; wherein the step of compressing the weight data of the machine learning model based on the distribution of the weight data of the machine learning model comprises: obtaining a plurality of weight data of a machine learning model, and equally dividing the plurality of weight data to obtain at least two weight division intervals; counting the number of weight data in each weight division interval, and performing numerical limit processing and normalization processing on the number of weight data corresponding to each weight division interval to obtain calibration quantity data corresponding to the weight division interval; based on the calibration quantity data corresponding to each weight division interval, re-dividing the plurality of weight data to obtain at least two weight quantization intervals, so that the number of weight data in each weight quantization interval is balanced; calculating a shared weight value of all weight data in each weight quantization interval, and taking the shared weight value as weight data of a compressed machine learning model to realize quantization processing of the weight data in the weight quantization interval, wherein the sum of bit numbers of all quantized weight data is less than the sum of bit numbers of all weight data before quantization.
10. The method of claim 9, wherein, The step of performing operations on the floating-point numbers comprises: performing model inference on the machine learning model based on the floating-point numbers. 11.A system for deploying a machine learning model, characterized by, An artificial intelligence open platform, a compression device and an embedded device are provided, the artificial intelligence open platform is connected with the compression device, and the artificial intelligence open platform is configured to train a machine learning model; the compression device is configured to obtain the machine learning model from the artificial intelligence open platform, compress weight data of the machine learning model based on a distribution of the weight data of the machine learning model, and obtain a compressed machine learning model; and send the compressed machine learning model to an embedded device to be deployed. The embedded device is connected with the compression device, and the embedded device is configured to receive the compressed machine learning model, convert weight data of the compressed machine learning model into floating-point numbers with the same value range as weight data of the machine learning model before compression, and perform operation on the floating-point numbers to complete lightweight deployment of the machine learning model on the embedded device. The step of compressing the weight data of the machine learning model based on the distribution of the weight data of the machine learning model comprises: obtaining a plurality of weight data of a machine learning model, and equally dividing the plurality of weight data to obtain at least two weight division intervals; counting the number of weight data in each weight division interval, and performing numerical limit processing and normalization processing on the number of weight data corresponding to each weight division interval to obtain calibration quantity data corresponding to the weight division interval; based on the calibration quantity data corresponding to each weight division interval, re-dividing the plurality of weight data to obtain at least two weight quantization intervals, so that the number of weight data in each weight quantization interval is balanced; calculating a shared weight value of all weight data in each weight quantization interval, and taking the shared weight value as weight data of a compressed machine learning model to realize quantization processing of the weight data in the weight quantization interval, wherein the sum of the bit numbers of all quantized weight data is less than the sum of the bit numbers of all weight data before quantization.
12. A computer readable storage medium for storing a computer program, characterized in that, The computer program, when executed by a processor, is configured to implement the machine learning model deployment method of any one of claims 1-10. The computer program, when executed by a processor, is configured to implement the machine learning model deployment method of any one of claims 1-10.
Citation Information
Patent Citations
Machine learning model training method and device and readable storage medium
CN114219095A
Method for increasing speed of running machine learning model by embedded device
CN114519432A