Neural network model deployment method and device, equipment, storage medium and computer program product

By performing structural analysis and singular value decomposition on the neural network model, and combining quantitative processing, the model is deployed to edge devices, solving the problem of resource constraints and achieving efficient edge device deployment and real-time response.

CN119990200APending Publication Date: 2025-05-13TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411889823.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to deploy complex neural network models to edge devices with limited storage space and computing resources, resulting in difficult to meet real-time response requirements.

Method used

By performing structural analysis of the neural network model to be deployed, the layer to be compressed is determined, and the weight matrix is singularly decomposed to obtain an approximate matrix, which is then quantized to reduce the number of parameters and bits, and finally deploy the model to the edge device.

Benefits of technology

It effectively reduces the number of parameters and computational complexity required by the model, allowing the neural network model to run efficiently on resource-constrained edge devices to meet the real-time response needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990200A_ABST
    Figure CN119990200A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of neural network model deployment, and discloses a neural network model deployment method and device, equipment, a storage medium and a computer program product, and the method comprises the steps: obtaining a to-be-deployed neural network model, carrying out the structural analysis of the to-be-deployed neural network model, and determining a to-be-compressed layer in the to-be-deployed neural network model; determining a weight matrix of the to-be-compressed layer, and performing singular value decomposition on the weight matrix to obtain an approximate matrix; compressing the to-be-deployed neural network model according to the approximate matrix; and performing quantification processing on the compressed to-be-deployed neural network model, and deploying the to-be-deployed neural network model after quantification processing to the target edge device. According to the method, the to-be-deployed model is subjected to structural analysis, the to-be-compressed layer is determined, the model is compressed by performing singular value decomposition on the weight matrix corresponding to the to-be-compressed layer, and the compressed model is quantified, so that the model can be deployed on the target edge equipment with limited storage and calculation capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of neural network model deployment technology, and in particular to a neural network model deployment method, apparatus, device, storage medium and computer program product. Background Art

[0002] With the rapid development of Industry 4.0 and IoT technology, various intelligent devices are widely used in industrial production, generating massive amounts of data with wide sources, large volumes and many types. Cloud computing has become the mainstream data processing method with its powerful computing power and storage resources. However, the surge in IoT devices has significantly increased the speed of data generation, and the demand for real-time response is growing. Since cloud computing relies on remote data centers, there is a high delay in data transmission, which makes it difficult to meet real-time requirements, and is also accompanied by problems such as high bandwidth consumption and privacy leakage. Based on this, edge intelligence has emerged as an emerging technology. It sinks data processing and computing power to edge devices close to the data source to achieve localized data processing, thereby effectively reducing transmission delays, reducing bandwidth consumption and ensuring data security.

[0003] The implementation of edge intelligence usually requires deploying models on edge devices. However, due to the limited storage space and computing resources of edge devices, it is often difficult to directly deploy complex models. Summary of the invention

[0004] The main purpose of this application is to provide a neural network model deployment method, apparatus, device, storage medium and computer program product, aiming to solve the problem of how to deploy a neural network model to an edge device with limited storage space and computing resources.

[0005] To achieve the above objectives, the present application proposes a neural network model deployment method, the method comprising:

[0006] Acquire a neural network model to be deployed, and perform structural analysis on the neural network model to be deployed to determine a layer to be compressed in the neural network model to be deployed;

[0007] Determine a weight matrix of the layer to be compressed, and perform singular value decomposition on the weight matrix to obtain an approximate matrix;

[0008] Compressing the neural network model to be deployed according to the approximate matrix;

[0009] The compressed neural network model to be deployed is quantized, and the quantized neural network model to be deployed is deployed to the target edge device.

[0010] In one embodiment, the step of performing singular value decomposition on the weight matrix to obtain an approximate matrix includes:

[0011] Decomposing the weight matrix to obtain a left singular matrix, a right singular matrix and a singular diagonal matrix;

[0012] Obtain a first singular value matrix based on a preset rank value, the left singular matrix, and the singular diagonal matrix;

[0013] Obtain a second singular value matrix based on the preset rank value, the right singular matrix, and the singular diagonal matrix;

[0014] An approximate matrix is ​​obtained according to the first singular value matrix and the second singular value matrix.

[0015] In one embodiment, before the step of obtaining a first singular value matrix based on a preset rank value, the left singular matrix and the singular diagonal matrix, the method further includes:

[0016] Determine the loss function and cost function corresponding to the neural network model to be deployed;

[0017] An auxiliary variable is determined according to the weight matrix, and a preset rank value is determined based on a preset regularization parameter, the auxiliary variable, the loss function, and the cost function.

[0018] In one embodiment, the step of quantizing the compressed neural network model to be deployed includes:

[0019] Obtain the target weight value and target activation value in the compressed neural network model to be deployed;

[0020] The compressed neural network model to be deployed is quantized based on the target weight value and the target activation value.

[0021] In one embodiment, the step of obtaining the target weight value and the target activation value in the compressed neural network model to be deployed includes:

[0022] Obtain the weight parameters of the compressed neural network model to be deployed;

[0023] Using a preset simulation quantization algorithm to simulate and quantize the compressed neural network model to be deployed based on a preset scaling factor and the weight parameter;

[0024] The target weight value and target activation value are obtained based on the simulation quantization results.

[0025] In one embodiment, the step of obtaining a target weight value and a target activation value based on the simulation quantization result includes:

[0026] Dequantizing the simulated quantization result to obtain an approximate model corresponding to the compressed neural network model to be deployed;

[0027] Sample data is obtained, and the approximate model is trained based on the sample data to obtain a target weight value and a target activation value.

[0028] In addition, to achieve the above purpose, the present application also proposes a neural network model deployment device, the device comprising:

[0029] A model acquisition module, used to acquire a neural network model to be deployed, and perform structural analysis on the neural network model to be deployed to determine a layer to be compressed in the neural network model to be deployed;

[0030] A data decomposition module, used to determine the weight matrix of the layer to be compressed, and perform singular value decomposition on the weight matrix;

[0031] A model compression module, used for compressing the neural network model to be deployed according to the approximate matrix;

[0032] The model deployment module is used to quantize the compressed neural network model to be deployed, and deploy the quantized neural network model to be deployed to the target edge device.

[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes a device, which system includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the neural network model deployment method as described above.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the neural network model deployment method described above are implemented.

[0035] In addition, in order to achieve the above-mentioned purpose, the present application also proposes a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the neural network model deployment method described above are implemented.

[0036] The present application proposes a neural network model deployment method, device, equipment, storage medium and computer program product, the method comprising: obtaining a neural network model to be deployed, and performing structural analysis on the neural network model to be deployed to determine the layer to be compressed in the neural network model to be deployed; determining the weight matrix of the layer to be compressed, and performing singular value decomposition on the weight matrix to obtain an approximate matrix; compressing the neural network model to be deployed according to the approximate matrix; quantizing the compressed neural network model to be deployed, and deploying the quantized neural network model to be deployed to the target edge device. The present application determines the layer to be compressed and the weight matrix corresponding to the layer to be compressed by performing structural analysis on the neural network model to be deployed, obtains an approximate matrix by using the weight matrix, and compresses the neural network model to be deployed by using the approximate matrix to reduce the number of parameters required for the model. The compressed neural network model to be deployed is then quantized to reduce the number of bits occupied by the parameters, so that the neural network model to be deployed can be deployed on the target edge device with limited storage space and computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0039] Figure 1 A flowchart of a first embodiment of the neural network model deployment method proposed in this embodiment;

[0040] Figure 2 This is an example diagram of singular value decomposition based on the LC algorithm in the neural network model deployment method proposed in this embodiment;

[0041] Figure 3 A flowchart of a second embodiment of the neural network model deployment method proposed in this embodiment;

[0042] Figure 4 A flowchart of quantization processing based on a pseudo quantizer in the neural network model deployment method proposed in this embodiment;

[0043] Figure 5 A structural diagram of a first embodiment of a neural network model deployment device provided in an embodiment of the present application;

[0044] Figure 6It is a schematic diagram of the structure of a device suitable for implementing the embodiments of the present application.

[0045] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0048] It should be noted that all directional indications in the embodiments of the present application (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0049] It is understandable that with the rapid development of Industry 4.0 and IoT technology, various intelligent devices are widely used in industrial production, generating massive data with wide sources, large volume and many types. Cloud computing has become the mainstream data processing method with its powerful computing power and storage resources. However, the surge in IoT devices has significantly increased the speed of data generation, and the demand for real-time response is growing. Since cloud computing relies on remote data centers, there is a high delay in the data transmission process, which makes it difficult to meet the real-time requirements, and is also accompanied by problems such as high bandwidth consumption and privacy leakage. Based on this, edge intelligence has emerged as an emerging technology. It realizes localized data processing by sinking data processing and computing power to edge devices close to the data source, thereby effectively reducing transmission delays, reducing bandwidth consumption and ensuring data security.

[0050] The implementation of edge intelligence usually requires deploying models on edge devices. However, due to the limited storage space and computing resources of edge devices, it is often difficult to directly deploy complex models.

[0051] Therefore, in order to solve the problem of how to deploy a neural network model to an edge device with limited storage space and computing resources, this embodiment proposes a neural network model deployment method, which includes: obtaining a neural network model to be deployed, and performing a structural analysis on the neural network model to be deployed to determine the layer to be compressed in the neural network model to be deployed; determining the weight matrix of the layer to be compressed, and performing singular value decomposition on the weight matrix to obtain an approximate matrix; compressing the neural network model to be deployed according to the approximate matrix; quantizing the compressed neural network model to be deployed, and deploying the quantized neural network model to be deployed to the target edge device. This embodiment determines the layer to be compressed and the weight matrix corresponding to the layer to be compressed by performing a structural analysis on the neural network model to be deployed, obtains an approximate matrix by using the weight matrix, and compresses the neural network model to be deployed by using the approximate matrix to reduce the number of parameters required for the model. Then, the compressed neural network model to be deployed is quantized to reduce the number of bits occupied by the parameters, so that the neural network model to be deployed can be deployed on the target edge device with limited storage space and computing resources.

[0052] For ease of understanding, the following combination Figures 1 to 6 The neural network model deployment method provided in the embodiments of the present application and the neural network model deployment method, apparatus, device, storage medium and computer program product provided in the following embodiments are introduced in detail.

[0053] The present application embodiment provides a neural network model deployment method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the neural network model deployment method proposed in this embodiment.

[0054] like Figure 1 As shown, the method includes:

[0055] Step S10: Obtain a neural network model to be deployed, and perform structural analysis on the neural network model to be deployed to determine a layer to be compressed in the neural network model to be deployed.

[0056] It should be noted that the execution subject of this embodiment can be a computing service system with neural network model deployment control, network communication and program running functions, such as a neural network model deployment device, or an electronic device capable of realizing the above functions. This embodiment uses a neural network model deployment device (hereinafter referred to as the device) for illustration, but does not specifically limit this embodiment.

[0057] It should also be noted that the above-mentioned neural network model to be deployed can be a neural network model trained in advance by the user. In this embodiment, the above-mentioned neural network model to be deployed can be a neural network model with a large number of parameters, such as a time series prediction model, etc. In this embodiment, the time series prediction model is used for explanation, but this embodiment is not specifically limited.

[0058] In addition, it can be understood that the neural network model is a neural network model composed of multiple layers, which can learn complex patterns of input data. For example, a convolutional neural network can be used for image classification, and a recurrent neural network can be used to process sequence data, such as text or time series data. The above-mentioned layer to be compressed can be a neural network that can be compressed in the above-mentioned neural network model to be deployed.

[0059] In a specific implementation, the user can load a trained neural network model into the above-mentioned device, which includes the structure of the neural network model (for example, how many layers the neural network model consists of) and the weight matrix corresponding to each layer in the neural network model. After obtaining the loaded neural network model, the above-mentioned device uses the neural network model as the neural network model to be deployed, and performs a structural analysis on the neural network model to be deployed: by analyzing the architecture of the model and the functions of the layers, determine which layers can be compressed. For example, the convolution layer usually has a large number of redundant parameters, which can be compressed by methods such as pruning or low-rank decomposition. The fully connected layer usually has a higher computational complexity and can be compressed by methods such as quantization. The above-mentioned device also needs to consider the impact of the selection of the compression layer on the performance of the model. For example, pruning the convolution layer may affect the feature extraction ability of the model, and quantizing the fully connected layer may cause the accuracy of the model to decrease. Therefore, the above-mentioned device needs to weigh the compression efficiency and model performance and select a suitable compression layer.

[0060] Step S20: Determine the weight matrix of the layer to be compressed, and perform singular value decomposition on the weight matrix to obtain an approximate matrix.

[0061] Step S30: compressing the neural network model to be deployed according to the approximate matrix.

[0062] It should be noted that the weight matrix can be a parameter matrix connecting different layers in a neural network, which determines the relationship between the input and output of the model. For example, in a convolution layer, each convolution kernel corresponds to a weight matrix, which is used to extract features in an image. The similarity matrix can be a low-rank approximation matrix obtained by performing singular value decomposition on the weight matrix. The similarity matrix can approximate the original matrix and reduce the number of parameters and computational complexity of the model.

[0063] In a specific implementation, after determining the layer to be compressed, the above-mentioned device obtains the corresponding weight matrix according to the layer to be compressed, and performs singular value decomposition on the corresponding weight matrix to obtain an approximate matrix of the weight matrix of the layer to be compressed, and uses the approximate matrix as the new weight matrix of the layer to be compressed to compress the above-mentioned model to be deployed. Since the similarity matrix is ​​a low-rank approximate matrix and approximates the original matrix, the number of parameters and the computational complexity of the model are reduced, and the function of the original model is retained.

[0064] Furthermore, in order to perform singular value decomposition on the weight matrix, the step of performing singular value decomposition on the weight matrix to obtain an approximate matrix includes:

[0065] Decomposing the weight matrix to obtain a left singular matrix, a right singular matrix and a singular diagonal matrix;

[0066] Obtain a first singular value matrix based on a preset rank value, the left singular matrix, and the singular diagonal matrix;

[0067] Obtain a second singular value matrix based on the preset rank value, the right singular matrix, and the singular diagonal matrix;

[0068] An approximate matrix is ​​obtained according to the first singular value matrix and the second singular value matrix.

[0069] It should be noted that the left singular matrix may be an orthogonal matrix whose column vectors are the left singular vectors of the weight matrix, the right singular matrix may be an orthogonal matrix whose column vectors are the right singular vectors of the weight matrix, and the singular vectors may be vectors corresponding to the singular values ​​of the matrix. The singular value is a kind of eigenvalue of the matrix, which describes the shape and size of the matrix and determines the most important direction of the matrix.

[0070] The preset rank value may be the rank of the approximate matrix, the singular diagonal matrix may be a diagonal matrix whose diagonal elements are the singular values ​​of the weight matrix and are arranged in descending order. The first singular matrix may be a matrix obtained by operating on the preset rank value, the left singular matrix and the singular diagonal matrix, and the second singular value matrix may be a matrix obtained by operating on the preset rank value, the right singular matrix and the singular diagonal matrix.

[0071] It should also be noted that singular value decomposition (SVD) is an important matrix decomposition technology that decomposes a matrix into the product of three matrices, as shown in the first decomposition algorithm as follows:

[0072] W=UΣV T ;

[0073] Wherein, W is the original matrix (i.e., the weight matrix in this embodiment); U is an m×m orthogonal matrix, which is the left singular matrix of W; V is an n×n orthogonal matrix, which is the right singular matrix of W; Σ is an m×n diagonal matrix (i.e., the singular diagonal matrix mentioned above), the non-diagonal elements are 0, and the diagonal elements are arranged from the singular values ​​of W in descending order, which can be expressed as Σ=diag(σ1,…,σ n ), σ1≥σ2≥…≥σ n ≥ 0. σ1,…,σ n are the singular values ​​in the original matrix.

[0074] When applying SVD to decompose a matrix, the original matrix is ​​usually decomposed into the product of two matrices. At this time, the singular value matrix is ​​no longer stored separately, but the singular values ​​are included in A (i.e., the first singular matrix mentioned above) and B (i.e., the second singular matrix mentioned above) respectively, so as to further reduce the parameter storage amount and computational complexity, as shown in the second decomposition algorithm as follows:

[0075]

[0076] In the formula, is the approximate matrix of the original matrix W, A is an m×r matrix, and B is an r×n matrix; is the truncated matrix of the original left singular matrix U, including only the first r columns of data of the original left singular matrix U; is the truncated matrix of the original right singular matrix V, including only the first r rows of data of the original right singular matrix V; is the truncated matrix of the original diagonal matrix Σ, which only includes the diagonal matrix composed of the first r singular values ​​arranged from large to small. In order to achieve efficient model compression, it is very important to choose a suitable r value (i.e., the rank of the weight matrix).

[0077] In a specific implementation, the device first obtains a left singular matrix, a right singular matrix and a singular diagonal matrix based on the weight matrix according to the first decomposition algorithm. The device then obtains a first singular value matrix based on a preset rank value, a left singular matrix and a singular diagonal matrix according to the second decomposition algorithm, and obtains a second singular value matrix based on the preset rank value, the right singular matrix and the singular diagonal matrix. Finally, the device obtains an approximate matrix based on the first singular value matrix and the second singular value matrix.

[0078] Furthermore, in order to obtain an accurate preset rank value, the weight matrix is ​​subjected to singular value decomposition, and before the step of obtaining a first singular value matrix based on the preset rank value, the left singular matrix and the singular diagonal matrix, the method further includes:

[0079] Determine the loss function and cost function corresponding to the neural network model to be deployed;

[0080] An auxiliary variable is determined according to the weight matrix, and a preset rank value is determined based on a preset regularization parameter, the auxiliary variable, the loss function, and the cost function.

[0081] It should be noted that the above-mentioned loss function can be a function that measures the difference between the predicted value and the true value of the above-mentioned neural network model to be deployed. The above-mentioned cost function can be a function that measures the complexity of the above-mentioned neural network model to be deployed. The above-mentioned auxiliary variables can be variables that help optimize the objective function. For example, in SVD, auxiliary variables can be used to represent low-rank approximate matrices. The above-mentioned regularization parameter can be a parameter for controlling the complexity of the model. In a specific implementation, the above-mentioned device can use a variety of methods to determine the preset rank value, such as the Bayesian Information Criterion, energy measurement, and Learning Compression (LC) algorithm. This embodiment uses the LC algorithm for explanation, but does not specifically limit this embodiment.

[0082] In a specific implementation, the above device needs to determine the loss function and cost function corresponding to the neural network model to be deployed. For example, assuming that the model to be deployed is a convolutional neural network (CNN) model for image classification, the loss function can select the cross entropy loss, and the cost function can select the number of model parameters. Next, the device determines the auxiliary variables according to the weight matrix, and determines the preset rank value according to the LC algorithm based on the preset regularization parameters, auxiliary variables, loss function and cost function.

[0083] refer to Figure 2 , Figure 2 This is an example diagram of singular value decomposition based on the LC algorithm in the neural network model deployment method proposed in this embodiment. In this example, the above-mentioned device uses the LC algorithm to obtain the preset rank value. The operation loss and calculation cost of the time series prediction model are used as the objective function, and the auxiliary variable Θ=(Θ1,…,Θ K ) is used to represent the low-rank approximation matrix of each layer, and the constraint condition W = Θ is imposed. The Lagrange multiplier with the same dimension as W is introduced to reduce the number of constraints. The objective function is written as follows: the rank of the weight matrix and the weight value are separable:

[0084]

[0085] Where L(W) represents the mean square error function between the predicted value and the true value of the Att-NBEATS model, λ is the regularization parameter, C(r) is the regularization term for rank selection, and represents the cost function of the model, which is usually the number of model parameters or the number of floating-point operations of the model, and can be written as C(r)=C1(r1)+…+C k (r k). From an optimization perspective, C(r) only depends on the rank of the weight matrix of each layer. Θ k represents the weight matrix corresponding to the k-th layer of the compressed neural network, ||W k -Θ k || 2 is a quadratic penalty parameter used to ensure that W and Θ are consistent. β=(β1,...,β k ) is the Lagrange product vector, μ is the penalty parameter, which punishes the deviation between W and Θ, gradually bringing W and Θ closer together so that they tend to satisfy the constraint of W = Θ.

[0086] Due to the separability of the rank r of the weight matrix and the weight value W in the objective function, the rank value determination in the LC algorithm can be divided into a learning (Learning, L) step (i.e. Figure 2 The L step in the code) and the compression (Compression, C) step (i.e. Figure 2 The C step in the diagram is as follows:

[0087]

[0088] The purpose of the L step is to ensure the performance of the model. The adaptive momentum stochastic optimization method is used to optimize W, and by gradually increasing the μ value, ||W k -Θ k || is forced to approach 0, so W k ≈Θ k , so that the low-rank weight matrix is ​​similar to the original matrix, and finally the updated weight matrix W is output.

[0089] The purpose of step C is to reduce the computational complexity of the model and to determine r by discrete gradient estimation. k , and select the first r of each layer matrix k Singular values ​​are obtained based on SVD to obtain a low-rank approximate matrix, thereby achieving model compression.

[0090] The L step and the C step are run alternately to finally determine the rank of each layer of the neural network and implement SVD-based model compression. The key parameter settings of the algorithm are shown in Table 1.

[0091] Table 1 Algorithm key parameter settings

[0092]

[0093] Step S40: quantize the compressed neural network model to be deployed, and deploy the quantized neural network model to be deployed to the target edge device.

[0094] It should be noted that the target edge device can be an edge device with storage space and computing resources. In a specific implementation, the device obtains an approximate matrix through the singular value decomposition and compresses the neural network model to be deployed to reduce the number of parameters required by the model by replacing the weight matrix of the layer to be compressed with the approximate matrix. Then, the compressed neural network model to be deployed needs to be quantized (such as linear quantization and symmetric quantization, etc.). This step involves converting the weights and activation values ​​of the model from floating point numbers to integers with low bit width to reduce the storage requirements of the model and improve the operating efficiency. After quantization, the device deploys the optimized neural network model to the target edge device, such as a smart camera or mobile sensor, which has limited resources but needs to process and analyze data in real time. In this way, the device not only reduces the complexity of the model, but also ensures the efficient operation and rapid response of the model in the edge environment.

[0095] This embodiment proposes a neural network model deployment method, the method comprising: obtaining a neural network model to be deployed, and performing a structural analysis on the neural network model to be deployed to determine the layer to be compressed in the neural network model to be deployed; determining the weight matrix of the layer to be compressed, and performing singular value decomposition on the weight matrix to obtain an approximate matrix; compressing the neural network model to be deployed according to the approximate matrix; quantizing the compressed neural network model to be deployed, and deploying the quantized neural network model to be deployed to the target edge device. This embodiment determines the layer to be compressed and the weight matrix corresponding to the layer to be compressed by performing a structural analysis on the neural network model to be deployed, obtains an approximate matrix by using the weight matrix, and compresses the neural network model to be deployed by using the approximate matrix to reduce the number of parameters required for the model. The compressed neural network model to be deployed is then quantized to reduce the number of bits occupied by the parameters, so that the neural network model to be deployed can be deployed on the target edge device with limited storage space and computing resources.

[0096] Based on the first embodiment, in the second embodiment, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be described in detail later. Figure 3 , Figure 3 This is a flow chart of the second embodiment of the neural network model deployment method proposed in this embodiment. Further, in order to quantize the compressed neural network model to be deployed, the step of quantizing the compressed neural network model to be deployed includes:

[0097] Step S41: Obtain the target weight value and target activation value in the compressed neural network model to be deployed;

[0098] Step S42: quantizing the compressed neural network model to be deployed based on the target weight value and the target activation value.

[0099] It should be noted that the above target weight value can be the connection strength between the layers in the neural network. The above target activation value can be the activeness of signal transmission. In a specific implementation, after the neural network model is compressed, the first task of the above device is to accurately obtain the target weight value and target activation value in the compressed model to be deployed. For example, in a convolutional neural network, the target weight value may represent the parameters of the convolution kernel, and the target activation value is the feature map generated after each convolution layer.

[0100] To achieve this goal, the device will traverse each layer of the model and extract the corresponding weight matrix and activation vector. Next, based on these target weight values ​​and target activation values, the device will quantize the compressed neural network model. Quantization involves converting floating-point parameters to a low-precision format, such as integers or low-bit-width floating-point numbers. This process involves selecting an appropriate quantization strategy, such as linear quantization or symmetric quantization, and determining the bit width of quantization, such as 8 bits or 16 bits.

[0101] During the quantization process, the above devices will use calibration techniques, such as minimum and maximum calibration or KL divergence calibration, to find the best quantization range and scaling factor. In this way, the weights and activation values ​​originally stored in 32-bit floating point numbers can be compressed to smaller data types, such as int8, thereby significantly reducing the storage requirements of the model and accelerating the inference speed. After quantization, the above devices will ensure that the neural network model can run efficiently on resource-constrained edge devices while maintaining acceptable prediction accuracy, such as real-time image recognition or speech processing tasks on IoT sensors or mobile devices.

[0102] Furthermore, in order to obtain the target weight value and target activation value of each layer of each neural network model to be deployed to quantize the neural network to be deployed, the step of obtaining the target weight value and target activation value in the compressed neural network model to be deployed includes:

[0103] Obtain the weight parameters of the compressed neural network model to be deployed;

[0104] Using a preset simulation quantization algorithm to simulate and quantize the compressed neural network model to be deployed based on a preset scaling factor and the weight parameter;

[0105] The target weight value and target activation value are obtained based on the simulation quantization results.

[0106] It should be noted that the above-mentioned weight parameters may be parameters in the weight matrix of each layer in the compressed neural network model to be deployed. The above-mentioned preset analog quantization algorithm may be an algorithm for converting floating-point weight parameters into low-precision representations. The above-mentioned preset scaling factor may be a coefficient preset by the user for adjusting the numerical range of the weight parameter.

[0107] In the specific implementation, after completing the compression step of the neural network model, the above-mentioned device will then focus on extracting the weight parameters in the compressed model to be deployed. These weight parameters are the key factors connecting the nodes of each layer in the neural network, and they determine the learning ability and prediction performance of the model. For example, in a deep learning model, the weight parameters may include the weight of the convolution kernel, the weight of the fully connected layer, etc.

[0108] The above-mentioned device will then use a preset simulated quantization algorithm to simulate quantize the compressed neural network model. Simulated quantization is a preprocessing technique that converts weight parameters from floating point numbers to low-precision representations, such as fixed point numbers, through a preset scaling factor. This process simulates the quantization effect of model parameters in actual deployment, allowing the device to evaluate the impact of quantization on model performance without actually changing the model parameters.

[0109] Specifically, the above device will scale and quantize each weight parameter according to a preset scaling factor. This process may involve finding the maximum absolute value of each weight parameter in order to determine the appropriate scaling factor, and then applying this factor to map the weight parameter to a more compact numerical range. For example, if the preset quantization bit width is 8 bits, the weight parameter will be quantized to an integer range of -128 to 127.

[0110] Based on the simulated quantization results, the above devices can obtain the target weight values ​​and target activation values. These values ​​are the parameters actually used by the quantized model during the inference process, and they are crucial to maintaining the prediction accuracy of the model. By simulating the quantization process, the above devices can not only verify the effectiveness of the quantization strategy, but also provide accurate parameters for subsequent actual quantization deployment, ensuring that the model can maintain performance while reducing computing resources and energy consumption when running on edge devices. For example, when deployed to mobile devices or embedded systems, this quantization process can significantly increase processing speed and reduce memory usage.

[0111] refer to Figure 4 , Figure 4 This is a flow chart of quantization processing based on a pseudo quantizer in the neural network model deployment method proposed in this embodiment. In one example, the above-mentioned device can insert a pseudo quantizer in each layer of the compressed neural network model to be deployed to simulate quantization of the compressed neural network model to be deployed, wherein the working principle of the pseudo quantizer, that is, the preset simulation quantization algorithm, is:

[0112]

[0113] x quantized =scaler·(q-zero_point);

[0114] Among them, x is the original floating point type data (that is, Figure 4 The initial weight in the quantized integer data, quantized is the dequantized floating-point data; scale is the scaling factor that maps the floating-point range to the quantized range; zero_point is the zero point that aligns the zero point of the original floating-point data with the zero point of the quantized integer data. scale and zero_point are determined by the distribution range of the weight values ​​and activation values ​​recorded for each layer.

[0115] Furthermore, in order to ensure the functional stability of the quantized neural network model to be deployed, the step of obtaining the target weight value and the target activation value based on the simulation quantization result includes:

[0116] Dequantizing the simulated quantization result to obtain an approximate model corresponding to the compressed neural network model to be deployed;

[0117] Sample data is obtained, and the approximate model is trained based on the sample data to obtain a target weight value and a target activation value.

[0118] It should be noted that the above approximate model can be a model obtained by analog quantization and inverse quantization that can be deployed to a model with limited computing resources and storage space, and has the same function as the above neural network to be deployed. The above sample data can be a subset of the data set used to train the model.

[0119] In the specific implementation, after completing the simulation quantization, the above-mentioned device will perform dequantization on the simulation quantization results in order to verify the accuracy of the quantization model. Dequantization refers to converting the quantized weight values ​​and activation values ​​back to floating-point representation to obtain the compressed approximate model corresponding to the neural network model to be deployed. This step helps to evaluate the errors introduced by quantization and ensure that the approximate model is as close as possible to the original model in performance.

[0120] Next, the above device will obtain a certain amount of sample data for training and verifying the model. Based on these sample data, the above device trains the approximate model with the goal of optimizing the model's predictive ability by adjusting weight parameters and activation values. During the training process, the device will use optimization algorithms such as gradient descent to update the model's parameters and ultimately obtain the target weight values ​​and target activation values.

[0121] refer to Figure 3 as well as Figure 4 , for ease of understanding, the following examples are given, but no specific limitation is made to this embodiment. In the process of predicting the motor state using the time series prediction model, in order to compress the above-mentioned time series prediction model and deploy it on the motor state prediction device. The above-mentioned device can select several sets of relatively complete full life cycle monitoring data from the motor state data set, arrange the data in chronological order, obtain time series data about the motor state, divide the training set and the test set according to the data ratio of 8:2, and after standardizing the data, use the sliding window method to construct a time window of length 8, and slide with a step size of 1 to obtain a time window of historical data. In each time window, take the sequence consisting of the first 7 data in the time window as the input sequence, and take the last data in the time window as the output sequence.

[0122] In this example, the C-MAPSS dataset is used for explanation. The C-MAPSS dataset is generated by simulation software developed by NASA (National Aeronautics and Space Administration) and is used to simulate the degradation process of aircraft engines. This dataset has 4 sub-datasets, each with a different number of operating conditions and fault conditions, as shown in Table 2, which is an introduction to the C-MAPSS dataset.

[0123] Table 2C-MAPSS dataset introduction

[0124]

[0125] After obtaining the data set, load the trained time series prediction model, including its structure and weight parameters. The specific structure is shown in Table 3, which is the model structure and parameter setting table.

[0126] Table 3 Model structure and parameter settings

[0127]

[0128] After obtaining the above-mentioned time series prediction model and its structure, the above-mentioned device analyzes the structure of the time series prediction model and finds that it mainly includes a fully connected layer and an attention layer. Among them, the fully connected layer has a strong representation ability due to its dense connection structure, but many connections are not necessary in practice, which leads to the fact that the fully connected layer usually has a high degree of redundancy. Therefore, it is necessary to decompose the fully connected layer. The attention layer mainly selects information that has a greater impact on the output result from the input features. Decomposing the selection mechanism therein may affect the weight distribution of the model, thereby weakening the effect of the attention mechanism. Therefore, it is unnecessary to decompose the attention layer. Therefore, the above-mentioned fully connected layer is used as the above-mentioned layer to be compressed. And the weight matrix corresponding to the fully connected layer is subjected to singular value decomposition by the method in the above-mentioned first embodiment. And the above-mentioned time series prediction model is compressed based on the approximate matrix obtained by the singular value decomposition.

[0129] The above device obtains the compressed time series prediction model (i.e. Figure 4 After the SVD compression model is obtained, the above-mentioned pseudo quantizer is inserted into each layer of the time series prediction model to simulate the quantization of the compressed time series prediction model, and the model is fine-tuned through the sample data obtained above and the above-mentioned simulated quantization results. Figure 4 In the forward propagation process, the weights and activation values ​​are represented by the data quantized by the pseudo quantizer. At this time, the model can perceive the error caused by quantization. Figure 4 When back propagating in the back propagation, the quantized integers are not updated, but the dequantized floating-point weights and activation values ​​are directly updated. By adjusting the distribution of weights and activation values, the quantized model can also maintain good performance. After fine-tuning is completed, the fake quantizer will be replaced by the real quantization operation, that is, the weights and activation values ​​are formally quantized to low-precision integers (i.e. Figure 4 The perceptual quantization fine-tuned weights are used to further compress the model memory and reduce computational complexity.

[0130] This embodiment also provides a first embodiment of a neural network model deployment device, please refer to Figure 5 , Figure 5 This is a structural diagram of a first embodiment of a neural network model deployment device provided in an embodiment of the present application, wherein the neural network model deployment device comprises:

[0131] A model acquisition module, used to acquire a neural network model to be deployed, and perform structural analysis on the neural network model to be deployed to determine a layer to be compressed in the neural network model to be deployed;

[0132] A data decomposition module, used to determine the weight matrix of the layer to be compressed, and perform singular value decomposition on the weight matrix;

[0133] A model compression module, used for compressing the neural network model to be deployed according to the approximate matrix;

[0134] A model deployment module is used to quantize the compressed neural network model to be deployed, and deploy the quantized neural network model to be deployed to the target edge device;

[0135] The data decomposition module is further used to decompose the weight matrix to obtain a left singular matrix, a right singular matrix and a singular diagonal matrix; obtain a first singular value matrix based on a preset rank value, the left singular matrix and the singular diagonal matrix; obtain a second singular value matrix based on the preset rank value, the right singular matrix and the singular diagonal matrix; obtain an approximate matrix according to the first singular value matrix and the second singular value matrix;

[0136] The data decomposition module is also used to determine the loss function and cost function corresponding to the neural network model to be deployed; determine the auxiliary variables according to the weight matrix, and determine the preset rank value based on the preset regularization parameter, the auxiliary variables, the loss function and the cost function.

[0137] Based on the above-mentioned first embodiment of the neural network model deployment device of the present application, a second embodiment of the neural network model deployment device of the present application is proposed.

[0138] In this embodiment, the model deployment module is further used to obtain a target weight value and a target activation value in the compressed neural network model to be deployed; and quantize the compressed neural network model to be deployed based on the target weight value and the target activation value;

[0139] The model deployment module is further used to obtain weight parameters in the compressed neural network model to be deployed; simulate and quantize the compressed neural network model to be deployed based on the preset scaling factor and the weight parameters using a preset simulation quantization algorithm; and obtain a target weight value and a target activation value based on the simulation quantization result;

[0140] The model deployment module is also used to dequantize the simulation quantization results to obtain an approximate model corresponding to the compressed neural network model to be deployed; obtain sample data, and train the approximate model based on the sample data to obtain target weight values ​​and target activation values.

[0141] The neural network model deployment device provided in this embodiment adopts the neural network model deployment method in the above embodiment, which can solve the problem of how to deploy the neural network model to edge devices with limited storage space and computing resources. Compared with the prior art, the beneficial effects of the neural network model deployment device provided in this embodiment are the same as the beneficial effects of the neural network model deployment method provided in the above embodiment, and the other technical features in the neural network model deployment device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0142] This embodiment provides a device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the neural network model deployment method in the above-mentioned embodiment one.

[0143] Reference below Figure 6 , Figure 6Schematic diagram of the structure of the device suitable for implementing the embodiment of the present application. The device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0144] like Figure 6 As shown, the device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0145] In particular, according to the present embodiment, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present embodiment includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present embodiment are executed.

[0146] The device provided in this embodiment adopts the neural network model deployment method in the above embodiment, which can solve the problem of how to deploy the neural network model to edge devices with limited storage space and computing resources. Compared with the prior art, the beneficial effects of the device provided in this embodiment are the same as the beneficial effects of the neural network model deployment method provided in the above embodiment, and the other technical features in the device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0147] It should be understood that the various parts disclosed in this embodiment can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0148] The above is only a specific implementation of this embodiment, but the protection scope of this embodiment is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in this embodiment, which should be included in the protection scope of this embodiment. Therefore, the protection scope of this embodiment should be based on the protection scope of the claims.

[0149] This embodiment provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the neural network model deployment method in the above-mentioned embodiment.

[0150] The computer-readable storage medium provided in this embodiment may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0151] The computer-readable storage medium may be included in the device; or may exist independently without being installed in the device.

[0152] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the device, the device: performs neural network model deployment.

[0153] The computer program code for performing the operation of the present embodiment can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0154] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present embodiment. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0155] The modules involved in the description of this embodiment may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0156] The readable storage medium provided in this embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned neural network model deployment method, and can solve the problem of how to deploy the neural network model to edge devices with limited storage space and computing resources. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as the beneficial effects of the neural network model deployment method provided in the above-mentioned embodiment, and will not be repeated here.

[0157] The above descriptions are only some embodiments, and are not intended to limit the patent scope of this embodiment. All equivalent structural changes made using the contents of the specification and drawings of this application under the technical concept of this application, or directly / indirectly applied in other related technical fields are included in the patent protection scope of this application.

Claims

1. A neural network model deployment method, characterized in that: The method comprises: Acquire a neural network model to be deployed, and perform structural analysis on the neural network model to be deployed to determine a layer to be compressed in the neural network model to be deployed; Determine a weight matrix of the layer to be compressed, and perform singular value decomposition on the weight matrix to obtain an approximate matrix; Compressing the neural network model to be deployed according to the approximate matrix; The compressed neural network model to be deployed is quantized, and the quantized neural network model to be deployed is deployed to the target edge device.

2. The method according to claim 1, characterized in that The step of performing singular value decomposition on the weight matrix to obtain an approximate matrix comprises: Decomposing the weight matrix to obtain a left singular matrix, a right singular matrix and a singular diagonal matrix; Obtain a first singular value matrix based on a preset rank value, the left singular matrix, and the singular diagonal matrix; Obtain a second singular value matrix based on the preset rank value, the right singular matrix, and the singular diagonal matrix; An approximate matrix is ​​obtained according to the first singular value matrix and the second singular value matrix.

3. The method according to claim 2, characterized in that Before the step of obtaining a first singular value matrix based on a preset rank value, the left singular matrix and the singular diagonal matrix, the method further includes: Determine the loss function and cost function corresponding to the neural network model to be deployed; An auxiliary variable is determined according to the weight matrix, and a preset rank value is determined based on a preset regularization parameter, the auxiliary variable, the loss function, and the cost function.

4. The method according to claim 1, characterized in that The step of quantizing the compressed neural network model to be deployed includes: Obtain the target weight value and target activation value in the compressed neural network model to be deployed; The compressed neural network model to be deployed is quantized based on the target weight value and the target activation value.

5. The method according to claim 4, characterized in that The step of obtaining the target weight value and the target activation value in the compressed neural network model to be deployed includes: Obtain the weight parameters of the compressed neural network model to be deployed; Using a preset simulation quantization algorithm to simulate and quantize the compressed neural network model to be deployed based on a preset scaling factor and the weight parameter; The target weight value and target activation value are obtained based on the simulation quantization results.

6. The method according to claim 5, characterized in that The step of obtaining a target weight value and a target activation value based on the simulation quantization result includes: Dequantizing the simulation quantization result to obtain an approximate model corresponding to the compressed neural network model to be deployed; Sample data is obtained, and the approximate model is trained based on the sample data to obtain a target weight value and a target activation value.

7. A neural network model deployment device, characterized in that: The device comprises: A model acquisition module, used to acquire a neural network model to be deployed, and perform structural analysis on the neural network model to be deployed to determine a layer to be compressed in the neural network model to be deployed; A data decomposition module, used to determine the weight matrix of the layer to be compressed, and perform singular value decomposition on the weight matrix; A model compression module, used for compressing the neural network model to be deployed according to the approximate matrix; The model deployment module is used to quantize the compressed neural network model to be deployed, and deploy the quantized neural network model to be deployed to the target edge device.

8. A neural network model deployment device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the neural network model deployment method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the neural network model deployment method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the steps of the neural network model deployment method according to any one of claims 1 to 6.