Model quantitative perceptual training system on edge device
Through the model quantization perception training system on edge devices, the problem of edge device resource limitations is solved, efficient deployment and optimization of model performance are achieved, adapting to different device states and meeting real-time application needs.
Patent Information
- Application Number
- CN202510845098.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Edge devices have limited computing and memory resources, making it difficult to effectively deploy and run high-precision deep learning models, and the quantization process may cause a significant decline in model performance.
A model quantization-aware training system is adopted, including data acquisition, preprocessing, quantization-aware training, model evaluation, deployment and dynamic scheduling modules. The model parameters are optimized through simulated quantization operations and gradient approximation methods to adapt to the hardware characteristics of edge devices.
This reduces the computational and memory requirements of the model on edge devices, maintains the high performance of the model, improves the inference speed, and meets real-time requirements.
Smart Images

Figure CN120745751A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge device technology, and in particular to a model quantization perception training system on an edge device. Background Art
[0002] In today's digital age, edge devices are increasingly used in various fields, such as smart homes, smart security, and industrial IoT. These devices often need to run deep learning models to achieve intelligent functions.
[0003] However, due to the limited hardware resources of edge devices (such as computing power, memory, storage, etc.), traditional high-precision deep learning models face many challenges when deployed on edge devices. Specifically, there are the following problems:
[0004] 1. Computing resource limitations: Edge devices have relatively weak computing capabilities, making them difficult to process large-scale, highly complex deep learning models. For example, in intelligent security monitoring, real-time object detection and recognition are required in the video streams captured by cameras. Using unoptimized high-precision models, such as those based on ResNet-101, when running on standard edge computing devices (such as the Raspberry Pi 4B), the computing speed falls far short of real-time requirements, resulting in significant delays in video analysis and the inability to detect anomalies in a timely manner.
[0005] 2. Limited memory resources: The weights and intermediate data of deep learning models typically occupy a large amount of memory. Edge devices with limited memory, such as some low-power IoT sensor nodes, may not be able to load the complete model. For example, the unquantized weight data of a simple speech recognition model may occupy tens of MB of memory, while some IoT sensor nodes only have a few MB of memory, making the model unusable on such devices.
[0006] 3. Model performance degradation: Directly quantizing a model without considering its impact on the model structure and parameters often leads to a significant drop in model performance. For example, in image classification tasks, directly quantizing a trained full-precision ResNet-50 model to INT8 format can cause the accuracy on the test set to drop sharply from 75% to around 50%, seriously affecting the model's effectiveness in practical applications.
[0007] Based on the above, a model quantization-aware training system on edge devices is invented. Summary of the Invention
[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0009] Model quantization-aware training system on edge devices, which includes:
[0010] Data acquisition module, used to collect data from various edge devices;
[0011] The data preprocessing module is used to clean, normalize, and enhance preprocessing operations on the collected data to improve the quality and diversity of the data and provide a better data foundation for subsequent model training;
[0012] The quantization-aware training module is responsible for applying quantization-aware training techniques during the training process to simulate quantization operations in the forward propagation of the model, quantize and dequantize weights and activation values, and use gradient approximation methods to calculate gradients in the backpropagation to optimize model parameters;
[0013] The model evaluation module is used to evaluate the model during or after training to determine whether the model performance meets the requirements by calculating the accuracy, recall rate, and F1 value indicators of the model on the validation set. If the model performance is poor, feedback will be given to the quantization-aware training module to adjust the training parameters or quantization strategy;
[0014] The model deployment module is used to deploy models that have undergone quantization-aware training and meet performance requirements to edge devices. During the deployment process, the model will be further optimized and adapted according to the hardware characteristics of the edge devices to ensure that the model can run efficiently on the edge devices.
[0015] The dynamic scheduling module is used to receive performance feedback from the model evaluation module, dynamically adjust the training parameters of the quantitative perception training module based on the feedback results, and determine the timing and method of model deployment based on the real-time resource status of the edge device.
[0016] As a preferred solution of the model quantization-aware training system on the edge device of the present invention, the quantization-aware training module includes:
[0017] The analog quantization configuration module is used to first determine the quantization parameters and then insert the quantization operation;
[0018] Model initialization and training parameter setting module, used to first initialize the model and then set the training parameters;
[0019] The forward propagation and simulation quantization execution module is used to first perform data input and forward calculation, and then simulate and quantize the effect;
[0020] Back propagation and gradient approximation processing module, which is used to first calculate the loss function and then perform gradient approximation calculation;
[0021] The model parameter optimization and update module is used to first update the parameters and then perform multiple iterative training.
[0022] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the simulation quantization configuration module includes:
[0023] Determine the quantization parameter module, which is used to specify the quantization bit width, quantization method, and quantization granularity based on the hardware capabilities of the edge device and the performance requirements of the model;
[0024] The quantization operation insertion module is used to insert quantization and dequantization operations during the forward propagation process of the model in the deep learning framework.
[0025] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the model initialization and training parameter setting module includes:
[0026] Model initialization module, used to initialize the selected deep learning model and set the model parameters and structure;
[0027] The training parameter setting module is used to first determine the key parameters in the training process, including learning rate, batch size, and number of iterations. Then, a dynamic learning rate adjustment strategy is adopted to use a larger learning rate in the early stage of training to accelerate convergence, and gradually reduce the learning rate as training progresses to improve model performance.
[0028] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the forward propagation and simulation quantization execution module includes:
[0029] The data input and forward calculation module is used to input the preprocessed data into the model and perform calculations according to the conventional forward propagation process. During the calculation process, the weights and activation values will undergo quantization and dequantization operations;
[0030] The module for simulating quantization effects is used to simulate the effects of quantization in actual deployment through quantization and dequantization operations. This allows the model to perceive the changes brought about by quantization during training, providing a basis for subsequent parameter optimization.
[0031] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the back propagation and gradient approximation processing module includes:
[0032] The loss function calculation module is used to calculate the loss function value based on the model output and the true label, and measure the difference between the model's current prediction result and the actual situation;
[0033] The gradient approximation calculation module is used to calculate the gradient using the gradient approximation technique. During backpropagation, the gradient is directly passed through the quantization operation without considering the non-differentiability of the quantization operation itself, thereby updating the model parameters.
[0034] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the model parameter optimization and update module includes:
[0035] The parameter update module is used to update the model parameters using an optimization algorithm based on the calculated gradients. During the update process, the model adjusts the parameters according to the quantized gradient information to adapt the model to the quantized representation and reduce the impact of quantization on model performance.
[0036] The multi-iteration training module is used to repeat the above-mentioned forward propagation, backpropagation and parameter update process for multiple iterative training. During the training process, the loss value and accuracy index are continuously monitored so that the training parameters can be adjusted in time according to the changes in the indicators until the model performance reaches the expected target or tends to be stable.
[0037] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the dynamic scheduling module includes:
[0038] The real-time equipment status monitoring module is used to first collect resource data, then monitor performance indicators, and finally perceive environmental parameters;
[0039] The training process dynamic adjustment module is used to first receive performance feedback, then optimize training parameters, and then fine-tune the quantization strategy;
[0040] The model deployment intelligent decision-making module is used to first determine the deployment timing, then select the deployment plan, and finally distribute the deployment tasks;
[0041] The continuous monitoring and feedback module of the operating status is used to first track the operating status, then handle abnormal situations, and then perform feedback optimization cycles.
[0042] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the device status real-time monitoring module includes:
[0043] The resource data collection module is used to collect key resource data of the device in real time through the built-in monitoring tools or system plug-ins of the edge device. The key resource data includes computing resources, memory usage, remaining storage space, and network bandwidth;
[0044] The performance indicator monitoring module is used to monitor the performance indicators of the model when running on edge devices, including inference time, response latency, and throughput. It can also analyze the model inference process with the help of performance analysis tools to obtain detailed performance data on the computation time and data transmission time of each layer to evaluate the current operating efficiency of the model;
[0045] Environmental parameter perception module, used to perceive the external environmental parameters of the edge device;
[0046] The training process dynamic adjustment module includes:
[0047] The performance feedback receiving module is used to establish a data interaction channel with the model evaluation module and receive performance feedback data during the model training process in real time;
[0048] A training parameter optimization module, which is used to dynamically adjust the training parameters of the quantization-aware training module based on performance feedback and device resource status;
[0049] The quantization strategy fine-tuning module is used to dynamically fine-tune the quantization strategy based on device resource conditions and model training performance.
[0050] As a preferred solution of the model quantization perception training system on the edge device of the present invention, the model deployment intelligent decision module includes:
[0051] The deployment timing judgment module is used to comprehensively consider the device resource status and model performance evaluation results to determine the optimal time for model deployment;
[0052] The deployment scheme selection module is used to select the most appropriate model deployment scheme based on the hardware characteristics and resource conditions of different edge devices;
[0053] The deployment task distribution module is used to reasonably allocate the selected model deployment tasks to the target edge devices, accurately transfer the deployment resources of the model files and configuration parameters to the corresponding devices through the device management system or distributed deployment tools, and start the deployment process to ensure that the model can be correctly installed and run on the edge devices;
[0054] The operating status continuous monitoring and feedback module includes:
[0055] The running status tracking module is used to continuously monitor the running status of the model after it is deployed to the edge device, collect resource usage and performance indicator changes in real time during the model operation, and track the accuracy of the model's inference results and resource usage fluctuations through logging and real-time data monitoring;
[0056] The abnormal situation processing module is used to take appropriate measures in time when abnormalities are detected in the model operation;
[0057] The feedback optimization loop module is used to feed back the model operation status and processing results to the dynamic scheduling module as a reference for the next round of scheduling decisions. By continuously collecting and analyzing model operation data, it continuously optimizes the training parameter adjustment strategy, quantization strategy and deployment plan, forming a closed-loop dynamic optimization process to improve the operating efficiency and model performance of the entire system.
[0058] Compared with existing technologies:
[0059] 1. Reduced computing resource consumption: Quantization-aware training converts the model's weights and activation values into low-precision representations, significantly reducing the computational effort required when running the model on edge devices. For example, running a quantized deep learning model on an ARM-based edge computing device significantly reduces computing resource consumption compared to an unquantized full-precision model, enabling complex models that would otherwise be impossible to run on such a device to run efficiently.
[0060] 2. Reduced memory usage: Quantized model weights and intermediate data significantly reduce the memory space occupied. For example, a typical image recognition model might use around 100MB of memory before quantization, but after quantization, the memory usage can be reduced to under 20MB. This allows more edge devices to easily load and run the model, improving the feasibility of model deployment on resource-constrained devices.
[0061] 3. Maintaining model performance: Despite quantization, the model maintains high performance after quantization due to the use of quantization-aware training techniques during training. In common tasks such as image classification and object detection, the accuracy of the quantized model is only moderately lower than that of the full-precision model, meeting the needs of most practical applications.
[0062] 4. Improved inference speed: Quantized models require less computation and fewer memory accesses, significantly improving inference speed on edge devices. In applications requiring high real-time performance, such as intelligent security monitoring and autonomous driving assistance, quantized models can process data more quickly and provide timely decision support. For example, in intelligent security monitoring, a quantized object detection model can detect a single image frame in milliseconds, meeting the needs of real-time monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic diagram of the overall framework of the present invention;
[0064] Figure 2 This is a schematic diagram of the framework of the quantitative perception training module of the present invention;
[0065] Figure 3 This is a schematic diagram of the dynamic scheduling module framework of the present invention. DETAILED DESCRIPTION
[0066] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0067] The present invention provides a model quantization perception training system on edge devices, please refer to Figure 1-Figure 3 ,include:
[0068] Data acquisition module, used to collect data from various edge devices;
[0069] The data preprocessing module is used to clean, normalize, and enhance preprocessing operations on the collected data to improve the quality and diversity of the data and provide a better data foundation for subsequent model training;
[0070] The quantization-aware training module is responsible for applying quantization-aware training techniques during the training process to simulate quantization operations in the forward propagation of the model, quantize and dequantize weights and activation values, and use gradient approximation methods to calculate gradients in the backpropagation to optimize model parameters;
[0071] The model evaluation module is used to evaluate the model during or after training to determine whether the model performance meets the requirements by calculating the accuracy, recall rate, and F1 value indicators of the model on the validation set. If the model performance is poor, feedback will be given to the quantization-aware training module to adjust the training parameters or quantization strategy;
[0072] The model deployment module is used to deploy models that have undergone quantization-aware training and meet performance requirements to edge devices. During the deployment process, the model will be further optimized and adapted according to the hardware characteristics of the edge devices to ensure that the model can run efficiently on the edge devices.
[0073] The dynamic scheduling module is used to receive performance feedback from the model evaluation module, dynamically adjust the training parameters of the quantitative perception training module based on the feedback results, and determine the timing and method of model deployment based on the real-time resource status of the edge device.
[0074] By setting up a dynamic scheduling module, it is possible to dynamically adjust training strategies and deployment plans based on real-time performance indicators during model training, as well as the computing resources and memory usage of edge devices. For example, when the current load on the edge device is high, the dynamic scheduling module can temporarily suspend model deployment or adjust the model's quantization strategy to reduce the resource demand for deployment. During training, if the model finds that certain indicators are improving slowly, it can automatically adjust training parameters such as learning rate and batch size, thereby improving training efficiency and model performance, and enhancing the system's adaptability to different scenarios and device states.
[0075] The quantization-aware training module includes:
[0076] The analog quantization configuration module is used to first determine the quantization parameters and then insert the quantization operation;
[0077] The analog quantization configuration module includes:
[0078] Determine the quantization parameter module, which is used to specify the quantization bit width, quantization method, and quantization granularity based on the hardware capabilities of the edge device and the performance requirements of the model;
[0079] The Insert Quantization Operation module is used to insert quantization and dequantization operations during the forward propagation of the model in the deep learning framework. Taking TensorFlow as an example, the QuantizeAwareTraining API is used to quantize the weights and activation values of each layer of the model. In the specific implementation, when quantizing the weight data, the maximum absolute value is first calculated to determine the quantization scale factor, and the weight data is mapped to the corresponding quantization range. For the activation value data, the appropriate quantization range is selected for mapping based on its distribution characteristics.
[0080] Model initialization and training parameter setting module, used to first initialize the model and then set the training parameters;
[0081] The model initialization and training parameter setting module includes:
[0082] Model initialization module, used to initialize the selected deep learning model and set the model parameters and structure;
[0083] The training parameter setting module is used to first determine the key parameters in the training process, including learning rate, batch size, and number of iterations. Then, a dynamic learning rate adjustment strategy is adopted. A larger learning rate is used in the early stage of training to accelerate convergence, and the learning rate is gradually reduced as training progresses to improve model performance.
[0084] The forward propagation and simulation quantization execution module is used to first perform data input and forward calculation, and then simulate and quantize the effect;
[0085] The forward propagation and analog quantization execution module includes:
[0086] The data input and forward calculation module is used to input the preprocessed data into the model and perform calculations according to the conventional forward propagation process. During the calculation process, the weights and activation values will undergo quantization and dequantization operations;
[0087] The module for simulating quantization effects is used to simulate the effects of quantization in actual deployments through quantization and dequantization operations. This allows the model to perceive the changes brought about by quantization during training, providing a basis for subsequent parameter optimization.
[0088] Back propagation and gradient approximation processing module, which is used to first calculate the loss function and then perform gradient approximation calculation;
[0089] The back propagation and gradient approximation processing module includes:
[0090] The loss function calculation module is used to calculate the loss function value based on the model output and the true label, and measure the difference between the model's current prediction result and the actual situation;
[0091] The gradient approximation calculation module is used to calculate the gradient using the gradient approximation technique. This allows the gradient to be directly passed through the quantization operation during backpropagation, without considering the non-differentiability of the quantization operation itself, thereby updating the model parameters.
[0092] Model parameter optimization and update module, which is used to first update the parameters and then perform multiple iterative training;
[0093] The model parameter optimization and update module includes:
[0094] The parameter update module is used to update the model parameters using an optimization algorithm based on the calculated gradients. During the update process, the model adjusts the parameters according to the quantized gradient information to adapt the model to the quantized representation and reduce the impact of quantization on model performance.
[0095] The multi-iteration training module is used to repeat the above-mentioned forward propagation, backpropagation and parameter update process for multiple iterative training. During the training process, the loss value and accuracy index are continuously monitored so that the training parameters can be adjusted in time according to the changes in the indicators until the model performance reaches the expected target or tends to be stable.
[0096] The dynamic scheduling module includes:
[0097] The real-time equipment status monitoring module is used to first collect resource data, then monitor performance indicators, and finally perceive environmental parameters;
[0098] The device status real-time monitoring module includes:
[0099] The resource data collection module is used to collect key resource data of the device in real time through the built-in monitoring tools or system plug-ins of the edge device. The key resource data includes computing resources, memory usage, remaining storage space, and network bandwidth;
[0100] The performance indicator monitoring module is used to monitor the performance indicators of the model when running on edge devices, including inference time, response latency, and throughput. It can also analyze the model inference process with the help of performance analysis tools to obtain detailed performance data on the computation time and data transmission time of each layer to evaluate the current operating efficiency of the model;
[0101] Environmental parameter perception module, used to perceive the external environmental parameters of the edge device;
[0102] The training process dynamic adjustment module is used to first receive performance feedback, then optimize training parameters, and then fine-tune the quantization strategy;
[0103] The training process dynamic adjustment module includes:
[0104] The performance feedback receiving module is used to establish a data interaction channel with the model evaluation module and receive performance feedback data during the model training process in real time;
[0105] The training parameter optimization module is used to dynamically adjust the training parameters of the quantization perception training module based on performance feedback and device resource status. If the current CPU load of the edge device is too high, the training batch size can be reduced to reduce the amount of data for each training session and reduce computing pressure. If the model converges slowly during training, the learning rate can be appropriately increased or the learning rate decay strategy can be adjusted to speed up the training process.
[0106] The quantization strategy fine-tuning module dynamically fine-tunes the quantization strategy based on device resources and model training performance. When device memory resources are limited, the quantization bit width can be further reduced or the quantization granularity can be adjusted to reduce model memory usage. If the model performance degrades significantly after quantization, the quantization accuracy of some key layers can be restored to strike a new balance between resource consumption and model performance.
[0107] The model deployment intelligent decision-making module is used to first determine the deployment timing, then select the deployment plan, and finally distribute the deployment tasks;
[0108] The model deployment intelligent decision module includes:
[0109] The deployment timing judgment module is used to comprehensively consider the device resource status and model performance evaluation results to determine the optimal time for model deployment. If the edge device is currently in a low-load state and the model's performance indicators on the test set meet the expected requirements, the model deployment process is triggered. If the device resources are insufficient or the model performance does not meet the requirements, the deployment is postponed and the device and model status are continuously monitored.
[0110] The deployment solution selection module is used to select the most appropriate model deployment solution based on the hardware characteristics and resource conditions of different edge devices. For devices with strong computing power but limited memory, a solution combining model compression and quantization is preferred to reduce memory usage. For devices with weak computing resources, a lightweight model is selected and the quantization strategy is optimized to ensure that the model can run efficiently on the device.
[0111] The deployment task distribution module is used to reasonably allocate the selected model deployment tasks to the target edge devices, accurately transfer the deployment resources of the model files and configuration parameters to the corresponding devices through the device management system or distributed deployment tools, and start the deployment process to ensure that the model can be correctly installed and run on the edge devices;
[0112] The continuous operation status monitoring and feedback module is used to first track the operation status, then handle abnormal situations, and then conduct feedback optimization cycles;
[0113] The operating status continuous monitoring and feedback module includes:
[0114] The running status tracking module is used to continuously monitor the running status of the model after it is deployed to the edge device, collect resource usage and performance indicator changes in real time during the model operation, and track the accuracy of the model's inference results and resource usage fluctuations through logging and real-time data monitoring;
[0115] The exception handling module is used to take appropriate measures in a timely manner when an exception is detected in the model operation; for example, it can automatically restart the model service, roll back to the previous version of the model, or adjust the model's operating parameters according to the exception to ensure stable model operation;
[0116] The feedback optimization loop module is used to feed back the model operation status and processing results to the dynamic scheduling module as a reference for the next round of scheduling decisions. By continuously collecting and analyzing model operation data, it continuously optimizes the training parameter adjustment strategy, quantization strategy and deployment plan, forming a closed-loop dynamic optimization process to improve the operating efficiency and model performance of the entire system.
[0117] When used, the specific steps are as follows:
[0118] S1: collects data from various edge devices through the data acquisition module;
[0119] S2: The data preprocessing module cleans, normalizes, and enhances preprocessing operations on the collected data to improve the quality and diversity of the data and provide a better data foundation for subsequent model training;
[0120] S3: By determining the quantization parameter module, the quantization bit width, quantization method and quantization granularity are clarified according to the hardware capabilities of the edge device and the performance requirements of the model; after clarification, the quantization operation module will be inserted into the deep learning framework to insert quantization and dequantization operations in the forward propagation process of the model; after insertion, the selected deep learning model will be initialized through the model initialization module to set the parameters and structure of the model; after initialization, the key parameters of the training process will be determined by setting the training parameter module, including learning rate, batch size, and number of iterations, and then a dynamic learning rate adjustment strategy will be adopted. A larger learning rate will be used to speed up the convergence speed at the beginning of training, and the learning rate will be gradually reduced as the training progresses to improve the model performance; then, the preprocessed data will be input into the model through the data input and forward calculation module, and calculated according to the conventional forward propagation process. During the calculation process, the weights and activation values will undergo quantization and dequantization operations; after calculation, the module will simulate the effect of quantization in actual deployment by simulating the quantization effect, so that the model The model perceives the changes brought about by quantization during the training process, providing a basis for subsequent parameter optimization; after simulation, the loss function value will be calculated according to the model output and the true label through the loss function calculation module to measure the difference between the current prediction result of the model and the actual situation; after calculation, the gradient will be calculated using the gradient approximation technology through the gradient approximation calculation module, so that the gradient can be directly passed through the quantization operation during back propagation, without considering the non-differentiability of the quantization operation itself, thereby realizing the update of the model parameters; after calculation, the model parameters will be updated using the optimization algorithm based on the calculated gradient through the parameter update module, and during the update process, the model will adjust the parameters according to the quantized gradient information, so that the model adapts to the quantized representation and reduces the impact of quantization on model performance; after the update, the above-mentioned forward propagation, back propagation and parameter update process will be repeated through multiple iterative training modules for multiple iterative training, and during the training process, the loss value and accuracy index will be continuously monitored so that the training parameters can be adjusted in time according to the changes in the index until the model performance reaches the expected target or tends to be stable;
[0121] S4: The model evaluation module evaluates the model during or after training to determine whether the model's performance meets the requirements by calculating the accuracy, recall, and F1 value indicators of the model on the validation set. If the model performance is poor, feedback will be given to the quantization-aware training module to adjust the training parameters or quantization strategy.
[0122] S5: The model deployment module deploys the quantization-aware trained model that meets performance requirements to the edge device. During the deployment process, the model is further optimized and adapted based on the hardware characteristics of the edge device to ensure that the model can run efficiently on the edge device.
[0123] S6: Through the resource data collection module, the built-in monitoring tools or system plug-ins of the edge device are used to collect the key resource data of the device in real time. The key resource data include computing resources, memory usage, remaining storage space, and network bandwidth. After collection, the performance indicators of the model running on the edge device will be monitored through the performance indicator monitoring module, including inference time, response delay, and throughput. The model inference process can be analyzed with the help of performance analysis tools to obtain detailed performance data of the computing time and data transmission time of each layer to evaluate the current operating efficiency of the model. After analysis, the external environmental parameters of the edge device will be perceived through the environmental parameter perception module. After perception, a data interaction channel will be established with the model evaluation module through the performance feedback receiving module to receive performance feedback data during the model training process in real time. After receiving, the training parameters of the quantization perception training module will be dynamically adjusted according to the performance feedback and device resource status through the training parameter optimization module. After adjustment, the quantization strategy fine-tuning module will dynamically fine-tune the quantization strategy according to the device resource situation and model training performance. After fine-tuning, the deployment timing judgment module will comprehensively consider the device resource status and model performance evaluation results to determine the best time to deploy the model. After judgment, the deployment timing judgment module will determine the best time to deploy the model. The deployment scheme selection module selects the most appropriate model deployment scheme based on the hardware characteristics and resource availability of different edge devices. Once selected, the selected model deployment task is rationally assigned to the target edge device through the deployment task distribution module. The deployment resources, including model files and configuration parameters, are accurately transferred to the corresponding device through the device management system or distributed deployment tool, and the deployment process is initiated to ensure that the model can be correctly installed and run on the edge device. After distribution, the operation status tracking module continuously monitors the model's operation status after the model is deployed to the edge device, collecting resource usage and performance indicator changes in real time during the model operation. It also tracks the accuracy of the model's inference results and resource usage fluctuations through logging and real-time data monitoring. After tracking, the exception handling module can take timely action when an anomaly in the model operation is detected. After processing, the model operation status and processing results are fed back to the dynamic scheduling module through the feedback optimization loop module as a reference for the next round of scheduling decisions. By continuously collecting and analyzing model operation data, the training parameter adjustment strategy, quantization strategy, and deployment scheme are continuously optimized, forming a closed-loop dynamic optimization process to improve the operation efficiency and model performance of the entire system.
[0124] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. Model quantization-aware training system on edge devices, characterized by: include: Data acquisition module, used to collect data from various edge devices; The data preprocessing module is used to clean, normalize, and enhance preprocessing operations on the collected data to improve the quality and diversity of the data and provide a better data foundation for subsequent model training; The quantization-aware training module is responsible for applying quantization-aware training techniques during the training process to simulate quantization operations in the forward propagation of the model, quantize and dequantize weights and activation values, and use gradient approximation methods to calculate gradients in the backpropagation to optimize model parameters; The model evaluation module is used to evaluate the model during or after training to determine whether the model performance meets the requirements by calculating the accuracy, recall rate, and F1 value indicators of the model on the validation set. If the model performance is poor, feedback will be given to the quantization-aware training module to adjust the training parameters or quantization strategy; The model deployment module is used to deploy models that have undergone quantization-aware training and meet performance requirements to edge devices. During the deployment process, the model will be further optimized and adapted according to the hardware characteristics of the edge devices to ensure that the model can run efficiently on the edge devices. The dynamic scheduling module is used to receive performance feedback from the model evaluation module, dynamically adjust the training parameters of the quantitative perception training module based on the feedback results, and determine the timing and method of model deployment based on the real-time resource status of the edge device.
2. The model quantization perception training system on the edge device according to claim 1, characterized in that The quantization-aware training module includes: The analog quantization configuration module is used to first determine the quantization parameters and then insert the quantization operation; Model initialization and training parameter setting module, used to first initialize the model and then set the training parameters; The forward propagation and simulation quantization execution module is used to first perform data input and forward calculation, and then simulate and quantize the effect; Back propagation and gradient approximation processing module, which is used to first calculate the loss function and then perform gradient approximation calculation; The model parameter optimization and update module is used to first update the parameters and then perform multiple iterative training.
3. The model quantization perception training system on the edge device according to claim 2, characterized in that The analog quantization configuration module includes: Determine the quantization parameter module, which is used to specify the quantization bit width, quantization method, and quantization granularity based on the hardware capabilities of the edge device and the performance requirements of the model; The quantization operation insertion module is used to insert quantization and dequantization operations during the forward propagation process of the model in the deep learning framework.
4. The model quantization perception training system on the edge device according to claim 2, characterized in that The model initialization and training parameter setting module includes: Model initialization module, used to initialize the selected deep learning model and set the model parameters and structure; The training parameter setting module is used to first determine the key parameters in the training process, including learning rate, batch size, and number of iterations. Then, a dynamic learning rate adjustment strategy is adopted to use a larger learning rate in the early stage of training to accelerate convergence, and gradually reduce the learning rate as training progresses to improve model performance.
5. The model quantization perception training system on the edge device according to claim 2, characterized in that The forward propagation and analog quantization execution module includes: The data input and forward calculation module is used to input the preprocessed data into the model and perform calculations according to the conventional forward propagation process. During the calculation process, the weights and activation values will undergo quantization and dequantization operations; The module for simulating quantization effects is used to simulate the effects of quantization in actual deployment through quantization and dequantization operations. This allows the model to perceive the changes brought about by quantization during training, providing a basis for subsequent parameter optimization.
6. The model quantization perception training system on the edge device according to claim 2, characterized in that The back propagation and gradient approximation processing module includes: The loss function calculation module is used to calculate the loss function value based on the model output and the true label, and measure the difference between the model's current prediction result and the actual situation; The gradient approximation calculation module is used to calculate the gradient using the gradient approximation technique. During backpropagation, the gradient is directly passed through the quantization operation without considering the non-differentiability of the quantization operation itself, thereby updating the model parameters.
7. The model quantization perception training system on the edge device according to claim 2, characterized in that The model parameter optimization and update module includes: The parameter update module is used to update the model parameters using an optimization algorithm based on the calculated gradients. During the update process, the model adjusts the parameters according to the quantized gradient information to adapt the model to the quantized representation and reduce the impact of quantization on model performance. The multi-iteration training module is used to repeat the above-mentioned forward propagation, backpropagation and parameter update process for multiple iterative training. During the training process, the loss value and accuracy index are continuously monitored so that the training parameters can be adjusted in time according to the changes in the indicators until the model performance reaches the expected target or tends to be stable.
8. The model quantization perception training system on the edge device according to claim 1, characterized in that The dynamic scheduling module includes: The real-time equipment status monitoring module is used to first collect resource data, then monitor performance indicators, and finally perceive environmental parameters; The training process dynamic adjustment module is used to first receive performance feedback, then optimize training parameters, and then fine-tune the quantization strategy; The model deployment intelligent decision-making module is used to first determine the deployment timing, then select the deployment plan, and finally distribute the deployment tasks; The continuous monitoring and feedback module of the operating status is used to first track the operating status, then handle abnormal situations, and then perform feedback optimization cycles.
9. The model quantization perception training system on the edge device according to claim 8, characterized in that The device status real-time monitoring module includes: The resource data collection module is used to collect key resource data of the device in real time through the built-in monitoring tools or system plug-ins of the edge device. The key resource data includes computing resources, memory usage, remaining storage space, and network bandwidth; The performance indicator monitoring module is used to monitor the performance indicators of the model when running on edge devices, including inference time, response latency, and throughput. It can also analyze the model inference process with the help of performance analysis tools to obtain detailed performance data on the computation time and data transmission time of each layer to evaluate the current operating efficiency of the model; Environmental parameter perception module, used to perceive the external environmental parameters of the edge device; The training process dynamic adjustment module includes: The performance feedback receiving module is used to establish a data interaction channel with the model evaluation module and receive performance feedback data during the model training process in real time; A training parameter optimization module, which is used to dynamically adjust the training parameters of the quantization-aware training module based on performance feedback and device resource status; The quantization strategy fine-tuning module is used to dynamically fine-tune the quantization strategy based on device resource conditions and model training performance.
10. The model quantization perception training system on the edge device according to claim 8, characterized in that The model deployment intelligent decision module includes: The deployment timing judgment module is used to comprehensively consider the device resource status and model performance evaluation results to determine the optimal time for model deployment; The deployment scheme selection module is used to select the most appropriate model deployment scheme based on the hardware characteristics and resource conditions of different edge devices; The deployment task distribution module is used to reasonably allocate the selected model deployment tasks to the target edge devices, accurately transfer the deployment resources of the model files and configuration parameters to the corresponding devices through the device management system or distributed deployment tools, and start the deployment process to ensure that the model can be correctly installed and run on the edge devices; The operating status continuous monitoring and feedback module includes: The running status tracking module is used to continuously monitor the running status of the model after it is deployed to the edge device, collect resource usage and performance indicator changes in real time during the model operation, and track the accuracy of the model's inference results and resource usage fluctuations through logging and real-time data monitoring; The abnormal situation processing module is used to take appropriate measures in time when abnormalities are detected in the model operation; The feedback optimization loop module is used to feed back the model operation status and processing results to the dynamic scheduling module as a reference for the next round of scheduling decisions. By continuously collecting and analyzing model operation data, it continuously optimizes the training parameter adjustment strategy, quantization strategy and deployment plan, forming a closed-loop dynamic optimization process to improve the operating efficiency and model performance of the entire system.
Citation Information
Patent Citations
Quantitative perception training method and device of neural network, and electronic equipment
CN113762061A
Deep learning system for edge device
CN118467164A
Edge device management system, management method and device
CN118587565A
Model quantification strategy determination method and device, model quantification method and device, medium and equipment
CN118690871A
Low-bit-width adaptive quantization method for image classification convolutional neural network
CN119312851A
Cited By
A progressive model inversion system and method for a data compression-free scenario
CN122635460A