Neural network model lightweight method and system for adaptive edge device
By adopting an adaptive lightweight approach for neural network models on edge devices, the model is tiered and lightweighted according to the computing power and storage space of the edge devices. Through periodic evaluation and optimization, the problem of balancing model inference speed and accuracy on edge devices is solved, and efficient model deployment and optimization are achieved.
Patent Information
- Application Number
- CN202511219586.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-28
AI Technical Summary
Existing deep neural network models are difficult to deploy efficiently on edge computing devices, and static lightweight methods cannot adapt to the diversity of different edge devices, making it difficult to achieve the optimal balance between model inference speed and accuracy.
An adaptive lightweight neural network model approach for edge devices is adopted. By constructing a lightweight model matching table, the model is lightweighted in stages according to the computing power and storage space of the edge device. The model is dynamically adjusted by periodically evaluating and providing feedback to optimize the model.
It improves the inference speed and accuracy of the model at the edge, meets the requirements of real-time performance and accuracy, achieves the optimal balance of the model at different working stages, and enhances the flexibility and applicability of the system.
Smart Images

Figure CN120851111A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and system for lightweighting neural network models for adaptive edge devices. Background Technology
[0002] In recent years, the Internet of Things (IoT) and edge computing technologies have developed rapidly, and the demand for intelligent applications running on edge devices is increasing daily. Deep neural networks have performed exceptionally well in fields such as image recognition and speech recognition, becoming a core technology driving the development of intelligent applications. However, traditional large-scale deep neural network models have a large number of parameters and computational demands, requiring significant resources. Edge devices, with their limited hardware capabilities, are difficult to deploy and run directly, leading to decreased inference performance and an inability to meet real-time and accuracy requirements. To address this, edge computing models have emerged. These models centralize data collection, computation, and processing locally, eliminating the need to transmit large amounts of data to the cloud. This reduces data transmission volume and network bandwidth pressure, while ensuring real-time and stable data processing, rapid response to local needs, and data security and privacy, effectively solving the problems inherent in cloud computing models.
[0003] Despite the numerous advantages of edge computing, the terminal processors used in edge computing are mostly embedded processors with relatively low performance, making it difficult to support complex data processing operations such as deep neural networks. Migrating complex computing tasks that originally ran on cloud servers to edge devices presents significant challenges. Therefore, researching the application and transformation of deep neural network models in edge computing scenarios has become a key task in promoting the development of edge intelligence.
[0004] Traditional cloud computing models require data to be transmitted to the cloud for centralized processing, and their limitations are becoming increasingly apparent, specifically: (1) data transmission is constrained by network conditions; (2) data processing is limited by the computing bottleneck of cloud servers; and (3) data transmission over the network poses security and privacy risks. Edge computing models can effectively solve the above problems of cloud computing models.
[0005] To improve the performance of deep neural network models, the model size is constantly increasing, but high performance is often accompanied by large size and high power consumption. However, the terminal processors of edge computing are mostly embedded processors with limited performance, making it difficult to support complex data processing operations such as deep neural networks. Migrating the complex computing tasks of cloud servers to edge devices faces huge challenges, which severely limits the application of deep neural networks at the edge.
[0006] To achieve the structural design and transformation of lightweight deep neural network models, it is necessary to explore a novel and efficient comprehensive compression and lightweight transformation scheme. Currently, deep neural network compression methods mainly include weight pruning, knowledge distillation, refined network design, weight quantization, and low-rank decomposition. These methods compress neural network models from different perspectives and each has its own characteristics.
[0007] While lightweight methods such as model quantization and model pruning exist, most are static optimization strategies, lacking flexibility and unable to dynamically adjust based on the storage space, computing power, and inference accuracy and recall parameters of different edge devices. In practical applications, edge devices exhibit significant hardware variations, and the resource requirements of the same device differ at different stages of operation. Existing static lightweight methods struggle to adapt to this diversity, failing to achieve an optimal balance between model inference speed and accuracy across various edge computing devices, thus limiting the widespread application of deep learning models at the edge. Summary of the Invention
[0008] To address these issues, this invention proposes a lightweight method and system for adaptive edge device neural network models.
[0009] According to one aspect of the present invention, a lightweight method for neural network models of adaptive edge devices is proposed, comprising the following steps:
[0010] S1. Construct and train a neural network model, and perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices. The index parameters include computing power index parameters and storage space index parameters, and obtain a matching table of typical edge computing power and lightweight model.
[0011] S2, perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table;
[0012] S3. Periodically evaluate the indicator parameters of the edge device to be deployed to obtain real-time evaluation results. Based on the real-time evaluation results and the actual usage effect indicator parameters of the lightweight model, determine whether to reload the new lightweight model or optimize the lightweight model.
[0013] Specifically, in S1, the neural network model is graded and lightweighted according to the index parameters of typical edge devices. This includes: classifying the computing power of the devices according to the computing power index parameters of the typical edge devices, performing a first graded lightweighting of the neural network model, and then performing a second graded lightweighting of the neural network model according to the storage space index parameters of the typical edge devices.
[0014] Specifically, the first-stage lightweighting of the neural network model involves classifying the computing power of the typical edge devices based on their computing power index parameters. This includes: weighting and summing the single-threaded computing speed and parallel computing capability of the typical edge devices to obtain a comprehensive computing power score, classifying the computing power level according to a certain threshold, and then performing lightweighting of the neural network model to suit different computing power levels for different classification results.
[0015] Specifically, the second-level lightweighting of the neural network model based on the storage space index parameters of the typical edge device includes: performing a second-level lightweighting based on the total storage space and remaining storage space of the typical edge device.
[0016] Specifically, in S2, an initial evaluation of the performance parameters of the edge device to be deployed is performed to obtain an initial evaluation result, which includes: obtaining the initial evaluation result based on the storage space performance parameters of the edge device to be deployed and the ratio of the computing power performance parameters of the edge device to the benchmark edge device, as measured by a benchmark test program.
[0017] Specifically, the actual performance metrics of the lightweight model include: inference speed, accuracy, recall, or mAP.
[0018] Specifically, S3 includes: calculating an initial comprehensive computing power score based on the computing power index parameters of the initial evaluation results; calculating a real-time comprehensive computing power score based on the computing power index parameters of the real-time evaluation results; setting a first comparison parameter and a second comparison parameter; when the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is less than the first comparison parameter, or the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is greater than the second comparison parameter and the actual usage effect index parameters of the lightweight model do not meet the usage requirements, then a new lightweight model is reloaded on the edge device to be deployed according to the real-time evaluation results and the typical edge computing power and lightweight model matching table; when the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is greater than or equal to the first comparison parameter and less than or equal to the second comparison parameter, and the actual usage effect index parameters of the lightweight model do not meet the usage requirements, the lightweight model is optimized.
[0019] Specifically, optimizing the lightweight model includes: when the actual performance metrics of the lightweight model do not meet the requirements, setting a set of optimization parameters, including a first optimization parameter, a second optimization parameter, and a third optimization parameter, and providing feedback on the error rate parameter of the lightweight model; when the error rate parameter is greater than the first optimization parameter, determining the cause of the errors in the lightweight model and adjusting the lightweight algorithm of the lightweight model according to the cause of the errors; when the error rate parameter is greater than the second optimization parameter, retraining the lightweight model; and when the error rate parameter is greater than the third optimization parameter, readjusting the architecture of the lightweight model.
[0020] According to one aspect of the present invention, a lightweight system for neural network models of adaptive edge devices is proposed, comprising the following modules according to any one of the first aspects:
[0021] The hierarchical lightweight module is configured to build and train a neural network model, and perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices. The index parameters include computing power index parameters and storage space index parameters, and obtain a matching table of typical edge computing power and lightweight model.
[0022] The model deployment module is configured to perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table.
[0023] The model update module periodically evaluates the indicator parameters of the edge device to be deployed, obtains real-time evaluation results, and determines whether to reload a new lightweight model or optimize the lightweight model based on the real-time evaluation results and the actual usage effect of the lightweight model.
[0024] According to one aspect of the present invention, a computer program product is provided having a computer program stored thereon, which, when executed by a processor, performs the method as described in the first aspect.
[0025] The advantages of this invention are:
[0026] (1) Strong adaptability: It can adaptively adjust according to the storage space, computing power and other parameters of different edge devices, and is applicable to various types of edge computing devices, which improves the versatility and applicability of the method.
[0027] (2) Performance optimization: Through adaptive model lightweighting adjustment, the inference speed of the model at the edge can be significantly improved while ensuring model accuracy. This reduces the decline in model inference performance caused by insufficient computing power and storage space of edge devices, and meets the real-time and accuracy requirements of practical applications.
[0028] (3) Dynamic balancing: It can achieve adaptive and optimal balance between model accuracy and inference speed at the edge according to the needs of different working stages, thereby improving the flexibility and efficiency of the system.
[0029] (4) Error feedback optimization: Based on the actual application of the model at the edge, error cases are fed back, and the error cases are analyzed to continuously optimize the model and lightweight algorithm. Attached Figure Description
[0030] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments and, together with the description, serve to explain the principles of the invention. Other embodiments and many anticipated advantages of the embodiments will be readily recognized as they become better understood through reference to the following detailed description. Elements in the drawings are not necessarily to scale. The same reference numerals refer to corresponding similar parts.
[0031] Figure 1 A flowchart illustrating a method for lightweighting a neural network model for an adaptive edge device according to the present invention is shown.
[0032] Figure 2 A detailed flowchart of a method for lightweighting a neural network model for an adaptive edge device according to the present invention is shown.
[0033] Figure 3 A schematic diagram of a lightweight neural network model system for an adaptive edge device according to the present invention is shown.
[0034] Figure 4 A schematic diagram of a computer system architecture suitable for implementing the embodiments of this application is shown. Detailed Implementation
[0035] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0036] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0037] Figure 1 A lightweight method for neural network models in adaptive edge devices is shown, comprising the following steps:
[0038] S1. Construct and train a neural network model, and perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices. The index parameters include computing power index parameters and storage space index parameters, and obtain a matching table of typical edge computing power and lightweight model.
[0039] S2, perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table;
[0040] S3. Periodically evaluate the indicator parameters of the edge device to be deployed to obtain real-time evaluation results. Based on the real-time evaluation results and the actual usage effect indicator parameters of the lightweight model, determine whether to reload the new lightweight model or optimize the lightweight model.
[0041] Specifically, in S1, the neural network model is graded and lightweighted according to the index parameters of typical edge devices. This includes: classifying the computing power of the devices according to the computing power index parameters of the typical edge devices, performing a first graded lightweighting of the neural network model, and then performing a second graded lightweighting of the neural network model according to the storage space index parameters of the typical edge devices.
[0042] Specifically, the first-stage lightweighting of the neural network model involves classifying the computing power of the typical edge devices based on their computing power index parameters. This includes: weighting and summing the single-threaded computing speed and parallel computing capability of the typical edge devices to obtain a comprehensive computing power score, classifying the computing power level according to a certain threshold, and then performing lightweighting of the neural network model to suit different computing power levels for different classification results.
[0043] Specifically, the second-level lightweighting of the neural network model based on the storage space index parameters of the typical edge device includes: performing a second-level lightweighting based on the total storage space and remaining storage space of the typical edge device.
[0044] Lightweighting methods for models include model pruning, quantization, knowledge distillation, and low-rank decomposition.
[0045] The benchmark program includes the following functions: ① Single-threaded computation speed evaluation: By running matrix multiplication (e.g., A×B) and convolution operations (e.g., sliding computation of a 3×3 convolution kernel on a 224×224 image), the time taken for a single operation is recorded. ② Parallel computing capability evaluation: The same test is run using multi-threading (e.g., OpenMP) or GPU acceleration (e.g., CUDA), and the speedup ratio after parallelization is recorded (e.g., single-threaded time / multi-threaded time). For example, if the single-threaded time is 100ms and the 4-threaded time is 30ms, the speedup ratio is 3.33, indicating higher parallel efficiency.
[0046] like Figure 2 As shown, in a specific embodiment, the model is built and trained in the cloud. Based on the computing power of typical edge devices, the model is tiered and lightweighted, establishing a "Matching Table of Typical Edge Computing Power and Lightweight Models." The single-threaded computing speed and parallel computing capability of the devices are evaluated through the preceding benchmark test program to tier the device computing power. A comprehensive computing power score is obtained by weighted summation of the single-threaded computing speed and parallel computing capability, and the score structure is tiered according to a certain threshold. For different tiering results, the neural network structure of the model is adjusted in the cloud, and lightweight models adapted to different computing power levels are trained in the cloud. The training results must ensure that the accuracy of the validation dataset is within the required range (different application scenarios have different requirements, such as recognition accuracy greater than 90%). The device's storage space parameters, including total storage space and remaining storage space, are obtained through the device's provided API interface.
[0047] After training, the model undergoes lightweight processing such as pruning or quantization. Similarly, the accuracy of the lightweight model needs to be evaluated using a validation dataset to ensure it meets the required range. If the model trained in the cloud fails to meet the required recognition accuracy when evaluated using a validation dataset, the model parameters and network structure will be adjusted before retraining. This ultimately results in a "Typical Edge Computing Power and Lightweight Model Matching Table," which is generated and maintained in the cloud before the edge device requests to load the model.
[0048] Specifically, in S2, an initial evaluation of the performance parameters of the edge device to be deployed is performed to obtain an initial evaluation result, which includes: obtaining the initial evaluation result based on the storage space performance parameters of the edge device to be deployed and the ratio of the computing power performance parameters of the edge device to the benchmark edge device, as measured by a benchmark test program.
[0049] When an edge device needs to apply a model, it sends a request to the cloud or server. A benchmark test program evaluates the edge device's computing power metrics and obtains the corresponding storage space metrics. Based on the evaluation results, a suitable lightweight model from the "Typical Edge Computing Power and Lightweight Model Matching Table" is selected and distributed to the edge device to be deployed. If a suitable lightweight model cannot be found in the "Typical Edge Computing Power and Lightweight Model Matching Table," the cloud / server performs adaptive lightweight model processing based on the evaluation results of the current edge device to be deployed and distributes it to the edge device.
[0050] In a specific embodiment, the computational load of the model is compressed and the inference speed of the model is adjusted based on the ratio of the computing power index parameters of the device to those of the baseline edge device.
[0051] By using the computing power of a selected edge device (usually a typical edge device used to test the model) as a benchmark and defining its computing power as 100%, the computing speed ratio or parallel acceleration ratio of the edge device to be deployed compared with the benchmark edge device is calculated, thereby defining whether the edge device to be deployed is a high-computing-power edge device or a low-computing-power edge device.
[0052] For example, in a target detection task, if the computing speed of the target device is less than 20% of that of the benchmark device, or the parallel speedup ratio is less than 2, it is defined as "weak computing power".
[0053] For devices with limited computing power, more efficient algorithms, such as depthwise separable convolutions, can be used to reduce the computational cost of the model. Simultaneously, the number of layers or neurons per layer can be reduced (e.g., reducing ResNet-18 from 18 layers to 12 layers) to speed up inference. For devices with greater computing power, the model complexity can be appropriately increased (e.g., expanding the number of channels in MobileNetV2 from 32 to 64) to improve model accuracy (model accuracy needs to be obtained through actual evaluation on a validation dataset during cloud training).
[0054] When the remaining storage space on the device is small (meaning the remaining storage space on the edge device is less than twice the size of the model, in which case the system cannot guarantee the storage of temporary files, the regular updating of the model, etc. If the edge device needs to run multiple models simultaneously, each model needs to reserve independent space), the compression ratio of the model parameter quantization can be appropriately increased, or the model pruning can be intensified to remove more unimportant neuron connections and parameters. When the device storage space is large, the compression ratio can be appropriately reduced to retain more model information and ensure the accuracy of the model.
[0055] In specific embodiments, unimportant neuron connections and parameters can be identified and removed using structured pruning and weight-based pruning methods. Weight-based pruning removes weights with small absolute values by calculating the absolute value or L1 / L2 norm of neuron connections (weights). For example, based on the target compression ratio, a threshold θ = 0.01 is set; if a connection in the weight matrix has a weight of 0.001, that connection is removed. Structured pruning includes channel pruning and neuron pruning. Channel pruning evaluates the importance of each channel using metrics such as L1 / L2 norm, gradient, and activation value, deletes the least important channels, and simultaneously deletes the corresponding rows / columns of the convolutional kernel. Then, the model accuracy is restored by fine-tuning the weights of the remaining channels. Neuron pruning targets fully connected neural networks, evaluates the importance of each neuron using metrics such as L1 norm and activation value, and deletes the least important neurons and all their input / output connections.
[0056] The overall computing power score for this test will be recorded as cp1.
[0057] Specifically, the actual performance metrics of the lightweight model include: inference speed, accuracy, recall, or mAP.
[0058] Specifically, S3 includes: calculating an initial comprehensive computing power score based on the computing power index parameters of the initial evaluation results; calculating a real-time comprehensive computing power score based on the computing power index parameters of the real-time evaluation results; setting a first comparison parameter and a second comparison parameter; when the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is less than the first comparison parameter, or the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is greater than the second comparison parameter and the actual usage effect index parameters of the lightweight model do not meet the usage requirements, then a new lightweight model is reloaded on the edge device to be deployed according to the real-time evaluation results and the typical edge computing power and lightweight model matching table; when the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is greater than or equal to the first comparison parameter and less than or equal to the second comparison parameter, and the actual usage effect index parameters of the lightweight model do not meet the usage requirements, the lightweight model is optimized.
[0059] During edge device inference, benchmark tests are run periodically (e.g., daily or during peak and off-peak periods) to assess computing power. Based on the assessment results, a decision is made on whether to load a more suitable lightweight model. The lightweighting strategy is dynamically adjusted based on monitoring results to achieve an adaptive and optimal balance between model accuracy and inference speed at the edge.
[0060] Specifically, the computing power benchmark test described above is performed on the edge devices at regular intervals to obtain a comprehensive computing power score. The score is then compared with the "Typical Edge Computing Power and Lightweight Model Matching Table" described above to determine whether the current comprehensive computing power score cp2 is compared with the initial comprehensive computing power score cp1.
[0061] Let ρ be the first comparison parameter and λ be the second comparison parameter. The function of ρ is to determine whether cp2 is significantly less than cp1. If this is the case, it indicates that a more lightweight model is needed to address the issue of slow inference speeds during peak computing power demands on edge devices in practical applications. The function of λ is to determine whether cp2 exceeds cp1 by a certain range. In this case, a more accurate model can be used instead. The value range of both coefficients is: 0 < ρ < 1, λ > 1.
[0062] In one embodiment, if cp2 / cp1 < ρ (e.g., ρ = 0.5), the edge device needs to reload the model from the cloud. If cp2 / cp1 > λ (e.g., λ = 1.5), and the actual recognition accuracy at the edge is lower than the required value, the edge device reloads a model with better recognition accuracy from the cloud using the current computing power score cp2. Otherwise, the original model remains unchanged.
[0063] If the change in cp2 compared to cp1 does not exceed the threshold (ρ≤cp2 / cp1≤λ), but the actual recognition accuracy on the edge device to be deployed is lower than the required value, then optimization is required. Otherwise, no adjustment is needed.
[0064] Specifically, optimizing the lightweight model includes: when the actual performance metrics of the lightweight model do not meet the requirements, setting a set of optimization parameters, including a first optimization parameter, a second optimization parameter, and a third optimization parameter, and providing feedback on the error rate parameter of the lightweight model; when the error rate parameter is greater than the first optimization parameter, determining the cause of the errors in the lightweight model and adjusting the lightweight algorithm of the lightweight model according to the cause of the errors; when the error rate parameter is greater than the second optimization parameter, retraining the lightweight model; and when the error rate parameter is greater than the third optimization parameter, readjusting the architecture of the lightweight model.
[0065] In actual use, the actual performance of the model is cumulatively evaluated. If the actual performance does not meet the requirements of the indicators (including but not limited to model inference speed, accuracy, recall, mAP, etc., the indicators to be focused on are different in different application scenarios), the error information is fed back to the cloud or server. The cloud or server performs statistical analysis on the error to determine whether it is necessary to re-lightweight the model, re-add error data for training, or readjust the model architecture.
[0066] Set a set of optimization parameters, including a first optimization parameter, a second optimization parameter, and a third optimization parameter, α, β, and γ, respectively. The settings of α, β, and γ need to be tested and configured based on the actual application scenario and specific business requirements. Several sets of parameters can be set initially (e.g., Group 1: α = 0.3; β = 0.3; γ = 0.3; Group 2: α = 0.2; β = 0.3; γ = 0.3, etc.; Group 3: α = 0.2; β = 0.2; γ = 0.3, etc.; and so on) for model training and testing, analyzing the model's performance under different parameter combinations. Compare the model's precision, recall, mAP, and other metrics under different parameter combinations. Based on the analysis results and specific business requirements (e.g., some application scenarios require high precision, while others prioritize recall), select the parameter configuration that optimizes the model's performance.
[0067] The optimization of the lightweight model can be divided into the following cases:
[0068] (1) Use error examples to evaluate the model before and after lightweighting. If there is a difference exceeding a certain proportion α (α can be 0.1, 0.2, 0.3, etc., which needs to be set for testing in combination with actual application scenarios and specific business needs), then the lightweighting method needs to be readjusted. Then, it is necessary to determine whether the proportion is caused by loss of precision or loss of key features, and make specific adjustments to the lightweighting algorithm.
[0069] Model precision loss refers to a decrease in overall performance (such as classification accuracy and mAP) after model weighting, but the model can still capture the key features of the input data. However, the reduced model complexity leads to blurred classification boundaries or decreased confidence. A more specific method for judgment is to perform inference on the model before and after weighting in the cloud. If the output confidence of the model generally decreases after weighting (e.g., from 0.9 to 0.6), but the classification result is still correct, then it is judged as precision loss.
[0070] Key feature loss in a model refers to the inability of a lightweight model to correctly extract or identify some key features in the input data (such as object edges in an image or semantic information in text), leading to a significant increase in the prediction error rate. A specific method for determining this is to use visualization tools (such as Grad-CAM) to show the model's region of interest in the input data. If the model's region of interest deviates from the key features (such as the main object in an image) after lightweighting, then key features may be lost.
[0071] If precision loss occurs, one of the following methods can be chosen to address it: ① Use the lightweighted model and fine-tune it on the original training data to restore some precision. ② Use the un-lightweighted model as the teacher model and the lightweighted model as the student model. Transfer the knowledge from the teacher model to the student model through knowledge distillation, optimize the distillation loss function (such as KL divergence), and improve the performance of the student model.
[0072] If the problem of key feature loss occurs, the following methods can be selected to handle it: ① Increase the number of layers or channels in the lightweight model to improve the model's feature extraction capability; ② Introduce attention mechanisms (such as SE module, CBAM module) to enhance the model's attention to key features; ③ Fuse the features of the lightweight model with those of the original model (such as feature concatenation, feature weighting) to retain key features.
[0073] (2) If more than a certain proportion of β (β can be 0.1, 0.2, 0.3, etc., and needs to be set according to the actual application scenario and specific business needs) of the error cases are significantly different from the training data on the cloud or server in terms of feature distribution, scenario type, etc., then the model needs to be retrained.
[0074] (3) If more than a certain proportion of errors (γ can be 0.1, 0.2, 0.3, etc., and needs to be set according to the actual application scenario and specific business requirements) are caused by model architecture errors, such as severe overfitting, underfitting, or inability to capture spatial hierarchical relationships and temporal information in the image, then the model architecture needs to be adjusted and retrained. Methods for adjusting the model architecture: ① For overfitting, reduce the number of network layers or the number of neurons per layer, or introduce regularization, or apply early stopping techniques (terminating training early when the validation set performance no longer improves). ② For underfitting, increase the number of network layers or the number of neurons per layer, or reduce the L1 / L2 regularization coefficient or Dropout rate, etc. ③ For cases where spatial hierarchical relationships and temporal information in the image cannot be captured, spatial or temporal modules need to be introduced, or multi-scale feature fusion methods need to be applied.
[0075] According to one aspect of the present invention, a lightweight system for neural network models of adaptive edge devices is proposed, comprising the following modules according to any one of the first aspects:
[0076] The hierarchical lightweight module 301 is configured to build and train a neural network model, and perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices. The index parameters include computing power index parameters and storage space index parameters, and obtain a matching table of typical edge computing power and lightweight model.
[0077] The model deployment module 302 is configured to perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table.
[0078] The model update module 303 periodically evaluates the indicator parameters of the edge device to be deployed, obtains real-time evaluation results, and determines whether to reload a new lightweight model or optimize the lightweight model based on the real-time evaluation results and the actual usage effect of the lightweight model.
[0079] The following is for reference. Figure 4 It shows a schematic diagram of the structure of a computer system 400 suitable for implementing electronic devices according to embodiments of the present application. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0080] like Figure 4 As shown, the computer system 400 includes a central processing unit (CPU) 401, which performs various appropriate actions and processes based on programs stored in read-only memory (ROM) 402 or programs loaded from storage section 409 into random access memory (RAM) 404. The RAM 404 also stores various programs and data required for the operation of the system 400. The CPU 401, ROM 402, ROM 403, and RAM 404 are interconnected via a bus 405. An input / output (I / O) interface 406 is also connected to the bus 405.
[0081] The following components are connected to I / O interface 406: an input section 407 including a keyboard, mouse, etc.; an output section 408 including a liquid crystal display (LCD) and speakers, etc.; a storage section 409 including a hard disk, etc.; and a communication section 410 including a network interface card such as a LAN card and a modem, etc. The communication section 410 performs communication processing via a network such as the Internet. A drive 411 is also connected to I / O interface 406 as needed. A removable medium 412, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 411 as needed so that computer programs read from it can be installed into storage section 409 as needed.
[0082] Specifically, according to embodiments of this disclosure, the processes described above with reference to the flowcharts are implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program is downloaded and installed from a network via communication section 410, and / or installed from removable medium 412. When the computer program is executed by central processing unit (CPU) 401, it performs the functions defined in the methods of this application.
[0083] It should be noted that the computer-readable storage medium of this application is a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium is, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium is any tangible medium containing or storing a program that is used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium includes a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium or any computer-readable storage medium other than a computer-readable storage medium may transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0084] Computer program code for performing the operations of this application is written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code executes entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer is connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or connected to an external computer (e.g., via the Internet using an Internet service provider).
[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram represents a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually execute substantially in parallel, and they may sometimes execute in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, is implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0086] The modules described in the embodiments of this application are implemented in software or hardware.
[0087] In another aspect, this application also provides a computer-readable storage medium included in the electronic device described in the above embodiments; or existing independently and not assembled into the electronic device. The computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: S1, construct and train a neural network model, perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices, the index parameters including computing power index parameters and storage space index parameters, and obtain a typical edge computing power and lightweight model matching table; S2, perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table; S3, periodically evaluate the index parameters of the edge device to be deployed to obtain a real-time evaluation result, and determine whether to reload a new lightweight model or optimize the lightweight model based on the real-time evaluation result and the actual usage effect index parameters of the lightweight model.
[0088] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A lightweight method for neural network models of adaptive edge devices, characterized in that, Includes the following steps: S1. Construct and train a neural network model, and perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices. The index parameters include computing power index parameters and storage space index parameters, and obtain a matching table of typical edge computing power and lightweight model. S2, perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table; S3. Periodically evaluate the indicator parameters of the edge device to be deployed to obtain real-time evaluation results. Based on the real-time evaluation results and the actual usage effect indicator parameters of the lightweight model, determine whether to reload the new lightweight model or optimize the lightweight model.
2. The method for lightweighting neural network models for adaptive edge devices according to claim 1, characterized in that, S1 performs hierarchical lightweighting of the neural network model based on the index parameters of typical edge devices. Specifically, this includes: classifying the computing power of the typical edge devices according to their computing power index parameters, performing a first-level lightweighting of the neural network model, and then performing a second-level lightweighting of the neural network model according to the storage space index parameters of the typical edge devices.
3. The method for lightweighting neural network models for adaptive edge devices according to claim 2, characterized in that, The computing power of the typical edge devices is classified according to their computing power index parameters. The first classification and lightweighting of the neural network model specifically includes: weighting and summing the single-threaded computing speed and parallel computing capability of the typical edge devices to obtain a comprehensive computing power score, classifying the computing power level according to a certain threshold, and performing lightweighting of the neural network model to suit different computing power levels for different classification results.
4. The method for lightweighting a neural network model for adaptive edge devices according to claim 2, characterized in that, The second level of lightweighting of the neural network model based on the storage space index parameters of the typical edge device specifically includes: performing a second level of lightweighting based on the total storage space and remaining storage space of the typical edge device.
5. The method for lightweighting a neural network model for adaptive edge devices according to claim 1, characterized in that, In S2, an initial evaluation is performed on the indicator parameters of the edge device to be deployed to obtain an initial evaluation result. Specifically, the initial evaluation result is obtained based on the storage space indicator parameters of the edge device to be deployed and the ratio of the computing power indicator parameters of the edge device to the computing power indicator parameters of the benchmark edge device, which are measured by a benchmark test program.
6. The method for lightweighting a neural network model for adaptive edge devices according to claim 1, characterized in that, The actual performance metrics of the lightweight model include: inference speed, accuracy, recall, or mAP.
7. The method for lightweighting neural network models for adaptive edge devices according to claim 1, characterized in that, S3 specifically includes: calculating an initial comprehensive computing power score based on the computing power index parameters of the initial evaluation results; calculating a real-time comprehensive computing power score based on the computing power index parameters of the real-time evaluation results; setting a first comparison parameter and a second comparison parameter; when the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is less than the first comparison parameter, or the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is greater than the second comparison parameter and the actual usage effect index parameters of the lightweight model do not meet the usage requirements, then a new lightweight model is reloaded on the edge device to be deployed according to the real-time evaluation results and the typical edge computing power and lightweight model matching table; when the ratio of the real-time comprehensive computing power score to the initial comprehensive computing power score is greater than or equal to the first comparison parameter and less than or equal to the second comparison parameter, and the actual usage effect index parameters of the lightweight model do not meet the usage requirements, the lightweight model is optimized.
8. The method for lightweighting a neural network model for adaptive edge devices according to claim 7, characterized in that, Optimizing the lightweight model specifically includes: when the actual performance metrics of the lightweight model do not meet the usage requirements, setting a set of optimization parameters, including a first optimization parameter, a second optimization parameter, and a third optimization parameter, and providing feedback on the error rate parameter of the lightweight model; when the error rate parameter is greater than the first optimization parameter, determining the cause of the errors in the lightweight model and adjusting the lightweight algorithm of the lightweight model according to the cause of the errors; when the error rate parameter is greater than the second optimization parameter, retraining the lightweight model; and when the error rate parameter is greater than the third optimization parameter, readjusting the architecture of the lightweight model.
9. A lightweight system for neural network models of adaptive edge devices, characterized in that, The method according to any one of claims 1 to 8 comprises the following modules: The hierarchical lightweight module is configured to build and train a neural network model, and perform hierarchical lightweighting of the neural network model according to the index parameters of typical edge devices. The index parameters include computing power index parameters and storage space index parameters, and obtain a matching table of typical edge computing power and lightweight model. The model deployment module is configured to perform an initial evaluation of the index parameters of the edge device to be deployed to obtain an initial evaluation result, and deploy the corresponding lightweight model on the edge device to be deployed according to the initial evaluation result and the typical edge computing power and lightweight model matching table. The model update module periodically evaluates the indicator parameters of the edge device to be deployed, obtains real-time evaluation results, and determines whether to reload a new lightweight model or optimize the lightweight model based on the real-time evaluation results and the actual usage effect of the lightweight model.
10. A computer program product, characterized in that, It stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.
Citation Information
Cited By
Control method and device of edge AI model, edge equipment and storage medium
CN121098718A
Lightweight model edge deployment optimization method based on incremental learning
CN121660113A
Agent AI-based edge gateway system
CN121814611A