Edge end AI equipment model training method

By performing preprocessing, selecting appropriate models, pruning, and knowledge distillation compression on edge devices, a distributed training environment is constructed. The optimized whale algorithm is used to allocate tasks, which solves the problem of resource constraints on edge devices and achieves efficient and accurate image processing.

CN120930724APending Publication Date: 2025-11-11CHINA NAT BUILDING MATERIALS TECH CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410969906.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Edge devices have limited computing power, storage resources, and energy supply, making it difficult to directly apply complex deep learning models, resulting in low image processing efficiency and the risk of data privacy leaks.

Method used

Preprocessing is performed on edge devices, a suitable pre-trained model is selected, the model is compressed through pruning and knowledge distillation, a distributed training environment is built, and the optimized whale algorithm is used to allocate training tasks for incremental learning and model fusion.

Benefits of technology

It achieves efficient and accurate image processing under limited resources, reduces resource waste, shortens the training cycle, and improves the model's adaptability and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930724A_ABST
    Figure CN120930724A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an edge end AI equipment model training method, which comprises the following steps of S1, preprocessing an original image; s2, selecting a proper pre-training model; s3, compressing and optimizing the model based on pruning and knowledge distillation operation; s4, constructing a distributed training environment, and distributing training tasks to a plurality of edge end AI devices for parallel operation; s5, adjusting parameters; s6, performing incremental learning on the edge end AI equipment by using the newly collected data; and S7, fusing each local model regularly, and updating the fused model to each edge end AI device. According to the method, the model is greatly compressed through pruning and knowledge distillation, a distributed training environment is constructed, the training speed is increased, the training period is shortened, training tasks are reasonably arranged by using the whale optimization algorithm according to the conditions of the edge end AI equipment, resource utilization is maximized, and resource waste is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method for training models for edge AI devices. Background Technology

[0002] With the rapid development of artificial intelligence technology, the demand for image processing in many fields is growing, such as security monitoring, autonomous driving, and medical diagnosis. However, traditional image processing methods usually rely on the powerful computing resources in the cloud for model training and inference, which not only leads to data transmission delays and bandwidth consumption, but may also cause the risk of data privacy leaks.

[0003] In practical applications, edge devices often need to process image data locally in real time and make fast and accurate decisions. However, these edge devices are limited in terms of computing power, storage resources and energy supply, making it difficult to directly apply complex deep learning models, especially image processing models such as neural networks, which have high accuracy but large storage resources. Therefore, there is a need for an edge AI device model training method that can achieve efficient and accurate image processing under limited resource conditions. Summary of the Invention

[0004] The purpose of this invention is to provide a method for training models for edge AI devices to solve the problems mentioned in the background art.

[0005] To address the aforementioned technical problems, the present invention aims to provide a method for training models for edge AI devices, comprising the following steps:

[0006] S1. Preprocess the original image on the edge AI device;

[0007] S2. Select a suitable pre-trained model based on the computing resources and task requirements of the edge AI device;

[0008] S3, a compression and optimization model based on pruning and knowledge distillation operations;

[0009] S4. Build a distributed training environment and distribute training tasks to multiple edge AI devices for parallel execution.

[0010] S5. Adjust parameters based on the real-time performance and resource status of edge AI devices;

[0011] S6. Use newly collected data to perform incremental learning on edge AI devices to update and optimize the model;

[0012] S7. Periodically merge the local models trained on each edge AI device and update the merged model on each edge AI device.

[0013] As a further improvement to this technical solution, the preprocessing in S1 includes image denoising, image enhancement, and image normalization.

[0014] As a further improvement to this technical solution, improved median filtering is used for image denoising in S1. Let the coordinates of a certain pixel in the image be (x, y), the size of the neighborhood window be m×n, and the pixel value in the neighborhood be f(i, j). For each pixel value f(i, j), check whether it is the maximum or minimum value in the neighborhood. If it is, it is marked as a noise point, and median filtering is applied to the current neighborhood window. If not, the pixel is not processed.

[0015] As a further improvement to this technical solution, the pruning operation in S3 is specifically as follows:

[0016] Step 1: Add an L1 regularization term to the loss function. For each weight in the neural network, the L1 regularization term is λ∑|w|, where λ is a hyperparameter that controls the strength of regularization, and ∑|w| is the L1 norm calculation for all weights w.

[0017] Step 2: Calculate the importance of the gradient evaluation weights, set a threshold, and prune the current weights when the gradient of a weight is less than the set threshold.

[0018] As a further improvement to this technical solution, the knowledge distillation operation in S3 is specifically as follows:

[0019] Step 1: Train the teacher model. Use rich data and powerful computing resources to train a large, high-performance teacher model. Teacher models usually have complex structures and a large number of parameters to achieve high accuracy.

[0020] Step 2: Define the distillation loss function. The distillation loss generally consists of two parts: the student model's prediction loss based on the true label and the student model's imitation loss based on the teacher model's output. Let the true label be y, and the student model's prediction be... The output of the teacher model is z t The student model's prediction of the teacher model's output is z. s ,have:

[0021]

[0022]

[0023] Among them, L prediction For the prediction loss function, y i For the i-th element in the actual tag, To predict the i-th element in the output of the student model, L distillationTo mimic the loss function, N is the total number of output elements, z s,i Let z be the i-th element of the student model's prediction of the teacher model's output. t,i This is the i-th element in the output of the teacher model;

[0024] Step 3: Train the student model using the distillation loss function from Step 2.

[0025] L tota =αL prediction +(1-α)L distillation

[0026] Among them, L total Let α be the total loss function, and α be the weight parameter where 0 < α < 1. By adjusting the value of α, the degree of emphasis of the student model on learning the real labels and imitating the output of the teacher model can be controlled.

[0027] Train the student model to minimize the total loss function L total .

[0028] As a further improvement to this technical solution, the specific steps of S4 are as follows:

[0029] S41. Assess the computing power, memory, and storage resources of each edge AI device;

[0030] S42. Based on the resource and task requirements of edge AI devices, allocate training tasks according to the optimized whale algorithm;

[0031] S43. Start the corresponding training process on each edge AI device, and exchange intermediate computing results and synchronization information between different edge AI devices through a distributed communication protocol.

[0032] As a further improvement to this technical solution, the specific steps in S42 for allocating training tasks based on the optimized whale algorithm are as follows:

[0033] S421. Initialization parameters: Set the whale population size, initial whale population position, and maximum number of iterations t. max The spatial dimension d of the whale, and the position of the whale represent the task allocation scheme;

[0034] S422. Boundary condition processing: Process each individual in the population separately and calculate the coefficient vector. and Generate uniformly distributed decision random numbers ρ.

[0035]

[0036]

[0037]

[0038] in, This is the control vector, and it decreases from 2 to 0 as the number of iterations increases throughout the iteration period. Let be a random vector on [0, 1], · represent element-wise multiplication, and t be the number of iterations. max Let w(t) be the maximum number of iterations, and w(t) be the inertia weight at the t-th iteration.

[0039] S423. Calculate the fitness value, using the total resource utilization of edge AI devices as the objective function;

[0040] S424, Random prey search: when ρ < 0.5 and hour,

[0041]

[0042] Where t is the current iteration number, This is the position vector of a randomly selected individual in the current population. Let be the position vector of the current whale individual in generation t. Let l be the position vector of the current whale individual in the (t+1)th generation, and l| represent the distance vector;

[0043] Shrink around prey: when ρ < 0.5 and At that time, a random disturbance term is introduced.

[0044]

[0045] in, This represents the optimal position vector for the objective function within the current population. To set parameters, To be a random increment that follows Brownian motion;

[0046] Spiral update position: when ρ≥0.5,

[0047]

[0048] Where s is a constant used to define the shape of the logarithmic spiral, and l is a uniformly distributed random number between [-1, 1].

[0049] As a further improvement to this technical solution, the real-time performance and resource status of the edge AI device in S5 include CPU utilization, memory usage, GPU computing power, network bandwidth, and storage capacity.

[0050] As a further improvement to this technical solution, S6 uses an online learning algorithm to input new data into the model batch by batch for training. At the same time, based on the concept drift detection method, when a significant change in data distribution is detected, the model is updated.

[0051] As a further improvement to this technical solution, S7 also includes monitoring the performance of each edge AI device after using the fusion model and collecting feedback information from the devices.

[0052] Compared with the prior art, the beneficial effects of the present invention are as follows: In this edge AI device model training method, the model is greatly compressed by pruning and knowledge distillation, a distributed training environment is constructed, the training speed is accelerated, the training cycle is shortened, and the training tasks are reasonably arranged by using the optimized whale algorithm according to the conditions of the edge AI device itself, so as to maximize resource utilization and reduce resource waste. Attached Figure Description

[0053] Figure 1 This is the overall flowchart of the present invention. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0055] like Figure 1 As shown, this embodiment provides a method for training an edge AI device model, including the following steps:

[0056] S1. Preprocess the original image on the edge AI device to improve accuracy;

[0057] S2. Based on the computing resources and task requirements of the edge AI device, select a suitable pre-trained model, usually a lightweight neural network.

[0058] S3. Based on pruning and knowledge distillation operations, the model is compressed and optimized to reduce the number of parameters and computational load, making it adaptable to the resource constraints of edge AI devices.

[0059] S4. Build a distributed training environment to distribute training tasks to multiple edge AI devices in parallel to improve training efficiency.

[0060] S5. Based on the real-time performance and resource status of edge AI devices, dynamically adjust parameters such as learning rate and training batch size to optimize the training process;

[0061] S6. Use newly collected data for incremental learning on edge AI devices to continuously update and optimize the model, thereby improving the model's adaptability and accuracy.

[0062] S7. Periodically merge the local models trained on various edge AI devices and update the merged model on each edge AI device to improve the model's generalization ability.

[0063] In this embodiment, the preprocessing in S1 includes image denoising, image enhancement, and image normalization. Image enhancement can use various algorithms such as sharpening filtering, Retinuex algorithm, and bilateral filtering. The formula for image normalization is:

[0064]

[0065] in, I represents the normalized pixel value, and I represents the pixel value of the original image. min I is the minimum pixel value in the original image. max This is the maximum pixel value in the original image.

[0066] Furthermore, in S1, an improved median filter is used for image denoising. Let the coordinates of a pixel in the image be (x, y), the size of the neighborhood window be m×n, and the pixel values ​​in the neighborhood be f(i, j). For each pixel value f(i, j), check whether it is the maximum or minimum value in the neighborhood. If it is, it is marked as a noise point, and the current neighborhood window is processed by median filtering. If not, the pixel is not processed. Noise points are basically extreme points of pixels in the neighborhood, while edge pixels do not have this characteristic. The improved median filter is used to distinguish edge points from noise, so that the edges are not blurred after filtering, and edge information is preserved, which helps image recognition and detection.

[0067] In this embodiment, the pruning operation in S3 is specifically as follows:

[0068] Step 1: Add an L1 regularization term to the loss function. For each weight in the neural network, the L1 regularization term is λ∑|w|, where λ is a hyperparameter that controls the strength of regularization, and ∑|w| is the L1 norm calculation for all weights w.

[0069] Step 2: Calculate the importance of the gradient evaluation weights, set a threshold, and prune the current weights when the gradient of a weight is less than the set threshold.

[0070] Furthermore, the knowledge distillation operation in S3 is specifically as follows:

[0071] Step 1: Train the teacher model. Use rich data and powerful computing resources to train a large, high-performance teacher model. Teacher models usually have complex structures and a large number of parameters to achieve high accuracy.

[0072] Step 2: Define the distillation loss function. The distillation loss generally consists of two parts: the student model's prediction loss based on the true label and the student model's imitation loss based on the teacher model's output. Let the true label be y, and the student model's prediction be... The output of the teacher model is z t The student model's prediction of the teacher model's output is z. s ,have:

[0073]

[0074]

[0075] Among them, L prediction For the prediction loss function, y i For the i-th element in the actual tag, To predict the i-th element in the output of the student model, L distillation To mimic the loss function, N is the total number of output elements, z s,i Let z be the i-th element of the student model's prediction of the teacher model's output. t,i This is the i-th element in the output of the teacher model;

[0076] Step 3: Train the student model using the distillation loss function from Step 2.

[0077] L total =αL prediction +(1-α)L distillation

[0078] Among them, l total Let α be the total loss function, and let α be the weight parameter where 0 < σ < 1. By adjusting the value of α, the emphasis of the student model on learning the real labels and imitating the output of the teacher model can be controlled.

[0079] Train the student model to minimize the total loss function L total .

[0080] In this embodiment, the specific steps of S4 are as follows:

[0081] S41. Assess the computing power, memory, and storage resources of each edge AI device;

[0082] S42. Based on the resource and task requirements of edge AI devices, allocate training tasks according to the optimized whale algorithm;

[0083] S421. Initialization parameters: Set the whale population size, initial whale population position, and maximum number of iterations t. max The spatial dimension d of the whale, and the position of the whale represent the task allocation scheme;

[0084] S422. Boundary condition processing: Process each individual in the population separately and calculate the coefficient vector. and Generate uniformly distributed decision random numbers ρ.

[0085]

[0086]

[0087]

[0088] in, This is the control vector, and it decreases from 2 to 0 as the number of iterations increases throughout the iteration period. Let be a random vector on [0, 1], · represent element-wise multiplication, and t be the number of iterations. max The maximum number of iterations is w(t), and the inertia weight at the t-th iteration is w(t). The adaptive inertia weight will gradually increase with the number of iterations, and enhance the influence of the optimal whale, so that other whales can quickly converge to the position of the optimal whale.

[0089] S423. Calculate the fitness value, using the total resource utilization of edge AI devices as the objective function;

[0090] S424, Random prey search: when ρ < 0.5 and hour,

[0091]

[0092] Where t is the current iteration number, This is the position vector of a randomly selected individual in the current population. Let be the position vector of the current whale individual in generation t. Let l represent the position vector of the current whale individual in the (t+1)th generation, and l·| represent the distance vector;

[0093] Shrink around prey: when ρ < 0.5 and At that time, a random disturbance term is introduced.

[0094]

[0095] in, This represents the optimal position vector for the objective function within the current population. To set parameters, To conform to the random increments of Brownian motion, it is easier to escape local optima;

[0096] Spiral update position: when ρ≥0.5,

[0097]

[0098] Where s is a constant used to define the shape of the logarithmic spiral, and l is a uniformly distributed random number between [-1, 1].

[0099] S43. Start the corresponding training process on each edge AI device, and exchange intermediate computing results and synchronization information between different edge AI devices through a distributed communication protocol.

[0100] In this embodiment, the real-time performance and resource status of the edge AI device in S5 include CPU utilization, memory usage, GPU computing power, network bandwidth, and storage capacity. When the device performance is high and resources are sufficient, the learning rate is appropriately increased to accelerate model convergence, and the training batch size is increased to make full use of hardware resources for efficient training. Conversely, if the device performance is limited or resources are scarce, the learning rate is reduced to avoid instability in the training process, and the training batch size is reduced to alleviate the device burden and ensure smooth training.

[0101] In this embodiment, S6 uses an online learning algorithm to input new data into the model batch by batch for training. At the same time, based on the concept drift detection method, when a significant change in data distribution is detected, the model is updated.

[0102] In this embodiment, S7 also includes monitoring the performance of each edge AI device after using the fusion model and collecting feedback information from the devices, mainly regarding the model's operating efficiency and accuracy, to further optimize and improve the update process.

[0103] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for training a model for an edge AI device, characterized in that, Includes the following steps: S1. Preprocess the original image on the edge AI device; S2. Select a suitable pre-trained model based on the computing resources and task requirements of the edge AI device; S3, a compression and optimization model based on pruning and knowledge distillation operations; S4. Build a distributed training environment and distribute training tasks to multiple edge AI devices for parallel execution. S5. Adjust parameters based on the real-time performance and resource status of edge AI devices; S6. Use newly collected data to perform incremental learning on edge AI devices to update and optimize the model; S7. Periodically merge the local models trained on each edge AI device and update the merged model on each edge AI device.

2. The edge AI device model training method according to claim 1, characterized in that: The preprocessing in S1 includes image denoising, image enhancement, and image normalization.

3. The edge AI device model training method according to claim 2, characterized in that: In step S1, an improved median filter is used for image denoising. Let the coordinates of a pixel in the image be (x, y), the size of the neighborhood window be m×n, and the pixel value in the neighborhood be f(i, j). For each pixel value f(i, j), check whether it is the maximum or minimum value in the neighborhood. If it is, it is marked as a noise point, and the median filter is applied to the current neighborhood window. If not, the pixel is not processed.

4. The edge AI device model training method according to claim 1, characterized in that, The pruning operation in S3 is specifically as follows: Step 1: Add an L1 regularization term to the loss function. For each weight in the neural network, the L1 regularization term is λ∑|w|, where λ is a hyperparameter that controls the strength of regularization, and ∑|wl is the L1 norm calculation for all weights w. Step 2: Calculate the importance of the gradient evaluation weights, set a threshold, and prune the current weights when the gradient of a weight is less than the set threshold.

5. The edge AI device model training method according to claim 1, characterized in that, The knowledge distillation operation in S3 is specifically as follows: Step 1: Train the teacher model. Use rich data and powerful computing resources to train a large-scale, high-performance teacher model. Step 2: Define the distillation loss function. The distillation loss generally consists of two parts: the student model's prediction loss based on the true label and the student model's imitation loss based on the teacher model's output. Let the true label be y, and the student model's prediction be... The output of the teacher model is z t The student model's prediction of the teacher model's output is z. s ,have: Among them, L prediction For the prediction loss function, y i For the i-th element in the actual tag, To predict the i-th element in the output of the student model, L distillation To mimic the loss function, N is the total number of output elements, z s,i Let z be the i-th element of the student model's prediction of the teacher model's output. t,i This is the i-th element in the output of the teacher model; Step 3: Train the student model using the distillation loss function from Step 2. L tota =αL prediction +(1-α)L distillation Among them, L total Let α be the total loss function, and let α be the weight parameter where 0 < α < 1. By adjusting the value of α, the emphasis of the student model on learning the real labels and imitating the output of the teacher model can be controlled. Train the student model to minimize the total loss function L total .

6. The edge AI device model training method according to claim 1, characterized in that, The specific steps of S4 are as follows: S41. Assess the computing power, memory, and storage resources of each edge AI device; S42. Based on the resource and task requirements of edge AI devices, allocate training tasks according to the optimized whale algorithm; S43. Start the corresponding training process on each edge AI device, and exchange intermediate computing results and synchronization information between different edge AI devices through a distributed communication protocol.

7. The edge AI device model training method according to claim 6, characterized in that, The specific steps in S42 for allocating training tasks based on the optimized whale algorithm are as follows: S421. Initialization parameters: Set the whale population size, initial whale population position, and maximum number of iterations t. max The spatial dimension d of the whale, and the position of the whale represent the task allocation scheme; S422. Boundary condition processing: Process each individual in the population separately and calculate the coefficient vector. and Generate uniformly distributed decision random numbers ρ. in, This is the control vector, and it decreases from 2 to 0 as the number of iterations increases throughout the iteration period. Let be a random vector on [0, 1], · represent element-wise multiplication, and t be the number of iterations. max Let w(t) be the maximum number of iterations, and w(t) be the inertia weight at the t-th iteration. S423. Calculate the fitness value, using the total resource utilization of edge AI devices as the objective function; S424, Random prey search: when ρ < 0.5 and hour, Where t is the current iteration number, For the position vector of a randomly selected individual in the current population, Let be the position vector of the current whale individual in generation t. Let be the position vector of the current whale individual in the (t+1)th generation, and || denote the distance vector; Shrink around prey: when ρ < 0.5 and At that time, a random disturbance term is introduced. in, This represents the optimal position vector for the objective function within the current population. To set parameters, To be a random increment that follows Brownian motion; Spiral update position: when ρ≥0.5, Where s is a constant used to define the shape of the logarithmic spiral, and l is a uniformly distributed random number between [-1, 1].

8. The edge AI device model training method according to claim 1, characterized in that: The real-time performance and resource status of edge AI devices in S5 include CPU utilization, memory usage, GPU computing power, network bandwidth, and storage capacity.

9. The edge AI device model training method according to claim 1, characterized in that: In step S6, an online learning algorithm is used to input new data into the model batch by batch for training. At the same time, based on the concept drift detection method, when a significant change in the data distribution is detected, the model is updated.

10. The edge AI device model training method according to claim 1, characterized in that: The S7 also includes monitoring the performance of each edge AI device after using the fusion model and collecting feedback information from the devices.