Edge real-time video analysis-oriented resource efficient continuous learning method and system
By designing an online model accuracy degradation predictor and a DRL controller to dynamically trigger retraining, the problem of limited edge server resources was solved, improving the accuracy and real-time performance of edge video analytics.
Patent Information
- Application Number
- CN202511101702.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
Smart Images

Figure CN120976829A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of edge video analysis, in particular to a resource-efficient continual learning method and system for edge real-time video analysis. BACKGROUND
[0002] With the rapid development of artificial intelligence and communication technology, real-time video analysis shows a wide application prospect in the fields of smart city, traffic management, industrial monitoring, etc. With the help of deep neural networks (DNNs), target detection, behavior recognition, video summarization, etc. can effectively improve the automation and intelligence level of video analysis system. In order to reduce video transmission overhead and protect data privacy, more and more video analysis systems are deployed on the network edge closer to the user, which not only reduces the data processing pressure of the cloud, but also improves the system response speed. However, due to the available resources such as computing, storage and bandwidth, the edge server is still difficult to efficiently support complex and computationally intensive high-precision DNN models.
[0003] In order to alleviate the contradiction between the limited resources of edge system and the high demand of real-time video analysis, some works try to compress the original complex DNN model to reduce its resource overhead while ensuring the accuracy of video analysis. The training method of these lightweight DNN models can achieve good results under specific data distribution. However, after the model is trained and deployed, the actual video stream data distribution it needs to process may be significantly different from the data distribution in the training phase (i.e. data drift), which leads to a serious decline in model performance. The occurrence of data drift is mainly due to: 1) the training data used by the cloud pre-trained model is usually from public data sets or data sets sampled from specific scenarios, which is difficult to cover all video stream data distributions in edge scenarios; 2) the video content captured by the camera (such as the target category appearing, the facial features needing to be recognized, etc.) will change over time. In particular, due to parameter compression and structure simplification, lightweight DNN models are more sensitive to changes in data distribution, thus causing more significant accuracy decline. In order to avoid the accuracy decline of the model caused by data drift, model retraining based on continual learning is gradually becoming an important way to improve edge video analysis. By collecting data online and continuously updating model parameters, the model can adapt to the changing data distribution.
[0004] To avoid the precision decrease of the lightweight model caused by data drift, a high-precision model in the cloud is usually used to guide the lightweight model to retrain to improve its precision. However, due to the calling cost of the high-precision model and the limited resources of the edge device, retraining can usually be performed only on part of the key samples. However, it is not feasible to continuously run the high-precision model or manually monitor the performance of the lightweight model. Therefore, it is necessary to design an online mechanism to realize real-time perception of the degree to which the lightweight model is affected by data drift, and then timely trigger retraining and accurately select the key training samples that cause the precision decrease.
[0005] The goal of retraining is to improve the precision of the model on the video frames that originally perform poorly, while avoiding the precision decrease on other frames. With the help of the key frame screening mechanism, although part of the potential key frames can be screened from all the video frames, it is still necessary to determine whether retraining is needed according to the actual degree of precision decrease. At the same time, by identifying and screening the region of interest (RoI) that causes the precision decrease in the video frames, the granularity of retraining can also be further refined, thereby improving the resource utilization efficiency. In addition, unlike the traditional training model from scratch, retraining is usually performed on a small number of samples and has a certain periodicity. This feature makes it easy for the lightweight model to forget old knowledge when learning new knowledge (i.e., catastrophic forgetting), thereby causing redundant retraining.
[0006] To avoid the interruption of real-time video analysis services, the generated retraining task is executed in parallel with the video analysis task, and the lightweight model used for video analysis is updated after the retraining is completed. Due to the limited resources available at the edge, the occupation of resources by retraining may affect the real-time performance of video analysis. If the delay of video analysis exceeds the specified service level objective (SLO), the analysis result of the current frame may be skipped, thereby causing a significant decrease in precision. Therefore, how to reasonably configure retraining to improve the precision of the model while avoiding affecting the real-time performance of video analysis has become a problem to be solved. SUMMARY
[0007] Therefore, the purpose of the present application is to provide a resource-efficient continuous learning method and system for edge real-time video analysis, which has important positive significance for improving the SP revenue in the MEC environment.
[0008] To achieve the above object, the application adopts the following technical scheme: a resource-efficient continuous learning method for edge-oriented real-time video analysis, including real-time video analysis and model retraining; an online model accuracy decline predictor oriented to RoI granularity is designed; the foreground contour features in the video frame are used to fit the RoI complexity, and a lightweight regression model is used to accurately estimate the RoI granularity-oriented data drift degree, thereby improving the resource efficiency of model retraining; at the same time, the accuracy decline degree of the lightweight model compared with the high-precision model is continuously collected; then a double-layer mixed sample pool is designed, and based on the designed online model accuracy decline predictor oriented to RoI granularity, the RoI potentially affected by data drift is continuously collected to the buffer pool, and the accumulated model accuracy decline degree is observed; if the decline degree exceeds the specified threshold, the sample batch in the buffer pool is input to the high-precision model to determine the final sample that needs to be retrained; at the same time, a certain proportion of RoI historical samples are randomly introduced to alleviate the catastrophic forgetting problem; finally, a retraining progress controller based on deep reinforcement learning DRL is developed; based on the evaluation results of each round, the model convergence state is dynamically estimated to determine the retraining termination time; by introducing the behavior cloning BC loss function, the controller converges to the optimal strategy faster in the offline training stage.
[0009] In a preferred embodiment, the real-time video analysis specifically includes: first, the video stream collected by the camera in real time is input to the background subtractor to extract the motion area and generate a candidate RoI set; then, the RoI set is input to the lightweight model for online analysis to obtain the real-time analysis result; at the same time, the CL4VA performs feature extraction on each RoI and inputs it to the accuracy decline predictor to predict the accuracy decline degree of the real-time analysis result compared with the Oracle model; subsequently, the CL4VA transmits the RoI whose accuracy decline degree exceeds the set threshold and its online analysis result to the buffer pool for further analysis, and continuously monitors the cumulative accuracy decline of the RoI in the buffer pool; when the total decline degree exceeds the threshold, the model retraining is adaptively triggered.
[0010] In a preferred embodiment, model retraining specifically includes: after generating the retraining task, firstly, the samples in the buffer pool are batch-input into the Oracle model to obtain high-precision analysis results; then, the true error between the online analysis results and the high-precision results of the Oracle model is evaluated, and the samples in the buffer pool are proportionally updated to the retraining sample pool in different modes according to the error threshold; simultaneously, the true error is also fed back to the accuracy degradation predictor to dynamically update the predictor; subsequently, retraining is performed on the lightweight model using the retraining sample pool; next, CL4VA designs a DRL-based progress controller to determine the appropriate time for retraining completion by interacting with the retraining process; the training status and evaluation information after each round are fed back to the controller to construct the proximal policy optimization PPO loss; simultaneously, expert actions based on offline simulation are used to construct the behavior cloning loss, thereby jointly training the policy network of the DRL agent; the trained policy network is used to interact with the retraining process to output the optimal time for retraining completion in real time; after retraining is completed, the online model is updated in real time to continuously improve the accuracy of subsequent video analysis.
[0011] In a preferred embodiment, the online model accuracy degradation predictor for RoI granularity specifically includes: RoI feature extraction, predictor construction, and dynamic updating;
[0012] RoI feature extraction includes: CL4VA uses the foreground contour information generated during the background modeling process to divide the video frame into multiple RoIs, and extracts statistical information reflecting their structural features and spatial distribution from the RoIs to represent their complexity; these foreground contours are obtained by extracting the boundaries after the background subtractor generates foreground masks in consecutive video frames;
[0013] The predictor construction includes: CL4VA takes the aforementioned RoI features as input, uses the accuracy error of the lightweight model compared to the Oracle model as a label, and introduces a stochastic gradient descent (SGD) regressor to fit the relationship between the two; after CL4VA starts running, it collects the contour features of all RoIs during the first retraining cycle and simultaneously inputs all RoIs into the lightweight model and the Oracle model to obtain recognition results with different accuracies; then, CL4VA calculates the accuracy error between the detection results of the two models, which is defined as...
[0014] error = 1 - F1(M small M oracle ), (1)
[0015] Among them, M small and M oracle These represent the detection results of the lightweight model and the Oracle model on RoI, respectively.
[0016] Subsequently, the collected contour features are used as input, and errors are used as labels to build a prediction model using an SGD regressor. To estimate the prediction error of input feature x In subsequent retraining cycles, CL4VA uses the trained regressor to perform precision error prediction on each selected RoI to determine whether it is a complex RoI, without calling the Oracle model to calculate the precision error; if the regressor's output is greater than the precision error threshold, CL4VA will treat it as a potential key sample and transfer it to the buffer sample pool along with the analysis results for further processing.
[0017] Dynamic updates include: once the retraining task is completed, the parameters and performance of the lightweight model are updated accordingly; for the collected potential key samples, CL4VA continues to collect the features of these samples and their accuracy errors on both models, and uses this newly collected information to dynamically update the regressor; CL4VA sets different regressor parameters for different datasets and lightweight models.
[0018] In a preferred embodiment, the two-layer hybrid sample pool specifically includes: a buffer sample pool storing key RoI samples predicted by the accuracy degradation predictor and a retraining sample pool maintaining RoI samples validated by the Oracle model and used for actual retraining; CL4VA uses the accuracy degradation predictor to calculate the features of each extracted RoI and determine whether it is a potential key sample; if the predicted accuracy degradation exceeds a threshold δ, CL4VA adds it to the buffer sample pool; CL4VA introduces cumulative accuracy degradation to measure the degree to which the model is affected by data drift, thereby adaptively triggering the retraining task; specifically, the cumulative accuracy degradation is defined as...
[0019]
[0020] in, Let N represent the prediction error of the i-th RoI, and N represent the number of RoIs in the current buffer sample pool.
[0021] like Exceeding the threshold δ max If the current model shows significant performance degradation, CL4VA will immediately construct the final retraining sample pool based on the current buffer pool and start the retraining task.
[0022] In a preferred embodiment, sample screening and strategy distillation specifically include: CL4VA first inputs batches of RoIs from the buffer pool into an Oracle model to obtain high-precision analysis results; then, it calculates the true error (error) for each RoI based on the analysis results of the Oracle model, and screens the RoI samples based on a threshold τ as follows:
[0023] (1) If error > τ, it indicates that the current RoI is not performing well on the lightweight model. The sample will be identified as a key sample affecting accuracy and added to the retraining sample pool. During the retraining process, the sample label is the recognition result of the Oracle model.
[0024] (2) If error≤τ, RoI will be considered a non-critical sample and added to the retraining sample pool with a certain probability p; the label of the non-critical sample is the recognition result of the lightweight model; during the retraining process, retraining on non-critical samples is regarded as a self-distillation process; self-distillation allows the model to use previous parameters for knowledge transfer.
[0025] After filtering all samples, CL4VA empties the buffer sample pool to collect potential key samples for the next cycle, while the retraining sample pool retains historical samples in the form of a queue with a fixed maximum capacity. If the number of samples to be added exceeds the remaining available space in the retraining sample pool, CL4VA will replace the samples in the old sample pool with a 50% probability to avoid drastic changes in sample features.
[0026] In a preferred embodiment, the retraining progress controller based on deep reinforcement learning (DRL) dynamically decides whether to continue retraining based on real-time states such as loss changes after each retraining round, thereby fully exploring the optimal balance between retraining time and accuracy improvement; the state space, action space, and reward function are defined as follows:
[0027] State space: Each retraining round is treated as a decision step, and the performance metrics of the lightweight model after each training round are encoded as a state vector s. t , to serve as the observation input for the DRL agent; specifically, s t Includes the number of samples N in the current retraining sample pool and the number of epochs that have been completed. t The model's loss in the previous training round. t The accuracy of the model on the validation set is acc. t Error between current accuracy and previous accuracy Accuracy growth rate of historical rounds Information; therefore, the state space is formally represented as
[0028]
[0029] in, and Defined respectively
[0030]
[0031]
[0032] Action Space: In each round of retraining, the DRL agent dynamically chooses to continue or terminate the training process based on the current system state; therefore, the action space is formally represented as...
[0033] a t ∈{0,1},(6)where, a t =0 indicates that retraining will continue in the next round, a t =1 indicates that retraining will terminate in the current round;
[0034] Reward function: The reward function is defined as follows:
[0035]
[0036] Among them, T retrain The retraining time is represented by α, and β represents the reward weights for improving control accuracy and the retraining time overhead, respectively.
[0037] In a preferred embodiment, the retraining progress control based on guided DRL is implemented as follows: using an expert action set... For input, where Represents state s i The algorithm outputs the converged DRL policy network by acquiring the optimal action offline; in each round, it constructs the s by reading the model accuracy change information in the evaluation results. t It is then fed into the policy network to determine whether to stop retraining.
[0038] In a preferred embodiment, the DRL agent, an expert experience pool for storing expert actions, and a local experience pool for storing local training samples are first initialized; at each time step t, it is determined whether to introduce expert actions based on the dynamic exploration progress λ (epoch). This allows the policy network to strike a balance between exploration and exploitation, while simultaneously sampling actions from the policy network and feeding them back to the system. If a = 1, retraining will be terminated, and the final model performance and retraining time cost will be recorded to calculate the global discounted reward. If retraining continues, the reward r for the current time slot will be returned after the next round of evaluation. t And construct the next time slot state s t+1 Next, the expert and online training samples are stored in the expert and local experience pools respectively, and the DRL policy network is updated. A discount factor γ is introduced, and after retraining, the cumulative discount reward is calculated, defined as...
[0039]
[0040] R t This is used to estimate the state value function V; then, the advantage function is estimated. Defined as
[0041]
[0042] δ t =r t +γV φ (s t+1 )-V φ (s t ), (10)
[0043] Where η represents the discount rate, δ t Represents temporal difference;
[0044] Subsequently, samples are taken from the local experience pool to construct the PPO loss function, which is defined as follows:
[0045]
[0046] in, This represents the probability ratio between the old and new strategies, and ε is used to control the magnitude of the strategy update.
[0047] Next, sampling is performed from the expert experience pool, and training is aided by the BC loss function, which is defined as follows:
[0048]
[0049] Combining the PPO and BC loss functions, a joint loss function is constructed, which is defined as follows:
[0050] L = L PPO +λ(epoch)·L BC (13)
[0051] Among them, the dynamic exploration progress λ (epoch) is relatively large in the initial stage, so as to approach expert behavior more quickly; then it gradually decays.
[0052] Finally, the policy network of the DRL agent is updated based on the joint loss, guiding it to stop retraining when the performance improvement tends to saturate. After each retraining round, CL4VA inputs the state information into the DRL agent and decides whether to continue retraining based on the output until the termination condition is triggered.
[0053] This invention provides a resource-efficient continuous learning system for edge real-time video analytics, which runs the aforementioned resource-efficient continuous learning method for edge real-time video analytics.
[0054] Compared with existing technologies, this invention has the following advantages: This invention proposes a novel resource-efficient continuous learning framework (CL4VA) for real-time edge video analytics. CL4VA improves the accuracy and real-time performance of edge video analytics by designing components such as a background subtractor, a RoI-granularity accuracy degradation predictor, a two-layer mixed sample pool, and a guided DRL-based retraining progress controller. Evaluation results based on various real-world video datasets show that compared to Vanilla without continuous learning, CL4VA can improve accuracy by up to 43.69% and by an average of 29.58%. Compared to advanced Ekya and AdaInf, CL4VA can also reduce latency by an average of 8.65% and improve accuracy by up to 5.57%. Furthermore, this invention verifies the effectiveness of CL4VA in different scenarios, and its core components introduce only very low overhead. By effectively combining LSTM and TD3 algorithms, a novel CONS method is proposed to solve the computational offloading problem for 5G network slices in MEC environments, aiming to improve the benefits of SPs and reduce their required costs. Based on real user communication traffic datasets, this invention verifies the effectiveness of the proposed CONS method in improving SP revenue through extensive experiments. Experimental results show that, compared to five other benchmark methods, the CONS method exhibits superior performance under different resource rental ratios. Compared to the advanced DDPG and TD3 methods, the CONS method also demonstrates faster and more stable convergence. Simulation results indicate that this method is of great significance for improving SP revenue in MEC environments. Attached Figure Description
[0055] Figure 1 This is an overview diagram of the CL4VA of a preferred embodiment of the present invention;
[0056] Figure 2 This is a schematic diagram illustrating the accuracy variations of different methods in a video frame according to a preferred embodiment of the present invention;
[0057] Figure 3 This is a schematic diagram of the ablation experiment of the core component in the CL4VA of the preferred embodiment of the present invention;
[0058] Figure 4 This is a schematic diagram illustrating the impact of the maximum capacity of the retraining sample pool on accuracy and retraining time in a preferred embodiment of the present invention. Detailed Implementation
[0059] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0060] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0061] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0062] refer to Figures 1-4 This paper presents a resource-efficient continuous learning system for real-time video analytics at the edge, proposing an online model accuracy degradation predictor at the RoI granularity. It utilizes foreground contour features in video frames to fit RoI complexity and employs a lightweight regression model to accurately estimate the degree of data drift at the RoI granularity, thereby improving the resource efficiency of model retraining. Simultaneously, it continuously collects the accuracy degradation of the lightweight model compared to the high-precision model to avoid predictor accuracy degradation caused by performance improvements after retraining. A two-layer hybrid sample pool is then designed, and based on the designed predictor, potentially data drift-affected RoIs are continuously collected into the buffer pool, with the cumulative model accuracy degradation observed. Once the degradation exceeds a specified threshold, samples from the buffer pool are batch-input into the high-precision model to determine the samples ultimately requiring retraining. Furthermore, a certain proportion of historical RoI samples are randomly introduced to mitigate the catastrophic forgetting problem, thereby improving the model's stability and generalization ability during the update process. Finally, a retraining progress controller based on deep reinforcement learning (DRL) is developed. Based on the evaluation results of each round, the model's convergence state is dynamically estimated to determine the timing for terminating retraining, thereby minimizing the impact of retraining on the real-time performance of video analytics while ensuring the effectiveness of retraining. Simultaneously, by introducing the Behavioral Cloning (BC) loss function, the controller can converge to the optimal policy more quickly during offline training.
[0063] like Figure 1 As shown, the CL4VA proposed in this invention consists of two parts: real-time video analysis and model retraining.
[0064] Real-time video analytics. First, the video stream captured in real-time by the camera is fed into a background subtractor to extract motion regions and generate a set of candidate RoIs. Next, the RoI set is fed into a lightweight model for online analysis to obtain real-time results. Simultaneously, CL4VA performs feature extraction on each RoI and feeds it into an accuracy degradation predictor to predict the degree of accuracy degradation of the real-time analysis results compared to the Oracle model. Subsequently, CL4VA transfers RoIs with accuracy degradation exceeding a set threshold, along with their online analysis results, to a buffer pool for further analysis, continuously monitoring the cumulative accuracy degradation of RoIs in the buffer pool. When the total degradation exceeds the threshold, adaptive model retraining is triggered.
[0065] Model Retraining. After generating the retraining task, samples from the buffer pool are first batch-input into the Oracle model to obtain high-precision analysis results. Next, the true error between the online analysis results and the high-precision results from the Oracle model is evaluated, and samples from the buffer pool are proportionally updated to the retraining sample pool in different modes according to the error threshold. Simultaneously, this true error is also fed back to the accuracy degradation predictor to dynamically update the predictor. Subsequently, retraining is performed on the lightweight model using the retraining sample pool. Next, CL4VA designed a DRL-based progress controller that interacts with the retraining process to determine the appropriate time to complete the retraining. To optimize the policy network, the training state and evaluation information after each round are fed back to the controller to construct the Proximal Policy Optimization (PPO) loss. Simultaneously, expert actions based on offline simulations are used to construct the behavior cloning loss, thereby jointly training the policy network of the DRL agent. The trained policy network is used to interact with the retraining process to output the optimal time to complete the retraining in real time. After retraining is completed, the online model is updated in real time to continuously improve the accuracy of subsequent video analysis.
[0066] Specifically, online accuracy degradation predictors at the RoI granularity include:
[0067] RoI Feature Extraction: When analyzing videos in real time, CL4VA first assesses whether the currently deployed lightweight model is affected by data drift, resulting in a decrease in accuracy, and quantitatively estimates the degree of decrease. To this end, CL4VA needs to accurately identify the set of video frames where the lightweight model has lower detection accuracy while the Oracle model has higher detection accuracy, and then use this set to retrain the lightweight model. However, in real-world scenarios, it is not possible to continuously run the Oracle model or manually monitor the accuracy of the lightweight model in real time. Typically, the lightweight model performs poorly in video frame regions with high complexity (e.g., regions containing small targets or targets similar to the background), and the degree of accuracy decrease is positively correlated with the complexity of the video frame region. Therefore, this invention indirectly infers the potential degree of accuracy decrease by estimating the complexity of the regions where different targets are located in the video. To avoid introducing excessive additional overhead, CL4VA uses the foreground contour information generated during background modeling to divide the video frame into multiple RoIs, and extracts statistical information reflecting their structural features and spatial distribution from the RoIs to represent their complexity. These foreground contours are obtained by extracting boundaries after the background subtractor generates foreground masks in consecutive video frames. Specifically, this invention considers five typical RoI features, the specific meanings and functions of which are shown in Table 1.
[0068] Table 1. RoI Feature Description
[0069]
[0070]
[0071] It is worth noting that the calculation of the above features is based on existing mainstream background subtractors, such as the Gaussian Mixture Model (GMM) and K-Nearest Neighbors (KNN) background subtractors provided by OpenCV. Therefore, the system only needs to calculate features based on the extracted foreground contours, without having to directly extract features from the image data.
[0072] Predictor Construction: To predict the degree of accuracy degradation at the RoI granularity, CL4VA uses the aforementioned RoI features as input and the accuracy error of the lightweight model compared to the Oracle model as a label. A Stochastic Gradient Descent (SGD) regressor is introduced to fit the relationship between the two. The predictor aims to accurately identify samples requiring retraining by using RoI features to predict the degree of accuracy degradation of the lightweight model. To better meet the real-time requirements of video analysis, the designed predictor adopts an online training method, without relying on pre-labeled datasets or offline training, thus possessing good practicality and deployment flexibility. Specifically, after CL4VA starts running, it collects the contour features of all RoIs in the first retraining cycle and synchronously inputs all RoIs into the lightweight model and the Oracle model to obtain recognition results with different accuracies. Then, CL4VA calculates the accuracy error between the detection results of the two models, which is defined as...
[0073] error = 1 - F1(M small M oracle ), (1)
[0074] Among them, M small and M oracle The results represent the detection results of the lightweight model and the Oracle model on RoI, respectively.
[0075] Subsequently, the collected contour features are used as input, and errors are used as labels to build a prediction model using an SGD regressor. To estimate the prediction error of input feature x In subsequent retraining cycles, CL4VA uses the trained regressor to perform precision error prediction on each selected RoI to determine whether it is a complex RoI, without calling the Oracle model to calculate the precision error. If the regressor's output is greater than the precision error threshold, CL4VA considers it a potentially critical sample and transfers it along with the analysis results to a buffer sample pool for further processing.
[0076] Dynamic updates: After the retraining task is completed, the parameters and performance of the lightweight model are updated accordingly. At this point, for the same RoI, the accuracy error of the lightweight model compared to the Oracle model will change (usually decrease). Therefore, to maintain good prediction accuracy, for the collected potential key samples, CL4VA continues to collect the features of these samples and their accuracy errors on both models, and uses this newly collected information to dynamically update the regressor. To achieve higher prediction accuracy, CL4VA sets different regressor parameters for different datasets and lightweight models, and the time cost of each update and prediction is very small (e.g., for 500 samples, an update can typically be completed within 17ms).
[0077] Specifically, the two-layer hybrid sample pool includes:
[0078] To improve the efficiency of key sample selection and mitigate the impact of catastrophic forgetting, CL4VA constructs a two-layer hybrid sample pool, designed to enhance model accuracy and stability while ensuring real-time video analysis. Specifically, the two-layer hybrid sample pool consists of two parts: a buffer sample pool stores key RoI samples predicted by the accuracy degradation predictor; and a retraining sample pool maintains RoI samples validated by the Oracle model and used for actual retraining.
[0079] Adaptive retraining is triggered. CL4VA uses a precision degradation predictor to compute features for each extracted RoI and determine whether it is a potential key sample. If the predicted precision degradation exceeds a threshold δ, CL4VA adds it to the buffer sample pool. To better balance the gain from retraining in terms of accuracy improvement with its resource overhead, CL4VA introduces cumulative precision degradation to measure the degree to which the model is affected by data drift, thus adaptively triggering the retraining task. Specifically, the cumulative precision degradation is defined as...
[0080]
[0081] in, Let represent the prediction error of the i-th RoI, and N represent the number of RoIs in the current buffer sample pool.
[0082] like Exceeding the threshold δ max If the current model shows significant performance degradation, CL4VA will immediately construct the final retraining sample pool based on the current buffer pool and start the retraining task.
[0083] Sample selection and strategy distillation. To accurately measure the actual accuracy degradation of each sample, CL4VA first inputs batches of RoIs from the buffer pool into an Oracle model to obtain high-precision analysis results. Then, based on the analysis results from the Oracle model, the true error (error) for each RoI is calculated, and the RoI samples are selected based on a threshold τ as follows:
[0084] (1) If error > τ, it indicates that the current RoI is not performing well on the lightweight model. In this case, the sample will be identified as a key sample affecting accuracy and added to the retraining sample pool. During the retraining process, the sample label is the recognition result of the Oracle model. Therefore, the retraining on these samples can be regarded as a local knowledge distillation process guided by the Oracle model, which aims to enable the lightweight model to learn the fitting ability of the Oracle in complex scenarios.
[0085] (2) If error ≤ τ, RoIs will be considered non-critical samples and added to the retraining sample pool with a certain probability p (e.g., 20%). The labels of non-critical samples are the recognition results of the lightweight model. Therefore, retraining on non-critical samples during retraining can be regarded as a self-distillation process. Self-distillation allows the model to utilize previous parameters for knowledge transfer, which helps to retain the memory of the original knowledge while learning knowledge from other models, thereby effectively suppressing catastrophic forgetting.
[0086] After filtering all samples, CL4VA clears the buffer sample pool to collect potential key samples for the next cycle, while the retraining sample pool retains historical samples in a queue with a fixed maximum capacity. If the number of samples to be added exceeds the remaining available space in the retraining sample pool, CL4VA will replace samples in the old sample pool with a 50% probability to avoid drastic changes in sample features. Through this design, CL4VA enables the lightweight model to efficiently learn Oracle model knowledge while avoiding forgetting old knowledge.
[0087] Specifically, DRL-based retraining progress control includes:
[0088] Due to the limited resources of edge systems, the additional overhead introduced by retraining must be carefully considered to avoid impacting the real-time performance of video analytics. Therefore, when generating and running retraining tasks, CL4VA determines the appropriate time to complete training to balance accuracy gain and resource overhead. Specifically, CL4VA designs a DRL-based progress controller that dynamically decides whether to continue retraining based on real-time states such as loss changes after each round of retraining, thereby fully exploring the optimal balance between retraining time and accuracy improvement.
[0089] Unlike training a model from scratch, retrained models tend to converge and stabilize quickly. However, during the retraining of lightweight models, the accuracy improvement typically decreases gradually, exhibiting a non-linear or even saturating trend. Therefore, identifying and terminating unnecessary retraining can effectively reduce its impact on the real-time performance of video analysis. Furthermore, completing retraining and updating the lightweight model at the appropriate time also helps improve the accuracy of subsequent video analysis, thereby effectively enhancing the overall system gain. Unlike existing works that use prior selection or offline analysis to pre-set retraining rounds, this invention models retraining as a Markov Decision Process (MDP) and designs a DRL-based retraining progress controller to determine the optimal time to terminate retraining, thus maximizing model accuracy while reducing the retraining time. Accordingly, the state space, action space, and reward function are defined as follows:
[0090] State space. Each retraining round is considered a decision step, and the performance metrics of the lightweight model after each training round are encoded as a state vector s. t , to serve as the observation input for the DRL agent. Specifically, s t Includes the number of samples N in the current retraining sample pool and the number of epochs that have been completed. t The model's loss in the previous training round. t The accuracy of the model on the validation set is acc. t Error between current accuracy and previous accuracy Accuracy growth rate of historical rounds Information such as these. Therefore, the state space can be formally represented as
[0091]
[0092] in, and Defined respectively
[0093]
[0094]
[0095] Action Space. In each round of retraining, the DRL agent dynamically chooses to continue or terminate the training process based on the current system state. Therefore, the action space can be formally represented as...
[0096] a t ∈{0,1},(6)where, a t =0 indicates that retraining will continue in the next round, a t =1 indicates that retraining will be terminated in the current round.
[0097] Reward function. To guide the DRL agent to maximize the gains from retraining, the reward function is defined as follows:
[0098]
[0099] Among them, T retrain The retraining time is represented by α, and β represents the reward weights for improving control accuracy and the retraining time overhead, respectively.
[0100] CL4VA utilizes the advanced PPO algorithm to train and optimize DRL agents. By introducing a policy update method based on shearing probability ratios, the PPO algorithm ensures stable policy updates while maintaining sample utilization efficiency. During the offline training phase, CL4VA sets a sufficient number of retraining rounds and collects the accuracy changes and time costs of each round to explore the optimal retraining progress control strategy. In particular, to improve policy performance and exploration efficiency, CL4VA constructs expert experience by simulating different training processes and recording the optimal termination timing, thereby leveraging imitation learning.
[25] The idea is to guide DRL agents to achieve fast convergence.
[0101] The key steps of the proposed retraining progress control based on guided DRL are shown in Algorithm 1. Algorithm 1 uses an expert action set... For input, where Represents state s i The algorithm outputs the converged DRL policy network by acquiring the optimal action offline. In each round, the algorithm constructs the s model by reading the model accuracy change information from the evaluation results. t The algorithm first initializes the DRL agent, an expert experience pool for storing expert actions, and a local experience pool for storing local training samples (lines 1-3). At each time step t, it determines whether to introduce expert actions based on the dynamic exploration progress λ (epoch). This aims to achieve a balance between exploration and exploitation in the policy network (lines 4-7), while sampling actions from the policy network and feeding them back to the system. If a=1, retraining will be terminated, and the final model performance and retraining time cost will be recorded to calculate the global discounted reward (lines 9-11). If retraining continues, the reward r of the current time slot will be returned after the next round of evaluation. t And construct the next time slot state s t+1(Lines 12-14). Next, the expert and online training samples are stored in the expert and local experience pools respectively (lines 15-16), and the DRL policy network is updated (lines 18-19). Specifically, to evaluate the cumulative effect of the reward function over time, this invention introduces a discount factor γ. After retraining, the cumulative discounted reward is calculated, defined as...
[0102]
[0103] R t This is used to estimate the state value function V. Next, the advantage function is estimated. Defined as
[0104]
[0105] δ t =r t +γV φ (s t+1 )-V φ (s t ), (10)
[0106] Where η represents the discount rate, δ t This represents the time difference.
[0107] Subsequently, samples are taken from the local experience pool to construct the PPO loss function, which is defined as follows:
[0108]
[0109] in, This represents the probability ratio between the old and new strategies, and ε is used to control the magnitude of the strategy update.
[0110] Next, sampling is performed from the expert experience pool, and training is aided by the BC loss function, which is defined as follows:
[0111]
[0112] Furthermore, by combining the PPO and BC loss functions, a joint loss function is constructed, which is defined as follows:
[0113] L = L PPO +λ(epoch)·L BC (13)
[0114] The dynamic exploration progress λ (epoch) is relatively large in the initial stage to more quickly approach expert behavior. It then gradually decays to encourage the policy to explore autonomously and learn better policies from non-expert samples.
[0115] Finally, the policy network of the DRL agent is updated based on the joint loss, guiding it to stop retraining when the performance improvement approaches saturation. After each retraining round, CL4VA inputs the state information into the DRL agent and decides whether to continue retraining based on the output, until the termination condition is triggered.
[0116]
[0117]
[0118] Method evaluation:
[0119] Datasets. This invention uses the following three real-world video datasets to comprehensively evaluate the performance of the proposed CL4VA:
[0120] (1) MOT: This dataset contains labeled data for large-scale multi-target tracking, consisting of 22 different indoor and outdoor scenes. The video data was recorded under different camera motion, angle, and imaging conditions for target tracking.
[0121] (2) UA-DETRAC: This dataset contains large-scale multi-target vehicle detection data, consisting of four types of vehicles (i.e., cars, buses, trucks and others) and four types of weather conditions (i.e., night, cloudy, sunny and rainy).
[0122] (3) Real-Camera: To further verify the performance of CL4VA in real-world video analysis scenarios, this invention collected videos recorded by real cameras deployed in a laboratory environment. Unlike the MOT and UA-DETRAC datasets, the target density in the collected videos is sparser. To mitigate interference from targetless video frames, video segments were collected during three time periods: 11:00-11:30 AM, 1:30-2:00 PM, and 5:00-5:30 PM. Furthermore, YOLOv8X was used to pre-detect the collected videos, filtering out targetless video frames to ensure that all frames contained at least one target. Simultaneously, the detection results of YOLOv8X in each frame were used as labels to evaluate the accuracy of different methods.
[0123] CL4VA Implementation. Based on a workstation equipped with an AMD Ryzen 99950x CPU, 64GB RAM, and an RTX 4090D, this invention uses Python and PyTorch to build a video analysis simulation system and implement the components in the proposed CL4VA. Video frames from the dataset are read using OpenCV, and a GMM-based background subtractor is created using `cv2.createBackgroundSubtractorMOG2()` to extract RoIs from the video frames. In the simulation experiments, the history length of the background subtractor is set to 200, and the threshold is set to 16. It maintains an estimate of the background image and uses subtraction to extract the foreground region, then performs blob detection to extract RoIs. An SGD regressor is implemented using the sklearn library, and the parameters are dynamically updated using the `partial_fit()` function. For the SGD regressor, the maximum number of iterations is set to 1000, and the initial learning rate is 0.01. YOLOv8S is used as a lightweight model to perform online video analysis, and YOLOv8X is used as an Oracle model to obtain high-precision analysis results. The precision drop threshold for identifying key samples was set to 0.2, the cumulative precision drop threshold for triggering retraining was set to 200, and the maximum capacity of the retraining sample pool was 1500. mAP (mean Average Precision) was used as the validation metric, which was obtained by calculating the area under the precision-recall curves for each category and averaging them.
[0124] Comparison Methods. To verify the superiority of the proposed CL4VA, this invention comprehensively compares it with the following benchmark methods:
[0125] (1) Vanilla: The lightweight YOLOv8s model is used for real-time video analysis without model retraining.
[0126] (2) Oracle: The high-precision model YOLOv8X was used for real-time video analysis to evaluate the difference between the retraining effect and its performance.
[0127] (3) Ekya: Constructs a configuration analyzer by performing several retraining sessions offline on a small number of samples. This configuration analyzer can be used to estimate the impact of training on all samples on the final accuracy and to select a training schedule that achieves the best balance between accuracy and overhead based on device resources before actual retraining.
[0128] (4) AdaInf: The cosine distance between the feature vector of YOLOv8s on each video frame and the average feature vector on the training set samples is calculated, and the video frame with the largest distance is selected as the key frame. The retraining progress is determined based on the key frame and the configuration analyzer in Ekya.
[0129] This invention compares the overall performance of different methods on different datasets. As shown in Table 2, all methods achieved the lowest and highest accuracy on the UA-DETRAC and Real-Camera datasets, respectively. This is because the UA-DETRAC dataset contains more densely packed targets in its video scenes compared to the other two datasets (especially at intersections, where dozens of vehicles may be present simultaneously). Furthermore, due to vehicle movement, it also contains more small targets. In contrast, most scenes in the MOT dataset are easier to identify. The Real-Camera dataset primarily contains single, relatively large targets, thus all methods generally achieve higher accuracy on this dataset.
[0130] Specifically, Vanilla achieves the lowest average latency across different datasets by using only a lightweight model for real-time video analysis without retraining. However, due to data distribution skew between the pre-trained lightweight model used by Vanilla and different datasets, its accuracy remains at the lowest level, making it difficult to guarantee the accuracy of video analysis. In contrast, Oracle represents the upper limit of detection accuracy. It uses YOLOv8X to analyze all video frames. Due to its more complex network structure and more model parameters, Oracle's average latency is significantly higher than other methods, making it difficult to guarantee the real-time performance of video analysis.
[0131] To achieve a better balance between accuracy and latency, Ekya, AdaInf, and the proposed CL4VA all introduce retraining mechanisms to dynamically adapt to environmental changes, thereby improving the performance of lightweight models. Ekya improves model accuracy to some extent through periodic retraining. Simultaneously, Ekya selects an appropriate number of training epochs based on device resources during the initialization phase to maintain a relatively low training load throughout the online phase. However, due to the lack of targeted sample selection, Ekya's accuracy improvement is limited.
[0132] AdaInf further introduces a key sample mining mechanism based on feature distance, building upon Ekya, which enhances the training efficiency for key samples. Therefore, its accuracy surpasses Ekya across multiple datasets. However, this key sample selection mechanism requires continuous calculation of the feature distance between online video frames and the retraining dataset, failing to fully capture data drift and incurring additional overhead. Consequently, its latency is higher than all other methods except Oracle.
[0133] Compared to other methods, the proposed CL4VA achieves higher detection accuracy than Ekya and AdaInf on multiple datasets, while also achieving lower latency, thus better balancing latency and accuracy. Notably, on multiple datasets, the average difference between CL4VA's accuracy and Oracle is less than 2.43%, while inference latency is reduced by more than 27.67%. This is because CL4VA's RoI-granularity accuracy degradation predictor accurately identifies RoIs where lightweight models perform poorly during video analysis, thereby improving the efficiency of utilizing key samples. Simultaneously, CL4VA enhances training diversity by mixing hard and random samples, effectively mitigating overfitting and catastrophic forgetting problems and overcoming the limitations of AdaInf and Ekya in sample construction strategies. Furthermore, CL4VA utilizes a DRL-based retraining progress controller to dynamically adjust the number of training epochs based on the model's convergence state, thereby minimizing the impact of retraining tasks on the real-time performance of video analysis.
[0134] Table 2 Comparison of overall performance of different methods on different datasets
[0135]
[0136]
[0137] Next, this invention compares the accuracy variations of different methods on video frame sequences. For example... Figure 2As shown, compared to Oracle, other methods exhibit lower accuracy in the early stages of the frame sequence. This is because these methods all employ lightweight models to ensure real-time video analysis. As the video analysis progresses, the accuracy of the three retraining methods (Ekya, AdaInf, and the proposed CL4VA) gradually approaches that of Oracle, while Vanilla remains at a lower accuracy. Specifically, Ekya and AdaInf trigger retraining at fixed time slots, and their accuracy increases at similar times with significant fluctuations. This is because they do not consider the model's original knowledge during retraining, leading to catastrophic forgetting. Compared to Ekya, AdaInf improves accuracy through a key sample selection mechanism based on feature distance. However, AdaInf only retrains on specified samples, causing more significant performance fluctuations. Unlike other methods, CL4VA designs a more precise key sample selection strategy and considers both random and historical samples to more efficiently learn from Oracle model knowledge while avoiding forgetting old knowledge, thus achieving higher accuracy and better stability. Notably, in the later stages of retraining, CL4VA's accuracy even surpasses that of Oracle. This is because in some video scenarios, the lightweight model may outperform Oracle in accuracy. In such cases, the accuracy degradation predictor treats these samples as non-critical and uses the analysis results of the lightweight model as training labels. Through this design, CL4VA effectively integrates the knowledge of both models in their high-accuracy scenarios, thereby maintaining good accuracy performance in the deployed edge video scenarios.
[0138] Next, to quantify the contribution of each core component to CL4VA performance, this invention removes the following components from CL4VA and observes their impact on video analysis accuracy and latency:
[0139] / AD (Accuracy Drop prediction): Removes feature predictions and randomly selects training samples.
[0140] / MS (Mixed Samples): Removes mixed samples and retrains using only the key samples selected from feature predictions.
[0141] / PC (Progress Control): Removes retraining progress control, and the number of retraining rounds is fixed at 20.
[0142] like Figure 3As shown, removing AD significantly impacts video analysis accuracy. This is because CL4VA cannot accurately identify RoI features and focus training on key samples, thus weakening the lightweight model's ability to identify complex video segments and regions. Similarly, removing MS also leads to a decrease in accuracy. This is because retraining only on key samples can easily cause the lightweight model to overfit, resulting in decreased accuracy in other non-key video scenes. Furthermore, when PC is removed from CL4VA, although accuracy only slightly decreases (0.23%), the average retraining time increases significantly. This is because while more training epochs can ensure model accuracy, they also affect the real-time performance of video analysis, causing some video frames to be unable to be analyzed in a timely manner. Completing retraining and updating the lightweight model faster can also improve the performance of subsequent video analysis more promptly. Therefore, compared to a fixed number of training epochs, the progress controller designed by CL4VA can further improve accuracy.
[0143] Next, this invention evaluated the additional time overhead introduced by the core components of CL4VA. As shown in Table 3, the time required to extract RoIs using the GMM-based background subtractor is 8.84 ms, indicating that it can quickly locate key regions from the original video frames for subsequent processing. Based on this, the time overhead for calculating RoI features is only 0.42 ms, which can be used to extract the RoI complexity information described in the foreground contours. For these feature inputs, the accuracy degradation prediction based on the SGD regressor can be quickly judged (only 0.20 ms), efficiently identifying complex RoIs. Finally, the DRL-based retraining progress controller can quickly determine the termination time of retraining based on historical information and the current environmental state, with a time overhead of only 1.50 ms. Therefore, CL4VA, composed of the above components, does not impose a significant burden on the video analysis system.
[0144] Table 3 Time overhead of core components in CL4VA
[0145]
[0146] Finally, this invention evaluates the impact of the maximum capacity of the retraining sample pool on the performance of different methods. For example... Figure 4As shown, the accuracy of all methods generally increases as the retraining sample pool size increases from 500 to 2500. This is because richer training samples help alleviate data drift, thereby improving the model's ability to fit complex samples. The proposed CL4VA exhibits the highest accuracy and the least retraining time at different sample pool sizes. Specifically, when the sample pool size is small, CL4VA shows a more significant accuracy improvement compared to AdaInf and Ekya. As the sample pool size increases, the accuracy gap between different methods also decreases. This is because when the available retraining samples are limited, the accuracy degradation predictor designed by CL4VA can more effectively select key samples that have a more significant impact on the accuracy of the lightweight model, thereby maximizing the improvement of video analysis accuracy. As the sample pool size increases, the knowledge learned by all methods from the Oracle model also increases, so the accuracy gap between the three gradually decreases. In addition, the retraining time of AdaInf and Ekya increases significantly with the increase of sample size. This is because they both have pre-defined training configurations, and the training cost increases with the increase of sample size. In contrast, CL4VA consistently maintains the lowest retraining time across different sample sizes. As the sample pool size increases, CL4VA's retraining time only rises slightly, demonstrating superior resource adaptability. This is because CL4VA's retraining progress controller effectively identifies and prematurely terminates redundant training, reducing its impact on the real-time performance and accuracy of video analytics.
[0147] The process or method of using this invention:
[0148] (1) CL4VA first performs foreground modeling on continuous video frames through the background modeling module, extracts the foreground contour and divides multiple regions of interest (RoI), and then extracts structural and spatial distribution related statistical features for each RoI, such as area, aspect ratio, convexity, and position, to characterize its complexity.
[0149] (2) The system uses a pre-trained accuracy degradation predictor to perform regression analysis on the complexity features of each RoI to predict the degree of accuracy degradation under the lightweight model. CL4VA automatically selects key samples (i.e. RoIs with significant accuracy degradation) from the video stream based on the prediction error value for subsequent retraining.
[0150] (3) Key samples enter the sample buffer. After the buffer reaches the set capacity, a mixed sample pool containing historical key samples and newly selected samples is constructed. If a simplified strategy is adopted, the construction of the mixed sample pool can be skipped, and key samples can be used directly for training.
[0151] (4) CL4VA has a built-in retraining scheduler based on deep reinforcement learning. It determines whether to continue retraining the model based on state information such as historical accuracy fluctuations, key sample distribution and retraining overhead, and dynamically judges when to end the training in order to maximize resource efficiency.
[0152] (5) Once retraining is complete, the system updates the lightweight model parameters and redeploys it to the edge to perform inference tasks, continuing to provide real-time video analytics services. Throughout the process, the system records the key states, accuracy changes, and resource overhead of each training round to optimize future strategies.
Claims
1. A resource-efficient and continuous learning method for real-time video analytics at the edge, characterized in that, This includes real-time video analysis and model retraining; designing an online model accuracy degradation predictor for RoI granularity; utilizing foreground contour features in video frames to fit RoI complexity, and using a lightweight regression model to accurately estimate the degree of data drift for RoI granularity, thereby improving the resource efficiency of model retraining; and continuously collecting the accuracy degradation of the lightweight model compared to the high-precision model. Then, a two-layer hybrid sample pool is designed, and based on the designed online model accuracy degradation predictor oriented towards RoI granularity, RoIs potentially affected by data drift are continuously collected into the buffer pool, and the cumulative model accuracy degradation is observed. If the degradation exceeds a specified threshold, samples in the buffer pool are batch-input into the high-precision model to determine the samples that finally need to be retrained. At the same time, a certain proportion of historical RoI samples are randomly introduced to alleviate the catastrophic forgetting problem. Finally, a retraining progress controller based on deep reinforcement learning (DRL) is developed. Based on the evaluation results of each round, the model convergence state is dynamically estimated to determine the retraining termination time. By introducing the Behavioral Cloning (BC) loss function, the controller converges to the optimal policy faster during the offline training phase.
2. The resource-efficient continuous learning method for real-time video analytics at the edge as described in claim 1, characterized in that, Real-time video analytics specifically includes: First, the video stream captured in real time by the camera is input into a background subtractor to extract motion regions and generate a candidate RoI set; then, the RoI set is input into a lightweight model for online analysis to obtain real-time analysis results; simultaneously, CL4VA performs feature extraction on each RoI and inputs it into an accuracy degradation predictor to predict the degree of accuracy degradation of the real-time analysis results compared to the Oracle model; subsequently, CL4VA transmits RoIs whose accuracy degradation exceeds a set threshold and their online analysis results to a buffer pool for further analysis, and continuously monitors the cumulative accuracy degradation of RoIs in the buffer pool; when the total degradation exceeds the threshold, adaptive model retraining is triggered.
3. The resource-efficient continuous learning method for edge real-time video analysis according to claim 1, characterized in that, The model retraining process specifically includes: after generating the retraining task, the samples in the buffer pool are first batch-input into the Oracle model to obtain high-precision analysis results; then, the true error between the online analysis results and the high-precision results of the Oracle model is evaluated, and the samples in the buffer pool are proportionally updated to the retraining sample pool in different modes according to the error threshold; simultaneously, this true error is also fed back to the accuracy degradation predictor to dynamically update the predictor; subsequently, retraining is performed on the lightweight model using the retraining sample pool; next, CL4VA designs a DRL-based progress controller, which interacts with the retraining process to determine the appropriate time for retraining to complete; the training status and evaluation information after each round are fed back to the controller to construct a proximal policy to optimize the PPO loss; at the same time, expert actions based on offline simulation are used to construct the behavior cloning loss, thereby jointly training the policy network of the DRL agent; the trained policy network is used to interact with the retraining process to output the optimal time for retraining to complete in real time; after retraining is completed, the online model will be updated in real time to continuously improve the accuracy of subsequent video analysis.
4. The resource-efficient continuous learning method for edge real-time video analysis according to claim 1, characterized in that, The online model accuracy degradation predictor for RoI granularity specifically includes: RoI feature extraction, predictor construction, and dynamic updating; RoI feature extraction includes: CL4VA uses the foreground contour information generated during the background modeling process to divide the video frame into multiple RoIs, and extracts statistical information reflecting their structural features and spatial distribution from the RoIs to represent their complexity; these foreground contours are obtained by extracting the boundaries after the background subtractor generates foreground masks in consecutive video frames; The predictor construction includes: CL4VA takes the aforementioned RoI features as input, uses the accuracy error of the lightweight model compared to the Oracle model as a label, and introduces a stochastic gradient descent (SGD) regressor to fit the relationship between the two; after CL4VA starts running, it collects the contour features of all RoIs during the first retraining cycle and simultaneously inputs all RoIs into the lightweight model and the Oracle model to obtain recognition results with different accuracies; then, CL4VA calculates the accuracy error between the detection results of the two models, which is defined as... error=1-F1(M small ,M oracle ), (1) Among them, M small and M oracle These represent the detection results of the lightweight model and the Oracle model on RoI, respectively. Subsequently, the collected contour features are used as input, and errors are used as labels to build a prediction model using an SGD regressor. To estimate the prediction error of input feature x In subsequent retraining cycles, CL4VA uses the trained regressor to perform precision error prediction on each selected RoI to determine whether it is a complex RoI, without calling the Oracle model to calculate the precision error; if the regressor's output is greater than the precision error threshold, CL4VA will treat it as a potential key sample and transfer it to the buffer sample pool along with the analysis results for further processing. Dynamic updates include: once the retraining task is completed, the parameters and performance of the lightweight model are updated accordingly; for the collected potential key samples, CL4VA continues to collect the features of these samples and their accuracy errors on both models, and uses this newly collected information to dynamically update the regressor; CL4VA sets different regressor parameters for different datasets and lightweight models.
5. The resource-efficient continuous learning method for edge real-time video analysis according to claim 1, characterized in that, The two-layer hybrid sample pool specifically includes: a buffer sample pool storing key RoI samples predicted by the accuracy degradation predictor, and a retraining sample pool maintaining RoI samples validated by the Oracle model and used for actual retraining; CL4VA uses the accuracy degradation predictor to calculate the features of each extracted RoI and determine whether it is a potential key sample; if the predicted accuracy degradation exceeds a threshold δ, CL4VA adds it to the buffer sample pool; CL4VA introduces cumulative accuracy degradation to measure the degree to which the model is affected by data drift, thereby adaptively triggering the retraining task; specifically, the cumulative accuracy degradation is defined as... in, Let N represent the prediction error of the i-th RoI, and N represent the number of RoIs in the current buffer sample pool. like Exceeding the threshold δ max If the current model shows significant performance degradation, CL4VA will immediately construct the final retraining sample pool based on the current buffer pool and start the retraining task.
6. The resource-efficient continuous learning method for edge real-time video analysis according to claim 5, characterized in that, The sample selection and strategy distillation process specifically includes: CL4VA first inputs batches of RoIs from the buffer pool into the Oracle model to obtain high-precision analysis results; then, based on the analysis results of the Oracle model, it calculates the true error (error) for each RoI, and performs the following selection on the RoI samples based on a threshold τ: (1) If error > τ, it indicates that the current RoI is not performing well on the lightweight model. The sample will be identified as a key sample affecting accuracy and added to the retraining sample pool. During the retraining process, the sample label is the recognition result of the Oracle model. (2) If error≤τ, RoI will be considered a non-critical sample and added to the retraining sample pool with a certain probability p; the label of the non-critical sample is the recognition result of the lightweight model; during the retraining process, retraining on non-critical samples is regarded as a self-distillation process; self-distillation allows the model to use previous parameters for knowledge transfer. After filtering all samples, CL4VA empties the buffer sample pool to collect potential key samples for the next cycle, while the retraining sample pool retains historical samples in the form of a queue with a fixed maximum capacity. If the number of samples to be added exceeds the remaining available space in the retraining sample pool, CL4VA will replace the samples in the old sample pool with a 50% probability to avoid drastic changes in sample features.
7. The resource-efficient continuous learning method for edge real-time video analysis according to claim 1, characterized in that, The retraining progress controller based on deep reinforcement learning (DRL) dynamically decides whether to continue retraining based on real-time states such as loss changes after each retraining round, thereby fully exploring the optimal balance between retraining time and accuracy improvement; the state space, action space, and reward function are defined as follows: State space: Each retraining round is treated as a decision step, and the performance metrics of the lightweight model after each training round are encoded as a state vector s. t , to serve as the observation input for the DRL agent; specifically, s t Includes the number of samples N in the current retraining sample pool and the number of epochs that have been completed. t The model's loss in the previous training round. t The accuracy of the model on the validation set is acc. t Error between current accuracy and previous accuracy Accuracy growth rate of historical rounds Information; therefore, the state space is formally represented as in, and Defined respectively Action Space: In each round of retraining, the DRL agent dynamically chooses to continue or terminate the training process based on the current system state; therefore, the action space is formally represented as a t ∈{0,1},(6)where, a t =0 indicates that retraining will continue in the next round, a t =1 indicates that retraining will terminate in the current round; Reward function: The reward function is defined as follows: Among them, T retrain The retraining time is represented by α, and β represents the reward weights for improving control accuracy and the retraining time overhead, respectively.
8. The resource-efficient continuous learning method for edge real-time video analysis according to claim 7, characterized in that, The implementation method of retraining progress control based on guided DRL is as follows: using expert action set For input, where Represents state s i The algorithm outputs the converged DRL policy network by acquiring the optimal action offline; in each round, it constructs the s by reading the model accuracy change information in the evaluation results. t It is then fed into the policy network to determine whether to stop retraining.
9. The resource-efficient continuous learning method for edge real-time video analysis according to claim 8, characterized in that, First, initialize the DRL agent, the expert experience pool for storing expert actions, and the local experience pool for storing local training samples; at each time step t, determine whether to introduce expert actions based on the dynamic exploration progress λ (epoch). This allows the policy network to achieve a balance between exploration and mining, while sampling actions from the policy network and feeding them back to the system; if a=1, retraining will be terminated, and the final model performance and retraining time cost will be recorded to calculate the global discount reward. If retraining continues, the reward r for the current time slot will be returned after the next round of evaluation. t And construct the next time slot state s t+1 Next, the expert and online training samples are stored in the expert and local experience pools respectively, and the DRL policy network is updated. A discount factor γ is introduced, and after retraining, the cumulative discount reward is calculated, defined as... R t This is used to estimate the state value function V; then, the advantage function is estimated. Defined as Where η represents the discount rate, δ t Represents temporal difference; Subsequently, samples are taken from the local experience pool to construct the PPO loss function, which is defined as follows: in, This represents the probability ratio between the old and new strategies, and ε is used to control the magnitude of the strategy update. Next, sampling is performed from the expert experience pool, and training is aided by the BC loss function, which is defined as follows: Combining the PPO and BC loss functions, a joint loss function is constructed, which is defined as follows: L=L PPO +λ(epoch)·L BC , (13) Among them, the dynamic exploration progress λ (epoch) is relatively large in the initial stage, so as to approach expert behavior more quickly; then it gradually decays. Finally, the policy network of the DRL agent is updated based on the joint loss, guiding it to stop retraining when the performance improvement tends to saturate. After each retraining round, CL4VA inputs the state information into the DRL agent and decides whether to continue retraining based on the output until the termination condition is triggered.
10. A resource-efficient continuous learning system for real-time edge video analytics, characterized in that, Run the resource-efficient continuous learning method for edge-oriented real-time video analytics as described in any one of claims 1-9.
Citation Information
Cited By
Intelligent agent continuous training and effect evaluation closed-loop method and system fusing work order feedback, and medium
CN121835735A