An online retraining method for multi-terminal video detection models

By using the online retraining method of multi-terminal video detection model of video detection model, data annotation, sample enhancement, mutual learning and FedAVG algorithm and other technologies, the problem of degradation of model accuracy and slow retraining speed caused by data drift is solved, and efficient model updates and hardware acceleration are achieved.

CN115587217BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211268163.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-06-27
Estimated Expiration
2042-10-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the data drift problem in video analysis, resulting in a decrease in the accuracy of shallow and lightweight models in real scenarios, and the online retraining speed of the model is slow, affecting system performance.

Method used

The online retraining method of multi-terminal video detection model is adopted. By performing data annotation, sample enhancement, mutual learning and FedAVG algorithm aggregation on edge servers, parallel updates of local models and intermediate models are realized, and optimizations are performed on digital accelerators and analog accelerators.

Benefits of technology

It significantly improves the accuracy of terminal models affected by data drift, shortens model update time, reduces hardware resource usage, and improves retraining efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115587217B_ABST
    Figure CN115587217B_ABST
Patent Text Reader

Abstract

The present invention relates to an online retraining method for a multi-terminal video detection model. The feature distributions of each terminal and the whole are obtained according to the sample labels and the input module of the intermediate model, and augmented samples are obtained by sampling on the overall feature distribution. Then, mutual learning is performed on the local model and the intermediate model corresponding to the new data set to update the local model and the intermediate model; multiple intermediate models are aggregated to generate a global model, and knowledge distillation is performed on the image set of the original data set by using the global model and the multiple local models corresponding thereto. The global model serves as a teacher model to impart knowledge to the local model serving as a student model; based on the characteristics of the model, the local model and the intermediate model are deployed on a digital accelerator, and certain processing is performed on the outputs of the two models, and the global model is deployed on a simulation accelerator. At the same time, weight noise is injected during the mutual learning process to enable the model to adapt to the influence of the noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of retraining of compressed deep models and hardware acceleration based on deep learning, and particularly to an online retraining method for multi-terminal video detection models. Background Art

[0002] With the improvement of people's living standards and the rapid development of technology, electronic devices such as smart phones have been widely popularized, and the deployment of deep learning on mobile devices has received more and more attention from researchers. However, limited by the huge neural network and the resource-limited mobile hardware platform, the deployment of neural networks on mobile devices has encountered many difficulties and obstacles. Researchers have started from model compression technology and deep learning hardware accelerators and made great progress in this field. However, video analysis based on deep learning will inevitably face the problem of data drift, that is, the video data in the real scene is different from the data during the training of the deep model. Under its influence, the accuracy of shallow lightweight models in the real scene may drop significantly, thus not meeting customer requirements, and online model retraining is an effective way to solve such problems.

[0003] Currently, most online model retraining relies on knowledge distillation technology to achieve. In this process, the terminal model, as the "student" model, learns from the "teacher" model located on the edge server, thereby realizing model update. However, due to limited resources on the edge server and the simple structure of the deployed "teacher" model, it will also face the problem of data drift, which affects the performance of the entire system. At the same time, the speed of online model retraining also affects the average inference accuracy of the terminal model. Therefore, accelerating the retraining process is also particularly important. Summary of the Invention

[0004] Technical Problems to be Solved

[0005] In order to avoid the deficiencies of the prior art, the present invention provides an online retraining method for multi-terminal video detection models.

[0006] Technical Solution

[0007] An online retraining method for multi-terminal video detection models, characterized by the following steps:

[0008] Step 1: Upload the data collected and screened by each terminal to the edge side, and use the global model located on the edge server to label the data without labels;

[0009] Step 2: Use the sample labels of the local model and the prediction module of the intermediate model to obtain the data set of each intermediate model before aggregation and the feature distribution of the overall data;

[0010] Step 3: All local models that share a global model sample the overall feature distribution in parallel according to the feature distribution to obtain augmented samples and update their original data sets;

[0011] Step 4: Based on the new data set, the corresponding local model and intermediate model are mutually learned to update the local model and the intermediate model. The loss functions corresponding to the training of the two models are rewritten as follows:

[0012] L local =αL Clocal +(1-α)D KL (P mid ||P local )

[0013] L mid =βL Cmid +(1-β)D KL (P mid ||P local )

[0014] where α and β are hyperparameters used to control the proportion of knowledge from data and other models, and L Clocal and L Cmid are the loss functions of the local model and the intermediate model based on data labels, P local and P mid are the inference results of the local model and the intermediate model, respectively;

[0015] Step 5: Use the FedAVG algorithm to aggregate the intermediate models of multiple terminal devices to generate a global model;

[0016] Step 6: Based on the data set originally collected and filtered by each terminal, the global model and its corresponding multiple local models are used for knowledge distillation. The global model acts as a "teacher" model to impart knowledge to the local model, which acts as a "student" model.

[0017] Step 7: The local model and the intermediate model are deployed on the digital accelerator, and the model input is processed to a certain extent. Specifically, after the global model completes the inference and annotation of the picture set uploaded by the terminal model, if there is no small target object in the picture, the picture pixel is reduced. If there is a small object in the picture, the picture pixel is not compressed to ensure the effect of retraining.

[0018] Step 8: Deploy the global model on the simulation accelerator and optimize the global model: On the one hand, the mutual learning used can make the model converge to a smoother minimum point, so that its robustness to noise is enhanced; on the other hand, weight noise is injected into the mutual learning process of the local model and the intermediate model to make the model adapt to the influence of noise. Specifically, in the forward propagation process of mutual learning, the weight of the lth layer is made to satisfy:

[0019] W l ∈N(W l0 ,σ N,l 2 )

[0020] σ N,l =μ(W lmax -W lmin )

[0021] And, in order to prevent the weight selection range from being too large and affecting the efficiency and accuracy of model training, further constraints are imposed on the weights:

[0022] W lmin ≤W l ≤W lmax

[0023] Wherein, W l0 is the original weight of the l-th layer, σ N,l is the injected noise, μ is the noise coefficient, W lmax and W lmin are respectively the maximum weight and the minimum weight of the l-th layer;

[0024] Step 9: The edge server transmits the updated model parameters back to the terminal, and the terminal immediately deploys the new model and continues with real-time video analysis.

[0025] A computer system, comprising: one or more processors, a computer-readable storage medium for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.

[0026] A computer-readable storage medium, storing computer-executable instructions, which are used to implement the above method when executed.

[0027] Beneficial effects

[0028] A multi-terminal video detection model online retraining method provided by the present invention uses a terminal compression model continuous evolution architecture based on mutual learning to complete the online update of the terminal model, significantly improving the accuracy of the compressed deep model affected by data drift. At the same time, through a retraining hardware acceleration method based on in-memory computing, the model update speed is greatly increased, and the hardware resources required for model retraining are significantly reduced, improving the retraining efficiency. Description of the drawings

[0029] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0030] Figure 1 Schematic diagram of the overall structure of the system for the online retraining method of the multi-terminal video detection model with low data transmission volume.

[0031] Figure 2 Data processing in the system. Specific implementation manners

[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0033] The present invention utilizes the following principles: by combining algorithms such as federated mutual learning and sample enhancement, retraining of the terminal model and the global model is achieved and model update is completed, and based on the characteristics of each model in the continuous evolution system, a suitable hardware acceleration scheme is selected for it, and the model deployed on it is optimized according to the hardware characteristics, so that the software and hardware are more adapted. The present invention can greatly improve the inference accuracy of the terminal model affected by data drift, and after hardware acceleration, the model accuracy, resource occupancy and acceleration effect have very significant improvements.

[0034] The present invention has a total of 2 attached drawings. Please refer to Figure 1 and Figure 2 as shown. The specific steps of the present invention are as follows:

[0035] Step 1: Upload the data collected and screened by each terminal to the edge side, and use the global model (uncompressed model) located on the edge server to label the data without labels.

[0036] Step 2: Use the sample labels of the local model (compressed model located at the terminal for video inference) and the prediction module of the intermediate model (the model is the same as the original global model and is used for mutual learning and model aggregation) to obtain the data set of each intermediate model before aggregation and the feature distribution of the overall data.

[0037] Step 3: All local models sharing a global model sample in parallel on the feature distribution according to the overall feature distribution to obtain augmented samples, update their original data sets, and with the help of sample enhancement, the influence of non-iid data of different terminal devices can be reduced, thereby improving the accuracy of the globally aggregated model and overcoming the heterogeneous problems of data and knowledge in federated learning.

[0038] Step 4: Based on the new data set, the corresponding local model and intermediate model are mutually learned to update the local model and the intermediate model. The loss functions corresponding to the training of the two models are rewritten as follows:

[0039] L local =αL Clocal +(1-α)D KL (P mid ||P local )

[0040] L mid =βL Cmid +(1-β)D KL (P mid ||P local )

[0041] where α and β are hyperparameters used to control the proportion of knowledge from data and other models, and L Clocal and L Cmid are the loss functions of the local model and the intermediate model based on data labels, P local and P mid These are the inference results of the local model and the intermediate model, respectively. In the mutual learning process, the accuracy of both the local model and the intermediate model will be improved, and the training effect of mutual learning is better than that of training the two models separately. There are three specific advantages: First, the category probability estimate output by the neural network will restore the connection information between different categories in the data to a certain extent. Therefore, the interaction of category estimates between networks can transmit and learn the data distribution characteristics, thereby improving the generalization ability of the model; secondly, mutual learning will also play a certain regularization role. The unique hot encoding of the true value label will make the model too certain about the prediction results during the training process, which is easy to cause the model to overfit. In the mutual learning process, the models learn each other's category probabilities, which can effectively prevent this phenomenon; finally, the network refers to the learning experience of other models to adjust its own learning process during the training process, which can make the results converge to a smoother minimum point, so it has better generalization ability and is less sensitive to noise, which gives us more choices for the system's hardware acceleration solution.

[0042] Step 5: Use the FedAVG algorithm to aggregate the intermediate models of multiple terminal devices to generate a global model.

[0043] Step 6: Based on the data set originally collected and filtered by each terminal, the global model and its corresponding multiple local models are used for knowledge distillation. The global model acts as a "teacher" model to impart knowledge to the local model, which acts as a "student" model.

[0044] Step 7: The local model and the intermediate model are deployed on a digital accelerator (such as a GPU), and certain processing is performed on the input of the model. Specifically, after the global model completes the inference annotation of the image set uploaded by the terminal model, if there are no small target objects in the image, the image pixels are reduced. If there are small objects in the image, in order to ensure the effect of retraining, the image pixels are not compressed. Under this optimization scheme, the accuracy of the model will not decrease significantly, while its video memory occupancy is significantly reduced. And due to the reduction of image pixels, the feature map input to the network becomes smaller, the intermediate results in model training decrease, reducing the computational volume and energy consumption of retraining, thereby accelerating the system retraining process and optimizing the system performance.

[0045] Step 8: The global model is deployed on a simulation accelerator, and we optimize the global model to increase its robustness to weight noise, so as to achieve better results on the simulation accelerator. On the one hand, the mutual learning we adopt can make the model converge to a smoother minimum point, strengthening its robustness to noise. On the other hand, weight noise is injected during the mutual learning process between the local model and the intermediate model to make the model adapt to the influence of noise. Specifically, during the forward propagation of mutual learning, the weights of the l-th layer satisfy:

[0046] W l ∈N(W l0 ,σ N,l 2 )

[0047] σ N,l =μ(W lmax -W lmin )

[0048] And, in order to prevent the weight selection range from being too large, affecting the efficiency and accuracy of model training, we further constrain the weights:

[0049] W lmin ≤W l ≤W lmax

[0050] where W l0 is the original weight of the l-th layer, σ N,L is the injected noise, μ is the noise coefficient, W lmax and W lmin are the maximum and minimum weights of the l-th layer respectively. Under this training method, the robustness of the aggregated global model to weight noise can be strengthened, the accuracy degradation during inference prediction on the simulation accelerator can be reduced, and the overall performance of the system can be improved.

[0051] Step 9: The edge server transmits the updated model parameters back to the terminal. The terminal immediately deploys the new model and continues with real-time video analysis.

[0052] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. An online retraining method for a multi-terminal video detection model, characterized in that Here are the steps: Step 1: Upload the data collected and filtered by each terminal to the edge, and use the global model located in the edge server to annotate the unlabeled data; Step 2: Use the sample labels of the local model and the prediction module of the intermediate model to obtain the data set of each intermediate model before aggregation and the feature distribution of the overall data; Step 3: All local models that share a global model sample the overall feature distribution in parallel according to the feature distribution to obtain augmented samples and update their original data sets; Step 4: Based on the new data set, the corresponding local model and intermediate model are mutually learned to update the local model and the intermediate model. The loss functions corresponding to the training of the two models are rewritten as follows: L local = αL Clocal + (1 - α)D KL (P mid ||P local ) It should be noted that there seems to be a possible error or ambiguity in the original formula in line 4 where it says "1 - αD" which might be intended as "1 - αD" or something else. The translation is based on the best understanding of the provided text. L mid = βL Cmid + (1 - β)D KL (P mid ||P local ) where α and β are hyperparameters used to control the proportion of knowledge from data and other models, L Clocal and L Cmid and L Cmid are the loss functions of the local model and the intermediate model based on data labels, respectively, and P local and P mid and P mid are the inference results of the local model and the intermediate model, respectively; Step 5: Use the FedAVG algorithm to aggregate the intermediate models of multiple terminal devices to generate a global model; Step 6: Based on the data set originally collected and filtered by each terminal, the global model and its corresponding multiple local models are used for knowledge distillation. The global model acts as a "teacher" model to impart knowledge to the local model, which acts as a "student" model. Step 7: The local model and the intermediate model are deployed on the digital accelerator, and the model input is processed to a certain extent. Specifically, after the global model completes the inference and annotation of the picture set uploaded by the terminal model, if there is no small target object in the picture, the picture pixel is reduced. If there is a small object in the picture, the picture pixel is not compressed to ensure the effect of retraining. Step 8: Deploy the global model on the simulation accelerator and optimize the global model: On the one hand, the mutual learning used can make the model converge to a smoother minimum point, so that its robustness to noise is enhanced; on the other hand, weight noise is injected into the mutual learning process of the local model and the intermediate model to make the model adapt to the influence of noise. Specifically, in the forward propagation process of mutual learning, the weight of the lth layer is made to satisfy: W l ∈ N(W l0 , σ N,l 2 ) σ N,l = μ(W lmax - W lmin ) In addition, in order to prevent the weight selection from being too large and affecting the efficiency and accuracy of model training, the weights are further constrained: W lmin ≤W l ≤W lmax Among them, W l0 is the weight of the original l-th layer, σ N,l is the injected noise, μ is the noise coefficient, W lmax and W lmin are the maximum weight and the minimum weight of the l-th layer respectively; Step 9: The edge server transmits the updated model parameters back to the terminal, which immediately deploys the new model and continues real-time video analysis.

2. A computer system, characterized in that include: One or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method of claim 1.

3. A computer-readable storage medium, characterized in that Computer executable instructions are stored, and when the instructions are executed, they are used to implement the method of claim 1.

Citation Information

Patent Citations

  • Federal mutual learning model training method for non-independent identically distributed data

    CN114091667A

  • Systems and methods for image-to-video re-identification

    WO2022134104A1