Context-aware online real-time video analysis method based on end-edge collaboration
Patent Information
- Application Number
- CN202310845052.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-07-11
AI Technical Summary
[0009]1.DNN模型的计算资源需求较高,需要特殊硬件(GPU,NPU等)才能达到实时要求;
[0066]1.基于强化学习,不需要先验的专家知识,并且算法的性能不会受到初始参数的影响,算法更加稳定。
Smart Images

Figure CN116866660B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge computing technology, and in particular to a context-aware online real-time video analysis method based on edge-end collaboration. Background Technology
[0002] Intelligent video analytics services have great application potential in fields such as smart parks and intelligent transportation. Meanwhile, with the development of science and technology, terminal devices such as smartphones and intelligent connected vehicles are ubiquitous in people's lives. However, challenges remain in meeting the real-time analysis needs arising from the massive amounts of video data generated by these terminal devices for real-time applications in decision support.
[0003] Traditional cloud computing architectures cannot meet the demands of large-scale real-time video analytics. With the widespread adoption of smartphones, connected cars, and surveillance cameras in daily life, the volume of real-time video stream data is growing exponentially. In industrial production environments, video analytics applications based on deep neural networks are often deployed on these terminal devices. However, the high computational resource consumption of deep neural networks and the low computational performance of terminal devices make it difficult to meet the high accuracy and low latency requirements of real-time inference tasks. Furthermore, large-scale video stream data relying on traditional cloud computing architectures is prone to processing queue congestion and high dependence on network bandwidth. In summary, the centralized processing approach of traditional cloud computing is increasingly unable to adapt to the explosive growth of data at the network edge and on terminals. Therefore, relying solely on terminal devices or cloud computing architectures cannot solve the ever-increasing demand for real-time video analytics.
[0004] Current edge-end collaborative architectures still exhibit a high dependence on network bandwidth. Traditional edge-end collaborative video analytics architectures rely solely on the data acquisition capabilities of the edge devices, still processing video analytics tasks at the edge. However, due to the highly computationally intensive nature of video stream analysis, this places even greater demands on network bandwidth during transmission. Furthermore, this architecture fails to fully utilize the performance of local terminal devices, resulting in a certain degree of waste of terminal computing resources.
[0005] In summary, designing an edge-to-edge collaborative video analytics method with low network persistence and high dependence, which fully utilizes local terminal computing resources, is an urgent problem to be solved.
[0006] Existing technology 1
[0007] Collaborative Analysis of Independent DNN Models Deployed at the Edge and End (Tan, Tianxiang et al.). Deep learning on mobile edges through neural processing units and edge computing[J].arXiv, 2021.
[0008] Disadvantages of existing technology 1
[0009] 1. DNN models have high computational resource requirements and require special hardware (GPU, NPU, etc.) to meet real-time requirements;
[0010] 2. It cannot adapt to dynamically changing task unloading environments. The algorithm needs to be rerun for each scheduling, and it cannot make a unified unloading decision for tasks with different needs.
[0011] Existing technology 2
[0012] Deploying DNN models at the edge and retraining terminal DNN models at the edge. Rivas, Daniel et al. Towards automatic model specialization for edge video analytics[J]. FUTURE GENERATI ON COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENC E, 2022, 134:399-413.
[0013] Disadvantages of existing technology 2
[0014] 1. Model retraining is time-consuming and may not converge; it also adapts slowly to dynamically changing video analysis environments.
[0015] 2. The terminal side requires high-performance hardware support, resulting in poor versatility;
[0016] 3. The data dependencies between the environment and system performance are not utilized, resulting in spatiotemporal redundancy in data analysis. Summary of the Invention
[0017] This invention addresses the shortcomings of existing technologies by providing a context-aware online real-time video analysis method based on edge-device collaboration. By establishing an adaptive mechanism for video frame parameters and offloading thresholds, after video data acquisition, the decision module adaptively sets the resolution and offloading threshold for the data. The local terminal then determines whether to offload the data to an edge device for processing based on the offloading threshold, thereby achieving real-time processing of video stream data.
[0018] To achieve the above-mentioned objectives, this invention discloses a context-aware online real-time video analysis method based on edge-end collaboration, specifically including the following steps:
[0019] Step 1: Acquire video frame data using a video capture device;
[0020] Step 2: Construct a decision module, which determines the resolution parameters and offloading threshold for processing video frame data;
[0021] Step 3: Preprocess the video frames according to the output of the decision module, including resolution processing of the video frame data according to the resolution parameters, and target tracking using the target tracking algorithm of the local terminal.
[0022] Step 4: The local terminal calculates the cumulative offset using a target tracking algorithm;
[0023] Step 5: The local terminal uses the cumulative offset to determine whether to offload the resolution-processed video frames to the edge device based on the offloading threshold according to the TCP / IP protocol;
[0024] Step 7: The local terminal receives the processing result and updates the target tracking result box on the local terminal.
[0025] Furthermore, the video frame data mentioned in step 1 includes real camera data and virtual camera data, and the camera must remain stationary during the video acquisition process.
[0026] Furthermore, the adaptive mechanism of the decision module for video frame parameters and offloading threshold in step 2 includes consideration of the current state of the video system. The current state of the video system is deployed on the local terminal, and the indicators considered by the adaptive mechanism include real-time network bandwidth, real-time analysis latency, and real-time analysis accuracy.
[0027] Furthermore, the algorithm for constructing the decision module in step 2 is based on UCB (Upper Confidence Bound) and Bayesian Optimization, as detailed below:
[0028] First, the time range is discretized into fixed round intervals, with rounds t = 1, 2, ... The decision module observes the system state, sets adaptive parameters for the video data of each round, and interacts with the environment to learn the optimal strategy.
[0029] The adaptive parameters include context, action, reward, and reward prediction (Bayesian optimization).
[0030] Context: Represents the current state of the system, including the system's analysis rate and accuracy. In round t, the context is:
[0031] c t =(P t-1 A t-1 (1-1)
[0032] Among them, P t-1 A represents the video frame analysis rate of the previous round. t-1 This indicates the accuracy of the video frame analysis in the previous round. The decision module is based on c. t Achieve adaptive video frame parameters and unloading threshold.
[0033] Action: The decision includes video frame resolution parameters and offload threshold parameters. The output decision action is defined as follows:
[0034] a t =(r t ,θ t (1-2)
[0035] Where r t This indicates the current video frame resolution parameter setting, θ. t This indicates the current video frame unloading threshold setting.
[0036] Rewards: Rewards are used to evaluate video frame inference performance; within each round t, based on a t The resolution parameter r in t The video frame data is preprocessed, and the cumulative offset d is calculated based on the optical flow algorithm. f,t and with a t The unloading threshold θ t The parameters are compared to determine whether to unload. A reward function is applied after all video frames within a round have been processed.
[0037] y t =P t +α·A t (1-3)
[0038] Where α is the weighting parameter.
[0039] Reward Prediction: Since the reward for the current round is unknown at the beginning of each round, Bayesian Optimization is used to predict the reward function; the reward function of Bayesian Optimization is defined as follows:
[0040] y t =f(z) t )+ε t (1-4)
[0041] Among them, z t =(a t ,c t ), which represents the state-action pair in each round, ε t ~N(0,σ 2 () represents random noise, describing the noise between the actual reward and the observed reward. The reward prediction function is defined as the mean function:
[0042] μ=E[f(z t (1-5)
[0043] Covariance (kernel) function
[0044] k z (z,z′)=E[(f(z)-μ(z))(f(z′)-μ(z′))] (1-6)
[0045] Here, the kernel function uses the Matérn kernel, specifically:
[0046]
[0047] Among them, parameter ν controls the sampling smoothness, B ν This represents the Bessel function, where the initial prior distribution of the reward function f(·) is set to N(0,k). z ).
[0048] Assume that the set of historical observation rewards at round t is Y. t =[y1,...,y t ] T For action-state pairs Z t =[z1,...,z t ] T Then when the new action state affects z t+1 Upon arrival, the observed distribution is updated as follows:
[0049]
[0050] Where, μ(Z) t )=[μ(z1),...,μ(z t )] T Kernel function K t =[k z (z,z′)]z,z′∈Z,k t (z)=[k z (z1,z),...,k z (z t ,z)] T Historical action state pairs and the set of observed rewards
[0051] Ht =(z1,y1),...,(z t ,y t If ), then for the predicted reward y t+1 The posterior distribution is:
[0052] y t+1 |H t ,z t+1 ~N(μ(z) t+1 |H t ),σ 2 (z t+1 |H t (1-9) Let μ t (z), k t (z,z′) and The mean, variance, and covariance of round t are expressed as follows:
[0053] μ t (z)=k t (z) T (K t +σ 2 I) -1 Y (1-10)
[0054] k t (z,z′)=k z (z,z′)-k t (z) T (K t +σ 2 I) -1 k t (z′) (1-11)
[0055]
[0056] To balance exploration and reward, UCB is used for action selection, specifically:
[0057]
[0058] Where D is the action space, β t Define a constant for the user. t These are the optimal decision parameters for the output.
[0059] Furthermore, the preprocessing described in step 3 involves scaling the video frame data according to the output resolution parameters of the decision module.
[0060] Furthermore, in step 4, the target tracking algorithm calculates the cumulative offset by transmitting the initial video frame to the edge device via TCP / IP protocol to generate a result box. After receiving this result box, the local terminal uses the target tracking algorithm to track the result box and calculate the cumulative coordinate offset between the tracking box and the result box.
[0061] Furthermore, step 5, which involves using the cumulative offset to determine whether to offload the resolution-processed video frames to the edge device based on the offloading threshold using the TCP / IP protocol, is as follows:
[0062] Cumulative offset > unload threshold → edge device inference;
[0063] Cumulative offset ≤ unload threshold → terminal device tracking.
[0064] Furthermore, the update of the target tracking result box of the local terminal in step 6 is based on the unloading result. If the video frame is unloaded to the edge device, the target tracking result box of the local terminal is updated to the edge device result box, and the cumulative offset is cleared to zero. If no unloading is performed, the target tracking result box of the local terminal is updated to the target tracking result box, and the cumulative offset is calculated.
[0065] Compared with the prior art, the advantages of the present invention are as follows:
[0066] 1. Based on reinforcement learning, it does not require prior expert knowledge, and the algorithm's performance is not affected by initial parameters, making the algorithm more stable.
[0067] 2. It can discover the data dependencies between the environment and system performance, thus achieving better uninstallation results.
[0068] 3. It has low requirements for the performance and bandwidth of local terminal equipment, and has good universal deployment of terminal equipment and high adaptability to network fluctuations. Attached Figure Description
[0069] Figure 1 This is a schematic diagram of the structure of a context-aware online real-time video analysis method based on edge-end collaboration according to an embodiment of the present invention.
[0070] Figure 2 This is a schematic diagram illustrating the application of a context-aware online real-time video analysis method based on edge-end collaboration according to an embodiment of the present invention. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and examples.
[0072] like Figure 1As shown, a context-aware online real-time video analysis method based on edge collaboration is presented. According to the system architecture implemented by this method, the system consists of a local terminal device and an edge device. First, the local terminal device acquires video stream data and uploads the first frame to the edge device for initializing the target detection box and sending back the detection result. The local terminal device extracts ORB feature points and performs Lucas-Kanade optical flow tracing, while simultaneously calculating the cumulative offset threshold in real time. The decision module outputs an unloading decision based on the cumulative offset threshold and environmental data.
[0073] like Figure 2 As shown, a context-aware online real-time video analysis method based on edge collaboration is proposed. The system architecture based on this method will be deployed on computer hardware, wherein the local terminal uses a Raspberry Pi 4B device and the edge device uses a Jetson Xavier.
[0074] The specific steps of this method are as follows:
[0075] Step 1: Acquire video frame data using a local terminal device. Local terminal device video frame acquisition can be achieved through various interfaces, including video file reading, virtual camera reading, and real camera reading. Simultaneously, local terminal device data acquisition should be performed when the camera is stationary. Specifically, the reading of video frame data involves:
[0076] (1) Reading video files. Video file formats include avi, mp4, etc.
[0077] (2) Virtual camera reading refers to video frame data from a camera implemented by software;
[0078] (3) Real camera reading refers to reading video frame data acquired by camera hardware in real time.
[0079] Step 2: Construct a decision module, which determines the resolution parameters and offloading threshold for processing video frame data through an adaptive mechanism. The adaptive mechanism of the decision module for video frame parameters and offloading threshold includes consideration of the current state of the video system, which is deployed on the local terminal. The indicators considered by the adaptive mechanism include real-time network bandwidth, real-time analysis latency, and real-time analysis accuracy.
[0080] The construction of the decision module in step 2 is mainly implemented by UCB (Upper Confidence Bound) and Bayesian Optimization, as detailed below:
[0081] First, the time range is discretized into fixed round intervals, rounds t = 1, 2, ... The decision module records the system state, then sets adaptive parameters for the video data of each round, and interacts with the environment to learn the optimal strategy;
[0082] Context: Represents the current state of the system, including the system's analysis rate and accuracy. In round t, the context is...
[0083] c t =(P t-1 A t-1 (1-1)
[0084] Among them, P t-1 A represents the video frame analysis rate of the previous round. t-1 This indicates the accuracy of the video frame analysis in the previous round. The decision module is based on c. t Achieve adaptive video frame parameters and unloading threshold.
[0085] Action: The decision-making module adaptively sets parameters for each round of video stream data based on the current system state. Decisions include video frame resolution parameters and offload threshold parameters. The output decision action is defined as follows:
[0086] a t =(r t ,θ t (1-2)
[0087] Where r t This indicates the current video frame resolution parameter setting, θ. t This indicates the current video frame unloading threshold setting.
[0088] Reward: The reward evaluates the performance of video frame inference; within each round t, based on a t The resolution parameter r in t The video frame data is preprocessed, and the cumulative offset d is calculated based on the target tracking algorithm. f,t and with a t The unloading threshold θ t The parameters are compared to determine whether to unload. A reward function is applied after all video frames within a round have been processed.
[0089] y t =P t +α·A t (1-3)
[0090] Where α is the weighting parameter.
[0091] Reward Prediction (Bayesian Optimization): Since the reward for the current round is unknown at the beginning of each round, Bayesian Optimization is used to predict the reward function; the reward function of Bayesian Optimization is defined as follows:
[0092] y t =f(z) t )+ε t (1-4)
[0093] Among them, z t =(a t ,c t ), which represents the state-action pair in each round, ε t ~N(0,σ 2 () represents random noise, describing the noise difference between the actual reward and the observed reward. The predicted reward function is defined as the mean function.
[0094] μ=E[f(z t (1-5)
[0095] Covariance (kernel) function:
[0096] k z (z,z′)=E[(f(z)-μ(z))(f(z′)-μ(z′))] (1-6)
[0097] Here, the kernel function uses the Matérn kernel, specifically:
[0098]
[0099] Among them, parameter ν controls the sampling smoothness, B ν This represents the Bessel function, where the initial prior distribution of the reward function f(·) is set to N(0,k). z ).
[0100] Assume that the set of historical observation rewards at round t is Y. t =[y1,...,y t ] T For action-state pairs Z t =[z1,...,z t ] T Then when the new action state affects z t+1 Upon arrival, the observed distribution is updated as follows:
[0101]
[0102] Where, μ(Z) t)=[μ(z1),...,μ(z t )] T Kernel function K t =[k z (z,z′)]z,z′∈Z,k t (z)=[k z (z1,z),...,k z (z t ,z)] T Historical action state pairs and the observed reward set H t =(z1,y1),...,(z t ,y t If ), then for the predicted reward y t+1 The posterior distribution is
[0103] y t+1 |H t ,z t+1 ~N(μ(z) t+1 |H t ),σ 2 (z t+1 |H t (1-9)
[0104] Let μ t (z), k t (z,z′) and Let the mean, variance, and covariance of round t be represented by the following formulas:
[0105] μ t (z)=k t (z) T (K t +σ 2 I) -1 Y (1-10)
[0106] k t (z,z′)=k z (z,z′)-k t (z) T (K t +σ 2 I) -1 k t (z′) (1-11)
[0107]
[0108] To balance exploration and reward, UCB (Upper Confidence Bound) is used for action selection, specifically...
[0109]
[0110] Where D is the action space, β t Define a constant for the user. t These are the optimal decision parameters for the output.
[0111] Step 3: The local terminal preprocesses the video frames according to the output of the decision module, including resolution processing of the video frame data according to the resolution parameters, and target tracking using the target tracking algorithm of the local terminal.
[0112] Step 4: The local terminal uses the accumulated offset to determine, based on the offloading threshold, whether to offload the resolution-processed video frames to the edge device according to the TCP / IP protocol; specifically: the system determines the offloading action a based on the selected action. t Scale the video frame resolution within round t to r. t The decision threshold is set to θ t .
[0113] Step 5, the local terminal uses the accumulated offset d f,t With decision threshold θ t The relationship determines whether the resolution-processed video frames are offloaded to the edge device based on the TCP / IP protocol; specifically, this relationship is:
[0114] d f,t >θ t →Edge device inference;
[0115] d f,t ≤θ t → Terminal device tracking.
[0116] Step 6: The local terminal receives the processing result and updates the target tracking result box of the local terminal. If the video frame is unloaded to the edge device, the target tracking result box of the local terminal is updated to the result box of the edge device, and the cumulative offset is cleared to zero; if no unloading is performed, the target tracking result box of the local terminal is updated to the target tracking result box, and the cumulative offset is calculated.
[0117] Finally, the context-aware online real-time video analysis method based on edge-end collaboration is described below. First, the video acquisition device acquires video frame data, the local terminal reads the video frames, and then the decision module selects the resolution parameter r using equation (1-13). t and unloading threshold θ tThe local terminal scales the video frames within the current round $t$ to a resolution $r_t$, sets the unloading threshold to $\theta_t$, and performs target tracking using a target tracking algorithm. Then, the local terminal calculates the cumulative offset using the tracking result box and the initial result box. Based on the cumulative offset and the unloading threshold, it determines whether to unload the resolution-processed video frames to the edge device using the TCP / IP protocol. If unloading to the edge device is required, the local terminal's target tracking result box is updated to the edge device's result box, and the cumulative offset is cleared to zero. If unloading is not required, the local terminal's target tracking result box is updated to the target tracking result box.
[0118] The methods described above according to the invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be stored as software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code. When said software or computer code is accessed and executed by the computer, processor, or hardware, it implements the context-aware online real-time video analytics method described herein based on edge-to-edge collaboration. Furthermore, when a general-purpose computer accesses the code used to implement the processing shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for performing the processing shown herein.
[0119] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the implementation methods of the present invention, and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of the present invention.
Claims
1. A context-aware online real-time video analysis method based on edge-end collaboration, characterized in that... Specifically, the following steps are included: Step 1: Acquire video frame data using a video capture device; Step 2: Construct a decision module, which determines the resolution parameters and offloading threshold for processing video frame data; The algorithm for constructing the decision module is based on UCB (Upper Confidence Bound) and Bayesian Optimization, as detailed below: First, the time range is discretized into fixed round intervals, rounds The decision-making module observes the system status, sets adaptive parameters for each round of video data, and interacts with the environment to learn the optimal strategy. The adaptive parameters include context, action, reward, and reward prediction (Bayesian optimization). Context: Represents the current state of the system, including the system's analysis rate and accuracy, within the round. The environment context is: (1-1); in, This indicates the video frame analysis rate of the previous round. This indicates the accuracy of the video frame analysis in the previous round; the decision module bases its decisions on... Achieve adaptive video frame parameters and offload threshold; Action: The decision includes video frame resolution parameters and offload threshold parameters. The output decision action is defined as follows: (1-2); in This indicates the current video frame resolution parameter setting. This indicates the current video frame unloading threshold setting; Rewards: Rewards are used to evaluate video frame inference performance; in each round Inside, according to resolution parameters in The video frame data is preprocessed, and the cumulative optical flow offset is calculated based on the optical flow algorithm. and with Unloading threshold in The parameters are compared to determine whether to unload. A reward function is applied after all video frames within a round have been processed. (1-3); in, These are weight parameters; Reward Prediction: Since the reward for the current round is unknown at the beginning of each round, Bayesian Optimization is used to predict the reward function; the reward function of Bayesian Optimization is defined as follows: (1-4); in, , represents the action-state pair in each round. Representing random noise, describing the noise between the actual reward and the observed reward; defining the reward prediction function as the mean function: (1-5); Covariance (kernel) function (1-6); Here, the kernel function uses the Matérn kernel, specifically: (1-7); Among them, parameters Controlling sampling smoothness, This represents the Bessel function; here, the reward function is set. The initial prior distribution is ; Assuming in round The following is a collection of historical observation rewards. For action state pairs Then when the new action state is Upon arrival, the observed distribution is updated as follows: (1-8); in, kernel function Historical action state pairs and the set of observed rewards For predicting rewards The posterior distribution is: (1-9); make , and Indicates a round The mean, variance, and covariance are given by the following formula: (1-10); (1-11); (1-12); To balance exploration and reward, UCB is used for action selection, specifically: (1-13); in, For the action space, Define constants for users. The optimal decision parameters are output. Step 3: Preprocess the video frames according to the output of the decision module, including resolution processing of the video frame data according to the resolution parameters, and target tracking using the target tracking algorithm of the local terminal. Step 4: The local terminal calculates the cumulative offset using a target tracking algorithm; Step 5: The local terminal uses the cumulative offset to determine whether to offload the resolution-processed video frames to the edge device based on the offloading threshold according to the TCP / IP protocol; Step 6: The local terminal receives the processing result and updates the target tracking result box on the local terminal.
2. The context-aware online real-time video analysis method based on edge-end collaboration according to claim 1, characterized in that: The video frame data mentioned in step 1 includes real camera data and virtual camera data, and the camera must remain stationary during the video acquisition process.
3. The context-aware online real-time video analysis method based on edge-end collaboration according to claim 1, characterized in that: The adaptive mechanism of the decision module for video frame parameters and offload threshold in step 2 includes consideration of the current state of the video system. The current state of the video system is deployed on the local terminal, and the indicators considered by the adaptive mechanism include real-time network bandwidth, real-time analysis latency, and real-time analysis accuracy.
4. The context-aware online real-time video analysis method based on edge-end collaboration according to claim 1, characterized in that: The preprocessing described in step 3 involves scaling the video frame data according to the output resolution parameters of the decision module.
5. The context-aware online real-time video analysis method based on edge-end collaboration according to claim 1, characterized in that: The target tracking algorithm described in step 4 calculates the cumulative offset by transmitting the initial video frame to the edge device via TCP / IP protocol to generate a result box. After receiving this result box, the local terminal uses the target tracking algorithm to track the result box and calculate the cumulative coordinate offset between the tracking box and the result box.
6. The context-aware online real-time video analysis method based on edge-end collaboration according to claim 1, characterized in that: Step 5, which involves using the cumulative offset to determine whether to offload the resolution-processed video frames to the edge device based on the offloading threshold using the TCP / IP protocol, is as follows: Cumulative offset > unload threshold → edge device inference; Cumulative offset Unload threshold → Terminal device tracking.
7. The context-aware online real-time video analysis method based on edge-end collaboration according to claim 1, characterized in that: The update of the target tracking result box of the local terminal in step 6 is based on the unloading result. If the video frame is unloaded to the edge device, the target tracking result box of the local terminal is updated to the edge device result box, and the cumulative offset is cleared to zero. If no unloading is performed, the target tracking result box of the local terminal is updated to the target tracking result box, and the cumulative offset is calculated.