Time-Varying Environment-Aware Video Analysis Task Offloading Method Based on Region of Interest
By collaborative decision-making between the local terminal and the edge server, the uninstallation strategy of the region of interest is determined, and the real-time and accuracy of small and medium-sized object detection in traditional methods is solved, real-time processing and resource optimization of high-definition video streams are achieved.
Patent Information
- Application Number
- CN202411273137.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-09-12
AI Technical Summary
Traditional video frame-level offloading methods have problems such as detection failure and high demand for network bandwidth resources when tracking and detecting small targets. They cannot meet the real-time and accuracy requirements of high-definition video analysis, and fail to make full use of heterogeneous resources of local terminals and edge servers.
The initial bounding box is generated by the edge server, and the local terminal decision module determines the packaging size and unloading threshold of the region of interest. Combining the confidence upper bound method, Bayesian optimization and Gaussian process, the cumulative offset and normalized interrelated values are calculated in real time, and adaptively decides whether to unload the region of interest to the edge server, and combines and packages using the two-dimensional packaging method and the Guillotine algorithm.
Real-time processing of small-target high-definition video streams in dynamic time-varying environments is realized, network resource dependence is reduced, end-edge heterogeneous resources are fully utilized, and processing efficiency and accuracy are improved.
Smart Images

Figure CN119135955B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge video analysis, and particularly relates to a time-varying environment perception video analysis task offloading method based on regions of interest. Background Art
[0002] With the rapid development of information technology, the generation and application of video data have increased explosively. In many fields, such as intelligent security, traffic monitoring, etc., video analysis plays a crucial role. However, there are still challenges in meeting the real-time analysis requirements for tracking and detecting small targets in high-definition videos.
[0003] When traditional video frame-level offloading is used to track and detect small targets, due to the small number of pixels occupied, unclear features, or easy interference by complex backgrounds, it is difficult to capture their movement trajectories during the process of target movement changes, and detection failures will occur. Secondly, when processing high-definition video data in the edge-cloud architecture, a large amount of video data needs to be transmitted to the edge server for processing and analysis, which will cause a huge pressure on the edge server for transmission and calculation, easily lead to data processing delays, and thus cannot meet the real-time analysis and rapid response requirements in practical applications. The traditional edge-cloud collaborative video frame-level offloading method cannot meet the real-time and accuracy requirements for high-definition video analysis with small targets.
[0004] In addition, the current edge-cloud collaborative architecture has a high dependence on network bandwidth. Its video analysis tasks are mainly processed by the edge side with high computing power, and video stream analysis has the characteristics of high computational intensity and extremely high requirements for network bandwidth in terms of transmission. At the same time, this form of edge-cloud architecture fails to fully utilize the heterogeneous resources of local terminals and edge servers, resulting in resource waste.
[0005] In summary, designing a low-network-dependence edge-cloud collaborative analysis method suitable for high-definition videos with small targets is an urgent problem to be solved.
[0006] Prior Art One
[0007] System based on on-device object tracking and edge-assisted analysis, Hanyao et al. Edge-assisted Online On-device Object Detection for Real-time Video Analytics [C]. IEEE INFOCOM 2021 - IEEE Conference on Computer Communications, 2021: 1 - 10.
[0008] Disadvantages of Prior Art Two
[0009] 1. There is a need for certain computing and network bandwidth resources to transmit video frames between the terminal and the edge, and there may still be a certain transmission delay.
[0010] 2. It depends on the initial offloading threshold setting, which affects the subsequent online learning time.
[0011] Prior Art II
[0012] Consider video frame-level offloading in an unknown and time-varying environment. A. Galanopoulos et al. AutoML for Video Analytics with Edge Computing [C]. IEEE INFOCOM 2021 - IEEE Conference on Computer Communications, 2021: 1 - 10.
[0013] Disadvantages of Prior Art II
[0014] 1. For video frame-level analysis tasks offloading of high-definition videos, there is still a problem of high demand for network bandwidth resources;
[0015] 2. Only relying on the DNN model deployed on the edge server to process video analysis tasks cannot fully utilize the resources of local terminals. Summary of the Invention
[0016] In view of the deficiencies of the prior art, the present invention provides a method for offloading video analysis tasks with time-varying environment awareness based on regions of interest. After capturing video data and obtaining video frames, an initial bounding box is generated by the edge server model, and the local terminal decision module determines the packing size of the region of interest and the offloading threshold. Based on the initial bounding box, the region of interest, its cumulative offset, and the normalized correlation value are obtained. Whether to offload the region of interest to the edge server is judged based on the offloading threshold, and the region of interest to be offloaded is combined and packed, so as to realize real-time processing of high-definition video stream data of small targets.
[0017] To achieve the above invention purpose, the present invention discloses a method for offloading video analysis tasks with time-varying environment awareness based on regions of interest, specifically including the following steps:
[0018] Step 1: The video acquisition device captures video data and obtains video frames through a decoder, and then an initial bounding box and the region of interest to be tracked are generated by the model deployed on the edge server;
[0019] Step 2: A decision module is built on the local terminal to determine the packing size of the region of interest, the cumulative pixel offset threshold of the region of interest, and the normalized correlation threshold of the region of interest;
[0020] Step 3: Use the target tracking module deployed on the terminal device for tracking, and calculate the cumulative pixel offset value and the normalized cross-correlation value of the region of interest in real time;
[0021] Step 4: The local terminal determines whether to offload the region of interest to the edge server through Socket communication based on the offloading threshold;
[0022] Step 5: The local terminal adaptively combines and packs the region of interest to be offloaded, and performs offloading according to the offloading judgment result;
[0023] Step 6: The local terminal receives the inference result from the edge server and updates the final inference result of the local terminal.
[0024] Furthermore, the video data in Step 1 comes from a real camera and a virtual camera, and the camera does not move during the acquisition process.
[0025] Furthermore, the decision-making module in Step 2 makes decisions based on time-varying environment perception, considering the current state of the system to set the packing size of the region of interest, the cumulative pixel offset threshold of the region of interest, and the normalized cross-correlation threshold, and finally balances the processing delay and the inference accuracy.
[0026] Furthermore, the implementation of the decision-making module of the local terminal in Step 2 combines the upper confidence bound method, Bayesian optimization, and Gaussian process based on a temporal kernel function, and its specific implementation is as follows:
[0027] First, decompose the time into multiple rounds t∈{1,2,...,T} with fixed temporal segmentation. The decision-making module determines the adaptive parameter settings and offloading thresholds for the set of regions of interest within each round of video frames by observing the environmental context, and interacts with the environment to obtain the optimal decision result;
[0028] The main elements of the decision-making problem include actions, environmental context, rewards, and reward predictions;
[0029] Actions: The decision-making actions include the packing size of the region of interest, the cumulative offset threshold of the region of interest, and the normalized cross-correlation threshold, which are defined as:
[0030] a t =(z t ,θ t ,n t ) (1)
[0031] where z t represents the suitable packing size of the region of interest combination for the current round, θ t represents the cumulative offset threshold setting for the current round, n t represents the normalized cross-correlation threshold for the current round, θ t and nt will jointly serve as the unloading threshold for subsequent determination of whether to unload.
[0032] Environmental context: The environmental context includes the processing delay and inference accuracy of the system, and is defined as:
[0033] c t =(D t-1 ,A t-1 ) (2)
[0034] where D t-1 represents the processing delay of the previous round, and A t-1 represents the inference accuracy of the previous round.
[0035] Reward: In each round, the local terminal packs, combines, and unloads the set of regions of interest generated from the video frame data according to the action a t . After completing the processing of the current round, the reward function is calculated to evaluate the current policy. The reward function is defined as:
[0036] y t =1 / D t +α·A t (3)
[0037] where α is the weight parameter.
[0038] Reward prediction: At the initial stage of decision-making, the environmental prior knowledge is unknown, and the relationship between the action context pair and the reward within each round is unknown. Therefore, Bayesian optimization is used to predict the reward function. The reward prediction function under Bayesian optimization is defined as:
[0039] y t =f(x t )+∈ t (4)
[0040] where x t =(a t ,c t ) represents the action context of d within each round, represents the random noise, that is, the difference between the actual reward and the observed reward. According to the Gaussian process, the mean function of the prior distribution of the reward prediction function is:
[0041]
[0042] Its covariance function is:
[0043]
[0044] where x and x′ are elements in the set X of action context pairs. Here, the initial prior distribution of the reward function f(·) is set to
[0045] For the update of the prior distribution of the reward prediction function, according to the Sherman-Morrison-Woodbury formula, the reward prediction function Y t = [y1,..., y t T The posterior distribution still follows a Gaussian process, and the mean of the posterior distribution after each round is:
[0046] μ t (x) = k t (x) T (K t + σ 2 Ι) -1 Y t (7)
[0047] The variance is:
[0048]
[0049] Among them, the kernel function K t = [k x (x, x′)], k t (x) = [k x (x1, x),..., k x (x t , x)] T , x, x′ ∈ X.
[0050] To solve the problem that the update effect of the reward function decreases due to "stale" historical data, first, introduce the time dimension into the kernel function k x and introduce the time series parameter t into the environmental context pair. Secondly, on the basis of the squared exponential kernel function k s , introduce the Ornstein-Uhlenbeck time covariance function k o , specifically:
[0051]
[0052] Among them, λ controls the decay rate of the importance of historical data samples. The larger λ is, the faster the importance of historical data samples decays, and the more important the new data samples will be; l represents the length scale, which determines the degree of mutual influence between x and x′.
[0053] If the experience time series information interval has uniformity, then for the kernel matrix is Among them, is the Hadamard product; Among them, u t With U t Indicates the reward prediction y t The weight of the corresponding data sample. Therefore, for The posterior distribution of is as follows:
[0054]
[0055] For the exploration-exploitation balance problem, the upper confidence bound method is used, and the action in each round is represented as:
[0056]
[0057] Among them, D is the action space; β t Is a custom constant used to control the weight ratio between exploration and utilization; the mean Represents the utilization value of each action in the action space; the variance Represents the importance of exploration in action selection to obtain more information.
[0058] Furthermore, the target tracking module described in step 3 will update the region of interest to be tracked obtained in step 1. The target tracking module calculates the cumulative offset by calculating the cumulative coordinate offset between the result box and the tracking box of each region of interest, and calculates the normalized mutual correlation value by calculating the pixel similarity degree of each region of interest in the previous frame and the current frame.
[0059] Furthermore, determining whether to unload the region of interest to the edge server based on the offloading threshold in step 4 is as follows:
[0060] Cumulative pixel offset > cumulative pixel offset threshold or normalized mutual correlation value < normalized mutual correlation threshold → unload to the edge server for inference;
[0061] Cumulative pixel offset ≤ cumulative pixel offset threshold and normalized mutual correlation value ≥ normalized mutual correlation threshold → unload to the local terminal;
[0062] Furthermore, adaptively combining and packing the regions of interest to be unloaded in step 5 is to combine and pack them according to the packing size output by the decision module. Its implementation combines the two-dimensional bin packing method and the Guillotine algorithm. The basic process is the placement of elements in the set of regions of interest to be unloaded, the division of blank rectangles, and the maintenance of the remaining rectangle list.
[0063] First, use the packing size z output by the decision module t Initialize the remaining rectangle list S.
[0064] For the placement of the element r in the set of regions of interest to be unloaded, first find the remaining rectangle s in S that can accommodate r and add it to the set S r .
[0065] For the splitting of the blank rectangle, split the blank rectangle in the direction with the smaller height and width to obtain s' and s''
[0066] For the maintenance of the list of remaining rectangles, add the set of remaining rectangles {s', s''} after removing the combined position s to S
[0067] Furthermore, the update of the final inference result of the local terminal described in step 6 is determined according to whether there is an inference box from the edge server. If there is an inference box, the final inference result is the composite result of the terminal box and the inference box; if there is no inference box, the final inference result is the target tracking result box of the local terminal
[0068] Compared with the prior art, the advantages of the present invention are as follows
[0069] 1. Based on the reinforcement learning method, it can learn unknown situations from historical data. By considering the environmental context information, the algorithm can make better decisions
[0070] 2. Analyze and utilize the dependency relationship between the environment and system performance, capture time-varying characteristics in a dynamic time-varying environment, and can obtain better offloading decision-making strategies
[0071] 3. It can reduce the dependence on network resources and utilize heterogeneous edge-terminal resources. Therefore, it has lower requirements for the performance of local terminal devices and has good general deployability and network fluctuation adaptability BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is a schematic structural diagram of a method for offloading time-varying environment-aware video analysis tasks based on regions of interest according to an embodiment of the present invention
[0073] Figure 2 is a schematic implementation diagram of a method for offloading time-varying environment-aware video analysis tasks based on regions of interest according to an embodiment of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples
[0075] As Figure 1As shown in the figure, a time-varying environment perception video analysis task offloading method based on regions of interest, and a system structure implemented according to this method. The system consists of a local terminal device and an edge server. First, the camera captures video stream data and obtains video frames through a decoder. The first frame is uploaded to the edge server to generate an initialized tracking bounding box and a set of regions of interest to be tracked, and the results are sent back to the local terminal. The local terminal performs OpticalFlow tracking, calculates the cumulative offset and normalized cross-correlation value of the regions of interest at the same time. The decision-making module determines the packing size and offloading threshold of the regions of interest, and then makes an offloading judgment according to the decision result and performs adaptive combination packing of the regions of interest.
[0076] As Figure 2 shown, the system structure of the present invention based on this method will be actually deployed. Among them, the local terminal uses Raspberry Pi 4B, and the edge server uses Jetson Xavier.
[0077] The specific steps for implementing this method are as follows:
[0078] Step 1, the local video acquisition device captures video stream data and obtains video frames. The local terminal device can obtain video frames through methods such as video file reading, virtual camera reading, and real camera reading. Among them, the formats of video file reading include avi, mp4, etc. Virtual camera reading refers to the video frame data in the camera implemented by software, and real camera reading refers to the video frame data obtained by real-time reading of the camera hardware.
[0079] Then, the video frames are uploaded to the edge server to initialize the bounding box and the set of regions of interest to be tracked through Socket communication.
[0080] Step 2, a decision-making module is built on the local terminal to determine the packing size of the regions of interest, the cumulative pixel offset threshold of the regions of interest, and the normalized cross-correlation threshold of the regions of interest. The decision-making module makes decisions on the settings of the packing size of the regions of interest, the cumulative pixel offset threshold of the regions of interest, and the normalized cross-correlation threshold based on time-varying environment perception, considering the current state of the system, and finally balances the processing delay and the inference accuracy.
[0081] For the implementation of the decision-making module in Step 2, it combines the upper confidence bound method, Bayesian optimization, and Gaussian process based on temporal kernel functions. The specific implementation is as follows:
[0082] First, the time is decomposed into multiple rounds t∈{1,2,...,T} with fixed temporal segmentation. The decision-making module determines the adaptive parameter settings and offloading thresholds for the set of regions of interest in each round of video frames by observing the environmental context, and interacts with the environment to obtain the optimal decision result;
[0083] The main elements of the decision-making problem include actions, environmental context, rewards, and reward prediction;
[0084] Actions: The decision-making actions include the packed size of the region of interest, the cumulative offset threshold of the region of interest, and the normalized correlation threshold, which are defined as:
[0085] a t =(z t ,θ t ,n t ) (1)
[0086] where z t represents the packed size of the region of interest combination suitable for the current round, θ t represents the cumulative offset threshold setting for the current round, and n t represents the normalized correlation threshold for the current round. θ t and n t will jointly serve as the offloading threshold for subsequent determination of whether to offload.
[0087] Environmental context: The environmental context includes the processing delay and inference accuracy of the system, which are defined as:
[0088] c t =(D t-1 ,A t-1 ) (2)
[0089] where D t-1 represents the processing delay of the previous round, and A t-1 represents the inference accuracy of the previous round.
[0090] Rewards: In each round, the local terminal packs and combines and offloads the set of regions of interest generated from the video frame data according to the action a t . After completing the processing of the current round, the reward function is calculated to evaluate the current strategy. The reward function is defined as:
[0091] y t =1 / D t +α·A t (3)
[0092] where α is the weight parameter.
[0093] Reward prediction: At the initial stage of decision-making, the environmental prior knowledge is unknown, and the relationship between the action context within each round and the reward is unknown. Therefore, Bayesian optimization is used to predict the reward function. The reward prediction function under Bayesian optimization is defined as:
[0094] y t =f(x t )+∈ t (4)
[0095] Among them, x t =(a t , c t ) represents the action context of d within each round. represents random noise, that is, the difference between the actual reward and the observed reward. According to the Gaussian process, the mean function of the prior distribution of the reward prediction function is:
[0096]
[0097] Its covariance function is:
[0098]
[0099] Among them, x and x' are elements in the set X of action context pairs. Here, the initial prior distribution of the reward function f(·) is set as
[0100] For the update of the prior distribution of the reward prediction function, according to the Sherman-Morrison-Woodbury formula, the posterior distribution of the reward prediction function Y t =[y1,..., y t T still follows the Gaussian process, then the mean of the posterior distribution after each round is
[0101] μ t (x)=k t (x) T (K t +σ 2 Ι) -1 Y t (7)
[0102] The variance is
[0103]
[0104] Among them, the kernel function K t =[k x (x, x')], k t (x)=[k x (x1, x),..., k x (x t , x)], x, x' ∈ X.
[0105] To solve the problem that the update effect of the reward function decreases due to historical "stale" data, first, introduce the time dimension into the kernel function k x , and introduce the time series parameter t into the environmental context pair. Secondly, in the squared exponential kernel function k s Introduce the Ornstein-Uhlenbeck time covariance function k on the basis of o , specifically:
[0106]
[0107] Among them, λ controls the decay rate of the importance of historical data samples. The larger λ is, the faster the importance of historical data samples decays, and new data samples will be more important; l represents the length scale, which determines the degree of mutual influence between x and x'.
[0108] If the experience time series information interval has uniformity, then for The kernel matrix is Among them, is the Hadamard product; Among them, u t and U t represent the weights of the data samples corresponding to the reward prediction y t . Therefore, for The posterior distribution is as follows:
[0109]
[0110] For the exploration and exploitation balance problem, the upper confidence bound method is used, and the action in each round is expressed as:
[0111]
[0112] Among them, D is the action space; β t is a user-defined constant used to control the weight ratio between exploration and exploitation; the mean represents the exploitation value of each action in the action space; the variance represents the importance of exploring to obtain more information in action selection.
[0113] Step 3, the local terminal updates the set of regions of interest to be tracked obtained in Step 1, and calculates the cumulative offset d f,t and the normalized mutual correlation value m f,t .
[0114] Step 4, the local terminal determines whether to offload the region of interest to the edge server through Socket communication based on the offloading threshold as:
[0115] d f,t > θ t or m f,t < n t → Offload to the edge server for inference;
[0116] d f,t ≤ θ t and m f,t ≥ n t → Unload to the local terminal;
[0117] Step 5, the local terminal performs adaptive combined packing according to the ROI packing size result output by the decision module. Its implementation combines the two-dimensional bin packing method and the Guillotine algorithm. The basic process is the placement of elements in the set of ROIs to be unloaded, the splitting of blank rectangles, and the maintenance of the remaining rectangle list.
[0118] First, use the packing size z output by the decision module t to initialize the remaining rectangle list S.
[0119] For the placement of element r in the set of ROIs to be unloaded, first find the remaining rectangle s in S that can hold r and put it into the set S r in.
[0120] For the splitting of blank rectangles, split the blank rectangle in the direction with the smaller height and width to obtain s' and s''.
[0121] For the maintenance of the remaining rectangle list, add the set of remaining rectangles {s', s''} after removing the combined position s to S.
[0122] Step 6, the local terminal receives the inference result from the edge server and updates the final inference result. It is judged according to whether there is an inference box from the edge server. If there is an inference box, the final inference result is the composite result of the terminal box and the inference box; if there is no inference box, the final inference result is the target tracking result box of the local terminal.
[0123] The method according to the present invention described above can be implemented in hardware, firmware, or can be implemented as software or computer code that can be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or can be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded via a network and to be stored in a local recording medium, so that the method described herein can be stored on such a recording medium for software processing using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, a method for offloading time-varying environment-aware video analysis tasks based on a region of interest described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the processing shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the processing shown herein.
[0124] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention according to the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A time-varying environment-aware video analysis task offloading method based on regions of interest, characterized in that: The following steps are involved: Step 1: The video acquisition device captures video data and obtains video frames through the decoder, and then generates the initial bounding box and the area of interest to be tracked through the model deployed on the edge server; Step 2: construct a decision module on the local terminal to determine the packing size of the region of interest, the cumulative pixel offset threshold of the region of interest, and the normalized correlation threshold of the region of interest; The decision module is based on time-varying environment perception and considers the current state of the system to make decisions on the setting of the packing size of the region of interest, the cumulative pixel offset threshold of the region of interest, and the normalized correlation threshold, and finally balances the processing delay and the inference accuracy; The implementation of the decision module of the local terminal combines the confidence upper bound method, Bayesian optimization, and Gaussian process based on the time series kernel function. The specific implementation is as follows: First, the time is decomposed into multiple rounds t∈{1,2,...,T} with fixed time sequence. The decision module determines the adaptive parameter settings and unloading thresholds for the set of regions of interest in each round of video frames by observing the environmental context, and interacts with the environment to obtain the optimal decision result. The main elements of a decision-making problem include actions, environmental context, rewards, and reward predictions; Action: The decision action includes the ROI packing size, the ROI cumulative offset threshold, and the normalized cross-correlation threshold, which are defined as: a t =(z t ,θ t ,n t ) (1) Among them, z t Indicates the combined packing size of the region of interest suitable for the current round, θ t Indicates the cumulative offset threshold setting for the current round, n t represents the normalized cross-correlation threshold of the current round, θ t and n t They will be used together as the uninstall threshold for subsequent determination of whether to uninstall; Environmental context: Environmental context includes the system’s processing latency and inference accuracy, which is defined as: c t =(D t-1 ,A t-1 ) (2) Among them, D t-1 represents the processing delay of the previous round, A t-1 Indicates the reasoning accuracy of the previous round; Reward: In each round, the local terminal takes action a t The set of regions of interest generated by the video frame data is packaged, combined and unloaded. After completing the processing of the current round, the reward function is calculated to evaluate the current strategy. The reward function is defined as: y t =1 / D t +α·A t (3) Among them, α is the weight parameter; Reward prediction: In the initial stage of decision making, the prior knowledge of the environment is unknown, and the relationship between the action-context pair and the reward in each round is unknown, so Bayesian optimization is used to predict the reward function; the reward prediction function under Bayesian optimization is defined as: y t =f(x t )+∈ t (4) Among them, x t =(a t ,c t ) represents the action context of d in each round, represents random noise, that is, the difference between the actual reward and the observed reward; according to the Gaussian process, the mean function of the prior distribution of the reward prediction function is: Its covariance function is: Where x and x′ are elements in the action-context pair set X; here, the initial prior distribution of the reward function f(·) is set to For the update of the prior distribution of the reward prediction function, according to the Sherman-Morrison-Woodbury formula, the reward prediction function Y t =[y1,...,y t ] T The posterior distribution of still obeys the Gaussian process, then the mean of the posterior distribution of each round is: m t (x)=k t (x) T (K t +s 2 I) -1 Y t (7) The variance is: Among them, the kernel function K t =[k x (x,x′)],k t (x) = [k x (x1,x),...,k x (x t ,x)] T ,x,x′∈X; First, in the kernel function k x The time dimension is introduced into the environment context pair, and the timing parameter t is introduced into the environment context pair. Secondly, in the square exponential kernel function k s Based on this, the Ornstein-Uhlenbeck time covariance function k is introduced o , specifically: Among them, λ controls the rate of decrease of the importance of historical data samples. The larger the λ is, the faster the importance of historical data samples decreases, and the new data samples will be more important. l represents the length scale, which determines the degree of mutual influence between x and x′. Since the time series information interval is uniform, The kernel matrix is in, ⊙ is the Hardman product; in, u t with U t represents the reward prediction y t The weight of the corresponding data sample; so for The posterior distribution of is as follows: For the exploration-exploitation balance problem, the confidence upper bound method is used, and the actions of each round are expressed as: Where D is the action space; β t is a custom constant used to control the weight ratio between exploration kernel utilization; mean Indicates the utilization value of each action in the action space; variance Indicates the importance of exploration to obtain more information in action selection; Step 3: Use the target tracking module deployed on the terminal device to track and calculate the cumulative pixel offset value and normalized correlation value of the area of interest in real time; Step 4: The local terminal determines whether to offload the area of interest to the edge server through Socket communication based on the offloading threshold; Step 5: The local terminal adaptively combines and packages the interested regions to be uninstalled, and performs uninstallation according to the uninstallation judgment result; Step 6: The local terminal receives the inference result from the edge server and updates the final inference result of the local terminal.
2. The method for offloading time-varying environment-aware video analysis tasks based on regions of interest according to claim 1 is characterized in that: The video data described in step 1 comes from a real camera and a virtual camera, and the camera does not move during the acquisition process.
3. The method for offloading time-varying environment-aware video analysis tasks based on regions of interest according to claim 1, characterized in that: The target tracking module described in step 3 will update the region of interest to be tracked obtained in step 1. The target tracking module calculates the cumulative offset as the cumulative coordinate offset between the result box and the tracking box of each region of interest, and calculates the normalized correlation value as the calculation of the pixel similarity between each region of interest in the previous frame and the current frame.
4. The method for offloading tasks of time-varying environment-aware video analysis based on regions of interest according to claim 1, characterized in that: The method described in step 4 for determining whether to offload the area of interest to the edge server through Socket communication based on the offloading threshold is: Cumulative pixel offset > cumulative pixel offset threshold or normalized correlation value < normalized correlation threshold → offload to edge server for inference; Cumulative pixel offset ≤ cumulative pixel offset threshold and normalized correlation value ≥ normalized correlation threshold → offload to the local terminal.
5. The method for offloading time-varying environment-aware video analysis tasks based on regions of interest according to claim 1, characterized in that: The adaptive combined packing of the interest areas to be unloaded described in step 5 is to combine and pack according to the packing size output by the decision module. Its implementation combines the two-dimensional packing method and the Guillotine algorithm. The basic process is the placement of elements in the interest area set to be unloaded, the segmentation of blank rectangles, and the maintenance of the remaining rectangle list; First, use the packing size z output by the decision module t Initialize the remaining rectangle list S; For the placement of element r in the set of interest areas to be unloaded, first find the remaining rectangle s in S that can hold r and put it into the set S. r middle; For the segmentation of the blank rectangle, the blank rectangle is segmented in the direction with smaller height and width to obtain s′ and s″; To maintain the remaining rectangle list, the remaining rectangle set {s′, s″} excluding the combined position s is added to S.
6. The method for offloading tasks of time-varying environment-aware video analysis based on regions of interest according to claim 1, characterized in that: The final reasoning result of updating the local terminal in step 6 is determined based on whether there is a reasoning box from the edge server. If there is a reasoning box, the final reasoning result is a composite result of the terminal box and the reasoning box. If there is no inference box, the final inference result is the target tracking result box of the local terminal.
Citation Information
Patent Citations
Method and device for transmitting information of region of interest in video
CN111193928A
End-side cooperation-based context-aware online real-time video analysis method
CN116866660A