Camera edge video analysis acceleration method and system based on area tracking
By optimizing edge video analysis through a region tracking method, and by merging and scheduling boundary candidate regions and tracking candidate regions, the problem of low resource utilization efficiency in edge video analysis is solved, achieving more efficient computation and lower latency.
Patent Information
- Application Number
- CN202511163265.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies have low resource utilization efficiency in edge video analysis and cannot effectively handle the continuity between video frames, resulting in increased computing requirements and latency sensitivity, especially in multi-camera scenarios where resource bottlenecks are severe.
By employing a region tracking method, we utilize boundary candidate region identification and candidate region construction, and combine Kalman filtering and the maximum bounding rectangle algorithm to optimize the merging and scheduling of candidate regions, thereby achieving efficient utilization of edge server resources.
It significantly reduced latency by an average of 37.59% and improved accuracy by up to 7.30%, demonstrating better performance, especially when the number of video streams increases, and significantly improving resource utilization efficiency.
Smart Images

Figure CN120953325A_ABST
Abstract
Description
Technical Field
[0001] This invention proposes a method and system for accelerating camera edge video analysis based on region tracking, which relates to the field of video analysis. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence, deep neural networks (DNNs) have been widely applied to video analytics to support scenarios such as traffic monitoring and anomaly detection. Considering the following two factors, an increasing number of video analytics applications are being deployed and run at the edge: 1) Limited upload bandwidth from the edge to the remote cloud makes data transmission a bottleneck in the video analytics pipeline, particularly detrimental to latency-sensitive video analytics tasks; 2) Videos often contain private information, and high reliability cannot be guaranteed for remote data transmission. However, the contradiction between the limited resources of edge servers and the ever-increasing computational demands of video analytics severely hinders the efficiency of edge video analytics, and this contradiction is intensifying: 1) DNN inference is a computationally intensive task, especially for target detection using complex convolutional layers; 2) As the number of camera video streams increases, the required computational cost also increases; 3) Existing cameras typically output 4-8K video, placing higher demands on processing power.
[0003] Some works attempt to alleviate the resource constraints of edge systems to accelerate video analytics. For example, Zhang et al. detected downsampled video frames and used a Long Short-Term Memory (LSTM) network to extract candidate regions, reducing the area required for inference. Wang et al. used lightweight features to determine the complexity of frames, deciding which frames to skip and the model size for detection. Furthermore, some works configure and schedule video frames to ensure analytical accuracy and reduce response latency. However, existing works typically transform video analytics into the independent recognition and judgment of each frame, then use DNNs to perform complete feature extraction and filtering operations. This approach essentially continues the image processing paradigm, ignoring the continuity between video frames, making it difficult to overcome system performance bottlenecks in terms of resource utilization efficiency. Summary of the Invention
[0004] In view of this, in order to fill the gaps and deficiencies in the existing technology, this invention proposes a method and system for accelerating camera edge video analysis based on region tracking, which can effectively improve the resource optimization granularity of video analysis in edge systems.
[0005] This invention proposes a method and system for accelerating camera edge video analysis based on region tracking, including the following:
[0006] This invention proposes a method for accelerating camera edge video analysis based on region tracking, characterized by the following:
[0007] Step S1: Obtain the regions where old targets have appeared through video pre-analysis, and use the boundary candidate region recognition method to pre-analyze the movement trajectory of old targets in the video stream to form a heat map, thereby constructing boundary candidate regions;
[0008] Step S2: Construct a tracking candidate region for the new target based on its potential movement range, and expand the range of the tracking candidate region; including a tracking candidate region merging mechanism based on the maximum bounding rectangle algorithm and an update mechanism for constructing the tracking candidate region based on the target position prediction model of Kalman filter;
[0009] Step S3: The candidate region merging and scheduling algorithm based on shelf strategy is used to merge the boundary candidate regions and tracking candidate regions generated in different video scenes into several fusion regions of different sizes and then distribute them to different servers to achieve efficient utilization of edge server resources.
[0010] Further, step S1 includes the following:
[0011] Step S11: Track and analyze the areas where historical targets have entered and exited, including the following:
[0012] Once the camera is deployed to the designated location, target detection is first performed on the entire frame of the captured video stream. The positions of the first and last appearances of all historical targets in the video frame are marked to construct a heat map representing the frequency of target entry and exit in different areas. In the heat map, the more overlapping areas of targets, the higher their entry and exit frequency. Furthermore, the video frame is divided into several grid cells according to the resolution, and the boundary candidate regions formed by the grid cells are determined according to the temperature distribution in the heat map.
[0013] Step S12: When a target is lost, threshold filtering is used to eliminate this noise phenomenon; this includes calculating the cross-union ratio (CUP) of the target box with each grid cell based on the target box in the heatmap. If the cumulative CUP of a certain grid cell exceeds a specified overlap threshold, the grid cell will be marked as a potential boundary candidate region.
[0014] Step S13: After identifying all mesh elements, construct a set of candidate boundary regions, including the following:
[0015] S1={(i,j)∣IoU(i,j)>θ}
[0016] Where (i,j) represents the grid cell in the i-th row and j-th column, θ represents the preset overlap threshold, and IoU represents the intersection-union ratio; further, all adjacent or contacting region blocks in set S1 are merged to obtain the set of candidate connected boundary regions:
[0017] S2={C1,C2,…,C k}
[0018] Where C 1≤l≤k = {(i1,j1),(i2,j2),…}, where k represents the number of candidate regions for connected boundaries;
[0019] Step S14: Automatic identification of boundary candidate regions is achieved by pre-analyzing the video stream captured by the camera, including running a complete target detection model on the edge server to capture the location where the target first appears and last disappears.
[0020] Further, step S2 includes the following:
[0021] Step S21: Construct a tracking candidate region for the new target based on its potential range of movement. When the new target appears in the boundary candidate region, create a tracking candidate region for the new target.
[0022] Step S21: When a new target is detected within the boundary candidate region, a tracking candidate region is created for the new target to achieve continuous tracking of the new target; firstly, the current location of the new target is determined based on the grid cells it covers when it is first detected, and then the tracking candidate region is expanded with the current detection box as the center based on the possible movement range of the target in the next frame.
[0023] Step S22: After creating a tracking candidate region for the new target, target detection can be performed only on the tracking candidate region to obtain the detection result of the new target in the next frame;
[0024] Step S23: When the position of the new target deviates from the center of the tracking candidate region in the next frame, move around the new target box and expand the tracking candidate region to achieve continuous tracking of the target;
[0025] Furthermore, firstly, the tracking candidate region is displaced according to the coordinate movement distance of the target box between two frames; then, the area ratio of the target box to the tracking candidate region is calculated to ensure that it is always less than the preset overlap threshold, so as to realize the function of adaptively adjusting the tracking candidate region to cope with the position movement and area change of the target in the video frame; finally, the above process is repeated until the target disappears into the boundary candidate region, and the corresponding tracking candidate region is deleted.
[0026] Furthermore, step S2 also includes the following:
[0027] Step S24: When expanding the tracking candidate region, a tracking candidate region merging mechanism based on the maximum bounding rectangle is adopted to solve the problem of overlap between the tracking candidate region and the boundary candidate region, as well as the problem of mutual interference in the detection of tracking candidate regions caused by multiple targets being close together. This includes the following:
[0028] Step S241: Construct a set R of candidate regions to store the regions for which merging and tracking candidate regions and boundary candidate regions have been completed;
[0029] Furthermore, define the tracking candidate region set B:
[0030] B = {b1, b2, ..., b} N}
[0031] Where b i = [x,y,w,h] represents the x-coordinate, y-coordinate, width, and height of a tracking candidate region;
[0032] Step S242: Copy the candidate region set B to obtain the set B to be processed. remian The tracking candidate region merging algorithm starts from the candidate set B each time. remian Take out an unmerged region b from the data. c As the starting point for the current merge, a merge queue Q and its circumscribed rectangle boundary are constructed as follows:
[0033] [x min ,y min ,x max ,y max ]
[0034] Where, x min The x-coordinate of the rectangle's boundary is the minimum value. min The minimum ordinate of the rectangle's boundary is represented by x. max The x-coordinate of the rectangle's boundary represents the maximum value of the rectangle's boundary. min This represents the maximum value of the ordinate of the rectangle's boundary.
[0035] Furthermore, in region b, which has not yet been merged... c Based on this, select those whose IoU exceeds the threshold δ in sequence. IoU Other regions are included in the merge set, and the minimum bounding rectangle boundary of the current merged region is updated, thereby maintaining the tracking of the overall coverage of the current merged region. The update process of the bounding rectangle boundary is defined as follows:
[0036] x min+2 ←min(x min+1 ,b x )
[0037] y min+2 ←min(ymin+1 ,b y )
[0038] x max+2 ←min(x max+1 ,b x )
[0039] y max+2 ←min(y max+1 ,b y )
[0040] Where x min+1 The x-coordinate represents the minimum value of the rectangle's boundary before the update. min+2 This represents the minimum x-coordinate of the rectangle boundary before the update.
[0041] Where y min+1 This represents the minimum y-coordinate of the rectangle boundary before the update. min+2 This represents the minimum ordinate of the rectangle's boundary before the update.
[0042] Where x max+1 This represents the maximum x-coordinate of the rectangle boundary before the update. max+2 This represents the maximum x-coordinate of the rectangle boundary before the update.
[0043] Where y max+1 This represents the minimum y-coordinate of the rectangle boundary before the update. max+2 This represents the minimum ordinate of the rectangle's boundary before the update.
[0044] Step S243: After processing the current region, from B remian Remove all merged regions and add the smallest bounding rectangle of that group of regions as a new candidate region to the candidate region set R; continue this process until all candidate regions have been processed. Finally, all merged regions are added to the candidate merge region set R, resulting in the following expression for R:
[0045] R = {r1, r2, ... r} n}
[0046] The current region includes region b, which has not yet been merged. c Based on this, and selecting those with an IoU exceeding the threshold δ IoU Other areas and b c The merged region;
[0047] Where r1, r2, ... r n This indicates the candidate region after merging.
[0048] Furthermore, step S2 also includes the following:
[0049] Step S25: Real-time prediction of the new target location using a Kalman filter-based target location prediction model includes the following:
[0050] Step S251: Initialize the Kalman filter, where the input of the Kalman filter is the target position z of the previous frame. t-1 and state vector x t-1 The output of the Kalman filter is the predicted position of the new target in the current frame; furthermore, prediction is performed based on the state vector of the previous time step to obtain the predicted state at the current time step. and predicting covariance
[0051] If a new target is detected at position z in the current frame t If the Kalman filter is updated, then if no new target location is detected in the current frame, then it is further determined whether the current target is in the boundary candidate region.
[0052] Step S252: Further, if the IoU between the tracking and boundary candidate regions where the new target is located is less than the threshold δ, then perform position prediction and return. If the IoU between the tracking and the boundary candidate region where the new target is located is greater than or equal to the threshold δ, then the new target is determined to have disappeared in the current video frame;
[0053] Define the cumulative number of undetected instances, 'a'; if 'a' exceeds a set threshold A... max If the target is lost, the tracking of the target will end, and the corresponding tracking candidate area will be deleted.
[0054] Furthermore, step S3 also includes the following:
[0055] Step S31: First, obtain the available edge server status, denoted as set S, which includes the following:
[0056] S = s1, s2, ... s M
[0057] Where s i ={cap i ,load i}, cap i and load i These represent the server's maximum computing power and real-time load, respectively; further, s is calculated based on the real-time status. i The computing power coefficient is defined as follows:
[0058]
[0059] Where r i For si The computing power coefficient is M, where M is the number of edge servers.
[0060] Furthermore, step S3 also includes the following:
[0061] Step S32: Calculate the total area s of all candidate regions and set the server's real-time computing power coefficient r. k As an area ratio constraint, the target area of each fusion region is determined; further, the set of fusion regions G for each group is initialized. k and cumulative area A k Then, all candidate regions are sorted in descending order of area to ensure that larger regions are allocated first; subsequently, the sorted regions I are traversed. i' For each region, calculate its relationship with the target T after being assigned to each group. k The area difference between them △ k It is then assigned to the group with the largest current area difference, and the cumulative area A of that group is updated. j .
[0062] Furthermore, step S3 also includes the following:
[0063] Step S33: After completing the region segmentation, merge the regions in each group; further, estimate the side length of the merged region in the current group, and sort the regions within the group from largest to smallest height; then, arrange the regions sequentially using a marching shelf strategy; and record their positions (x, y) in the merged region for subsequent merging, and calculate the blank rate α of the splicing result; if α exceeds a set threshold... Then, the area of the candidate region with the largest area in the current fusion region is reduced to 90% of its original size. Finally, the set of all fusion regions that satisfy the ratio constraint is returned.
[0064] According to a second aspect of the present invention, a camera edge video analysis acceleration system based on region tracking includes an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements a camera edge video analysis acceleration method based on region tracking as described in any one of the present invention.
[0065] According to a third aspect of the present invention, a camera edge video analysis acceleration system based on region tracking includes a computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program implements a camera edge video analysis acceleration method based on region tracking as described in any one of the present invention.
[0066] The present invention has the following advantages:
[0067] (1) A boundary candidate region recognition method was designed. DNN inference oriented towards candidate regions in multi-camera scenes was explored and used as a new dimension to optimize real-time edge video analysis. By pre-analyzing the heatmap of the frequency of target entry and exit in the video scene, RegTrack can quickly and accurately capture new targets in the scene by running DNN inference in the boundary candidate region instead of the whole frame.
[0068] (2) A tracking candidate region construction and update mechanism was designed. RegTrack constructs tracking candidate regions for new targets based on their potential movement range and updates the candidate regions quickly by utilizing the target's positional movement between consecutive frames. To avoid target overlap and loss during tracking, a tracking candidate region merging mechanism based on the maximum bounding rectangle and a target position prediction model based on Kalman filtering were developed to ensure the accuracy of the candidate regions.
[0069] (3) A candidate region merging and scheduling algorithm based on a shelf strategy was designed. Specifically, based on available edge system resources and real-time load, all candidate regions are efficiently merged and then scheduled to the most suitable server for DNN inference. Simultaneously, to avoid potential blank areas consuming inference resources, significant target regions are identified and reasonably scaled during the merging process to reduce the actual inference area while maintaining video analysis accuracy.
[0070] (4) Based on real surveillance video datasets and camera video streams, extensive experiments have verified the effectiveness of the proposed RegTrack. Compared to benchmark methods, RegTrack reduces latency by an average of 37.59% and improves accuracy by up to 7.30%. In particular, RegTrack exhibits a more significant performance advantage as the number of video streams increases, and its components require very low time overhead. Attached Figure Description
[0071] Figure 1 This is a schematic diagram of the method steps of the present invention.
[0072] Figure 2 This is a schematic diagram of the system of the present invention.
[0073] Figure 3 This is a heat map of the target entry and exit frequency and a schematic diagram of the boundary candidate region for this invention.
[0074] Figure 4 This is a schematic diagram showing the overlap of different candidate regions in this invention.
[0075] Figure 5 This is a schematic diagram of the merging of candidate regions in this invention.
[0076] Figure 6This is a comparative diagram showing the different methods of the present invention using different pre-trained models.
[0077] Figure 7 This diagram illustrates the performance comparison of different methods of the present invention under different input frame sizes.
[0078] Figure 8 This is a schematic diagram illustrating the performance of the present invention at different expansion ratios.
[0079] Figure 9 This is a schematic diagram illustrating the performance of the present invention under different numbers of video streams. Detailed Implementation
[0080] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.
[0081] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0082] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0083] like Figure 1 and Figure 2 As shown, this invention proposes a method and system for accelerating camera edge video analysis based on region tracking, including the following:
[0084] like Figure 1 As shown, this invention proposes a method for accelerating camera edge video analysis based on region tracking, characterized by the following:
[0085] Step S1: Obtain the regions where old targets have appeared through video pre-analysis, and use the boundary candidate region recognition method to pre-analyze the movement trajectory of old targets in the video stream to form a heat map, thereby constructing boundary candidate regions;
[0086] Step S2: Construct a tracking candidate region for the new target based on its potential movement range, and expand the range of the tracking candidate region; including a tracking candidate region merging mechanism based on the maximum bounding rectangle algorithm and an update mechanism for constructing the tracking candidate region based on the target position prediction model of Kalman filter;
[0087] Step S3: The candidate region merging and scheduling algorithm based on shelf strategy is used to merge the boundary candidate regions and tracking candidate regions generated in different video scenes into several fusion regions of different sizes and then distribute them to different servers to achieve efficient utilization of edge server resources.
[0088] In one embodiment of the present invention, step S1 includes the following:
[0089] Step S11: Track and analyze the areas where historical targets have entered and exited, including the following:
[0090] Once the camera is deployed to the designated location, target detection is first performed on the entire frame of the captured video stream. The positions of the first and last appearances of all historical targets in the video frame are marked to construct a heat map representing the frequency of target entry and exit in different areas. In the heat map, the more overlapping areas of targets, the higher their entry and exit frequency. Furthermore, the video frame is divided into several grid cells according to the resolution, and the boundary candidate regions formed by the grid cells are determined according to the temperature distribution in the heat map.
[0091] like Figure 3 As shown, in one embodiment of the present invention, as Figure 3 As shown in (a), the positions of the first and last appearances of all targets in the video frame are marked to construct a heatmap representing the frequency of target entry and exit in different regions. In the heatmap, the more overlapping regions a target is in, the higher its entry and exit frequency. Figure 3 As shown in (b), the dashed box illustrates the constructed boundary candidate region. This invention divides the video frame into several grid cells according to its resolution and determines the boundary candidate region composed of these grid cells based on the temperature distribution in the heatmap. Since 100% model accuracy cannot be guaranteed, individual targets may suddenly appear or disappear during target detection and tracking. The method described in this invention uses threshold filtering to eliminate this noise phenomenon.
[0092] Step S12: When a target is lost, threshold filtering is used to eliminate this noise phenomenon; this includes calculating the cross-union ratio (CUP) of the target box with each grid cell based on the target box in the heatmap. If the cumulative CUP of a certain grid cell exceeds a specified overlap threshold, the grid cell will be marked as a potential boundary candidate region.
[0093] Step S13: After identifying all mesh elements, construct a set of candidate boundary regions, including the following:
[0094] S1={(i,j)∣IoU(i,j)>θ}
[0095] Where (i,j) represents the grid cell in the i-th row and j-th column, θ represents the preset overlap threshold, and IoU represents the intersection-union ratio; further, all adjacent or contacting region blocks in set S1 are merged to obtain the set of candidate connected boundary regions:
[0096] S2={C1,C2,…,C k}
[0097] Where C 1≤l≤k = {(i1,j1),(i2,j2),…}, where k represents the number of candidate regions for connected boundaries;
[0098] Furthermore, l represents the l-th connected boundary candidate region, which is the l-th element in S2, and k represents the number of connected boundary candidate regions, which is the number of elements in S2.
[0099] Step S14: Automatic identification of boundary candidate regions is achieved by pre-analyzing the video stream captured by the camera, including running a complete target detection model on the edge server to capture the location where the target first appears and last disappears.
[0100] In one embodiment of the present invention, step S2 includes the following:
[0101] Step S21: Construct a tracking candidate region for the new target based on its potential range of movement. When the new target appears in the boundary candidate region, create a tracking candidate region for the new target.
[0102] Step S21: When a new target is detected within the boundary candidate region, a tracking candidate region is created for the new target to achieve continuous tracking of the new target; firstly, the current location of the new target is determined based on the grid cells it covers when it is first detected, and then the tracking candidate region is expanded with the current detection box as the center based on the possible movement range of the target in the next frame.
[0103] Step S22: After creating a tracking candidate region for the new target, target detection can be performed only on the tracking candidate region to obtain the detection result of the new target in the next frame;
[0104] Step S23: When the position of the new target deviates from the center of the tracking candidate region in the next frame, move around the new target box and expand the tracking candidate region to achieve continuous tracking of the target;
[0105] like Figure 4 As shown, in one embodiment of the present invention, when expanding the tracking candidate region, the following two overlap cases need to be considered:
[0106] (1) When the tracking candidate region overlaps with the boundary candidate region, the tracking candidate region may be repeatedly created for the existing target during the new target capture process. For example... Figure 4 As shown in (a), the dashed box represents the constructed tracking candidate region, and the solid box represents the final target detection result. When a target first appears in the boundary candidate region, the method described in this invention creates a tracking candidate region for it. In subsequent video frames, since the target is still within the boundary candidate region, the system attempts to repeatedly create tracking candidate regions for it. This can cause the tracking candidate region to overlap with the boundary candidate region, resulting in the target being detected in both regions simultaneously. Ultimately, this can lead to the same target having two different detection boxes in the original video frame. If this overlap is not handled properly, new tracking candidate regions will be continuously created during the new target capture process until the target completely leaves the boundary candidate region. This repeated detection not only causes unnecessary system overhead but also seriously affects the accuracy of region tracking.
[0107] (2) When multiple targets are close together, their respective tracking candidate regions may detect each other. For example... Figure 4 As shown in (b), when the method described in this invention creates a tracking candidate region for target A, it overlaps to some extent with the tracking candidate region for target B. This causes the target B's detection result to be obtained when the tracking candidate region is fed into the target detection model for inference. Similarly, the target A's detection result may also be output when the tracking candidate region of target B is inferred. The above situation leads to repeated detection of the same target, and this situation becomes more and more serious as the number of overlapping targets increases.
[0108] Furthermore, in one embodiment of the present invention, the tracking candidate region is first displaced according to the coordinate movement distance of the target box between two frames; then, the area ratio of the target box to the tracking candidate region is calculated to ensure that it is always less than a preset overlap threshold, thereby realizing the function of adaptively adjusting the tracking candidate region to cope with the position movement and area change of the target in the video frame as described in the present invention; finally, the method of the present invention repeats the above process until the target disappears in the boundary candidate region, and then the corresponding tracking candidate region is deleted.
[0109] In one embodiment of the present invention, step S2 further includes the following:
[0110] Step S24: When expanding the tracking candidate region, a tracking candidate region merging mechanism based on the maximum bounding rectangle is adopted to solve the problem of overlap between the tracking candidate region and the boundary candidate region, as well as the problem of mutual interference in the detection of tracking candidate regions caused by multiple targets being close together. This includes the following:
[0111] In one embodiment of the present invention, the algorithm 1 of the tracking candidate region merging mechanism of the maximum bounding rectangle is shown in Table 1 below.
[0112] Table 1. Algorithm 1: Merging of Tracking Candidate Regions Based on the Maximum Bounding Rectangle.
[0113]
[0114] Furthermore, the specific implementation steps of Algorithm 1 based on the merging of tracking candidate regions according to the maximum bounding rectangle, as shown in Table 1, include the following:
[0115] Step S241: Construct a set R of candidate regions to store the regions for which merging and tracking candidate regions and boundary candidate regions have been completed;
[0116] In one embodiment of the present invention, a set of tracking candidate regions B is defined:
[0117] B = {b1, b2, ..., b} N}
[0118] Where b i = [x,y,w,h], representing the location, width, and height of a tracking candidate region;
[0119] Step S242: Copy the candidate region set B to obtain the set B to be processed. remian The tracking candidate region merging algorithm starts from the candidate set B each time. remian Take out an unmerged region b from the data. c As the starting point for the current merge, a merge queue Q and its circumscribed rectangle boundary are constructed as follows:
[0120] [x min ,y min ,x max ,y max ]
[0121] Where, x min The x-coordinate of the rectangle's boundary is the minimum value. min The minimum ordinate of the rectangle's boundary is represented by x. max The x-coordinate of the rectangle's boundary represents the maximum value of the rectangle's boundary. min This represents the maximum value of the ordinate of the rectangle's boundary.
[0122] Furthermore, in region b, which has not yet been merged... c Based on this, select those whose IoU exceeds the threshold δ in sequence. IoU Other regions are included in the merge set, and the minimum bounding rectangle boundary of the current merged region is updated, thereby maintaining the tracking of the overall coverage of the current merged region. The update process of the bounding rectangle boundary is defined as follows:
[0123] x min+2 ←min(x min+1 ,b x )
[0124] y min+2 ←min(y min+1 ,b y )
[0125] x max+2 ←min(x max+1 ,b x )
[0126] y max+2 ←min(y max+1 ,b y )
[0127] Where x min+1 The x-coordinate represents the minimum value of the rectangle's boundary before the update. min+2 This represents the minimum x-coordinate of the rectangle boundary before the update.
[0128] Where y min+1 This represents the minimum y-coordinate of the rectangle boundary before the update. min+2 This represents the minimum ordinate of the rectangle's boundary before the update.
[0129] Where x max+1 This represents the maximum x-coordinate of the rectangle boundary before the update. max+2 This represents the maximum x-coordinate of the rectangle boundary before the update.
[0130] Where y max+1 This represents the minimum y-coordinate of the rectangle boundary before the update. max+2 This represents the minimum ordinate of the rectangle's boundary before the update.
[0131] Step S243: After processing the current region, from B remian Remove all merged regions and add the smallest bounding rectangle of that group of regions as a new candidate region to the candidate region set R; continue this process until all candidate regions have been processed. Finally, all merged regions are added to the candidate merge region set R, resulting in the following expression for R:
[0132] R = {r1, r2, ... r} n}
[0133] The current region includes region b, which has not yet been merged. c Based on this, and selecting those with an IoU exceeding the threshold δ IoU Other areas and b c The merged region;
[0134] Where r1, r2, ... r n This indicates the candidate region after merging.
[0135] In one embodiment of the present invention, the effect of merging candidate regions is as follows: Figure 5 As shown.
[0136] In one embodiment of the present invention, step S2 further includes the following:
[0137] Step S25: Real-time prediction of the new target location using a Kalman filter-based target location prediction model includes the following:
[0138] The target position prediction algorithm based on Kalman filtering is shown in Table 2.
[0139] Table 2 Algorithm 2 Target Position Prediction Based on Kalman Filtering
[0140]
[0141]
[0142] Furthermore, the specific implementation of Algorithm 2 for target position prediction based on Kalman filtering in this invention, according to Table 2, includes the following:
[0143] Step S251: Initialize the Kalman filter, where the input of the Kalman filter is the target position z of the previous frame. t-1 and state vector x t-1 The output of the Kalman filter is the predicted position of the new target in the current frame; furthermore, prediction is performed based on the state vector of the previous time step to obtain the predicted state at the current time step. and predicting covariance
[0144] If a new target is detected at position z in the current frame t If the Kalman filter is updated, then if no new target location is detected in the current frame, then it is further determined whether the current target is in the boundary candidate region.
[0145] Step S252: Further, if the IoU between the tracking and boundary candidate regions where the new target is located is less than the threshold δ, then perform position prediction and return. If the IoU between the tracking and the boundary candidate region where the new target is located is greater than or equal to the threshold δ, then the new target is determined to have disappeared in the current video frame;
[0146] Define the cumulative number of undetected instances, 'a'; if 'a' exceeds a set threshold A... max If the target is lost, the tracking of the target will end, and the corresponding tracking candidate area will be deleted.
[0147] In one embodiment of the present invention, step S3 further includes the following:
[0148] Step S31: RegTrack first obtains the available edge server status, denoted as set S, which includes the following:
[0149] S = s1, s2, ... s M
[0150] Where s i ={cap i ,load i}, cap i and load i These represent the server's maximum computing power and real-time load, respectively; further, s is calculated based on the real-time status. i The computing power coefficient is defined as follows:
[0151]
[0152] Where r i For s i The computing power coefficient is M, where M is the number of edge servers.
[0153] The candidate region merging and scheduling algorithm based on shelf strategy is shown in Table 3.
[0154] Table 3. Candidate Region Merging and Scheduling Based on Shelf Strategy in Algorithm 3
[0155]
[0156]
[0157]
[0158] As shown in Table 3, the specific implementation of Algorithm 3 for candidate region merging and scheduling based on shelf strategy in this invention includes the following:
[0159] Step S32: Calculate the total area s of all candidate regions and set the server's real-time computing power coefficient r. k As an area ratio constraint, the target area of each fusion region is determined; further, the set of fusion regions G for each group is initialized. k and cumulative area A k Then, all candidate regions are sorted in descending order of area to ensure that larger regions are allocated first; subsequently, the sorted regions I are traversed. i' For each region, calculate its relationship with the target T after being assigned to each group. k The area difference between them △ kIt is then assigned to the group with the largest current area difference, and the cumulative area A of that group is updated. j .
[0160] In one embodiment of the present invention, step S3 further includes the following:
[0161] Step S33: After completing the regional division, merge the regions in each group; further, estimate the current...
[0162] The side length of the merged region is determined, and the regions within the group are sorted from largest to smallest height. Then, a progressive shelf strategy is used to arrange the regions sequentially. The position (x, y) of each region within the merged region is recorded for subsequent merging, and the blanking rate α of the spliced result is calculated. If α exceeds a set threshold... Then the area of the candidate region with the largest area in the current fusion region is reduced to 90% of its original size. Finally, the set of all fusion regions that meet the proportional constraints is returned.
[0163] This invention proposes a camera edge video analysis acceleration system based on region tracking, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements a camera edge video analysis acceleration method based on region tracking as described in any one of the present invention.
[0164] This invention proposes a camera edge video analysis acceleration system based on region tracking, comprising a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements a camera edge video analysis acceleration method based on region tracking as described in any one of the present invention.
[0165] In addition to the above, the present invention also includes some embodiments for verifying the technical solutions proposed in the present invention, including the following:
[0166] To verify the feasibility of the proposed method and system for accelerating camera edge video analysis based on region tracking, RegTrack is constructed as the implementation framework for the method and system. RegTrack includes the following:
[0167] RegTrack is an implementation framework for a region-tracking-based camera edge video analysis acceleration method and system described in this invention, which can improve the resource optimization granularity of video analysis in edge systems. First, by pre-analyzing the heatmap of target entry and exit frequencies in the video stream, a boundary candidate region identification method is designed to achieve efficient capture of new targets. Next, tracking candidate regions are constructed for the captured targets to continuously track them and minimize the area requiring inference. Subsequently, a tracking candidate region merging mechanism based on the maximum bounding rectangle and a target position prediction model based on Kalman filtering are developed to effectively handle target overlap and loss problems. Finally, a candidate region merging and scheduling algorithm based on a shelf strategy is designed to fully utilize edge system resources and minimize video analysis response latency. Extensive experiments based on real surveillance video datasets and camera video streams validate the effectiveness of the proposed RegTrack. Compared with benchmark methods, RegTrack reduces latency by an average of 37.59% and improves accuracy by up to 10.10%. In particular, RegTrack's performance can be further improved with the increase in the number of devices, and its components incur only very low time overhead. Therefore, RegTrack is completely different from existing tracking technologies.
[0168] This invention uses the following three experimental datasets to comprehensively evaluate the performance of the proposed RegTrack:
[0169] (1) MOT: This dataset contains large-scale labeled data for multi-object tracking, in which 22 different indoor and outdoor scenes were recorded under different camera motions, camera angles and imaging conditions for pedestrian tracking.
[0170] (2) UA-DETRAC: This dataset contains large-scale multi-target vehicle detection data, taking into account 4 types of vehicles and 4 weather conditions.
[0171] (3) Real-Camera: To further test the effectiveness of RegTrack in actual video surveillance, this invention collected videos recorded by real cameras deployed in a laboratory environment. Unlike the previous two datasets, the target density in this dataset is sparser. To avoid unnecessary interference caused by targetless video frames, videos were collected for three time periods: 11:00-11:30 AM, 1:30-2:00 PM, and 5:00-5:30 PM. Furthermore, YOLOv8X was used to preprocess the collected videos, filtering out targetless video frames to ensure that all frames contained at least one target. Simultaneously, the detection results of YOLOv8X on each frame were used as labels to evaluate detection accuracy.
[0172] This invention uses three pre-trained DNN models—YOLOv8S, YOLOv8L, and ResDet50 (i.e., EfficientDet with ResNet50 as the backbone network)—to perform object detection. By default, the evaluation scheme uses the MOT dataset and YOLOv8S as the object detection dataset and model, respectively. Inference latency is measured using average frame inference latency, and the following three key metrics are used to evaluate the accuracy of object detection and tracking:
[0173] (1) Precision: Measures the target detection accuracy of the model, which is defined as follows:
[0174]
[0175] Here, TP and FP represent the number of correctly and incorrectly detected targets, respectively.
[0176] (2) Recall: Measures the model's ability to detect targets, defined as follows:
[0177]
[0178] Where FN represents the number of targets missed.
[0179] (3) MOTA (Multiple Object Tracking Accuracy): Measures the accuracy of multi-object tracking, considering three types of errors: missed detection, false detection, and ID switching. It is defined as follows:
[0180]
[0181] Here, IDSW represents the number of times the target identity changes during the tracking process, and GT represents the total number of real targets.
[0182] This invention uses PyTorch and OpenCV to build the proposed video analysis system and implements the components in RegTrack. Two servers were used to build the edge system: a host machine with an Intel i5 12400KF CPU, 32GB RAM, and an RTX 4060, and a workstation with an AMD Ryzen 9 9950x CPU, 64GB RAM, and an RTX 4090D. For the parameters in RegTrack, the expansion ratio of the tracking candidate region was set to 20%, the maximum forget time for target loss was 15 frames, the maximum IoU for metrics such as Precision was 0.5, the blank rate detection trigger threshold was 0.3, the other IoU overlap thresholds were 0.3, and the minimum allowable scaling area was 15,000 pixels. The GPU utilization and available video memory during server runtime were tracked using Nvidia's pynvml.
[0183] This invention uses PyTorch and OpenCV to build the proposed video analysis system and implement the components in RegTrack. Two servers were used to build the edge system, including a host with an Intel i5 12400KF CPU, 32GB RAM, and an RTX 4060, and a workstation with an AMD Ryzen 9 9950x CPU, 64GB RAM, and an RTX 4090D. For the parameters in RegTrack, the expansion ratio of the tracking candidate region was set to 20%, the maximum forgetting time for target loss was 15 frames, the maximum IoU for metrics such as Precision was 0.5, the blank rate detection trigger threshold was 0.3, the other IoU overlap thresholds were 0.3, and the minimum allowable scaling area was 15,000 pixels. The GPU utilization and available video memory of the server were tracked using Nvidia's pynvml. Performance comparisons of different methods on different datasets are shown in Table 4.
[0184] Table 4 Performance comparison of different methods on different datasets
[0185]
[0186]
[0187] To verify the superiority of the proposed RegTrack, this invention provides a comprehensive comparison with the following benchmark methods:
[0188] (1) Full-Frame Detection (FFD): Detects all video frames completely.
[0189] Input into the object detection model for video analysis.
[0190] (2) GMM-Guided Detection (GGD): The background remover based on the Gaussian Mixture Model (GMM) of OpenCV is used to extract RoIs from the video frames, and then all RoIs are input into the target detection model in the output order for video analysis.
[0191] (3) Elf: Performs a downsampling-based full-frame detection every few frames to capture newly emerging targets, and uses LSTM to predict the target's movement trajectory to expand the tracking candidate region. At the same time, the tracking candidate region is scheduled to different servers for load balancing based on the number and performance of edge servers.
[0192] (4) Gecko: Detects the pixel variation between two frames and uses a feature extractor to select the optimal model for inference for different frames. When the variation between two frames is small, the frame skip controller adjusts the detection interval of the video stream and uses a lightweight model for inference. When the variation between two frames is large, the model with the highest accuracy is selected for inference. YOLOv8S and YOLOv8L are used to implement the lightweight and highest accuracy models, respectively.
[0193] First, this invention evaluates the latency and accuracy of different methods on different datasets. As shown in Table I, all methods achieved the lowest and highest accuracy on the UADET and Real-Camera datasets, respectively. This is because the UADET dataset has a denser number of targets in the video scenes compared to the other two datasets, especially at intersections where there may be dozens of vehicles. Additionally, due to vehicle movement, there are more small targets. In contrast, targets in most scenes of the MOT dataset are easier to identify. The Real-Camera dataset collected in this invention mainly contains single targets, and these targets are relatively large; therefore, different methods can achieve higher accuracy on this dataset. The accuracy and latency of FFD are both at the median of the five methods, which this invention uses as a benchmark to analyze other methods.
[0194] Specifically, GGD exhibits the highest latency because its background removal algorithm performs indiscriminate pixel comparisons for each frame and performs serial inference on the RoI sequence. Even at low target densities, this RoI extraction strategy still incurs high overhead during RoI extraction. At high target densities, this serial inference strategy cannot fully utilize the parallel capabilities of the GPU, resulting in high latency. Furthermore, GGD is directly affected by the accuracy of the background removal algorithm; if it fails to capture target changes between two frames or filters out important target contours, it will directly lead to the loss of subsequent target detection. Therefore, GGD demonstrates lower accuracy across multiple metrics.
[0195] Besides the RegTrack proposed in this invention, Elf achieves the lowest latency because it uses LSTM to predict target movement trajectories to construct tracking candidate regions, thereby enabling frame content filtering. Simultaneously, Elf can schedule different tracking candidate regions based on edge server status, thus efficiently utilizing available edge resources. However, this interval sampling and prediction-based region extraction strategy struggles to respond promptly to sudden events in video frames, which are often more critical during target monitoring. Therefore, although Elf achieves lower latency, its accuracy is significantly lower than FFD in most scenarios, making it difficult to meet the high accuracy requirements of practical video analytics systems.
[0196] Gecko achieved accuracy second only to RegTrack, and all three metrics showed high accuracy on the UADET dataset, thanks to its dynamic model selection mechanism. When complex scenes are detected, Gecko adjusts the precision model for video analysis to achieve better performance. To balance latency, Gecko uses a lightweight model and dynamic frame intervals for object detection in scenes with fewer targets. However, the switching between high-precision and lightweight models introduces latency, especially in scenes with rapidly changing target density, resulting in higher latency for Gecko compared to most other methods. Furthermore, similar to the interval downsampling strategy used by Elf, dynamic frame interval inference may not respond promptly to fast-moving targets, potentially leading to a decrease in accuracy.
[0197] Among all methods, the proposed RegTrack achieves the lowest latency and exhibits the highest accuracy across various datasets. This is because RegTrack employs a shelf-based candidate region merging and scheduling algorithm, which efficiently merges candidate regions from multiple video streams and fully utilizes edge server resources for inference, thus achieving the lowest latency. Furthermore, RegTrack outperforms FFD in most scenarios. This is because FFD may experience accuracy degradation for targets with small areas or similar to the background, as it may fail to detect these targets. In contrast, RegTrack, through its designed Kalman filter-based target location prediction algorithm, can avoid accuracy degradation even in cases of target loss by predicting target trajectories. Moreover, compared to Elf and Gecko, RegTrack also achieves efficient adaptation to fast-moving targets, thus avoiding accuracy degradation caused by the sudden appearance or rapid movement of targets.
[0198] Next, the present invention comprehensively evaluates the generalization of the proposed RegTrack by changing the model type, model size, and input frame size. Figure 6We compared the mean time-of-flight (MOTA) and latency of different methods on pre-trained models such as YOLOv8S, YOLOv8L, and ResDet50. Since Gecko employs a dynamic model selection strategy, its performance was not evaluated in this set of experiments. RegTrack significantly outperformed GGD and Elf in MOTA, was comparable to FFD, and surpassed FFD in YOLOv8L. Meanwhile, RegTrack achieved the lowest latency, particularly significantly lower than FFD in YOLOv8L and ResDet50. This indicates that for more complex models, the latency reduction effect brought by the region area saved by RegTrack is more significant. Furthermore, different model sizes have the greatest impact on GGD, and the least impact on RegTrack and Elf. This is because GGD executes inference of multiple candidate regions serially, and the increase in model complexity has the most significant impact on it. In contrast, RegTrack and Elf have better adaptability by scheduling candidate regions to different servers.
[0199] Figure 7 The impact of different input frame sizes on the performance of various methods was compared. As the frame size increased, the MOTA and latency of all methods increased. This is because a larger input resolution allows the model to capture more details, thus more accurately capturing the target, but it also requires more computational resources for inference, leading to higher latency. Specifically, when the input frame size increased to 720, FFD, Gecko, and the proposed RegTrack achieved comparable accuracy, significantly higher than GGD and Elf. However, due to the use of a more complex model for video analysis and inadequate candidate region scheduling, Gecko's latency was significantly higher than other methods except GGD, and continued to increase with increasing frame size. Thanks to efficient scheduling strategies, RegTrack and Elf maintained consistently low latency, showing only a slight increase with increasing frame size.
[0200] Figure 8 The impact of the tracking candidate region expansion ratio on RegTrack performance was evaluated. When the expansion ratio increased from 10% to 20%, all accuracy metrics and latency increased. This is because for RegTrack, a larger expansion ratio allows for more effective target acquisition, but also leads to increased latency due to the larger inference area. When the expansion ratio increased from 20% to 50%, all accuracy metrics showed a slight increase, while latency continued to increase proportionally. This is because at this point, the tracking candidate region can fully cover the target's possible range of movement, and further expansion of its area does not bring a significant improvement in accuracy. However, latency continues to increase due to the expansion of the area.
[0201] Figure 9The performance of RegTrack was evaluated under different numbers of video streams to verify its good system scalability in large-scale scenarios. As the number of video streams increases, the latency of both RegTrack and FFD increases accordingly. However, the latency required by RegTrack is consistently lower than that of FFD, and the gap between the two continues to widen. Meanwhile, the accuracy error between the two remains less than 2%. This is because RegTrack can effectively filter candidate regions from different video streams, then merge these regions into multiple whole frames and schedule them to different servers to accelerate inference. Experimental results verify the scalability of the proposed RegTrack in large-scale video stream scenarios and its effectiveness in accelerating video analysis.
[0202] Table 5 tests the additional time overhead introduced by the RegTrack core components. First, the average time overhead for filtering overlapping regions based on the maximum bounding rectangle is only 0.01ms, which can quickly identify overlapping regions and merge dense target regions, thereby avoiding the construction of duplicate tracking candidate regions. Second, target position prediction takes only 0.51ms. The Kalman filter is updated when there is a target box output, and the target trajectory is predicted when there is no target box output, so as to avoid the decrease in accuracy caused by target loss. Finally, the multi-scale probing progressive merging merges candidate regions from all cameras and efficiently distributes them to different devices for analysis, which only incurs a time overhead of 5.07ms. Therefore, the time overhead introduced by the RegTrack core components will not place a significant burden on the video analysis system.
[0203] Table 5 Time Cost of RegTrack Core Components
[0204] Tracking candidate region merging Target location prediction Candidate region merging and scheduling 0.01ms 0.51ms 5.07ms
[0205] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for accelerating camera edge video analysis based on region tracking, characterized in that, Includes the following: Step S1: Obtain the regions where old targets have appeared through video pre-analysis, and use the boundary candidate region recognition method to pre-analyze the movement trajectory of old targets in the video stream to form a heat map, thereby constructing boundary candidate regions; Step S2: Construct a tracking candidate region for the new target based on its potential movement range, and expand the range of the tracking candidate region; including a tracking candidate region merging mechanism based on the maximum bounding rectangle algorithm and an update mechanism for constructing the tracking candidate region based on the target position prediction model of Kalman filter; Step S3: The candidate region merging and scheduling algorithm based on shelf strategy is used to merge the boundary candidate regions and tracking candidate regions generated in different video scenes into several fusion regions of different sizes and then distribute them to different servers to achieve efficient utilization of edge server resources.
2. The method for accelerating camera edge video analysis based on region tracking according to claim 1, characterized in that, Step S1 includes the following: Step S11: Track and analyze the areas where historical targets have entered and exited, including the following: Once the camera is deployed to the designated location, target detection is first performed on the entire frame of the captured video stream. The positions of the first and last appearances of all historical targets in the video frame are marked to construct a heat map representing the frequency of target entry and exit in different areas. In the heat map, the more overlapping areas of targets, the higher their entry and exit frequency. Furthermore, the video frame is divided into several grid cells according to the resolution, and the boundary candidate regions formed by the grid cells are determined according to the temperature distribution in the heat map. Step S12: When a target is lost, threshold filtering is used to eliminate this noise phenomenon; this includes calculating the cross-union ratio (CUP) of the target box with each grid cell based on the target box in the heatmap. If the cumulative CUP of a certain grid cell exceeds a specified overlap threshold, the grid cell will be marked as a potential boundary candidate region. Step S13: After identifying all mesh elements, construct a set of candidate boundary regions, including the following: S1={(i,j)∣IoU(i,j)>θ} Where (i,j) represents the grid cell in the i-th row and j-th column, θ represents the preset overlap threshold, and IoU represents the intersection-union ratio; further, all adjacent or contacting region blocks in set S1 are merged to obtain the set of candidate connected boundary regions: S2={C1,C2,…,C k } Where C 1≤l≤k = {(i1,j1),(i2,j2),…}, where k represents the number of candidate regions for connected boundaries; Step S14: Automatic identification of boundary candidate regions is achieved by pre-analyzing the video stream captured by the camera, including running a complete target detection model on the edge server to capture the location where the target first appears and last disappears.
3. The method for accelerating camera edge video analysis based on region tracking according to claim 2, characterized in that, Step S2 includes the following: Step S21: Construct a tracking candidate region for the new target based on its potential range of movement. When the new target appears in the boundary candidate region, create a tracking candidate region for the new target. Step S21: When a new target is detected within the boundary candidate region, a tracking candidate region is created for the new target to achieve continuous tracking of the new target; firstly, the current location of the new target is determined based on the grid cells it covers when it is first detected, and then the tracking candidate region is expanded with the current detection box as the center based on the possible movement range of the target in the next frame. Step S22: After creating a tracking candidate region for the new target, target detection can be performed only on the tracking candidate region to obtain the detection result of the new target in the next frame; Step S23: When the position of the new target deviates from the center of the tracking candidate region in the next frame, move around the new target box and expand the tracking candidate region to achieve continuous tracking of the target; Furthermore, firstly, the tracking candidate region is displaced according to the coordinate movement distance of the target box between two frames; then, the area ratio of the target box to the tracking candidate region is calculated to ensure that it is always less than the preset overlap threshold, so as to realize the function of adaptively adjusting the tracking candidate region to cope with the position movement and area change of the target in the video frame; finally, the above process is repeated until the target disappears into the boundary candidate region, and the corresponding tracking candidate region is deleted.
4. The method for accelerating camera edge video analysis based on region tracking according to claim 3, characterized in that, Step S2 also includes the following: Step S24: When expanding the tracking candidate region, a tracking candidate region merging mechanism based on the maximum bounding rectangle is adopted to solve the problem of overlap between the tracking candidate region and the boundary candidate region, as well as the problem of mutual interference in the detection of tracking candidate regions caused by multiple targets being close together. This includes the following: Step S241: Construct a set R of candidate regions to store regions that have completed merging and tracking candidate regions and boundary candidate regions; Furthermore, define the tracking candidate region set B: B={b1,b2,…,b N } Where b i = [x,y,w,h] represents the x-coordinate, y-coordinate, width, and height of a tracking candidate region; Step S242: Copy the candidate region set B to obtain the set B to be processed. remian The tracking candidate region merging algorithm starts from the candidate set B each time. remian Take out an unmerged region b from the data. c As the starting point for the current merge, a merge queue Q and its circumscribed rectangle boundary are constructed as follows: (x min ,and min ,x max ,and max ] Where, x min The x-coordinate of the rectangle's boundary is the minimum value. min The minimum ordinate of the rectangle's boundary is represented by x. max The x-coordinate of the rectangle's boundary represents the maximum value of the rectangle's boundary. min This represents the maximum value of the ordinate of the rectangle's boundary. Furthermore, in region b, which has not yet been merged... c Based on this, select those whose IoU exceeds the threshold δ in sequence. IoU Other regions are included in the merge set, and the minimum bounding rectangle boundary of the current merged region is updated, thereby maintaining the tracking of the overall coverage of the current merged region. The update process of the bounding rectangle boundary is defined as follows: x min+2 ←min(x min+1 ,b x ) and min+2 ←min(y min+1 ,b y ) x max+2 ←min(x max+1 ,b x ) and max+2 ←min(y max+1 ,b y ) Where x min+1 The x-coordinate represents the minimum value of the rectangle's boundary before the update. min+2 This represents the minimum x-coordinate of the rectangle boundary before the update. Where y min+1 This represents the minimum y-coordinate of the rectangle boundary before the update. min+2 This represents the minimum ordinate of the rectangle's boundary before the update. Where x max+1 This represents the maximum x-coordinate of the rectangle boundary before the update. max+2 This represents the maximum x-coordinate of the rectangle boundary before the update. Where y max+1 This represents the minimum y-coordinate of the rectangle boundary before the update. max+2 This represents the minimum ordinate of the rectangle's boundary before the update. Step S243: After processing the current region, from B remian Remove all merged regions and add the smallest bounding rectangle of that group of regions as a new candidate region to the candidate region set R; continue this process until all candidate regions have been processed. Finally, all merged regions are added to the candidate merge region set R, resulting in the following expression for R: R={r1,r2,…r n} The current region includes region b, which has not yet been merged. c Based on this, and selecting those with an IoU exceeding the threshold δ IoU Other areas and b c The merged region; Where r1, r2, ... r n This indicates the candidate region after merging.
5. The method for accelerating camera edge video analysis based on region tracking according to claim 4, characterized in that, Step S2 also includes the following: Step S25: Real-time prediction of the new target location using a Kalman filter-based target location prediction model includes the following: Step S251: Initialize the Kalman filter, where the input of the Kalman filter is the target position z of the previous frame. t-1 and state vector x t-1 The output of the Kalman filter is the predicted position of the new target in the current frame; furthermore, prediction is performed based on the state vector of the previous time step to obtain the predicted state at the current time step. and predicting covariance If a new target is detected at position z in the current frame t If the Kalman filter is updated, then if no new target location is detected in the current frame, then it is further determined whether the current target is in the boundary candidate region. Step S252: Further, if the IoU between the tracking and boundary candidate regions where the new target is located is less than the threshold δ, then perform position prediction and return. If the IoU between the tracking and the boundary candidate region where the new target is located is greater than or equal to the threshold δ, then the new target is determined to have disappeared in the current video frame; Define the cumulative number of undetected instances, 'a'; if 'a' exceeds a set threshold A... max If the target is lost, the tracking of the target will end, and the corresponding tracking candidate area will be deleted.
6. The method for accelerating camera edge video analysis based on region tracking according to claim 5, characterized in that, Step S3 also includes the following: Step S31: First, obtain the available edge server status, denoted as set S, which includes the following: S=s1,s2,…s M Where s i ={cap i ,load i }, cap i and load i These represent the server's maximum computing power and real-time load, respectively; further, s is calculated based on the real-time status. i The computing power coefficient is defined as follows: Where r i For s i The computing power coefficient is M, where M is the number of edge servers.
7. The method for accelerating camera edge video analysis based on region tracking according to claim 6, characterized in that, Step S3 also includes the following: Step S32: Calculate the total area s of all candidate regions and set the server's real-time computing power coefficient r. k As an area ratio constraint, the target area of each fusion region is determined; further, the set of fusion regions G for each group is initialized. k and cumulative area A k Then, all candidate regions are sorted in descending order of area to ensure that larger regions are allocated first; subsequently, the sorted regions I are traversed. i' For each region, calculate its relationship with the target T after being assigned to each group. k The area difference between them △ k It is then assigned to the group with the largest current area difference, and the cumulative area A of that group is updated. j .
8. The method for accelerating camera edge video analysis based on region tracking according to claim 7, characterized in that, Step S3 also includes the following: Step S33: After completing the region segmentation, merge the regions in each group; further, estimate the side length of the merged region in the current group, and sort the regions within the group from largest to smallest height; then, arrange the regions sequentially using a marching shelf strategy; and record their positions (x, y) in the merged region for subsequent merging, and calculate the blank rate α of the splicing result; if α exceeds a set threshold... Then, the area of the candidate region with the largest area in the current fusion region is reduced to 90% of its original size. Finally, the set of all fusion regions that satisfy the ratio constraint is returned.
9. A camera edge video analysis acceleration system based on region tracking, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a camera edge video analysis acceleration method based on region tracking as described in any one of claims 1 to 8.
10. A camera edge video analysis acceleration system based on region tracking, comprising a computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a camera edge video analysis acceleration method based on region tracking as described in any one of claims 1 to 8.