A method, system and medium for accelerating video analysis
By dividing the video image areas into easy-to-detection and difficult-to-detection areas, and selecting appropriate frame rate and configuration combinations, the problems of low video analysis efficiency and waste of computing resources in the prior art are solved, and efficient and accurate video analysis is achieved.
Patent Information
- Application Number
- CN202210077582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-24
AI Technical Summary
The prior art is inefficient in video analysis, is wasted computing resources and cannot be applied to dynamic video scenarios, making it difficult to take into account both the efficiency and accuracy of video analysis.
By adaptively dividing the video image area into easy-to-detection areas and difficult-to-detection areas according to easy-to-detection, and selecting appropriate frame rates and configuration combinations in different areas, video analysis is performed using the object detection model.
While ensuring the accuracy of video analysis, it significantly improves video analysis efficiency, reduces calculation overhead and reasoning time overhead, and can be suitable for dynamic video scenarios.
Smart Images

Figure CN114511525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video detection technology, and in particular to a method, system and medium for accelerating video analysis. Background Art
[0002] With the development of the times, cameras have entered thousands of households. Whether in family life, cities or enterprises, a large number of cameras are deployed. For the Skynet project alone, it is estimated that the number of surveillance cameras will be around 560 million by the end of 2021. A large number of cameras monitor in real time 24 hours a day, generating large amounts of historical video data. Through video analysis, these historical video data can meet a variety of needs such as traffic control, security monitoring, factory workshop monitoring, and urban security prevention and control.
[0003] With the rapid development of computer vision technology, deep neural networks have been widely used in video analysis. Video analysis technologies such as image classification, target detection, and abnormal behavior detection are used in various video surveillance scenarios. For example, through abnormal behavior detection, traffic management departments can automatically identify illegal vehicles on the road; through target detection technology, analyzing security surveillance videos can help public security departments quickly capture criminals. It can be seen that historical video data has great value, and analyzing historical video data is of great significance.
[0004] With the rapid development of computer vision technology, deep neural networks have been widely used in various video surveillance scenarios for video analysis. Object detection is a basic task in computer vision. A reliable object detection algorithm is the basis for understanding and analyzing complex scenes. The performance of the object detection algorithm will directly affect the performance of subsequent high-level tasks in computer vision. However, using deep neural networks for reasoning requires high computational costs, and as the requirements for model detection capabilities and reasoning accuracy increase, models with stronger representation capabilities and larger computational workloads are increasingly needed. Even the configuration that meets the accuracy threshold will have many orders of magnitude changes in its resource requirements. When conducting large-scale video data analysis, the high computational cost and processing time overhead become problems that need to be solved urgently.
[0005] Video data has three characteristics of big data: huge data volume, low value density, and fast growth rate. On the one hand, data processing efficiency determines whether we can make full use of the rapidly growing video data and mine the data value; on the other hand, surveillance video data is mainly used for post-query, and we often need to obtain query results quickly, so it is also crucial to achieve low latency in video analysis. At the same time, as a basic task of computer vision, the performance of target detection will directly affect the performance of subsequent high-level computer vision tasks such as action recognition, target tracking, and behavior understanding, and thus determine the availability of artificial intelligence applications. Therefore, how to balance the efficiency of video analysis detection and reasoning accuracy, and speed up the efficiency of video analysis while ensuring reasoning accuracy, is of great significance for understanding surveillance video content.
[0006] In the prior art, there are the following related technical solutions:
[0007] (1) Reduce the set of videos to be detected by filtering out frames that do not contain relevant information for the current query, thereby improving the efficiency of video analysis. Filtering out a frame requires understanding how the frame will affect the query results, and existing systems mainly use frame differencing. Video frames whose low-level features (e.g., pixel values) have not changed significantly (based on static thresholds) are expected to produce the same detection results. Therefore, frame differencing is used to determine whether the video content has changed significantly to decide whether to filter.
[0008] (2) Detection is performed while the image is being captured, that is, the real-time video being captured is processed and the detection results are indexed. When a query is made for a specific category (such as an ambulance), an index search is performed directly to directly locate the video frame containing the queried object category.
[0009] (3) Selecting the optimal configuration combination for the current video clip by adjusting the configuration parameters of the video analysis system. The selected video configuration parameter combination includes the target detection model, image resolution, and detection step size, and the video analysis system configuration combination is updated regularly.
[0010] However, the prior art has the following disadvantages:
[0011] 1. Low efficiency of video analysis. The parameter configuration combination search of the video analysis system is based on the configuration combination selection of the entire video. Often, due to the existence of a small part of the clip or area, the search results of the entire video configuration are increased, and a high-computation configuration combination is selected, which reduces the efficiency of video analysis.
[0012] 2. Waste of computing resources. When performing detection while ingesting video, since usually only a small portion of the recorded frames are queried, most of the ingestion time may be wasted.
[0013] 3. Cannot be applied to dynamic video scenes. The frame filtering-based method cannot be used in dynamic video scenes due to its principle limitation. Summary of the invention
[0014] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to provide a method, system and medium for accelerating video analysis, thereby improving the efficiency of large-scale video data analysis. Compared with other solutions, the present invention can improve the efficiency of video analysis while ensuring the accuracy of video analysis, without limiting the application scenarios of the video analysis system, and reduce the computational overhead and the inference time overhead.
[0015] To achieve the above-mentioned object, the present invention provides a system for accelerating video analysis, comprising: a partition selector, a plurality of detection frame rate selectors, a plurality of configuration search modules and a target detection model, wherein the output end of the partition selector is connected to the plurality of detection frame rate selectors, the plurality of detection frame rate selectors are respectively connected to the plurality of configuration search modules, and the output ends of the plurality of configuration search modules are connected to the target detection model, wherein:
[0016] A partition selector adaptively divides the video image structure into an easy-to-detect area and a hard-to-detect area according to the detectability of the video image area;
[0017] A detection frame rate selector is used to select a frame rate for the easy-to-detect area and the difficult-to-detect area using a frame rate selector;
[0018] A configuration search module is used to search for configuration information of the image after the frame rate is selected until the target accuracy is reached and a complete combination of parameter configuration information is obtained; the configuration information includes but is not limited to image size and target detector;
[0019] The target detection model is used to combine the searched parameter configuration information and apply it to queries of other times of the video image.
[0020] A method for accelerating video analysis comprises the following steps:
[0021] Each video is segmented according to a specific duration, and the image structure of the specific segmented video is adaptively divided into an easy-to-detect area and a difficult-to-detect area according to the detectability of the video image area;
[0022] Using a frame rate selector to select a frame rate for the easy-to-detect area and the difficult-to-detect area;
[0023] Searching for configuration information of the image after the frame rate selection until the target accuracy is reached, and obtaining a complete combination of parameter configuration information; the configuration information includes but is not limited to image size and target detector;
[0024] The searched parameter configuration information is combined and applied to the query of other times of the video image.
[0025] Furthermore, the specific segmented video is adaptively divided into an easy-to-detect area and a difficult-to-detect area according to the detectability of the video image area. Specifically, the threshold is set to T, and the detection result of configuration combination 1 in the image block of the i-th row and the j-th column is Det1 ij , the detection result of configuration combination 2 is Det2 ij ,if:
[0026] Det1 ij -Det2 ij >T,(i,j)∈difficult-to-detect area;
[0027] Det1 ij -Det2 ij <T, (i, j)∈ easy-to-detect area.
[0028] Furthermore, the easy-to-detect area and the difficult-to-detect area use an exhaustive method to select the frame rate when using the minimum model (yolo v5s), so that the detection frame rate is 1. When F1>(1+a)*f, a model search is performed, where F1 represents the comprehensive detection value. The F1 value is used to measure the detection effect of the video analysis system. The F1 value is defined as follows:
[0029]
[0030]
[0031]
[0032] Among them, TP refers to the positive samples predicted by the model as positive, FP refers to the negative samples predicted by the model as positive, FN refers to the positive samples predicted by the model as negative, precision is the accuracy rate, and recall is the recall rate.
[0033] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-mentioned method for accelerating video analysis.
[0034] The beneficial effects of the present invention are:
[0035] 1. High efficiency of video analysis. The present invention partitions the image according to whether the video image area is easy to detect and divides it according to the spatial dimension. In addition, the configuration combination and detection frame rate suitable for different partitions are selected between different partitions, and the method of differential detection of different partitions divides the image according to the time and space dimensions, which fully demonstrates the idea of divide and conquer of video analysis tasks, avoids the general application of large computational models to full image and full frame rate detection, thereby achieving the reduction of computational overhead and time overhead during video analysis query while ensuring a certain inference accuracy.
[0036] 2. Save computing resources. This solution only needs to perform detection on the video data to be queried for post-query of the video data set, which greatly reduces the retrieval scope and does not need to detect all the ingested videos when ingesting them.
[0037] 3. Widely applicable video scenes. This solution can also be well applied in dynamic video scenes.
[0038] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a system model diagram of the present invention.
[0040] Figure 2 It is a diagram of dividing the video of the present invention into segments according to specific durations.
[0041] Figure 3 It is the video image partition diagram of the present invention.
[0042] Figure 4 It is a flow chart of the video analysis system of the present invention.
[0043] Figure 5 It is the image area division diagram of the present invention.
[0044] Figure 6 It is a frame rate selection and configuration search flow chart of the present invention. DETAILED DESCRIPTION
[0045] like Figure 1 As shown, the present invention provides a system for accelerating video analysis, including: a partition selector, multiple detection frame rate selectors, multiple configuration search modules and a target detection model, the partition selector output end is connected to the multiple detection frame rate selectors, the multiple detection frame rate selectors are respectively connected to the multiple configuration search modules, and the multiple configuration search module output ends are connected to the target detection model, wherein:
[0046] A partition selector adaptively divides the video image structure into an easy-to-detect area and a hard-to-detect area according to the detectability of the video image area;
[0047] A detection frame rate selector is used to select the frame rate for the easy-to-detect area and the difficult-to-detect area using the frame rate selector;
[0048] A configuration search module is used to search for configuration information of the image after the frame rate is selected until the target accuracy is reached and a complete combination of parameter configuration information is obtained; the configuration information includes but is not limited to image size and target detector;
[0049] The target detection model is used to combine the searched parameter configuration information and apply it to queries of other times of the video image. Figure 1 There is no fixed model for the target detection model, which can be yolo v5s in this embodiment, or other target detection models in computer vision such as Faster RCNN.
[0050] For query-related video datasets, such as Figure 2 As shown, each video is segmented into segments of specific duration ( Figure 2 The specific segmented video is input into the partition detector, which adaptively divides the image structure into easy-to-detect areas and difficult-to-detect areas according to the detectability of the video image area. Figure 3 As shown. Use the frame rate selector to select the frame rate for the two regions respectively, and then continue to search for other configurations, including image size and target detection model, to obtain a complete video analysis system parameter configuration combination. Finally, apply the searched parameter configuration combination to other time periods of the video ( Figure 2 The query is performed on the entire video (part 2 of the video clip). The video is divided into several clips, and each video clip is divided into two parts. The first part is used for video configuration parameter search, and the configuration parameter results searched in the first part are used for video analysis in the second part.
[0051] like Figure 4 As shown, a method for accelerating video analysis comprises the following steps:
[0052] Each video is segmented according to a specific duration, and the image structure of a specific segmented video is adaptively divided into an easy-to-detect area and a difficult-to-detect area according to the detectability of the video image area;
[0053] Use the frame rate selector to select the frame rate for easy-to-detect areas and difficult-to-detect areas;
[0054] Searching for configuration information of the image after the frame rate selection until the target accuracy is achieved, and obtaining a complete combination of parameter configuration information; the configuration information includes but is not limited to image size and target detector;
[0055] The searched parameter configuration information is combined and applied to the query of the video image at other times.
[0056] For the partition selector, the candidate parameter set for the detection frame rate is [1, 2, 3, 4, 5], and the candidate models for the target detector are [yolo v5s, yolo v5 m, yolo v5 l, yolo v5 x]. The lower the detection frame rate and the larger the candidate target detector model, the higher the detection accuracy and the greater the time overhead. When the detection frame rate is 1, the target detector selects the model with the largest computational effort and the highest detection accuracy (here is yolo v5 x) as the optimal configuration combination 1, and the detection result of this configuration combination is used as the basic truth (Ground Truth). In addition, when the detection frame rate is 1, the target detector selects the model with the smallest computational effort and the lowest detection accuracy (here is yolo v5 s) as the control configuration combination 2. Divide the image in the video into m*n image blocks ( Figure 5 In the example, m=4, n=5), this partitioning method is only an example, and other partitioning methods are also possible.
[0057] For a frame of image, use the strong target detection capability model and the weak target detection capability model for detection respectively. In each region block, the detection result of the strong target detection capability model is taken as the basic fact. When the number of missed objects by the weak target detection capability model exceeds the specified threshold, the target area is a difficult detection area. Otherwise, it is an easy detection area. Set the threshold to T, and assume that the detection result of configuration combination 1 in the i-th row and j-th column image block is Det1 ij , the detection result of configuration combination 2 is Det2 ij ,if:
[0058] Det1 ij -Det2 ij >T,(i,j)∈difficult-to-detect area;
[0059] Det1 ij -Det2 ij <T, (i, j)∈ easy-to-detect area.
[0060] The frame rate selection and configuration search flow chart is as follows Figure 6As shown. In this embodiment, the easy-to-detect area and the difficult-to-detect area use an exhaustive method to select the frame rate when using the minimum model (Yolo V5 S), where when the target detector candidate model set is [Yolo V5 S, Yolo V5 M, Yolo V5 L, Yolo V5 X], the lightest target detection model is Yolo V5 S. When the detection frame rate is 1, when F1> (1+a)*f, a model search is performed, where F1 represents the comprehensive detection value, and the F1 value is used to measure the detection effect of the video analysis system. The F1 value is defined as follows:
[0061]
[0062]
[0063]
[0064] Among them, TP refers to the positive samples predicted by the model as positive, FP refers to the negative samples predicted by the model as positive, FN refers to the positive samples predicted by the model as negative, precision is the accuracy rate, and recall is the recall rate.
[0065] The threshold setting here is only an example, and other settings are also possible. The model search strategy is to prioritize the configuration of the difficult-to-detect area, and improve the configuration of the easy-to-detect area when the combination of the difficult-to-detect area configuration reaches the highest. This strategy can maximize the improvement of the comprehensive F1 value of the detection while ensuring the lowest final comprehensive configuration overhead.
[0066] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of the above-mentioned method for accelerating video analysis when executed by a processor.
[0067] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A system for accelerating video analysis, characterized in that: The system selects configuration combinations and detection frame rates suitable for different partitions between different partitions, and detects differential speeds of different partitions; the system includes: a partition selector, multiple detection frame rate selectors, multiple configuration search modules and a target detection model, the output end of the partition selector is connected to multiple detection frame rate selectors, the multiple detection frame rate selectors are respectively connected to multiple configuration search modules, and the output ends of the multiple configuration search modules are connected to the target detection model, wherein: A partition selector adaptively divides the video image structure into an easy-to-detect area and a hard-to-detect area according to the detectability of the video image area; A detection frame rate selector is used to select a frame rate for the easy-to-detect area and the difficult-to-detect area using a frame rate selector; A configuration search module is used to search for configuration information of the image after the frame rate is selected until the target accuracy is reached and a complete combination of parameter configuration information is obtained; the configuration information includes image size and target detector; The target detection model is used to combine the searched parameter configuration information and apply it to queries of other times of the video image.
2. A method for accelerating video analysis, wherein the method uses the system of claim 1 to accelerate video analysis, characterized in that: The method described herein selects configuration combinations and detection frame rates that are suitable for different partitions, and detects differential speeds of different partitions; The specific steps include: Each video is segmented according to a preset duration, and the segmented video is adaptively divided into an easy-to-detect area and a difficult-to-detect area according to the detectability of the video image area; Using a frame rate selector to select a frame rate for the easy-to-detect area and the difficult-to-detect area; Searching for configuration information of the image after the frame rate selection until the target accuracy is reached, and obtaining a complete combination of parameter configuration information; the configuration information includes image size and target detector; The searched parameter configuration information is combined and applied to the query of other times of the video image.
3. A method for accelerating video analysis as claimed in claim 2, characterized in that: The segmented video is adaptively divided into an easy-to-detect area and a difficult-to-detect area according to the detectability of the video image area, specifically: For a frame of image, the strong target detection capability model and the weak target detection capability model are used for detection respectively. In each area block, the detection result of the strong target detection capability model is taken as the basic fact. When the number of missed objects of the weak target detection capability model exceeds the specified threshold, the target area is a difficult-to-detect area, otherwise, it is an easy-to-detect area. Set the threshold to T, and assume that the detection result of configuration combination 1 in the image block in the i-th row and j-th column is Det1 ij , the detection result of configuration combination 2 is Det2 ij ,if: Det1 ij -Det2 ij >T,(i,j)∈difficult-to-detect area; Det1 ij -Det2 ij <T, (i, j)∈ easy-to-detect area.
4. A method for accelerating video analysis as claimed in claim 2, characterized in that: The frame rates of the easy-to-detect area and the difficult-to-detect area are selected by exhaustive method, and the detection frame rate is 1. When F1>(1+a)*f, a model search is performed; Where f represents the target measurement indicator threshold, a represents the degree to which the target measurement indicator threshold is allowed to be exceeded, and F1>(1+a)*f means that when the detection effect of the current configuration combination exceeds the target detection effect f measured by a, a parameter configuration search is performed to select a lighter configuration combination. On the premise of achieving the target detection effect f, a smaller configuration combination is selected; F1 represents the harmonic mean of precision and recall, which is a measure of the detection results. The F1 value is used to measure the detection effect of the video analysis system. The F1 value is defined as follows: Among them, TP refers to the positive samples predicted by the model as positive, FP refers to the negative samples predicted by the model as positive, FN refers to the positive samples predicted by the model as negative, precision is the accuracy rate, and recall is the recall rate.
5. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for accelerating video analysis according to any one of claims 2 to 4 are implemented.
Citation Information
Patent Citations
Monitoring-video-oriented masked face detection method
CN104866843A
Abnormal behavior detection method, device, electronic device, and storage medium
CN109086696A