A method for improving real-time video target detection performance

By combining real-time video transmission and keyframe judgment modules with optical flow algorithms and detection box stability optimization, the problems of frame discontinuity and detection box jitter in video detection are solved, achieving high efficiency and stability in real-time video target detection.

CN117095323BActive Publication Date: 2025-11-04GUANGZHOU TIANYUE ELECTRONICS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210515438.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2025-11-04
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

Existing deep learning object detection technologies suffer from problems such as discontinuous detection between frames, abrupt changes in detection boxes, and the inability to simultaneously achieve high accuracy and speed in video detection. This is especially true on edge devices such as mobile phones, where computational demands are high and memory consumption is significant. Existing keyframe extraction technologies require prior knowledge and cannot meet the needs of real-time video.

Method used

The system employs a real-time video transmission module, a keyframe judgment module, and a detection box stability optimization module, using optical flow and exponential moving weighted average algorithms to optimize detection performance. The real-time video transmission module is responsible for video stream data acquisition and transmission, the keyframe judgment module determines keyframes based on optical flow thresholds, and the detection box stability optimization module uses an exponential moving weighted average algorithm to smooth the detection box position.

Benefits of technology

It improves the speed and stability of real-time video target detection, reduces detection box jitter and discontinuity, and optimizes the processing performance of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095323B_ABST
    Figure CN117095323B_ABST
Patent Text Reader

Abstract

The application provides a method for improving real-time video target detection performance.The method for improving real-time video target detection performance comprises the following steps: S1: a real-time video transmission module calls a mobile phone camera to collect video stream data, and transmits the video stream data according to a transmission protocol after compression and coding.The method for improving real-time video target detection performance provided by the application prevents the detection frame from discontinuous, jittery and other abrupt changes, improves the speed of video target detection, and performs weighted smoothing on the detection result to increase the stability of the inter-frame detection frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video processing algorithm, and particularly relates to a method for improving real-time video target detection performance. BACKGROUND

[0002] At present, the deep learning target detection technology has high precision and speed when detecting targets on static images, and is widely applied in various industries. However, when it is applied to continuous detection of videos, problems such as discontinuous inter-frame detection, abrupt change of detection frame, and inability to simultaneously satisfy high precision and speed may occur. Since the deep learning technology has large calculation amount and consumes a lot of memory, the above problems are more obvious when it is applied to edge devices such as mobile phones.

[0003] Since there is a large amount of redundant information between continuous frames of a video, it is unnecessary to perform high-precision detection on each frame. Therefore, in the actual application of video detection, a method of extracting key frames first and then performing target detection is usually adopted to optimize the speed and precision of detection. Common key frame extraction techniques include interval method, frame difference method, and optical flow estimation. The interval method mainly sets a frame interval. A video usually contains 30 frames of images per second. If a frame interval of 5 is set, one frame is taken as a key frame every 5 frames, and the remaining frames are taken as non-key frames. The interval method is very simple, but the frames obtained by the interval method are not representative. When no motion occurs in the video, the key frames obtained by the interval method still have a large amount of redundancy. The frame difference method commonly used includes two-frame difference method and three-frame difference method. The core of the frame difference method is to calculate the difference value of the global gray scale of adjacent frames, and compare the difference value with a pre-set threshold value. If the difference value is greater than the threshold value, it indicates that there are more key features, and the frame is taken as a key frame. The frame difference method is simple and effective, but it is based on global difference value and may ignore local information. In addition, the threshold value needs to be changed according to the image content. The value of the threshold value determines the quality of the obtained key frames. The core of the optical flow method is to calculate the optical flow vector of the pixels of adjacent frames. However, the calculation amount of the optical flow method is large, and there are dense optical flow and sparse optical flow.

[0004] Therefore, it is necessary to provide a method for improving real-time video target detection performance to solve the above technical problems. SUMMARY

[0005] The present application provides a method for improving real-time video target detection performance, which solves the problems of large calculation overhead, and inability to simultaneously satisfy precision and speed in video target detection. The existing key frame extraction technology for videos needs prior knowledge and cannot meet the key frame extraction of real-time videos.

[0006] To solve the above technical problems, the present application provides a method for improving real-time video target detection performance, which comprises the following steps:

[0007] S1: a real-time video transmission module, which calls a mobile phone camera to collect video stream data, and transmits the data according to a transmission protocol after compression and coding;

[0008] S2: key frame judgment module, S21, when the video transmission module inputs the frame in real time, first judge whether it is the first frame input by the video detection this time, if it is the first frame, it is directly judged as a key frame, then input the target detection module for detection, record the detection result returned by the module, and continue to input the next frame; if it is not the first frame, the next step is performed;

[0009] S22, get the current input frame fn, record the last key frame as fn-1, and get the target detection result of fn-1, the result mainly includes n [class, confidence, [x, y, w, h]], that is, the detection of n targets [category, category confidence, [center point x coordinate, center point y coordinate, detection frame width, detection frame height]], traverse the target detection result of fn-1, extract the detection information of each detection frame, extract the image information of fn-1 frame and fn frame detection frame position according to fn-1 frame detection frame position, and correspond fn-1 frame detection frame and fn frame preset frame image information one by one;

[0010] S23, use optical flow algorithm to calculate fn-1 frame detection frame and fn frame preset frame image, get the optical flow vector of the image;

[0011] S24, after extracting the detection frame of the corresponding position of the adjacent frames and completing the optical flow calculation, the total sum of the optical flow vectors of all pixels in each detection frame is calculated, and the optical flow vectors of all regions are sorted to find the maximum optical flow vector and its corresponding detection frame region, which are respectively recorded as max_flow and max_bbox;

[0012] S25, set the optical flow threshold, the threshold is 1 / 4 of the length of max_bbox detection frame x, y, calculate the optical flow component of max_flow, Mx represents the component of optical flow on x axis, and My represents the component of optical flow on y axis, and compare the component values with the threshold value respectively: when the calculated value is greater than the threshold value, the current frame is judged as a key frame, and the target detection module is input for detection; when the calculated value is less than the threshold value, the current frame is judged as a non-key frame, and the target detection result of fn-1 frame is used;

[0013] S26, the next input judgment is performed, and the key frame judgment module is completed;

[0014] S3: target detection module, use YOLOV5s target detection model for training and optimization, and convert the model weight file to tflite format, and deploy detection on mobile phone side;

[0015] S4: detecting frame stability optimization module, when the real-time video input continuous frame, the pixels between adjacent frames may have changed due to light brightness change, codec noise, etc., but we can't distinguish the difference on these micro pixels visually, and the picture we see is still the same, since the target detection is a convolutional processing of image pixels, and the final regression of the detection frame and the post-processing operation such as nms, which causes the detection frame jitter, changes abruptly, and discontinuous problems between adjacent frames, specifically, taking the target detection result of fn-1 frame, fn-1=[cls, conf, [x, y, w, h]], wherein the detection frame information is fn-1_bbox=[x, y, w, h]; taking the target detection result of fn frame, fn=[cls, conf, [x, y, w, h]], wherein the detection frame information is fn_bbox=[x, y, w, h], due to the change of pixels between frames, and the nms processing after the target detection for single frame image, resulting in fn-1_bbox and fn_bbox information in the two frames of detection results being completely irrelevant, which is reflected on the image as jitter and discontinuity of the detection frame on the continuous frame;

[0016] Therefore, the present scheme adds a detection frame stability optimization module, uses an exponential moving weighted average algorithm, combines the detection frame results of key frames and non-key frames, and processes them to optimize the stability of the terminal application. Specifically, the actual detection result fn-1=[cls, conf, [x, y, w, h]] of fn-1 frame is obtained, and the actual target detection result fn=[cls, conf, [x, y, w, h]] of fn frame is obtained, and the bbox information [x, y, w, h] in the fn-1 and fn detection results is processed by exponential moving weighted average, the exponential weighted moving average is to smooth the current value by the current actual value and the previous period (the average of how many previous data is determined by the weight), to generate a smooth trend curve, the specific formula is as follows:

[0017]

[0018] Where Vt is the moving average prediction value at time t; θt is the true value at time t; β is the weight, which determines the t value of the average; 1-β^t is the deviation correction term, as t increases, β^t will gradually approach to 0, and 1-β^t will gradually approach to 1, solving the problem of inaccurate estimation in the early stage;

[0019] The above steps complete the entire process of real-time video target detection;

[0020] Preferably, the main algorithm steps of S22 to S24 are:

[0021] Traverse the target detection result of fn-1 frame;

[0022] extract the detection information of the ith detection box, [class, confidence, [x, y, w, h]];

[0023] According to the position information [x, y, w, h] of the ith detection box, extract the pixels in the corresponding coordinates in the fn-1 frame and fn frame images, and mark them as the region image In-1 and the region image In;

[0024] Calculate the dense optical flow of the region In-1 and the region In, and record the result as M(x, y, n), which represents the pixel optical flow vector at (x, y) in the nth frame;

[0025] Calculate the sum of the optical flow vectors in the region, and record it as Mi, which represents the optical flow vector of the ith region;

[0026] After traversing all the detection boxes, calculate the optical flow vectors and find the largest optical flow vector and its corresponding region, and record the value as max_flow and the region as max_bbox.

[0027] Preferably, the main algorithm steps of S3 to S4 are:

[0028] Obtain the actual target detection information of the fn-1 frame, and record it as the list fn-1;

[0029] Obtain the region optical flow vector calculated in the key frame judgment module fn-1->fn;

[0030] Obtain the actual target detection information of the fn frame, and record it as the list fn;

[0031] Traverse the list fn-1:

[0032] Extract the detection information of the ith detection box, and record it as fn-1_i = [class, confidence, [x, y, w, h]];

[0033] Extract the optical flow vector of the ith region, and record it as the vector M;

[0034] Use Kalman filtering to predict and update the extracted detection box using the position of the ith detection box and the optical flow vector M as prior information, and record it as fn_i_prediction;

[0035] Set the search area with the fn_i_prediction center point [x, y], radius M;

[0036] Traverse the actual target detection information of the fn frame:

[0037] When the detection box center point [x, y] falls within the search area of step iv, and the IOU value with fn_i_prediction box is the largest, record the actual detection information at this time as fn_i_actual;

[0038] The actual detection frame [x, y, w, h] of fn-1_i and fn_i_ is weighted using an exponential moving weighted average, and the output fn_i_ average is output; in the scheme, the beta weight is 0.7, representing the data of the average of 3 detection results, and [x, y, w, h] is substituted into the above formula respectively, and the weighted average result is obtained, that is

[0039]

[0040] The prediction and update are completed, and the weighted smoothed result is output.

[0041] Compared with the related art, the method for improving real-time video target detection performance provided by the application has the following beneficial effects:

[0042] The application provides a method for improving real-time video target detection performance, and a real-time video transmission module is mainly responsible for pushing and pulling a stream of a mobile terminal real-time video, and inputting stream data into a key frame judgment module; the key frame judgment module is responsible for extracting a key frame from a real-time transmission video frame, inputting the key frame into a target detection module for detection, and not detecting a non-key frame, so that the processing performance of the target detection model on the real-time video is improved; the target detection module is responsible for using a neural network to perform convolution identification on a picture to identify the category of each target on the picture and detect the specific position of the target, and then inputting the detection result into a detection frame stability optimization module; the detection frame stability optimization module is responsible for counting the detection results of a previous frame and a current frame, and optimizing the position of the detection frame to prevent the detection frame from appearing discontinuous, jitter and other abrupt changes, improve the speed of video target detection, and perform weighted smoothing on the detection result to increase the stability of the inter-frame detection frame. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 The application provides a method for improving real-time video target detection performance.

[0044] Figure 2 The application provides a method for improving real-time video target detection performance. DETAILED DESCRIPTION

[0045] The application will be further described below in combination with the drawings and embodiments.

[0046] Please refer to Figure 1 , Figure 2 , wherein Figure 1 The application provides a method for improving real-time video target detection performance. Figure 2 The application provides a method for improving real-time video target detection performance.

[0047] S1: Real-time video transmission module, call the mobile phone camera to collect video stream data, after compression and coding, transmit according to the transmission protocol; the real-time video transmission module is mainly responsible for the push stream and pull stream of the mobile phone real-time video, and inputs the stream data into the key frame judgment module,

[0048] S2: Key frame judgment module, the key frame judgment module is responsible for extracting key frames from the real-time transmitted video frames, inputting the target detection module for detection, and not detecting non-key frames, so as to improve the processing performance of the target detection model on real-time video, S21, when the video transmission module inputs the frame in real time, first judge whether it is the first frame input for this time video detection, if it is the first frame, then directly judge it as a key frame, then input the target detection module for detection, record the detection result returned by the module, and continue to input the next frame; if it is not the first frame, then proceed to the next step;

[0049] S22, get the current input frame fn, record the last key frame as fn-1, and get the target detection result of fn-1, the result mainly includes n [class, confidence, [x, y, w, h]], that is, the detection of n targets [category, category confidence, [center point x coordinate, center point y coordinate, detection frame width, detection frame height]]. Traverse the target detection result of fn-1, extract the detection information of each detection frame, extract the image information of fn-1 frame and fn frame detection frame position according to fn-1 frame detection frame position, and correspond fn-1 frame detection frame and fn frame preset frame image information one by one;

[0050] S23, use optical flow algorithm to calculate fn-1 frame detection frame and fn frame preset frame image, get the optical flow vector of the image. Optical flow method is a method of finding the corresponding relationship between the last frame and the current frame by using the change of pixels in time domain and the correlation between adjacent frames, so as to calculate the motion information of objects between adjacent frames.

[0051] For example, the position of point A on the n-1th frame is (x1, y1), and the position of point A on the nth frame is (x2, y2), that is, In-1(x1, y1) = In(x2, y2) = In-1(x1 + ux, x1 + vy), then the optical flow of In-1->In is (ux, vy), where u and v represent the velocity components on the x and y axes, and ux and vy represent the offsets on the x and y axes. Therefore, given a pair of pictures (fn-1->fn), the optical flow graph between the pair of pictures can be calculated, which has the same size as the two frames of pictures. Optical flow algorithms can be divided into sparse optical flow and dense optical flow. The sparse optical flow algorithm first extracts feature points in the image, then calculates the optical flow vectors of the feature points in the two images, which has a small amount of calculation but is not very accurate. The dense optical flow algorithm directly calculates the optical flow vector of each pixel point in the image, which has a large amount of calculation but high accuracy. The present scheme mainly detects target information in the video, and does not pay attention to the rest of the information. Therefore, for the whole image optical flow estimation, irrelevant information interference will be introduced, and the amount of calculation will be large. For real-time video frames, attention should be focused on the part of the adjacent frame that is the target and moves. Therefore, the present scheme only performs dense optical flow calculation on the detected results of the target, avoids the interference of the rest of the information, reduces the amount of calculation, and improves the accuracy.

[0052] S24, after extracting the detection frame of the corresponding position of the adjacent frame and completing the optical flow calculation, the sum of the optical flow vectors of all pixels in each detection frame is counted. The optical flow vectors of all regions are sorted to find the largest optical flow vector and its corresponding detection frame region, which are denoted as max_flow and max_bbox, respectively.

[0053] S25, set the optical flow threshold, which is 1 / 4 of the length of the max_bbox detection frame x and y. Calculate the optical flow components of max_flow, Mx represents the component of the optical flow on the x axis, and My represents the component of the optical flow on the y axis. Compare the component values with the threshold value respectively: when the calculated value is greater than the threshold value, the current frame is determined as a key frame, which is input into the target detection module for detection; when the calculated value is less than the threshold value, the current frame is determined as a non-key frame, and the target detection result of the fn-1 frame is used.

[0054] S26, the next input judgment is performed, and the key frame judgment module is completed.

[0055] S3: target detection module, using YOLOV5s target detection model for training and optimization, and converting the model weight file to tflite format for deployment and detection on the mobile phone side; the target detection module is responsible for using a neural network to convolve and identify the categories of each target on the picture, and detect the specific position, and then input the detection result into the detection frame stability optimization module

[0056] S4: a detection box stability optimization module, the detection box stability optimization module is responsible for counting the detection results of the previous frame and the current frame, and optimizing the position of the detection box, to prevent the detection box from appearing discontinuous, jitter and other abrupt changes. When the real-time video input continuous frame, the pixels between adjacent frames may have changed due to light intensity changes, coding noise, etc., but we cannot distinguish the difference in these microscopic pixels visually, and the picture is still the same. Since target detection is a convolutional processing of image pixels, the final regression detection box, and nms post-processing operation, these cause the detection box to jitter, change abruptly, and be discontinuous between adjacent frames. Specifically, the target detection result of fn-1 frame is taken, which is abbreviated as fn-1=[cls, conf, [x, y, w, h]], wherein the detection box information is fn-1_bbox=[x, y, w, h]; the target detection result of fn frame is taken, which is abbreviated as fn=[cls, conf, [x, y, w, h]], wherein the detection box information is fn_bbox=[x, y, w, h]. Due to the change of pixels between frames, and the nms processing after target detection for single frame image detection, the information of fn-1_bbox and fn_bbox in the detection results of two frames is completely irrelevant, which is reflected in the image as jitter and discontinuity of the detection box on the continuous frames.

[0057] Therefore, the present scheme adds a detection box stability optimization module, uses an exponential moving weighted average algorithm, combines the detection box results of key frames and non-key frames, and processes them to optimize the stability of the terminal application. Specifically, the actual detection result fn-1=[cls, conf, [x, y, w, h]] of fn-1 frame is obtained, the actual target detection result fn=[cls, conf, [x, y, w, h]] of fn frame is obtained, and the bbox information [x, y, w, h] in the fn-1 and fn detection results is respectively processed by exponential moving weighted average. Exponential weighted moving average is to smooth the current value by the current actual value and the previous period (determined by the weight value how many previous data is averaged) to generate a smooth trend curve. The specific formula is as follows:

[0058]

[0059] Wherein Vt is the moving average prediction value at time t; θt is the true value at time t; β is the weight, which determines the t value of the average; 1-β^t is the deviation correction term, as t increases, β^t will gradually approach to 0, and 1-β^t will gradually approach to 1, solving the problem of inaccurate estimation in the early stage;

[0060] The above steps complete the entire process of real-time video target detection;

[0061] The main algorithm steps of S22 to S24 are:

[0062] Iterate through the target detection results of frame fn-1;

[0063] Extract the detection information of the i-th detection box: [class, confidence, [x, y, w, h]].

[0064] Based on the position information [x, y, w, h] of the i-th detection box, extract the pixels at the corresponding coordinates in the fn-1 frame and fn frame images, and label them as region image In-1 and region image In;

[0065] Dense optical flow is calculated for regions In-1 and In, and the result is denoted as M(x, y, n), which represents the pixel optical flow vector at (x, y) in the nth frame;

[0066] The sum of optical flow vectors within the statistical region is denoted as Mi, representing the optical flow vector of the i-th region;

[0067] After traversing all detection boxes, the optical flow vectors are counted, and the largest optical flow vector and its corresponding region are found. The value is denoted as max_flow, and the region is denoted as max_bbox.

[0068] The main algorithm steps from S3 to S4 are as follows:

[0069] Obtain the actual target detection information for frame fn-1, and denote it as list fn-1;

[0070] Obtain the region optical flow vector fn-1->fn calculated in the keyframe determination module;

[0071] Obtain the actual target detection information of frame fn, and denote it as list fn;

[0072] Traverse list fn-1:

[0073] Extract the detection information of the i-th detection box, denoted as fn-1_i = [class, confidence, [x, y, w, h]].

[0074] Extract the optical flow vector of the i-th region, denoted as vector M;

[0075] Using the position of the i-th detection box and the optical flow vector M as prior information, Kalman filtering is used to predict and update the extracted detection boxes, denoted as fn_i_prediction;

[0076] Set the search area with the predicted center point [x, y] and radius M;

[0077] Actual target detection information for traversing frame fn:

[0078] When the center point [x, y] of the detection frame falls into the search area of step iv, and the IOU value with fn_i_prediction frame is the maximum, the actual detection information at this time is recorded as fn_i_actual;

[0079] The detection frame [x, y, w, h] of fn-1_i and fn_i_actual is weighted using exponential moving weighted average, and fn_i_average is output; in the scheme, the beta weight value is 0.7, representing the data of the average of 3 detection results, [x, y, w, h] is substituted into the above formula respectively, and the weighted average result is obtained, that is

[0080]

[0081] The prediction and update are completed, and the weighted smoothed result is output.

[0082] Business logic: use the real-time video transmission module to call the mobile phone camera to obtain the video stream, compress and encode the video stream, and then perform stream data transmission; use the key frame judgment module to judge the input data, and extract the key frame of the real-time stream data; use the target detection model to perform convolution processing on the data, detect the category and position of the target, use the detection frame stability optimization module to optimize the target detection result, use the real-time video transmission module to return the data, and display the data on the user side.

[0083] The above only describes the embodiments of the present application, and does not limit the patent range of the present application, any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection range of the present application.

Claims

1. A method for improving real-time video target detection performance, characterized in that, Comprise the following steps: S1: Real-time video transmission module, call mobile phone camera for video stream data acquisition, after compression coding according to transmission protocol transmission; S2: Key frame judgment module, S21, when the video transmission module real-time input frame, first judge whether it is the first frame of this video detection input, if it is the first frame, then directly determine the key frame, then input target detection module for detection, record module returns the detection result, continue to input the next frame;If not the first frame, then proceed to the next step; S22, get the current frame as fn frame, record the last key frame as fn-1 frame, and get the target detection result of fn-1 frame, the result mainly contains n [class, confidence, [x, y, w, h]], that is, the detection of n targets [category, category confidence, [center point x coordinate, center point y coordinate, detection frame width, detection frame height]], traverse the target detection result of fn-1 frame, extract the detection information of each detection frame, extract the image information of fn-1 frame and fn frame detection frame position according to fn-1 frame detection frame position, and correspond fn-1 frame detection frame and fn frame preset frame image information one by one; S23, use optical flow algorithm to calculate fn-1 frame detection frame and fn frame preset frame image, get the optical flow vector of image; S24, after extracting the detection frame of corresponding position and completing the optical flow calculation of adjacent frames, statistics the sum of optical flow vector of all pixels in each detection frame, sort the optical flow vector of all regions, find out the maximum optical flow vector and its corresponding detection frame area, respectively recorded as max_flow and max_bbox; S25, set the optical flow threshold, the threshold is 1 / 4 of the length of max_bbox detection frame x, y, calculate the optical flow component of max_flow, Mx represents the component of optical flow on x axis, My represents the component of optical flow on y axis, and compare the component values with the threshold value respectively: when the calculated value is greater than the threshold value, the current frame is determined as key frame, input target detection module for detection;When the calculated value is less than the threshold value, the current frame is determined as non-key frame, and the target detection result of fn-1 frame is used; S26, the next input judgment is carried out, and the key frame judgment module is completed; S3: Target detection module, use YOLOV5s target detection model for training and optimization, and convert the model weight file to tflite format, and deploy detection on mobile phone side; S4: detecting frame stability optimization module, when the real-time video input continuous frame, the pixels between adjacent frames have changed due to light brightness change, codec noise, but we can't distinguish the difference on these micro pixels visually, the picture we see is still the same, since the target detection is a convolution operation on image pixels, and the final regression is a detection frame, and there is nms post-processing operation, which causes the detection frame jitter, changes abruptly, and discontinuous problems between adjacent frames, specifically, the target detection result of fn-1 frame is taken, which is abbreviated as fn-1=[cls, conf, [x, y, w, h]], wherein the detection frame information is fn-1_bbox=[x, y, w, h]; the target detection result of fn frame is taken, which is abbreviated as fn=[cls, conf, [x, y, w, h]], wherein the detection frame information is fn_bbox=[x, y, w, h], due to the change of pixels between frames, and the nms processing after the target detection for single frame image, the information of fn-1_bbox and fn_bbox in the two frame detection results is completely irrelevant, which is reflected in the image as jitter and discontinuity of the detection frame on the continuous frame; Therefore, the scheme adds a detection frame stability optimization module, uses an exponential moving weighted average algorithm, combines the detection frame results of key frames and non-key frames, and processes them to optimize the stability of the terminal application, specifically, the actual detection result fn-1=[cls, conf, [x, y, w, h]] of fn-1 frame is obtained, the actual target detection result fn=[cls, conf, [x, y, w, h]] of fn frame is obtained, the bbox information [x, y, w, h] in the fn-1 and fn detection results is respectively processed by exponential moving weighted average, the exponential weighted moving average is to modify the current value by smoothing the current value and the previous period to generate a smooth trend curve, the specific formula is as follows: t = 1, 2, 3...n Wherein is the moving average prediction value at Vt time; θt is the true value at t time; β is the weight, which determines the t value of the average; 1-β^t is the deviation correction term, as t increases, β^t will gradually approach to 0, and 1-β^t will gradually approach to 1, solving the problem of inaccurate estimation in the early stage; The whole process of real-time video target detection is completed through the above steps.

2. The method of claim 1, wherein, The main algorithm steps of S22 to S24 are: Traverse the target detection result of fn-1 frame; Extract the detection information of the i-th detection frame, [class, confidence, [x, y, w, h]]; According to the position information [x, y, w, h] of the i-th detection frame, the pixels in the corresponding coordinates of fn-1 frame and fn frame image are extracted, which are marked as region image In-1 and region image In; Calculate the dense optical flow of the region In-1 and the region In, and get the result M(x, y, n), which represents the pixel optical flow vector at (x, y) of the n-th frame; Statistical sum of optical flow vectors in the region, denoted as Mi, representing the optical flow vector of the i-th region; After traversing all detection boxes, the optical flow vectors are counted, and the largest optical flow vector and its corresponding region are found. The value is denoted as max_flow, and the region is denoted as max_bbox.

3. The method of claim 1, wherein, The main algorithm steps from S3 to S4 are as follows: Obtain the actual target detection information for frame fn-1, and denote it as list fn-1; Obtain the region optical flow vector fn-1->fn calculated in the keyframe determination module; Obtain the actual target detection information of frame fn, and denote it as list fn; Traverse list fn-1: Extract the detection information of the i-th detection box, denoted as fn-1_i=[class, confidence, [x, y, w, h]]; Extract the optical flow vector of the i-th region, denoted as vector M; Using the position of the i-th detection box and the optical flow vector M as prior information, Kalman filtering is used to predict and update the extracted detection boxes, denoted as fn_i_prediction; Set the search area with the predicted center point [x, y] and radius M; Actual target detection information for traversing frame fn: When the center point [x, y] of the detection box falls into the search area of ​​step iv and has the maximum IOU value with the fn_i_predicted box, the actual detection information at this time is recorded as fn_i_actual. The exponentially moving weighted average is used to weight both fn-1_i and fn_i_the actual detection boxes [x, y, w, h], and output the average fn_i_. In this scheme, the β weight is 0.7, representing the average of 3 detection results. Substituting [x, y, w, h] into the above formula, the weighted average result is obtained. Complete the prediction and update, and output the weighted smoothed result.

Citation Information

Patent Citations

  • An improved method for improving the stability of video object detection

    CN109902620A

  • Target detection method, system and device, storage medium and computer device

    CN109978756A