Quick video auditing method and system based on time sequence prediction

Through a fast video auditing method based on timing prediction, the location and status of the target in the video are predicted, and the computing resources are concentrated in the area of ​​interest are solved, and the problems of slow processing speed and high-resolution video in the existing technology are solved, and high-efficiency and low-resource consumption video auditing is achieved.

CN120220035AActive Publication Date: 2025-06-27海看网络科技(山东)股份有限公司

Patent Information

Application Number
CN202510694953.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

When existing video auditing methods deal with high-resolution and high frame rate video streams, computing resources and bandwidth resources consume huge amounts, resulting in slow processing speed and high cost, making it difficult to meet the requirements of real-time or large-scale processing.

Method used

Using a fast video auditing method based on timing prediction, the predicted location and state of the target may appear in subsequent frames is used to centralize the computing resources to scan and identify the predicted area of ​​interest (ROI) instead of global scanning.

Benefits of technology

It significantly reduces the computational complexity and resource consumption, improves the processing speed and efficiency of video audits, is more adaptable than high-resolution videos, and continuously optimizes prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220035A_ABST
    Figure CN120220035A_ABST
Patent Text Reader

Abstract

The invention discloses a quick video auditing method and system based on time sequence prediction, and mainly relates to the technical field of computer vision. Comprising the following steps: constructing a time sequence state prediction model, performing global scanning on an initial frame of a video stream, and calling a target detection algorithm to identify and position a target; extracting state information of the target in the initial frame to form an initial state vector # imgabs0 #, and initializing a time sequence state prediction model; and performing iterative loop processing on each frame after the initial frame by using the initialized time sequence state prediction model to complete video auditing. The method has the advantages that computing resources are concentrated in a predicted region of interest (ROI) to be scanned and recognized by predicting the possible position and state of a target in a subsequent frame instead of global scanning, so that the computing resources are greatly saved, and the auditing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically to a fast video review method and system based on temporal prediction. Background Art

[0002] With the popularization of applications such as video surveillance, content review, and behavior analysis, efficient and accurate processing and analysis of massive video data have become an important requirement. Traditional video review or analysis methods usually involve global scanning and object detection for each frame of the video. Although this method can obtain comprehensive information, when processing high-resolution and high-frame-rate video streams, it consumes a large amount of computing resources (such as CPU and GPU processing time) and bandwidth resources, resulting in slow processing speed and high cost, and it is difficult to meet the requirements of real-time or large-scale processing.

[0003] To improve efficiency, researchers have proposed some methods based on object tracking. However, existing methods may not fully utilize the motion state information of the object to actively predict its future position, or are limited to specific tracking algorithms.

[0004] Therefore, there is an urgent need for a technical solution that can significantly reduce the computational complexity and improve the video review processing speed. Summary of the Invention

[0005] The purpose of the present invention is to provide a fast video review method and system based on temporal prediction, which predicts the possible positions and states of the object in subsequent frames, and focuses the computing resources on scanning and identifying the predicted region of interest (ROI), rather than performing global scanning, thereby greatly saving computing resources and improving the review efficiency.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: On the one hand, the present invention provides a fast video review method based on temporal prediction, including the following steps: Step S1: Construct a temporal state prediction model, perform a global scan on the initial frame of the video stream, and call an object detection algorithm to identify and locate the object; Step S2: Extract the state information of the object in the initial frame to form an initial state vector , and initialize the temporal state prediction model; Step S3: Use the initialized temporal state prediction model to perform iterative loop processing on each frame after the initial frame to complete the video review.

[0007] Preferably, in the step S2, the state information of the object in the initial frame includes but is not limited to: the position information of the target of interest , size information and estimating motion parameters from multi-frame data, where the motion parameters include but are not limited to velocity information and acceleration information .

[0008] Preferably, step S3 includes the following steps: S31: Define each frame after the initial frame as the current frame , and predict the state of the current frame through a sequential state prediction model based on the correction state of the previous frame ; S32: Generate a region of interest in the current frame according to the position information and size information of the predicted state, where the size of the region of interest is the product of the predicted size and the magnification factor , where ; S33: Perform target analysis within the region of interest, where the target analysis includes but is not limited to: precise positioning, feature extraction, and classification recognition; S34: If a target is detected within the region of interest, obtain the measurement state of the target, and use to correct the sequential state prediction model, and input the posterior state as the input for the next frame prediction.

[0009] Preferably, the generation of the region of interest satisfies the following conditions: The width of the region of interest is: , and the height is: , where .

[0010] Preferably, the target analysis includes but is not limited to: Precisely positioning the target within the region of interest using a deep learning-based detector; Extracting the texture, color, or shape features of the target to enhance the tracking robustness; Identifying the class attributes of the target through a classification model.

[0011] Preferably, the model correction uses the Kalman gain to calculate the weighted fusion of the predicted state and the measurement state, or updates the weight parameters of the deep learning model through the backpropagation algorithm.

[0012] Preferably, step S3 further includes: If consecutive ​​​​If no target is detected in the frame or the prediction deviation exceeds the threshold, the analysis of the region of interest is paused and a global scan is triggered. The target is recaptured or a new target is discovered through the global scan, and the tracking is reinitialized or terminated according to the scan results.

[0013] Preferably, the conditions for triggering a global scan include, but are not limited to: Triggering at a fixed time interval; triggering by a target loss event; triggering by a scene mutation event.

[0014] On the other hand, an audit system based on the fast video audit method based on temporal prediction as described above is provided, including: An initialization module for performing a global scan, target detection, and initial state extraction; A temporal prediction module for predicting the target state of the current frame based on historical states; An ROI generation module for dynamically delimiting the region of interest; A local analysis module for performing target analysis tasks within the ROI; A state correction module for fusing prediction and measurement data to update the model; A fault tolerance control module for handling target loss events and triggering a global scan.

[0015] Preferably, the temporal prediction module supports multi-model switching, including but not limited to: Kalman filtering, particle filtering, and LSTM networks, and dynamically selects the optimal model according to the scene complexity.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Greatly improve the audit efficiency: By concentrating the analysis and calculation on the predicted region of interest (ROI) instead of full-frame processing, the amount of calculation required for each frame is significantly reduced, thus accelerating the overall speed of video audit; 2. Significantly reduce the consumption of computing resources: Since only complex analyses such as target detection and recognition are performed on the ROI, redundant calculations for a large amount of background or irrelevant regions in the video frame are avoided, effectively saving hardware resources such as processors and memory; 3. Have stronger adaptability and more obvious advantages for high-resolution videos: As the video resolution increases, the total number of pixels increases multiplicatively, resulting in a sharp increase in the computational burden of traditional full-frame analysis. By focusing on the ROI, the amount of calculation is basically related to the size of the target and has little relationship with the growth of the overall resolution. Therefore, the higher the resolution, the more prominent the advantages of the present method in terms of efficiency improvement and resource saving; 4. Continuously optimize the prediction accuracy: After successfully finding the target within the ROI, use its latest measurement status to update the time-series state prediction model, achieving continuous tracking of the target state and iterative optimization of the prediction model, which helps to more accurately predict the target position and define the ROI in subsequent frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the flowchart of the method of the present invention; Figure 2 , where A is an example diagram of the initial detection frame of the present invention around the detected human face; B is an example diagram of the ROI generated by using Alpha-Beta prediction based on Figure 2 A in; C is an example diagram of the expanded ROI intercepted based on Figure 2 B in; D is an example diagram of the result obtained by face recognition within C of the present invention and the face position is framed; Figure 2 is an example diagram of the complete process visualization of prediction-identification-prediction based on the Alpha-Beta model of the present invention, where A, B, C, D, E, F are Figure 3 the specific process diagrams of; Figure 3 ; Figure 4 , where A is an example diagram of initializing the particle filter of the present invention; B is the region of interest ROI expanded based on Figure 4 A in; C is an example diagram of intercepting the region of interest ROI based on Figure 4 B in; D is an example diagram of the result obtained by face recognition within C of the present invention and the face position is framed; Figure 4 is an example diagram of the complete process visualization of prediction-identification-prediction based on the particle filter model of the present invention, where A, B, C, D, E, F are Figure 5 the specific process diagrams of; Figure 5 ; Figure 6 is an example diagram of the complete process visualization of the particle filter model of the present invention losing the target and recapturing it, where A, B, C, D are Figure 6 the specific process diagrams of; Figure 7 is the schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.

[0019] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationships of various components or elements of the present invention, and do not specifically refer to any component or element in the present invention, and should not be construed as a limitation to the present invention.

[0020] Embodiment 1: As Figure 1 shown, this embodiment provides a fast video review method based on time series prediction using an Alpha-Beta filter as a time series state prediction model, including the following steps: Step S1: Construct a time series state prediction model, and perform a global scan on the initial frame of the video stream, and call a target detection algorithm to identify and locate the target; Step S2: Extract the state information of the target in the initial frame to form an initial state vector , and initialize the time series state prediction model; Step S3: Use the initialized time series state prediction model to perform iterative loop processing on each frame after the initial frame to complete video review.

[0021] Among them, steps S1-S2 in this embodiment can be regarded as the "initialization" stage (cold start process). Step S1 can be regarded as the "target capture and initial state extraction" stage, and step S2 can be regarded as the "prediction model initialization" stage. In the "target capture and initial state extraction" stage and the "prediction model initialization" stage, the following is specifically executed: First, detect the target of interest through a global image scanning method (such as a conventional target detection algorithm). Once the target is detected, as Figure 2 shown in A, the initial detection box surrounds the detected face. In the first frame of the video, its initial measurement state includes the center position and size , which are used to initialize the state vector of the Alpha-Beta filter. In this embodiment, the state vector contains the position, size, and their respective rates of change (speeds) of the target. Initially, the position and size components are set to the measured values, while all rate of change components (speeds , , and the rate of change of size , ) are usually set to zero.

[0022] Step S3 includes: S31: Define each frame after the initial frame as the current frame , based on the corrected state of the previous frame , predict the current frame through the temporal state prediction model 's state ; ("Temporal State Prediction" stage) S32: According to the predicted state 's position information and size information, generate a region of interest in the current frame, and the size of the region of interest is the product of the predicted size and the magnification factor , where ; ("Region of Interest (ROI) Generation" stage) S33: Perform target analysis within the region of interest, and the target analysis includes but is not limited to: precise positioning, feature extraction, classification and recognition; ("Target Analysis within ROI" stage) S34: If the target is detected within the region of interest, obtain the measurement state of the target , use to correct the temporal state prediction model, input the posterior state , as the input for the next frame prediction. ("State Update and Iteration Preparation" stage) In the "Temporal State Prediction" stage, specifically execute: For each subsequent frame , use the corrected state at time , and the Alpha-Beta filter predicts the target state of the current frame . The prediction process is based on a simplified dynamic model, such as a uniform motion and uniform size change model: Predicted position: ; ; Predicted size: ; The speed and size change rate are usually predicted to remain constant: ; where is the time interval between frames.

[0023] In the "Region of Interest (ROI) Generation" stage, specifically execute: According to the predicted target state 's predicted position and predicted size , combined with a preset or adaptive magnification factor (such as 1.8), generate a predicted region of interest (ROI). This ROI is expected to contain the actual position of the target in the current frame. (As Figure 2As shown in B, the purple dashed box is the ROI generated based on Alpha-Beta prediction, and the yellow dashed line is the ROI after expansion using the expansion coefficient.) In the "Target Analysis within ROI" stage, the following is specifically executed: As Figure 2 shown in C, face recognition will scan the face within this area, that is, the target detection algorithm is only executed within the predicted ROI generated in the "Region of Interest (ROI) Generation" stage, rather than scanning the entire image, significantly reducing the computational amount; If a target is successfully detected within the ROI (as Figure 2 shown in D, the green box is the target detected within the ROI), then the current measurement state is obtained .

[0024] In the "State Update and Iteration Preparation" stage, the following is specifically executed: If the target is found within the ROI, then the measured value and the predicted value are used to correct (update) the state of the Alpha-Beta filter. The update process is as follows: Calculate the measurement residual (the difference between the observed value and the predicted value): (Calculated separately for each component); Update the state: ; ; (Similarly, use and to update the size and its change rate), where , , , are preset filter parameters, and the corrected state will be used for the prediction of the next frame .

[0025] Finally, for the remaining video frames in the video, repeat the above "Initialization" stage to "State Update and Iteration Preparation" stage, and perform loop processing until the video ends. As Figure 3 shown in A, B, C, D, E, F, they are visualization examples of some of the frames.

[0026] Example 2: This example provides a fast video review method based on temporal prediction using a particle filter as a temporal state prediction model, specifically including the following steps: In the "Initialization" stage, the following is specifically executed: Similar to Embodiment 1, the initial target is detected through global scanning (such as Figure 4 the initial detection box shown as A in ). After obtaining the initial measurement state the system initializes a number of particles (200 particles are initialized in this example). Each particle represents a state hypothesis of the initial state of the target. These particles are usually generated by sampling according to a certain initial uncertainty distribution (such as a Gaussian distribution) around Figure 4 (as shown as A in the appendix, a large number of particles are scattered around the initial target in the figure, and the purple box is the position of the prediction box). The initial weights of all particles are usually set to equal values .

[0027] In the "temporal state prediction" stage, the following is specifically executed: For subsequent frames , each particle (from the set of particles after update and resampling in the previous frame) is propagated (evolved) according to a preset dynamic model to obtain its predicted state in the current frame : ; The dynamic model is a simple random walk (i.e., the state at the previous moment plus a random perturbation).

[0028] In the "region of interest (ROI) generation" stage, the following is specifically executed: The predicted ROI is determined based on the collective positions and distributions of all predicted particles . For example, the weighted average state of these predicted particles can be calculated as the predicted center position and size of the target in the current frame , and then combined with an amplification factor to generate the predicted ROI (as shown as B in Figure 4 , the yellow dashed box is the ROI predicted based on the particle cloud distribution, and the current state of the particle cloud may also be shown in the figure).

[0029] In the "target analysis within ROI" stage, the following is specifically executed: The system only performs target detection within the predicted ROI (as shown as C in Figure 4 , face recognition will scan for faces within this area), and if a target is detected within this area, its measurement state is obtained; If a target is successfully detected within the ROI (as shown as D in Figure 4 , the green box is the target detected within the ROI).

[0030] In the "Status Update and Iteration Preparation" stage, the following operations are specifically performed: Weight calculation: Based on the measurement value and the similarity (or "likelihood") with each predicted particle to update the weight of the particle , the closer the particle state is to the measurement value, the higher its weight. For example, the similarity can be calculated based on the exponential decay function of the difference between the two state vectors. All weights are then normalized so that the sum is 1; Resampling: To prevent particle degradation (i.e., a few particles occupy almost all the weights), a number of particles are redrawn from the current particle set according to the updated weights to form a new particle set , particles with high weights have a higher probability of being selected multiple times. After resampling, the weights of the new particle set are usually reset to equal values , this step enables the particles to better predict the actual state of the target.

[0031] Finally, for the remaining video frames , repeat the above "Initialization" stage to "Status Update and Iteration Preparation" stage in a loop until the video ends. As Figure 5 shown in A, B, C, D, E, F, they are visualization examples of some of the frames.

[0032] In addition, this embodiment also provides fault tolerance processing: If the particle weights cannot be effectively updated within the predicted ROI (for example, all particles do not match any potential detections within the ROI), and this situation persists for M (preferably 7 in this example) frames, a global scan will be performed to attempt to re-capture the target, and the particle filter will be re-initialized based on the new strong detection results. As Figure 6 shown in the images of A, B, C, D, if no valid target can be detected continuously in the ROI, a global scan is triggered to re-determine the target position and re-initialize the particle weights.

[0033] This preferred embodiment based on the particle filter can provide more robust ROI prediction for target tracking in complex scenarios through its probabilistic modeling of state uncertainty, thus supporting an efficient video review process.

[0034] As Figure 7 shown, this embodiment also provides an audit system based on the above fast video audit method based on temporal prediction, including: An initialization module for performing global scanning, target detection, and initial state extraction; A temporal prediction module for predicting the target state of the current frame based on historical states; An ROI generation module for dynamically demarcating an area of interest; A local analysis module for performing a target analysis task within the ROI; A state correction module for fusing prediction and measurement data to update the model; A fault tolerance control module for handling target loss events and triggering a global scan.

[0035] The time series prediction module supports multi-model switching, including but not limited to: Kalman filtering, particle filtering, and LSTM networks, and dynamically selects the optimal model according to the scene complexity.

[0036] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A fast video review method based on time series prediction, characterized in that, It includes the following steps: Step S1: Construct a temporal state prediction model, perform a global scan on the initial frame of the video stream, and call the target detection algorithm to identify and locate the target; Step S2: Extract the state information of the target in the initial frame to form an initial state vector , and initialize the temporal state prediction model; Step S3: Use the initialized temporal state prediction model to perform iterative loop processing on each frame after the initial frame to complete video review.

2. The fast video review method based on time series prediction according to claim 1, wherein, In the step S2, the state information of the target in the initial frame includes but is not limited to: the position information of the target of interest , the size information , and the motion parameters estimated by multi-frame data, where the motion parameters include but are not limited to the velocity information , the acceleration information .

3. A fast video review method based on time series prediction according to claim 1, characterized in that The said Step S3 includes the following steps: S31: Define each frame after the initial frame as the current frame , based on the correction status of the previous frame , predict the status of the current frame through the timing status prediction model ; ; S32: According to the predicted state 's position information and size information, generate a region of interest in the current frame, and the size of the region of interest is the product of the predicted size and the magnification factor , where ; S33: Perform target analysis within the region of interest, and the target analysis includes but is not limited to: precise positioning, feature extraction, and classification recognition; S34: If the target is detected within the region of interest, obtain the measurement state of the target and use the modified timing state prediction model to input the posterior state as the input for the next frame prediction.

4. A fast video review method based on time series prediction according to claim 2, characterized in that, The generation of the region of interest meets the following conditions: The width of the region of interest is: , and the height is: , where .

5. A fast video review method based on time series prediction according to claim 3, characterized in that, The target analysis includes but is not limited to: Precisely locate the target within the region of interest using a deep learning-based detector; Extract the texture, color, or shape features of the target to enhance the tracking robustness; Identify the category attributes of the target through a classification model.

6. The fast video review method based on time series prediction according to claim 1, wherein The model correction adopts the weighted fusion of the predicted state and the measured state by calculating the Kalman gain, or updates the weight parameters of the deep learning model through the backpropagation algorithm.

7. A fast video review method based on time series prediction according to claim 3, characterized in that The said Step S3 also includes: If consecutive frames do not detect the target or the prediction deviation exceeds the threshold, pause the region of interest analysis and trigger a global scan to recapture the target or discover a new target through the global scan, and re-initialize or terminate the tracking according to the scan result.

8. A fast video review method based on time series prediction according to claim 7, characterized in that The conditions for triggering a global scan include but are not limited to: Triggered at a fixed time interval; triggered by a target loss event; triggered by a scene mutation event.

9. An auditing system based on the fast video auditing method based on time series prediction as described in claim 1, characterized in that, It includes: An initialization module for performing global scan, target detection, and initial state extraction; A temporal prediction module for predicting the target state of the current frame based on the historical state; An ROI generation module for dynamically demarcating the region of interest; A local analysis module for performing target analysis tasks within the ROI; A state correction module for fusing prediction and measurement data to update the model; A fault tolerance control module for handling target loss events and triggering a global scan.

10. An auditing system according to claim 9, characterized in that, The said temporal prediction module supports multi-model switching, including but not limited to: Kalman filter, particle filter, and LSTM network, and dynamically selects the optimal model according to the scene complexity.

Citation Information

Patent Citations

  • Video processing method and device, video detection model training method and device and medium

    CN117274851A

  • Video behavior recognition method and device, equipment and medium

    CN118823622A

  • Intelligent bad behavior auditing method based on attitude estimation and time sequence prediction

    CN119942651A

Cited By

  • Unmanned aerial vehicle real-time detection tracking method and device based on YOLOv8 and dynamic ROI

    CN121280954A