A method for extracting key video data based on multi-dimensional semantic information

By employing time-domain sampling, Gaussian mixture models, and multi-modal Transformer models, the method addresses storage and retrieval inefficiencies in video processing by extracting and representing key data in a multi-dimensional format, reducing redundancy and improving retrieval efficiency.

CN116994176BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310883076.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2025-07-15
Estimated Expiration
2043-07-18

AI Technical Summary

Technical Problem

The existing video key data extraction methods have a lot of redundant background data and coarse characterization granularity, and lack of systematic multi-dimensional information extraction and characterization, resulting in large storage space requirements and low retrieval efficiency.

Method used

Through time domain sampling, video background is constructed based on Gaussian hybrid model, key targets are extracted using single-stage object detection and target tracking algorithms, image quality scores are calculated, text summary is generated in combination with Transformer, and multi-dimensional target representation structure is constructed.

Benefits of technology

Significantly reduce storage space requirements, improve data information density, and meet application needs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994176B_ABST
    Figure CN116994176B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting key video data based on multi-dimensional semantic information. First, time-domain sampling and preprocessing are performed on the input video; then a video background is constructed based on the Gaussian mixture model; then a single-stage object detection network is used to extract and screen key objects in the non-background area of the video frames; an object tracking algorithm is used to track the key objects to obtain a sequence of object bounding boxes; the object motion information is calculated, the quality score of the image patch within each tracked bounding box is calculated, and the image patch with the largest quality score is selected as the typical object image; an object fine-grained attribute extraction model is used to extract the color and model subclass information of the object; a video description generation model based on Transformer is used to generate a text summary of the key object; finally, a multi-dimensional representation structure of the key object is constructed, and the video background and all object multi-dimensional representations are stored as key data. The present invention can greatly reduce the required storage space and improve the data information density.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video processing, and particularly relates to a method for extracting key data of a video. Background Art

[0002] In recent years, with the rapid development of video acquisition methods such as fixed sensors, smart phones, and aerial photography by drones, and the wide application of video sharing websites, the number of videos is increasing at an explosive rate. Videos record the process of social and life changes in an intuitive and vivid way. Therefore, it is of great significance to process and analyze videos and extract key data, which has great application value in many people's livelihood and economic fields such as intelligent security and criminal detection.

[0003] However, the explosively growing videos pose higher requirements for storage, transmission, and processing. For example, limited by the capacity of storage devices, surveillance videos recorded continuously for 24 hours can generally be stored for about one week, or even less time. In the face of abnormal situations, most of them retrieve videos by manually playing them back to determine whether relevant people or objects appear. This method is not only time-consuming and laborious, but also prone to problems such as missed detection and missed effective time. Therefore, how to intelligently extract key video data, greatly reduce the required storage space, and improve the retrieval efficiency is of great significance in many aspects.

[0004] The intelligent extraction of key video data aims to extract key people and objects that appear in the video on the basis of semantic analysis of the video, so as to delete irrelevant data and improve the data information density. Its main difficulties are the extraction and discovery of key targets, multi-dimensional information extraction and representation. There are a large number of static and dynamic targets in the video, plus noise interference. How to extract targets from them is a key point. On the other hand, after determining the key targets, how to mine multi-dimensional information and construct a suitable representation to improve the representation accuracy and user retrieval efficiency is another difficulty. Most of the existing methods for extracting key video data focus on key frame extraction and have not yet paid attention to the target level; or focus on single modules such as target detection and target tracking, lacking a systematic method. Summary of the Invention

[0005] To overcome the deficiencies of the prior art, the present invention provides a method for extracting key video data based on multi-dimensional semantic information. First, temporal sampling is performed on the input video for preprocessing; then, a video background is constructed based on the Gaussian mixture model; next, a single-stage object detection network is used to extract and screen key objects in the non-background regions of the video frames; an object tracking algorithm is used to track the key objects to obtain a sequence of object bounding boxes; the object motion information is calculated, the quality score of the image patch within each tracked bounding box is calculated, and an image patch with the maximum quality score is selected as the typical object image; an object fine-grained attribute extraction model is used to extract the color and model subclass information of the object; a video description generation model based on Transformer is used to generate a text summary of the key object; finally, a multi-dimensional representation structure of the key object is constructed, and the video background and all object multi-dimensional representations are stored as key data. The present invention can greatly reduce the required storage space and improve the data information density.

[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0007] Step 1: For the input video, perform temporal sampling to reduce the video frame rate to 2 FPS, and perform preprocessing operations such as white balance and color correction.

[0008] Step 2: Based on the video frame sequence obtained in Step 1, construct a video background based on the Gaussian mixture model.

[0009] Step 3: Based on the video background obtained in Step 2, use a single-stage object detection network to extract and screen key objects in the non-background regions of the video frames; use an object tracking algorithm to track the key objects to obtain a sequence of object bounding boxes.

[0010] Step 4: According to the sequence of object bounding boxes obtained in Step 3, calculate the object motion information, including the object appearance time, disappearance time, and spatio-temporal motion coordinate trajectory.

[0011] Step 5: Based on the sequence of object bounding boxes obtained in Step 3, calculate the quality score of the image patch within each bounding box, and select an image patch with the maximum quality score as the typical object image.

[0012] Step 6: Use an object fine-grained attribute extraction model to extract the color and model subclass information of the object.

[0013] Step 7: Use a video description generation model based on Transformer to generate a text summary of the key object.

[0014] Step 8: Based on the results of Steps 3 to 7, construct a multi-dimensional representation structure of the key object, and finally store the video background and all object multi-dimensional representations as key data.

[0015] Preferably, step 2 is specifically as follows:

[0016] Step 2-1: The Gaussian mixture model consists of K single Gaussian models and is used to describe the brightness distribution of a pixel point at different times through weighted summation; the video background extraction process based on the Gaussian mixture model is as follows:

[0017] Step 2-2: Gaussian mixture model initialization: Randomly initialize the means of K Gaussian distributions, set the variance to 9, and assign the weight to 1 / K;

[0018] Step 2-3: Take one frame of image at a time, compare each pixel value in the image with the means of each single Gaussian model. If the deviation is less than 2.5 times the variance, update the mean μ, standard deviation σ, and weight w of the matching Gaussian model; if neither is satisfied, discard the model with the smallest weight and newly establish a Gaussian model with its mean being the current pixel value, standard deviation being 9, and weight being the smallest weight;

[0019] Step 2-4: Arrange the Gaussian models in descending order according to the w i / σ i value; Select the first B Gaussian distributions as the background mode according to the following formula:

[0020]

[0021] where w i is the weight of the i-th Gaussian model, and the parameter T represents the proportion of the background;

[0022] Step 2-5: Judging pixel by pixel whether the pixel value matches the selected B Gaussian distributions. If it matches, the pixel is a background point; otherwise, it is a foreground;

[0023] Step 2-6: Perform a loop of steps 2-3 to 2-5 for all frames to obtain the background sequence of the current sequence, and take the average to obtain the background of this scene.

[0024] Preferably, K = 5 and T = 0.6.

[0025] Preferably, step 3 is specifically as follows:

[0026] Step 3-1: Use the single-stage object detection model yolo v8 for object detection and output the detection boxes of the objects; use the pre-trained model parameters on the COCO dataset to obtain the optimal object detection model;

[0027] Step 3-2: Apply the non-maximum suppression algorithm to filter the detection boxes obtained in the previous step to avoid generating multiple overlapping detection boxes for the same target. In non-maximum suppression, first select the detection box with the highest predicted score by the target detection model, then judge the overlap degree between other candidate boxes and the selected target box. If it is greater than the threshold T = 0.5, delete the detection box with a smaller score. Then, sequentially select the candidate box with the highest predicted score from the remaining set of detection boxes until all target boxes are traversed.

[0028] Step 3-3: Use the Sort object tracking algorithm to generate the target sequence image. The specific process is as follows: for the detection boxes at time t-1, predict their positions in the t-th frame based on the linear Kalman filter. Then, perform Hungarian matching on the position prediction results and the target detection boxes through the IoU value to obtain the tracking boxes at time t.

[0029] Preferably, the image quality score is calculated in the following way:

[0030] Step 5-1: Send the image into the neural network model of Resnet-50 pre-trained on ImageNet, and use the output of the penultimate layer as the image semantic feature vector.

[0031] Step 5-2: Convert the image into a grayscale image, and then calculate the image entropy representing the image information content:

[0032]

[0033] where p(n) represents the probability that the pixel gray value is n.

[0034] Step 5-3: Concatenate the obtained semantic feature vector and the image entropy into a vector, and then send it into an MLP network with 2 hidden layers to predict the image quality score. The hidden layer dimensions of this MLP network are 64 and 32 respectively, and the parameters of this network are trained using the image quality evaluation dataset LIVE.

[0035] Preferably, the target fine-grained attribute extraction model is a multi-classification network that fuses a 3-layer perceptron with LSTM. Send the image semantic feature vector obtained in step 5 into a 3-layer MLP network for feature refinement, then integrate the obtained result into the LSTM network, and finally output a target attribute vector att = [col, vel_cat]; col represents the target color attribute, and its values include red, white, black, blue, silver, other; vel_cat represents the fine-grained type attribute of the vehicle, including bicycle, sedan, SUV, bus, truck. When the target is not a vehicle, its value is None.

[0036] Preferably, the Transformer-based video description generation model includes a video feature extraction module and a cross-modal interaction module:

[0037] The video feature extraction module is a video Swin Transformer, and its model parameters adopt the pre-trained parameters on the Kinetics action recognition task. First, it divides the image sequence of T×H×W×3 into multiple 3D blocks of 2×4×4×3. Each 3D block passes through a linear embedding layer, multiple block merging layers, and Swin Transformer blocks in sequence, and finally generates an 8C-dimensional feature vector, where C is 96.

[0038] The cross-modal interaction module performs a sequence-to-sequence cross-modal translation task to generate the final natural language description. Its input is the vector sequence obtained in the previous step, and it generates the text description of the video in an autoregressive manner.

[0039] Preferably, the multi-dimensional representation of the target is expressed as:

[0040] data = [cat, st, et, loc, att, img, txt]

[0041] where cat represents the target category, st and et are the appearance and disappearance times of the target respectively, loc is the spatio-temporal motion trajectory, img represents the typical target image obtained in step 5, att is the target attribute vector obtained in step 6, and txt is the text summary obtained in step 7.

[0042] The beneficial effects of the present invention are as follows:

[0043] (1) By extracting key data such as the background and target in the video, the present invention finally represents the video as a multi-dimensional representation of the background image and the target, which can greatly reduce the required storage space and improve the data information density.

[0044] (2) The present invention designs a multi-dimensional target representation method. Through multiple modules such as detection, tracking, and text generation, a multi-modal target representation method containing attributes, images, texts, etc. is constructed, which can meet the application requirements of different scenarios and different users. Description of the Drawings

[0045] Figure 1 is the flowchart of the method of the present invention. Detailed Embodiments

[0046] The present invention will be further described below with reference to the drawings and embodiments.

[0047] The present invention proposes a method for extracting key video data based on multi-dimensional semantic information, which solves the defects of existing video key frame extraction methods, such as a large amount of redundant background data and coarse representation granularity. The present invention extracts key data at the target granularity and constructs a multi-dimensional information extraction and representation method, which can not only reduce redundant background data, but also effectively represent multi-dimensional information such as the motion, attributes, and appearance of key data.

[0048] To achieve the above object, the specific steps of the present invention are as follows:

[0049] Step 1: For the input video, perform temporal sampling to reduce the video frame rate to 2 FPS, and perform preprocessing operations such as white balance and color correction.

[0050] Step 2: Based on the Gaussian mixture model, construct the video background for the video frame sequence obtained in Step 1.

[0051] The Gaussian mixture model consists of K single Gaussian models and is used to describe the brightness distribution of a pixel point at different times through weighted summation. Specifically, K is set to 5. The process of extracting the video background based on the Gaussian mixture model is as follows:

[0052] 1) Initialization. Randomly initialize the means of the K Gaussian distributions, set the variance to 9, and assign the weight to 1 / K;

[0053] 2) Take a single frame of image at a time, compare each pixel value with the means of the single Gaussian models. If the deviation is less than 2.5 times the variance, update the mean μ, standard deviation σ, and weight w of the matching Gaussian model; if neither is satisfied, discard the model with the smallest weight and newly establish a Gaussian model with its mean as the current pixel value, standard deviation of 9, and weight as the smallest weight;

[0054] 3) Sort the Gaussian models in descending order according to the value of w i / σ i Select the first B Gaussian distributions as the background model according to the following formula

[0055]

[0056] where w i is the weight of the i-th Gaussian model, and the parameter T = 0.6 represents the proportion of the background.

[0057] 4) Judging pixel by pixel whether the pixel value matches the selected B Gaussian distributions. If it matches, the pixel is a background point; otherwise, it is a foreground;

[0058] 5) Loop Steps 2 - 4 for all frames to obtain the background sequence of the current sequence, and take the average to obtain the background of the scene.

[0059] Step 3: Based on the video background obtained in Step 2, use a single-stage object detection network to extract and screen key objects in the non-background area of the video frames; use an object tracking algorithm to track the key objects and obtain a sequence of bounding boxes of the objects.

[0060] 1) Use the single-stage object detection model YOLO v8 for object detection and output the detection boxes of the objects. Further, use the pre-trained model parameters on the COCO dataset to obtain the optimal object detection model.

[0061] 2) Adopt the non-maximum suppression algorithm to screen the detection boxes obtained in the previous step to avoid generating multiple overlapping detection boxes for the same object. In non-maximum suppression, first select the detection box with the largest prediction score of the object detection model, and then judge the overlap degree between other candidate boxes and the selected object box. If it is greater than the threshold T = 0.5, delete the object box with a smaller score; then sequentially select the candidate box with the largest prediction score from the remaining set of detection boxes until all object boxes are traversed.

[0062] 3) Adopt the Sort object tracking algorithm to generate a sequence of object images. The specific process is to predict the position in the t-th frame based on the linear Kalman filter for the detection boxes at the (t - 1)-th moment. Then, perform Hungarian matching on the position prediction result and the object detection box through the IoU value to obtain the tracking box at the t-th moment.

[0063] Step 4: According to the sequence of object bounding boxes obtained in Step 3, calculate the object motion information, including the appearance time, disappearance time, and spatio-temporal motion coordinate trajectory of the object.

[0064] Step 5: Based on the sequence of object bounding boxes obtained in Step 3, calculate the quality score of the image patch within each bounding box, and select an image patch with the largest quality score as the typical object image.

[0065] The image quality score is calculated in the following way:

[0066] 1) Send the image into the neural network model of ResNet-50 pre-trained on ImageNet, and use the output of the penultimate layer as the image semantic feature vector.

[0067] 2) Convert the image into a grayscale image. Then calculate the image entropy representing the amount of information in the image

[0068]

[0069] where p(n) represents the probability that the pixel grayscale value is n.

[0070] 3) Concatenate the obtained semantic feature vector and image entropy into a vector, and then feed it into an MLP network with two hidden layers to predict the quality score of the image. The hidden layer dimensions of this MLP network are 64 and 32 respectively, and the parameters of the network are trained using the LIVE image quality assessment dataset.

[0071] Step 6: Use the target fine-grained attribute extraction model to extract the color and model subclass information of the target.

[0072] The target fine-grained attribute extraction model is a multi-classification network that fuses a 3-layer perceptron with LSTM. Feed the image semantic feature vector obtained in Step 5 into a 3-layer MLP network for feature refinement, then integrate the obtained result into the LSTM network, and finally output a multi-dimensional target attribute vector att = [col, vel_cat]. col represents the target color attribute, and its values include red, white, black, blue, silver, and others. vel_cat represents the fine-grained type attribute of the vehicle, including bicycle, car, SUV, bus, truck. When the target is not a vehicle, its value is None.

[0073] Step 7: Use the Transformer-based video description generation model to generate a text summary of the key target.

[0074] The Transformer-based video description generation model mentioned above includes a video feature extraction module and a cross-modal interaction module.

[0075] 1) The video feature extraction module is a video Swin Transformer, and the model parameters adopt the pre-trained parameters on the Kinetics action recognition task. First, it divides the image sequence of T×H×W×3 into multiple 3D blocks of 2×4×4×3. Each 3D block passes through a linear embedding layer, multiple block merge layers, and Swin Transformer blocks in sequence, and finally generates an 8C-dimensional feature vector. Generally, C is taken as 96.

[0076] 2) The cross-modal interaction module performs a sequence-to-sequence cross-modal translation task to generate the final natural language description. Its input is the vector sequence obtained in the previous step, and it generates the text description of the video in an autoregressive manner.

[0077] Step 8: Based on the results of Steps 3-7, construct a multi-dimensional representation structure of the key target, and finally store the video background and all target multi-dimensional representations as key data.

[0078] The multi-dimensional representation of the target can be expressed as

[0079] data = [cat, st, et, loc, att, img, txt]

[0080] Among them, cat represents the target category, st and et are the target appearance and disappearance times respectively, loc is the spatio-temporal motion trajectory, img represents the target typical image obtained in step 5, att is the target attribute vector obtained in step 6, and txt is the text summary obtained in step 7.

Claims

1. A method for extracting key video data based on multi-dimensional semantic information, characterized in that, It includes the following steps: Step 1: For the input video, perform temporal sampling to reduce the video frame rate to 2 FPS, and perform preprocessing operations such as white balance and color correction; Step 2: Based on the video frame sequence obtained in Step 1, construct a video background using a Gaussian mixture model; Step 3: Based on the video background obtained in Step 2, use a single-stage object detection network to extract and screen key objects in the non-background area of the video frame; Use an object tracking algorithm to track the key objects to obtain a sequence of object bounding boxes; Step 4: According to the sequence of object bounding boxes obtained in Step 3, calculate the object motion information, including the object appearance time, disappearance time, and spatio-temporal motion coordinate trajectory; Step 5: Based on the sequence of object bounding boxes obtained in Step 3, calculate the quality score of the image patch within each bounding box, and select an image patch with the maximum quality score as the typical object image; Step 5-1: Send the image into the neural network model of Resnet-50 pre-trained on ImageNet, and use the output of the penultimate layer as the image semantic feature vector; Step 5-2: Convert the image to a grayscale image, and then calculate the image entropy representing the amount of information in the image: wherein represents the probability that the pixel gray value is n; Step 5-3: Concatenate the obtained semantic feature vector and image entropy into a vector, and then send it into an MLP network with 2 hidden layers to predict the quality score of the image; the hidden layer dimensions of this MLP network are 64 and 32 respectively, and the parameters of this network are trained using the image quality assessment dataset LIVE; Step 6: Use an object fine-grained attribute extraction model to extract the color and model subclass information of the object; Step 7: Use a video description generation model based on Transformer to generate a text summary of the key object; Step 8: Based on the results of Steps 3 to 7, construct a multi-dimensional representation structure of the key object, and finally store the video background and all object multi-dimensional representations as key data.

2. The video key data extraction method based on multi-dimensional semantic information according to claim 1, wherein The specific content of Step 2 is as follows: Step 2-1: The Gaussian mixture model consists of K single Gaussian models, and is used to describe the brightness distribution of a pixel point at different times through weighted summation; the video background extraction process based on the Gaussian mixture model is as follows: Step 2-2: Gaussian mixture model initialization: Randomly initialize the means of K Gaussian distributions, set the variance to 9, and assign the weight to 1 / K; Step 2-3: Take one frame of image at a time, compare each pixel value in the image with the mean of each single Gaussian model. If the deviation is less than 2.5 times the variance, update the mean μ, standard deviation σ and weight of the matching Gaussian model. If neither condition is satisfied, discard the model with the smallest weight and newly establish a Gaussian model with its mean being the current pixel value, standard deviation being 9, and weight being the smallest weight. Step 2-4: Arrange each Gaussian model in descending order according to the value; Select the first B Gaussian distributions as the background mode according to the following formula: Among them, is the weight of the i-th Gaussian model, and the parameter T represents the proportion of the background; Step 2-5: Judging pixel by pixel whether the pixel value matches the selected B Gaussian distributions. If it matches, the pixel is a background point, otherwise it is a foreground; Step 2-6: Perform the loop of Steps 2-3 to 2-5 for all frames to obtain the background sequence of the current sequence, and calculate the average to obtain the background of this scene.

3. A method for extracting key video data based on multi-dimensional semantic information according to claim 2, characterized in that The said K = 5, T = 0.

6.

4. A method for extracting key video data based on multi-dimensional semantic information according to claim 2, characterized in that The specific content of Step 3 is as follows: Step 3-1: Use the single-stage object detection model yolo v8 for object detection and output the detection boxes of the objects; use the pre-trained model parameters on the COCO dataset to obtain the optimal object detection model; Step 3-2: Apply the non-maximum suppression algorithm to filter the detection boxes obtained in the previous step to avoid generating multiple overlapping detection boxes for the same target. In non-maximum suppression, first select the detection box with the highest predicted score by the target detection model, then judge the overlap degree between other candidate boxes and the selected target box. If it is greater than the threshold T = 0.5, delete the target box with a smaller score. Then, sequentially select the candidate box with the highest predicted score from the remaining detection box set until all target boxes are traversed. Step 3-3: Generate the target sequence image using the Sort object tracking algorithm; the specific process is to predict the position in the frame for the detection box at time t-1 based on the linear Kalman filter; then perform Hungarian matching on the position prediction result and the target detection box through the IoU value to obtain the tracking box at time t.

5. A method for extracting key video data based on multi-dimensional semantic information according to claim 1, characterized in that, The target fine-grained attribute extraction model is a multi-classification network that combines a 3-layer perceptron with LSTM. The image semantic feature vector obtained in Step 5 is sent into a 3-layer MLP network for feature refinement, and then the obtained result is integrated into the LSTM network. Finally, a target attribute vector att = [col, vel_cat] is output. col represents the target color attribute, and its values include red, white, black, blue, silver, and others; vel_cat represents the fine-grained type attribute of the vehicle, including bicycle, sedan, SUV, bus, truck. When the target is not a vehicle, its value is None.

6. A method for extracting key video data based on multi-dimensional semantic information according to claim 1, characterized in that, The Transformer-based video description generation model includes a video feature extraction module and a cross-modal interaction module: The video feature extraction module is a video Swin Transformer, and its model parameters adopt the pre-trained parameters on the Kinetics action recognition task. First, it divides the T×H×W×3 image sequence into multiple 2×4×4×3 3D blocks. Each 3D block passes through a linear embedding layer, multiple block merging layers, and Swin Transformer blocks in sequence, and finally generates an 8C-dimensional feature vector, where C is 96. The cross-modal interaction module performs a sequence-to-sequence cross-modal translation task to generate the final natural language description. Its input is the vector sequence obtained in the previous step, and it generates the text description of the video in an autoregressive manner.

7. A method for extracting key video data based on multi-dimensional semantic information according to claim 1, characterized in that, The multi-dimensional representation of the target is expressed as: Among them, cat represents the target category, st and et are the target appearance and disappearance times respectively, loc is the spatio-temporal motion trajectory, img represents the typical target image obtained in step 5, att is the target attribute vector obtained in step 6, txt is the text summary obtained in step 7.

Citation Information

Patent Citations

  • Video monitoring method and device

    CN107872644A

  • Live broadcast method, electronic device and readable storage medium

    CN111556332A