Monitoring video frame extraction system based on cloud computing

Through multimodal data reception and cloud background purification technology, combined with optical flow tracking and hierarchical clustering, keyframe extraction and background removal of monitoring video frames are optimized, and the problems of false detection and missed detection in complex scenarios in traditional methods are solved, achieving high-precision keyframe extraction and foreground image purification.

CN120279465AInactive Publication Date: 2025-07-08昆明峰辉科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510425186.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional surveillance video frame extraction methods rely on a single feature for keyframe extraction, resulting in false detection or missed detection in complex scenarios, and the inability to effectively utilize spatial and temporal information, affecting the accuracy of keyframe extraction.

Method used

The multi-modal data receiving module is used to collect multiple monitoring video streams in real time, extract keyframes through the dynamic screening module of space-time features, and use the cloud background purification module to remove background interference. Combined with optical flow tracking, hierarchical clustering and GAN background removal technologies, keyframe selection and background elimination are optimized.

Benefits of technology

It improves the screening accuracy of keyframes in complex scenarios, reduces the false detection and missed detection rates, and improves the purity and integrity of the foreground image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279465A_ABST
    Figure CN120279465A_ABST
Patent Text Reader

Abstract

The invention discloses a monitoring video frame extraction system based on cloud computing, and the system comprises a multi-modal data receiving module which is used for collecting multiple paths of monitoring video streams in real time, and transmitting the multiple paths of monitoring video streams to a spatial-temporal feature dynamic screening module; relates to the technical field of video data processing. Through the arrangement of a spatial-temporal feature dynamic screening module, the spatial-temporal distance between any two video frames is calculated; the Euclidean distance of target centroid coordinates in two video frames, the timestamp difference value in the two video frames and the normalized difference value of target movement speeds in the two video frames are comprehensively considered in the space-time distance, the limitation of traditional single feature analysis is broken through in combination with hierarchical clustering and feature activeness evaluation, and the accuracy of feature activeness evaluation is improved. Through hierarchical clustering, representativeness and feature activeness of the key frames are ensured, and multi-dimensional changes of colors, textures and motion are synthesized, so that the screening precision of the key frames in a complex scene is improved; and through contour semantic analysis and secondary screening, false detection or missing detection of the key frame is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data processing, and particularly to a monitoring video frame extraction system based on cloud computing. Background Art

[0002] Monitoring video refers to a video stream that is collected in real time by security monitoring devices for security protection, behavior monitoring, or data recording.

[0003] Traditional methods mostly rely on single features such as optical flow method or color histogram for key frame extraction. Publication No. CN111723713B discloses a video key frame extraction method based on the optical flow method, which triggers key frame extraction by tracking the number of feature points through the optical flow method, but does not comprehensively consider multi-dimensional features of color, texture, and semantics. When the video scene contains complex motion or illumination changes, relying solely on the number of feature points is prone to false detection or missed detection. And Publication No. CN111723713B only processes a single video stream and cannot utilize spatio-temporal correlation information, which limits the key frame extraction accuracy in complex scenarios. Summary of the Invention

[0004] To solve the technical problems existing in the background art, the present invention proposes a monitoring video frame extraction system based on cloud computing.

[0005] A monitoring video frame extraction system based on cloud computing proposed by the present invention includes:

[0006] A multi-modal data receiving module: for collecting multiple monitoring video streams in real time and transmitting the multiple monitoring video streams to the spatio-temporal feature dynamic screening module;

[0007] A spatio-temporal feature dynamic screening module: for extracting key frames from multiple monitoring video streams and transmitting the key frames to the cloud background purification module;

[0008] A cloud background purification module: for removing the background interference of the key frames.

[0009] Preferably, in the spatio-temporal feature dynamic screening module, the multi-modal data receiving module transmits the collected multiple monitoring video streams to the spatio-temporal feature dynamic screening module;

[0010] For each video frame in the multiple monitoring video streams, optical flow tracking and optical flow field consistency verification are performed;

[0011] Extract the spatio-temporal features of each video frame in the multiple monitoring video streams and construct a multi-dimensional feature vector;

[0012] The spatio-temporal features include spatial position, motion speed, and timestamp;

[0013] Each video frame forms a class by itself, and the spatio-temporal distance between any two video frames is calculated;

[0014] Iteratively merge the frame classes with the closest distance through the hierarchical clustering algorithm, terminate the clustering when the inter-class distance exceeds the preset threshold, and the average spatio-temporal distance is less than the preset threshold of the spatio-temporal distance;

[0015] Within each cluster, calculate the feature activity scores of all frames;

[0016] Select the frame with the highest feature activity score from each cluster as the key frame, and the remaining video frames as redundant frames;

[0017] The redundant frames enter the secondary judgment through contour semantic analysis, and the key frames are selected again. The remaining video frames after the secondary judgment are the frames to be deleted;

[0018] Transmit the key frames and the frames to be deleted to the cloud background purification module.

[0019] Preferably, in the spatio-temporal feature dynamic screening module, for each video frame in each multi-channel monitoring video stream, optical flow tracking and optical flow field consistency verification are performed as follows:

[0020] Adopt the improved Lucas-Kanade algorithm combined with ORB feature detection to extract feature points from the first frame in the video frames of the multi-channel monitoring video stream;

[0021] Track the feature points in the subsequent video frames of the first frame of the multi-channel monitoring video stream through the improved Lucas-Kanade algorithm, and calculate the motion trajectories of the feature points between adjacent frames;

[0022] By continuously updating the positions of the feature points, obtain the optical flow vectors of each feature point;

[0023] Adopt the contour-optical flow joint matching algorithm to perform consistency verification on each frame of video data in the multi-channel monitoring video stream.

[0024] Preferably, in the spatio-temporal feature dynamic screening module, calculate the spatio-temporal distance between any two video frames as follows:

[0025] Let (x i , y i ) and (x j , y j ) be the centroid coordinates of the targets in two video frames respectively;

[0026] Calculate the Euclidean distance of the centroid coordinates of the targets in the two video frames, denoted as D 空间 ;

[0027] Calculate the time stamp difference in the two video frames, denoted as D 时间 ;

[0028] Calculate the normalized difference in the target motion speed between two video frames, denoted as D 运动 ;

[0029] For D 空间 , D 时间 , D 运动 Perform normalization processing, and after normalization, they are respectively denoted as D 归一 空间 , D 归一 时间 , D 归一 运动 ;

[0030] The spatio-temporal distance calculation formula D is:

[0031] D = α × D 归一 空间 + β × D 归一 时间 + γ × D 归一 运动 , and α + β + γ = 1.

[0032] Preferably, in the spatio-temporal feature dynamic screening module, within each cluster, calculate the feature activity scores of all frames as follows:

[0033] Quantify the color, texture, and motion features in the video frame to generate corresponding feature vectors;

[0034] For the color feature, use a color histogram to statistically analyze the pixel distribution of different color channels;

[0035] For the texture feature, use a gray-level co-occurrence matrix to extract the contrast and correlation features of the texture;

[0036] For the motion feature, represent it through the statistical information of the optical flow vector;

[0037] Calculate the change rate of each feature vector between adjacent frames as the activity index of the feature vector;

[0038] Assign weights to each feature vector;

[0039] Sum the activity indices of each feature vector multiplied by the corresponding weights to obtain the feature activity score of the video frame.

[0040] Preferably, in the spatio-temporal feature dynamic screening module, redundant frames enter the secondary judgment through contour semantic analysis and key frames are selected again as follows:

[0041] Extract the contour of each video frame of the redundant frame, and use an edge detection algorithm combined with morphological processing to obtain the contour of the target;

[0042] Then convert the contour to a chain code;

[0043] Classify and segment different objects in the video frames through semantic segmentation technology;

[0044] Adopt a contour chain code matching algorithm to calculate the similarity of object contours between adjacent frames;

[0045] When the contour similarity of the video frame is lower than the set threshold, this video frame is a key frame.

[0046] Preferably, the cloud background purification module includes:

[0047] Multi-reference frame fusion unit: Receive the frames to be deleted and key frames in the spatio-temporal feature dynamic screening module, and select multiple frames from the frame queue to be deleted to generate a dynamic background model;

[0048] GAN background generator: Based on the CycleGAN architecture, remove the background from the key frames to generate foreground images;

[0049] Background residual correction unit: Calculate the difference between the original key frame and the foreground image generated by GAN to obtain a residual image, perform edge detection and morphological processing on the residual image, extract the background residual area, and use bidirectional interpolation or texture synthesis to repair the background residual area, eliminate the background residual, and generate the final foreground image.

[0050] A method for extracting monitoring video frames based on cloud computing, including the following steps:

[0051] S1. Real-time collect multiple monitoring video streams, and upload them to the cloud after preprocessing by edge nodes;

[0052] S2. For each video frame in each multiple monitoring video stream, perform optical flow tracking and optical flow field consistency verification;

[0053] Extract the spatio-temporal features of each video frame in the multiple monitoring video streams, and construct a multi-dimensional feature vector;

[0054] The spatio-temporal features include spatial position, motion speed, and timestamp;

[0055] Each video frame forms a class by itself, and calculate the spatio-temporal distance between any two video frames;

[0056] Iteratively merge the frame classes with the closest distance through a hierarchical clustering algorithm, terminate the clustering when the inter-class distance exceeds the preset threshold, and the average spatio-temporal distance is less than the preset threshold of the spatio-temporal distance;

[0057] Within each cluster, calculate the feature activity scores of all frames;

[0058] Select the frame with the highest feature activity score from each cluster as the key frame, and the remaining video frames as redundant frames;

[0059] The redundant frames enter the secondary judgment through contour semantic analysis, and key frames are selected again. The video frames remaining after the secondary judgment are frames to be deleted;

[0060] S3. Select multiple frames from the queue of frames to be deleted to generate a dynamic background model;

[0061] Based on the CycleGAN architecture, remove the background from the key frames to generate foreground images;

[0062] Calculate the difference between the original key frames and the foreground images generated by the GAN to obtain the residual images. Perform edge detection and morphological processing on the residual images, extract the background residue areas, and use bidirectional interpolation or texture synthesis to repair the background residue areas, eliminate the background residue, and generate the final foreground images.

[0063] In the present invention, the proposed monitoring video frame extraction system based on cloud computing has the following beneficial technical effects:

[0064] 1. Through the setting of the spatio-temporal feature dynamic screening module, by calculating the spatio-temporal distance between any two video frames, when screening the key frames for the first time, the spatio-temporal distance comprehensively considers the Euclidean distance of the target centroid coordinates in the two video frames, the time stamp difference in the two video frames, and the normalized difference of the target motion speed in the two video frames. Combining hierarchical clustering and feature activity evaluation, it breaks through the limitations of traditional single-feature analysis. Through hierarchical clustering, it ensures the representativeness of the key frames. The feature activity synthesizes multi-dimensional changes in color, texture, and motion. The selection of representative frames combines the activity score and the spatio-temporal distance. While ensuring the similarity of frames within the cluster, it is easier to capture significant change points and improve the screening accuracy of key frames in complex scenarios; and through contour semantic analysis for secondary screening, it effectively responds to scene changes and reduces the false detection or missed detection of key frames.

[0065] 2. In the cloud background purification module, when using GAN for background removal, there are still some background residue problems. Through the background residual correction unit, calculate the residual images, extract the background residue areas, and use bidirectional interpolation or texture synthesis to repair the background residue areas, eliminate the background residue, and improve the purity of the foreground images.

[0066] The additional aspects and advantages of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is the principle block diagram of the system of the present invention;

[0068] Figure 2 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0069] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0070] As Figure 1 shown, a monitoring video frame extraction system based on cloud computing includes:

[0071] Multi-modal data receiving module: used to collect multiple monitoring video streams in real time and transmit the multiple monitoring video streams to the spatio-temporal feature dynamic screening module;

[0072] Spatio-temporal feature dynamic screening module: used to extract key frames from multiple monitoring video streams and transmit the key frames to the cloud background purification module;

[0073] Cloud background purification module: used to remove the background interference of the key frames.

[0074] The cloud background purification module includes:

[0075] Multi-reference frame fusion unit: receives the frames to be deleted and key frames in the spatio-temporal feature dynamic screening module, and selects multiple frames from the frame queue to be deleted to generate a dynamic background model;

[0076] GAN background generator: based on the CycleGAN architecture, removes the background from the key frames to generate a foreground image;

[0077] Background residual correction unit: calculates the difference between the original key frame and the foreground image generated by GAN to obtain a residual image, performs edge detection and morphological processing on the residual image, extracts the background residual area, and uses bidirectional interpolation or texture synthesis to repair the background residual area, eliminates the background residual, and generates the final foreground image. Makes the foreground image more complete and pure, and further eliminates the background residual.

[0078] The GAN background generator is an image processing module based on a generative adversarial network, used to generate a pure foreground image without background from video frames with background interference;

[0079] In the cloud background purification module, when using GAN for background removal, there are still some background residual problems. Through the background residual correction unit, the residual image is calculated, the background residual area is extracted, and bidirectional interpolation or texture synthesis is used to repair the background residual area, eliminates the background residual, and improves the purity of the foreground image.

[0080] In an alternative embodiment, in the cloud background purification module, texture features of the foreground image generated by the GAN are extracted to obtain a foreground texture feature map; the texture features of the background in the original video frame are analyzed to construct a background texture model; according to the foreground texture feature map and the background texture model, a texture synthesis algorithm is used to adjust the texture of the foreground image to make it more consistent with the texture of the surrounding environment.

[0081] In an alternative embodiment, in the cloud background purification module, based on cosine similarity, multiple frames that are the most similar are selected from the queue of frames to be deleted as background references;

[0082] An initial background model is generated through median filtering, and then combined with a Gaussian mixture model to dynamically update the background statistical characteristics, adapt to illumination changes and dynamic backgrounds, and realize the modeling of a dynamic background model.

[0083] In an alternative embodiment, in the spatio-temporal feature dynamic screening module, the multi-modal data receiving module transmits the collected multiple monitoring video streams to the spatio-temporal feature dynamic screening module;

[0084] For each video frame in each multiple monitoring video stream, optical flow tracking and optical flow field consistency verification are performed;

[0085] The spatio-temporal features of each video frame in the multiple monitoring video streams are extracted, including spatial position, motion speed, and timestamp, to construct a multi-dimensional feature vector;

[0086] The spatial position includes the centroid coordinates of the target;

[0087] The motion speed is calculated through the optical flow vector;

[0088] Each video frame forms a class by itself, and the spatio-temporal distance between any two video frames is calculated;

[0089] Through the hierarchical clustering algorithm, the frame classes with the closest distance are iteratively merged, and the clustering is terminated when the inter-class distance exceeds a preset threshold. The frames within the generated clustering clusters ensure overall spatio-temporal similarity through multiple merging processes, and the average spatio-temporal distance is less than the preset threshold of the spatio-temporal distance;

[0090] Within each cluster, the feature activity scores of all frames are calculated;

[0091] The frame with the highest feature activity score is selected from each cluster as the key frame, and the remaining video frames are redundant frames;

[0092] The redundant frames are subjected to secondary judgment through contour semantic analysis, and key frames are selected again. The remaining video frames after the secondary judgment are frames to be deleted;

[0093] The key frames and the frames to be deleted are transmitted to the cloud background purification module;

[0094] Through the above detailed process, the spatio-temporal feature dynamic screening module can extract key frames more accurately in complex surveillance video scenarios.

[0095] In an optional embodiment, in the spatio-temporal feature dynamic screening module, for each video frame in each multi-channel surveillance video stream, optical flow tracking and optical flow field consistency verification are performed as follows:

[0096] An improved Lucas-Kanade algorithm combined with ORB feature detection is used to extract feature points from the first frame of the video frames in the multi-channel surveillance video stream;

[0097] The Lucas-Kanade algorithm is an optical flow calculation method based on differences, used to estimate the pixel motion between adjacent image frames. The improved Lucas-Kanade algorithm adds an anti-occlusion mechanism on the basis of the Lucas-Kanade algorithm. For example, by setting a tracking confidence threshold for feature points, when the tracking confidence is lower than this threshold, it is considered that the feature point may be occluded and the tracking of it is paused;

[0098] ORB feature detection is an image feature detection and description algorithm, and ORB feature detection is used to supplement feature point information and enhance the robustness of the improved Lucas-Kanade algorithm under different illuminations and viewpoints;

[0099] The improved Lucas-Kanade algorithm is used to track the feature points in the subsequent video frames of the first frame of the multi-channel surveillance video stream, and calculate the motion trajectory of the feature points between adjacent frames;

[0100] By continuously updating the positions of the feature points, the optical flow vector of each feature point is obtained;

[0101] The optical flow vector reflects the motion speed and direction of the feature points;

[0102] A contour-optical flow joint matching algorithm is used to perform consistency verification on each frame of video data in the multi-channel surveillance video stream;

[0103] Specifically, the optical flow vector is combined with the motion information of the target contour to determine whether the optical flow field conforms to the motion pattern of the target. When the motion direction of the target contour is inconsistent with the directions of most optical flow vectors, it is determined that there is mis-tracking or background interference, and at this time, the optical flow vector is corrected or removed;

[0104] In an optional embodiment, in the spatio-temporal feature dynamic screening module, the spatio-temporal distance between any two video frames is calculated as follows:

[0105] Let (x i , y i ) and (x j , yj ) are the centroid coordinates of the target in two video frames respectively;

[0106] Calculate the Euclidean distance between the centroid coordinates of the target in two video frames, denoted as D 空间 ;

[0107] Calculate the time stamp difference between two video frames, denoted as D 时间 ;

[0108] Calculate the normalized difference of the target motion speed in two video frames, denoted as D 运动 ;

[0109] For D 空间 and D 时间 and D 运动 Perform normalization processing, and after normalization, they are denoted as D 归一 空间 and D 归一 时间 and D 归一 运动 ;

[0110] The spatio-temporal distance calculation formula D is:

[0111] D = α × D 归一 空间 + β × D 归一 时间 + γ × D 归一 运动 , and α + β + γ = 1;

[0112] In an optional embodiment, α = 0.4, β = 0.3, γ = 0.3;

[0113] This method comprehensively considers the influence of spatial distance and time interval. For example, when the spatial distance is similar, the shorter the time interval, the smaller the spatio-temporal distance;

[0114] In an optional embodiment, in the spatio-temporal feature dynamic screening module, within each cluster, calculate the feature activity scores of all frames as follows:

[0115] Quantify the color, texture, and motion features in the video frame to generate corresponding feature vectors;

[0116] For color features, use color histograms to statistically analyze the pixel distribution of different color channels;

[0117] For texture features, use gray-level co-occurrence matrices to extract texture contrast and correlation features;

[0118] For motion features, represent them through the statistical information of optical flow vectors, and the statistical information such as average speed and speed variance;

[0119] Calculate the change rate of each eigenvector between adjacent frames as the activity index of the eigenvector;

[0120] For example, the activity of color features can be calculated through the difference in color histograms, and the activity of texture features can be measured through the change in gray-level co-occurrence matrices.

[0121] Assign weights to each eigenvector, and adjust the size of the weights according to different scenarios and requirements. For example, in a scenario where moving objects are of concern, the weight of motion features can be appropriately increased.

[0122] Sum the activity indices of each eigenvector multiplied by the corresponding weights to obtain the feature activity score of the video frame;

[0123] In an optional embodiment, in the spatio-temporal feature dynamic screening module, redundant frames enter the secondary judgment through contour semantic analysis, and key frames are selected again as follows:

[0124] Extract the contour of each video frame of the redundant frames, and use an edge detection algorithm combined with morphological processing to obtain the contour of the target;

[0125] Then convert the contour into a chain code;

[0126] The chain code can concisely represent the shape information of the contour, facilitating subsequent matching and analysis;

[0127] Through semantic segmentation technology, different targets in the video frame are classified and segmented;

[0128] Identify different types of targets such as people, vehicles, and objects, and assign corresponding semantic labels to each target. Through contour semantic segmentation, the motion and changes of the target can be analyzed more accurately.

[0129] Adopt a contour chain code matching algorithm to calculate the similarity of the target contours between adjacent frames;

[0130] When using the contour chain code matching algorithm to calculate the similarity of the target contours between adjacent frames, use a time warping algorithm to handle the stretching and deformation of the contour sequence on the time axis and find the optimal matching path;

[0131] When the contour similarity of the video frame is lower than the set threshold, the video frame is a key frame;

[0132] The edge detection algorithm includes the Canny edge detection algorithm.

[0133] Through the setting of the spatio-temporal feature dynamic screening module, by calculating the spatio-temporal distance between any two video frames, when screening key frames for the first time, the spatio-temporal distance comprehensively considers the Euclidean distance of the target centroid coordinates in the two video frames, the time stamp difference in the two video frames, and the normalized difference of the target motion speed in the two video frames. Combining hierarchical clustering and feature activity evaluation, it breaks through the limitations of traditional single-feature analysis. By hierarchical clustering, it ensures the representativeness of key frames. The feature activity synthesizes multi-dimensional changes in color, texture, and motion. The selection of representative frames combines the activity score and spatio-temporal distance. While ensuring the similarity of frames within the cluster, it is easier to capture significant change points and improve the screening accuracy of key frames in complex scenarios; and through contour semantic analysis for secondary screening, it effectively deals with scene changes and reduces the false detection or missed detection of key frames.

[0134] Such as Figure 2 A method for extracting surveillance video frames based on cloud computing shown below includes the following steps:

[0135] S1. Real-time collect multiple surveillance video streams and upload them to the cloud after preprocessing by edge nodes;

[0136] S2. For each video frame in each multiple surveillance video stream, perform optical flow tracking and optical flow field consistency verification;

[0137] Extract the spatio-temporal features of each video frame in the multiple surveillance video streams and construct a multi-dimensional feature vector;

[0138] The spatio-temporal features include spatial position, motion speed, and time stamp;

[0139] Each video frame forms a class by itself, and calculate the spatio-temporal distance between any two video frames;

[0140] Iteratively merge the frame classes with the closest distance through the hierarchical clustering algorithm. When the inter-class distance exceeds the preset threshold, terminate the clustering, and the average spatio-temporal distance is less than the preset threshold of the spatio-temporal distance;

[0141] Within each cluster, calculate the feature activity scores of all frames;

[0142] Select the frame with the highest feature activity score from each cluster as the key frame, and the remaining video frames as redundant frames;

[0143] The redundant frames enter the secondary judgment through contour semantic analysis, and select key frames again. The remaining video frames after the secondary judgment are frames to be deleted;

[0144] S3. Select multiple frames from the queue of frames to be deleted to generate a dynamic background model;

[0145] Based on the CycleGAN architecture, remove the background from the key frames to generate foreground images;

[0146] Calculate the difference between the original key frame and the foreground image generated by the GAN to obtain a residual image. Perform edge detection and morphological processing on the residual image to extract the background residual area. Use bidirectional interpolation or texture synthesis to repair the background residual area, eliminate the background residual, and generate the final foreground image.

[0147] At the same time, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0148] In the embodiments provided by the present invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the above-described invention embodiments are merely illustrative. For example, the division of modules is only a logical function division, and there can be other division methods in actual implementation.

[0149] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0151] For those skilled in the operation and maintenance in this field, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the basic characteristics of the present invention.

[0152] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A monitoring video frame extraction system based on cloud computing, characterized in that, Including: Multimodal data receiving module: It is used to collect multiple monitoring video streams in real time and transmit the multiple monitoring video streams to the spatio-temporal feature dynamic screening module; Spatio-temporal feature dynamic screening module: It is used to extract key frames from multiple monitoring video streams and transmit the key frames to the cloud background purification module; Cloud background purification module: It is used to remove the background interference of key frames.

2. The monitoring video frame extraction system based on cloud computing according to claim 1, characterized in that, In the spatio-temporal feature dynamic screening module, the multimodal data receiving module transmits the collected multiple monitoring video streams to the spatio-temporal feature dynamic screening module; For each video frame in each multiple monitoring video stream, optical flow tracking and optical flow field consistency verification are performed; Extract the spatio-temporal features of each video frame in the multiple monitoring video streams and construct a multi-dimensional feature vector; The spatio-temporal features include spatial position, motion speed, and timestamp; Each video frame forms a separate class, and the spatio-temporal distance between any two video frames is calculated; Iteratively merge the frame classes with the closest distance through the hierarchical clustering algorithm. When the inter-class distance exceeds the preset threshold, the clustering is terminated, and the average spatio-temporal distance is less than the preset threshold of the spatio-temporal distance; Within each cluster, calculate the feature activity scores of all frames; Select the frame with the highest feature activity score from each cluster as the key frame, and the remaining video frames are redundant frames; The redundant frames enter the secondary judgment through contour semantic analysis, and the key frames are selected again. The remaining video frames after the secondary judgment are frames to be deleted; Transmit the key frames and the frames to be deleted to the cloud background purification module.

3. The monitoring video frame extraction system based on cloud computing according to claim 2, characterized in that, In the spatio-temporal feature dynamic screening module, for each video frame in each multiple monitoring video stream, optical flow tracking and optical flow field consistency verification are performed as follows: Adopt the improved Lucas-Kanade algorithm combined with ORB feature detection to extract feature points from the first frame in the video frames of the multiple monitoring video streams; Track the feature points in the subsequent video frames of the first frame of the multiple monitoring video streams through the improved Lucas-Kanade algorithm, and calculate the motion trajectories of the feature points between adjacent frames; Obtain the optical flow vector of each feature point by continuously updating the position of the feature point; Adopt the contour-optical flow joint matching algorithm to perform consistency verification on each frame of video data in the multiple monitoring video streams.

4. The monitoring video frame extraction system based on cloud computing according to claim 2 or 3, characterized in that In the spatio-temporal feature dynamic screening module, calculate the spatio-temporal distance between any two video frames as follows: Let (x i , y i ) and (x j , y j ) be the centroid coordinates of the target in two video frames respectively; Calculate the Euclidean distance between the centroid coordinates of the targets in two video frames, denoted as D 空间 ; Calculate the time stamp difference between two video frames, denoted as D 时间 ; Calculate the normalized difference in the target motion speed between two video frames, denoted as D 运动 ; For D 空间 、D 时间 、D 运动 perform normalization processing, and after normalization, they are respectively counted as D 归一 空间 、D 归一 时间 、D 归一 运动 ; The spatio-temporal distance calculation formula D is: D = α × D 归一 空间 + β × D 归一 时间 + γ × D 归一 运动 , and α + β + γ = 1.

5. The monitoring video frame extraction system based on cloud computing according to claim 4, characterized in that, In the spatio-temporal feature dynamic screening module, within each cluster, calculate the feature activity scores of all frames as follows: Quantify the color, texture, and motion features within the video frame to generate corresponding feature vectors; For the color feature, use the color histogram to statistically analyze the pixel distribution of different color channels; For the texture feature, use the gray-level co-occurrence matrix to extract the contrast and correlation features of the texture; For the motion feature, it is represented by the statistical information of the optical flow vector; Calculate the change rate of each feature vector between adjacent frames as the activity index of the feature vector; Assign weights to each feature vector; Sum the activity indices of each feature vector multiplied by the corresponding weights to obtain the feature activity score of this video frame.

6. The monitoring video frame extraction system based on cloud computing according to claim 5, characterized in that, In the spatio-temporal feature dynamic screening module, the redundant frames enter the secondary judgment through contour semantic analysis, and the key frames are selected again as follows: For each video frame of the redundant frames, contour extraction is performed, and the contour of the target is obtained by combining edge detection algorithms with morphological processing; Then the contour is converted into a chain code; Through semantic segmentation technology, different targets in the video frame are classified and segmented; Using the contour chain code matching algorithm, calculate the similarity of the target contours between adjacent frames; When the contour similarity of the video frame is lower than the set threshold, the video frame is a key frame.

7. The monitoring video frame extraction system based on cloud computing according to claim 1, wherein The cloud background purification module includes: Multi-reference frame fusion unit: Receive the frames to be deleted and key frames in the spatio-temporal feature dynamic screening module, and select multiple frames from the queue of frames to be deleted to generate a dynamic background model; GAN background generator: Based on the CycleGAN architecture, remove the background from the key frame to generate a foreground image; Background residual correction unit: Calculate the difference between the original key frame and the foreground image generated by the GAN to obtain a residual image, perform edge detection and morphological processing on the residual image, extract the background residual area, and use bidirectional interpolation or texture synthesis to repair the background residual area, eliminate the background residue, and generate the final foreground image.

8. The method for extracting surveillance video frames based on cloud computing according to any one of claims 1-7, characterized in that Including the following steps: S1. Real-time collect multiple monitoring video streams, and upload them to the cloud after preprocessing by the edge node; S2. For each video frame in each multiple monitoring video stream, perform optical flow tracking and optical flow field consistency verification; Extract the spatio-temporal features of each video frame in the multiple monitoring video streams, and construct a multi-dimensional feature vector; The spatio-temporal features include spatial position, movement speed, and timestamp; Each video frame forms a class by itself, and calculate the spatio-temporal distance between any two video frames; Iteratively merge the frame classes with the closest distance through the hierarchical clustering algorithm, terminate the clustering when the inter-class distance exceeds the preset threshold, and the average spatio-temporal distance is less than the preset threshold of the spatio-temporal distance; Within each cluster, calculate the feature activity scores of all frames; Select the frame with the highest feature activity score from each cluster as the key frame, and the remaining video frames are redundant frames; The redundant frames enter the secondary judgment through contour semantic analysis, select the key frames again, and the remaining video frames after the secondary judgment are the frames to be deleted; S3. Select multiple frames from the queue of frames to be deleted to generate a dynamic background model; Based on the CycleGAN architecture, remove the background from the key frame to generate a foreground image; Calculate the difference between the original key frame and the foreground image generated by the GAN to obtain a residual image, perform edge detection and morphological processing on the residual image, extract the background residual area, and use bidirectional interpolation or texture synthesis to repair the background residual area, eliminate the background residue, and generate the final foreground image.

Citation Information

Patent Citations

  • A video keyframe extraction method and system based on optical flow

    CN111723713B