A key frame extraction method, device, equipment and medium
By extracting motion and color features from video frame sequences, adaptively setting keyframe thresholds, and using a target detection model to accurately extract keyframes, the problems of time-consuming, labor-intensive, and inaccurate methods in existing technologies are solved, enabling accurate localization and automatic extraction of moving targets in video frame sequences.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AGRICULTURAL BANK OF CHINA
- Filing Date
- 2022-09-16
- Publication Date
- 2026-07-03
AI Technical Summary
In existing technologies, keyframe extraction from video suffers from problems such as being time-consuming and labor-intensive, being limited by human physiological characteristics, having inflexible thresholds, and not fully utilizing video frame information, resulting in inaccurate and incomplete extraction.
By acquiring the video frame sequence of the moving target, motion and color features are extracted, the inter-frame difference index is determined, and the key frame threshold is adaptively set according to the inter-frame difference index and the total number of frames. The key frames are then accurately extracted using the target detection model.
It achieves precise positioning and capture of moving targets in video frame sequences, and automatically and flexibly extracts key frames, solving the problems of inflexibility and inaccuracy in existing technologies.
Smart Images

Figure CN115471772B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and medium for extracting keyframes. Background Technology
[0002] With the development of internet and IoT technologies, the number of videos is growing exponentially, estimated to increase 50-fold every 10 years. However, videos suffer from drawbacks such as high variability, large data volume, and low abstraction, making keyframe extraction technology increasingly important.
[0003] Currently, for videos containing moving targets (people, vehicles, etc.), keyframe extraction is generally performed manually, i.e., by manually identifying keyframes. However, manual keyframe identification is time-consuming and labor-intensive, and is limited by human physiological characteristics, leading to errors or omissions in keyframe extraction. Furthermore, some non-manual keyframe extraction methods suffer from inflexible keyframe thresholds and insufficient utilization of video frame information, resulting in incomplete or inaccurate keyframe extraction. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and medium for extracting keyframes, which can adaptively set keyframe thresholds and flexibly and accurately filter out keyframes.
[0005] According to one aspect of the present invention, a method for extracting keyframes is provided, comprising:
[0006] The process involves acquiring a sequence of video frames containing moving targets, extracting motion and color features from the sequence, and obtaining multidimensional feature extraction results.
[0007] The inter-frame difference index is determined based on the multi-dimensional feature extraction results and the number of video shots that match the video frame sequence to be processed.
[0008] The keyframe threshold is determined based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed.
[0009] The target key video frames in the video frame sequence to be processed are determined by using the target detection model, the inter-frame difference index, and the key frame threshold.
[0010] According to another aspect of the present invention, a keyframe extraction apparatus is provided, comprising:
[0011] The feature extraction module is used to acquire the video frame sequence of the moving target and extract motion and color features from the video frame sequence to obtain multi-dimensional feature extraction results.
[0012] The inter-frame difference index determination module is used to determine the inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots that match the video frame sequence to be processed.
[0013] The keyframe threshold determination module is used to determine the keyframe threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed.
[0014] The target key video frame determination module is used to determine the target key video frames in the video frame sequence to be processed by using the target detection model, the inter-frame difference index, and the key frame threshold.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the keyframe extraction method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the keyframe extraction method described in any embodiment of the present invention.
[0020] The technical solution of this invention acquires a sequence of video frames containing a moving target, extracts motion and color features from the sequence to obtain multi-dimensional feature extraction results, and then determines an inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots matching the sequence. Based on the inter-frame difference index and the total number of frames in the sequence, a keyframe threshold is determined. Furthermore, a target detection model, the inter-frame difference index, and the keyframe threshold are used to identify the target's key video frames in the sequence. This solution extracts motion and color features to uncover potential key features of moving targets in the sequence and adaptively generates keyframe thresholds based on the degree of difference between video frames and the total number of frames in the sequence. This achieves accurate localization and capture of moving targets in the sequence, enabling automatic and flexible extraction of target key video frames. It solves the problems of inflexible and inaccurate keyframe extraction in existing technologies, and allows for adaptive setting of keyframe thresholds to flexibly and accurately select key frames.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart illustrating a keyframe extraction method provided in Embodiment 1 of the present invention;
[0024] Figure 2 A flowchart illustrating a keyframe extraction method provided in Embodiment 2 of the present invention;
[0025] Figure 3 This is a flowchart of another keyframe extraction method provided in Embodiment 2 of the present invention;
[0026] Figure 4 This is a schematic diagram of a keyframe extraction device provided in Embodiment 3 of the present invention;
[0027] Figure 5 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1 This is a flowchart of a keyframe extraction method provided in Embodiment 1 of the present invention. This embodiment is applicable to keyframe extraction from videos containing moving targets. The method can be executed by a keyframe extraction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0032] S110. Obtain the video frame sequence of the moving target to be processed, and extract motion features and color features from the video frame sequence to be processed to obtain multi-dimensional feature extraction results.
[0033] The moving target can be any moving object, determined according to actual tracking needs. Optionally, the moving target can be a person and / or a vehicle, etc. The video frame sequence to be processed can be a sequence of video frames associated with the moving target, used to extract the key video frames where the moving target is located. The video frame sequence to be processed can include multiple video frames. Motion features can be behavioral features occupying time and space. Color features can be visual features in color space. The multidimensional feature extraction result can be the feature extraction result obtained after extracting motion and color features from the video frame sequence to be processed.
[0034] In this embodiment of the invention, a video can be acquired according to the need for key video frame detection of a video including a moving target, and then the acquired video can be converted to obtain a video frame sequence to be processed. Motion features and color features can then be extracted from the video frame sequence to be processed to obtain multi-dimensional feature extraction results.
[0035] S120. Determine the inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots that match the video frame sequence to be processed.
[0036] The number of video shots can be the total number of shots used to complete video recording. The inter-frame difference index can be used to characterize the degree of difference between two consecutive video frames in the video frame sequence to be processed.
[0037] In this embodiment of the invention, the dimensions of the multidimensional feature extraction results can be unified, and then a feature vector can be created based on the multidimensional feature extraction results with unified dimensions. The number of video shots of the aforementioned acquired video (i.e. the number of video shots that match the video frame sequence to be processed) can be further determined. Then, based on the feature vector and the number of video shots that match the video frame sequence to be processed, the inter-frame difference index of two consecutive video frames to be processed (hereinafter referred to as consecutive video frames to be processed) in the video frame sequence to be processed can be determined sequentially.
[0038] S130. Determine the key frame threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed.
[0039] The keyframe threshold can be a threshold determined based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed, and is used to filter the video frames to be processed in the video frame sequence.
[0040] In this embodiment of the invention, the mean value of the inter-frame difference index can be determined based on the inter-frame difference index of the sequentially determined consecutive video frames to be processed and the total number of frames in the video frame sequence to be processed, and the obtained mean value can be used as the key frame threshold.
[0041] S140. Using the target detection model, inter-frame difference index, and key frame threshold, determine the target key video frames in the video frame sequence to be processed.
[0042] The object detection model can be a network used to identify target objects in an image.
[0043] In this embodiment of the invention, the video frames to be processed in the video frame sequence to be processed can be screened for the first time according to the inter-frame difference index and the key frame threshold. The screened consecutive video frames to be processed with an inter-frame difference index greater than the key frame threshold are input into the target detection model. Then, according to the output result of the target detection model, the video frames to be processed input into the target detection model are screened for the second time to obtain the target key video frames in the video frame sequence to be processed.
[0044] The technical solution of this invention acquires a sequence of video frames containing a moving target, extracts motion and color features from the sequence to obtain multi-dimensional feature extraction results, and then determines an inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots matching the sequence. Based on the inter-frame difference index and the total number of frames in the sequence, a keyframe threshold is determined. Furthermore, a target detection model, the inter-frame difference index, and the keyframe threshold are used to identify the target's key video frames in the sequence. This solution extracts motion and color features to uncover potential key features of moving targets in the sequence and adaptively generates keyframe thresholds based on the degree of difference between video frames and the total number of frames in the sequence. This achieves accurate localization and capture of moving targets in the sequence, enabling automatic and flexible extraction of target key video frames. It solves the problems of inflexible and inaccurate keyframe extraction in existing technologies, and allows for adaptive setting of keyframe thresholds to flexibly and accurately select key frames.
[0045] Example 2
[0046] Figure 2 This is a flowchart of a keyframe extraction method provided in Embodiment 2 of the present invention. This embodiment is a specific embodiment based on the above embodiment, providing specific optional implementation methods for extracting motion features and color features from the video frame sequence to be processed, and obtaining multi-dimensional feature extraction results. Figure 2 As shown, the method includes:
[0047] S210. Obtain the video frame sequence of the moving target and extract the first motion feature based on the Gaussian mixture background model to obtain the moving target area of each video frame.
[0048] Gaussian mixture background model (GMM) is a model that describes the variation of each pixel using a Gaussian distribution. Generally, the changes of each pixel in a video frame can be viewed as a random process that continuously generates pixel values; the GMM can be used to represent the background of a video frame.
[0049] The first motion feature can be the motion feature extracted after background subtraction between the video frame to be processed and the Gaussian mixture background model. The moving target area can be the area occupied by the moving target in the background subtraction result between the video frame to be processed and the Gaussian mixture background model.
[0050] In this embodiment of the invention, after obtaining the sequence of video frames to be processed for the moving target, a Gaussian mixture background model can be used to model the target. Then, the current video frame to be processed is subjected to background subtraction with the modeling result. Based on the background subtraction result, the first motion feature is extracted to determine the area occupied by the moving target in the background subtraction result, thereby obtaining the area of the moving target in the current video frame to be processed. Similarly, the area of the moving target in each video frame to be processed can be obtained.
[0051] S220. Using the diamond search method, the second motion feature extraction is performed on the continuous video frames to be processed in the video frame sequence to obtain the motion vectors of the continuous video frames to be processed.
[0052] The diamond search method is an algorithm that uses a diamond shape as a search template to determine the region with the highest similarity between two image frames. The second motion feature can be the motion features extracted from the video frame to be processed based on the diamond search method. The motion vector can be a vector formed between regions in two image frames that meet a similarity threshold (which can be set as needed).
[0053] In this embodiment of the invention, the diamond search method can be used to perform similar region matching on the continuous video frames to be processed in the video frame sequence, and the regions that meet the similarity threshold can be used to extract the second motion feature to obtain the vector formed between the regions that meet the similarity threshold in the continuous video frames to be processed, that is, to obtain the motion vector of the continuous video frames to be processed.
[0054] S230. Extract color features from the video frame sequence to be processed to obtain the color entropy of each video frame to be processed.
[0055] Color entropy can be used to describe the richness of colors in an image.
[0056] In this embodiment of the invention, color features can be extracted from each video frame in the video frame sequence to be processed based on the principle of information entropy calculation, so as to obtain the color entropy of each video frame to be processed.
[0057] Information entropy is a concept derived from thermodynamics and applied to information science. It represents the probability of the occurrence of a specific type of information. The more ordered a system is, the lower its information entropy; conversely, the more chaotic a system is, the higher its information entropy.
[0058] S240. Determine the multidimensional feature extraction results based on the moving target area of each video frame to be processed, the motion vector of the continuous video frames to be processed, and the color entropy of each video frame to be processed.
[0059] In this embodiment of the invention, the dimensions of the moving target area and the color entropy of each video frame to be processed can be adjusted according to the dimensions of the motion vector of the continuous video frames to be processed, so as to use the feature extraction result with unified dimensions as the multidimensional feature extraction result.
[0060] S250. Based on the multidimensional feature extraction results and the number of video shots that match the video frame sequence to be processed, determine the inter-frame difference index.
[0061] In an optional embodiment of the present invention, determining the inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots matching the video frame sequence to be processed may include: determining the current continuous inter-frame feature vector based on the multi-dimensional feature extraction results matching the current continuous video frames to be processed; performing clustering processing on the video frame sequence to be processed to obtain the number of video shots matching the video frame sequence to be processed; calculating the target weight coefficient based on the number of video shots; and determining the inter-frame difference index of the current continuous video frames to be processed based on the current continuous inter-frame feature vector and the target weight coefficient.
[0062] Here, the current consecutive frame feature vector can be a vector describing the features of each dimension of the current consecutive video frames to be processed. The target weight coefficient can be determined by the number of video shots and is a weight coefficient that matches the number of feature dimensions in the multi-dimensional feature extraction result.
[0063] In this embodiment of the invention, based on the multidimensional feature extraction results of the current continuous video frames to be processed, the feature extraction results of the current continuous video frames to be processed in each dimension can be determined. Then, based on the feature extraction results of the current continuous video frames to be processed in each dimension and the sum of the feature extraction results of all the continuous video frames to be processed in each dimension, the inter-frame feature vector of the current continuous video frames can be determined. Thus, based on the color entropy of the current continuous video frames to be processed, the sequence of video frames to be processed is clustered to obtain the number of video shots matching the sequence of video frames to be processed. Further, based on the number of video shots, a weight coefficient matching the number of feature dimensions of the multidimensional feature extraction results is determined to obtain the target weight coefficient. Then, based on the inter-frame feature vector of the current continuous video frames and the target weight coefficient, a vector multiplication operation is performed to obtain the inter-frame difference index of the current continuous video frames to be processed.
[0064] S260. Determine the key frame threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed.
[0065] In an optional embodiment of the present invention, determining the keyframe threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed may include: summing the inter-frame difference indices of all consecutive video frames to be processed to obtain a target sum value; determining the target number of all consecutive video frames to be processed based on the total number of frames in the video frame sequence to be processed; and determining the keyframe threshold based on the quotient of the target sum value and the target number.
[0066] The target sum can be the sum of the inter-frame difference indices of all consecutive video frames to be processed. The target number can be the total number of consecutive video frames to be processed.
[0067] In this embodiment of the invention, the inter-frame difference index of all consecutive video frames to be processed can be summed first to obtain the target sum value. Then, the difference between the total number of frames in the sequence of video frames to be processed and 1 is used as the target number of all consecutive video frames to be processed. Furthermore, the quotient of the target sum value and the target number is used as the key frame threshold.
[0068] S270. Using the target detection model, inter-frame difference index, and key frame threshold, determine the target key video frames in the video frame sequence to be processed.
[0069] In an optional embodiment of the present invention, determining the target key video frame in the video frame sequence to be processed by using a target detection model, an inter-frame difference index, and a key frame threshold may include: determining the initial screening key video frames of the video frame sequence to be processed based on the key frame threshold and the inter-frame difference index; identifying moving targets in the initial screening key video frames based on the target detection model to obtain moving target association data of the initial screening key video frames; and determining the target key video frame based on the moving target association data and the intersection-union ratio (IU / R) threshold.
[0070] The initial screening of key video frames can be the result of filtering a sequence of video frames to be processed based on key frame thresholds and inter-frame difference indices, used as input to the target detection model. Moving target association data can be the model output data obtained after inputting the initial screening of key video frames into the target detection model. Moving target association data can include the target category and location of the detection boxes identified by the target detection model. The intersection-over-union (IoU) threshold can be a pre-set threshold for filtering the initial screening of key video frames. IoU can be understood as the ratio of the overlapping area of two targets to the non-overlapping area of two targets.
[0071] In this embodiment of the invention, the inter-frame difference index can be grouped according to the order of the video frames to be processed in the video frame sequence, and the minimum value in each group can be deleted. The steps of grouping the inter-frame difference index and deleting the minimum value in each group are performed iteratively until the results of the iteration are all greater than the key frame threshold. The video frames to be processed that match the iteration results are used as the initial key video frames of the video frame sequence. The initial key video frames are then input into the target detection model. The target detection model identifies the moving targets in the initial key video frames and obtains the moving target association data. Based on the moving target association data, the intersection-union ratio (CIRR) of the detection boxes of the moving targets in two consecutive video frames in the initial key video frames is calculated. Based on the CIRR of the detection boxes of the moving targets in two consecutive video frames and the CIRR threshold, the target key video frames are determined.
[0072] In an optional embodiment of the present invention, determining the initial screening key video frames of the video frame sequence to be processed based on the key frame threshold and the inter-frame difference index may include: grouping the inter-frame difference index according to the continuity relationship of each consecutive video frame to be processed to obtain each current inter-frame difference index group; performing minimum value deletion processing on each current inter-frame difference index group to obtain the processing result of each group; comparing the processing result of each group with the key frame threshold, and if there is at least one group processing result with data less than the key frame threshold, then updating each current inter-frame difference index group according to the processing result of each group, and returning to perform the minimum value deletion processing on each current inter-frame difference index group until the processing result of each group is greater than the key frame threshold.
[0073] The current inter-frame difference index grouping can be the result of grouping the inter-frame difference index according to the continuity of the consecutive video frames to be processed. The grouping processing result can be the result after deleting the minimum value in the current inter-frame difference index group.
[0074] In this embodiment of the invention, based on the continuity relationship of each consecutive video frame in the video frame sequence to be processed, two consecutive inter-frame difference indices can be grouped together to obtain each current inter-frame difference index group. Further, a minimum value deletion process is performed on each current inter-frame difference index group, that is, the minimum value in each current inter-frame difference index group is deleted to obtain the processing result of each group. Then, the processing result of each group is compared with the key frame threshold. If there is at least one group processing result with data less than the key frame threshold, then each current inter-frame difference index group is updated according to the processing result of each group, and the operation of performing minimum value deletion on each current inter-frame difference index group is returned until the processing result of each group is greater than the key frame threshold.
[0075] In an optional embodiment of the present invention, determining the target key video frame based on moving target association data and cross-union ratio (CURBR) threshold may include: when it is determined that there is a moving target in the initial screening key video frames based on moving target association data, redundant video frames in the initial screening key video frames are removed according to the NMS algorithm and the CURBR threshold to obtain the target key video frame.
[0076] Redundant video frames can be those video frames that need to be removed from the initial screening of key video frames. These are video frames to be processed that have an intersection-over-union (IoU) ratio greater than a threshold with the redundant video frames in the initial screening of key video frames.
[0077] In this embodiment of the invention, it can be determined whether the target category within the detection box of the initial screening key video frame is a moving target category based on the moving target association data. If so, it is determined that there is a moving target in the initial screening key video frame. Then, the cross-union ratio (CUI) threshold of the NMS (Non-Maximum Suppression) algorithm is set. Based on the position of the moving target in the detection box in the moving target association data and the NMS algorithm, a group of video frames to be processed (the group of video frames to be processed includes two consecutive video frames in the initial screening key video frames) with a CUI less than the CUI threshold is determined. One of the video frames to be processed in the group of video frames to be processed with a CUI less than the CUI threshold is taken as a redundant video frame. Then, the redundant video frames in the initial screening key video frames are removed to obtain the target key video frame.
[0078] In a specific example, the keyframe extraction steps are as follows:
[0079] Step 1: Convert the monitoring video of the moving target into a sequence of video frames to be processed.
[0080] Step 2: Convert the video frames fi (i = 1, 2, ..., M) in the video frame sequence to be processed from the RGB (red, green, blue) space to the HSV (hue, saturation, brightness) space, thus converting them from RGB to the HSV space, which is more consistent with human visual perception. Here, M is the total number of frames in the video frame sequence. Standardize the range of the HSV components and obtain a histogram. The standardized ranges are as follows:
[0081]
[0082] Where z > y > x > 0, for example, z = 22, y = 17, x = 12. Furthermore, the standardized calculation method in this scheme is as follows: H represents hue, S represents saturation, and V represents brightness.
[0083] The formula for calculating color entropy is as follows:
[0084]
[0085] Where λ1+λ2+λ3=1. For example, λ1=0.5, λ2=0.3, λ3=0.2. h(i), s(i), v(i) represent the statistics of the standardized HSV histogram, and log is the logarithm to the base 2.
[0086] f k and the next video frame to be processed f k+1 Compare and extract f k to f k+1 The motion vector. A diamond search method is used for each 16*16 (representing the number of pixels) search block C. j (j=1,2,…,N) performs motion vector estimation to obtain the x-axis direction component X. j and the y-axis component Y j The sum of the norms of the coordinate axes is obtained, and the motion vector is calculated as follows:
[0087]
[0088] f is calculated using a Gaussian mixture background model and background subtraction. k The area of the moving target. Let f k Background subtraction is performed using a Gaussian mixture background model. After subtraction, the pixels containing the moving target are marked as 255 (white), and the background is marked as 0 (black). The area of the moving target is A(f). k The number of white pixels in the image after background subtraction is calculated.
[0089] Color entropy and moving target area are features extracted from each video frame to be processed, while motion vectors are the difference features between two consecutive video frames to be processed. The color entropy and moving target area of HSV are subtracted from each consecutive frame to maintain the consistency of the units. The representation is as follows:
[0090] E(f k ,f k+1 )=|E(f k )-E(f k+1 )|
[0091] A(f k ,f k+1 )=|A(f k )-A(f k+1 )|
[0092] Normalize the eigenvalues to construct the feature vector for each video frame to be processed, that is, determine the feature vector between consecutive frames. The representation of the feature vector between consecutive frames is as follows:
[0093]
[0094] Step 3: After obtaining the feature vectors between consecutive frames, further determine the number of video shots. The specific process is as follows: Compare the similarity of the HSV color entropy of each video frame to be processed and the cluster center frame. If the similarity is greater than the threshold of 0.85, the two frames are judged to be similar. Then, the video frame to be processed that matches the current cluster center frame is merged into the cluster that matches the current cluster center frame, and the video frame to be processed in the newly added cluster is used as the new cluster center frame. If no cluster center frame similar to the current video frame to be processed is found, a new cluster is created for the current video frame to be processed, and the current video frame to be processed is added to the new cluster. The HSV color entropy similarity is calculated as follows:
[0095] Similarity=s1×a1+s2×a2+s3×a3
[0096] Where s1 is the sum of the minimum values of the current video frame to be processed and the cluster center frame in the HSV histogram H dimension, S2 is the sum of the minimum values of the current video frame to be processed and the cluster center frame in the HSV histogram S dimension, and S3 is the sum of the minimum values of the current video frame to be processed and the cluster center frame in the HSV histogram V dimension. For example, a1 can be set to 0.5, a2 to 0.3, and a3 to 0.2.
[0097] The number of video shots η is obtained based on similarity and clustering rules, that is, the number of clusters is taken as the number of video shots η.
[0098] Step 4: Dynamically sum the features of each feature vector in the continuous inter-frame feature vectors according to the number of video shots η, and obtain the inter-frame difference index.
[0099] The inter-frame difference index is calculated as follows:
[0100]
[0101] The target weight coefficients ω1, ω2, and ω3 are calculated as follows:
[0102]
[0103] Step 5: Arrange the inter-frame difference indices sequentially by frame number: {d(f1,f2),d(f2,f3),...,d(f...} M-1 ,f M The keyframe threshold is calculated based on the following formula:
[0104]
[0105] Grouping every two adjacent inter-frame difference indices in the inter-frame difference index into a group, and removing the minimum value in each group, this process is repeated iteratively until all inter-frame difference values are greater than the keyframe threshold. Then, each remaining inter-frame difference index d(f) is taken. k ,f k+1 The frame corresponding to the index k+1 of the video is the key frame, and the initial key video frames are obtained.
[0106] Step 6: Use the object detection model to obtain the object category and location in the initial screening key video frames.
[0107] Step 7: Filter out redundant detection boxes using the NMS algorithm. If the number of remaining detection boxes is greater than half the number of detection boxes before filtering, it is determined that there are redundant video frames between the two video frames to be processed. Specifically, this includes: comparing the motion target association data in the initially screened key video frames; if no motion target appears, the process ends; if a motion target appears, the position and category of the motion target detection boxes in the two video frames to be processed are filtered using the NMS algorithm. The cross-union ratio (CUI) threshold of the NMS algorithm is set to 0.50. Then, it is determined whether the CUI of the two detection boxes is greater than the CUI threshold. If the CUI is greater than 0.50, one of the two frames is deleted; if the CUI is less than 0.50, both frames are retained to obtain the target key video frame.
[0108] In summary, a flowchart of steps 1 to 7 can be found here. Figure 3 .
[0109] The technical solution of this invention involves acquiring a sequence of video frames containing moving targets, and performing a first motion feature extraction based on a Gaussian mixture background model to obtain the moving target area of each video frame. Then, using a diamond search method, a second motion feature extraction is performed on consecutive video frames to obtain their motion vectors. Next, color feature extraction is performed on the video frame sequence to obtain the color entropy of each frame. Based on the moving target area, motion vectors, and color entropy of each frame, a multi-dimensional feature extraction result is determined. Further, based on the multi-dimensional feature extraction result and the number of video shots matching the video frame sequence, an inter-frame difference index is determined. Then, based on the inter-frame difference index and the total number of frames in the video frame sequence, a keyframe threshold is determined. Finally, using a target detection model, the inter-frame difference index, and the keyframe threshold, the target key video frames in the video frame sequence are identified. This solution extracts motion and color features to uncover potential key features of moving targets in the video frame sequence to be processed. It can also adaptively generate key frame thresholds based on the degree of difference between video frames and the total number of frames in the video frame sequence to be processed, thereby achieving accurate localization and capture of moving targets in the video frame sequence to be processed. In other words, it achieves automatic, flexible and accurate extraction of key video frames of targets, solving the problems of inflexible and inaccurate key frame extraction in existing technologies. It can adaptively set key frame thresholds and flexibly and accurately filter out key frames.
[0110] Example 3
[0111] Figure 4 This is a schematic diagram of a keyframe extraction device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: a feature extraction module 310, an inter-frame difference index determination module 320, a key frame threshold determination module 330, and a target key video frame determination module 340, wherein,
[0112] The feature extraction module 310 is used to acquire the video frame sequence of the moving target and extract motion and color features from the video frame sequence to obtain multi-dimensional feature extraction results.
[0113] The inter-frame difference index determination module 320 is used to determine the inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots that match the video frame sequence to be processed.
[0114] The keyframe threshold determination module 330 is used to determine the keyframe threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed.
[0115] The target key video frame determination module 340 is used to determine the target key video frames in the video frame sequence to be processed by using the target detection model, the inter-frame difference index, and the key frame threshold.
[0116] The technical solution of this invention acquires a sequence of video frames containing a moving target, extracts motion and color features from the sequence to obtain multi-dimensional feature extraction results, and then determines an inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots matching the sequence. Based on the inter-frame difference index and the total number of frames in the sequence, a keyframe threshold is determined. Furthermore, a target detection model, the inter-frame difference index, and the keyframe threshold are used to identify the target's key video frames in the sequence. This solution extracts motion and color features to uncover potential key features of moving targets in the sequence and adaptively generates keyframe thresholds based on the degree of difference between video frames and the total number of frames in the sequence. This achieves accurate localization and capture of moving targets in the sequence, enabling automatic and flexible extraction of target key video frames. It solves the problems of inflexible and inaccurate keyframe extraction in existing technologies, and allows for adaptive setting of keyframe thresholds to flexibly and accurately select key frames.
[0117] Optionally, the feature extraction module 310 includes a first feature extraction unit, a second feature extraction unit, a third feature extraction unit, and a multi-feature combination unit. The first feature extraction unit is used to perform first motion feature extraction on the video frame sequence to be processed based on a Gaussian mixture background model to obtain the motion target area of each video frame to be processed. The second feature extraction unit is used to perform second motion feature extraction on consecutive video frames to be processed in the video frame sequence to be processed using a diamond search method to obtain the motion vector of each consecutive video frame to be processed. The third feature extraction unit is used to extract color features from the video frame sequence to be processed to obtain the color entropy of each video frame to be processed. The multi-feature combination unit is used to determine the multi-dimensional feature extraction result based on the motion target area of each video frame to be processed, the motion vector of each consecutive video frame to be processed, and the color entropy of each video frame to be processed.
[0118] Optionally, the inter-frame difference index determination module 320 includes a feature vector determination unit, a video shot number determination unit, and an inter-frame difference index calculation unit. The feature vector determination unit is used to determine the current continuous inter-frame feature vector based on the multi-dimensional feature extraction results matching the current continuous video frames to be processed. The video shot number determination unit is used to perform clustering processing on the video frame sequence to be processed to obtain the number of video shots matching the video frame sequence to be processed. The inter-frame difference index calculation unit is used to calculate a target weight coefficient based on the number of video shots, and determine the inter-frame difference index of the current continuous video frames to be processed based on the current continuous inter-frame feature vector and the target weight coefficient.
[0119] Optionally, the keyframe threshold determination module 330 includes a target sum determination unit, a target quantity determination unit, and a keyframe threshold calculation unit. The target sum determination unit is used to sum the inter-frame difference indices of all the continuous video frames to be processed to obtain a target sum. The target quantity determination unit is used to determine the target quantity of all the continuous video frames to be processed based on the total number of frames in the video frame sequence. The keyframe threshold calculation unit is used to determine the keyframe threshold based on the quotient of the target sum and the target quantity.
[0120] Optionally, the keyframe threshold determination module 330 includes a preliminary key video frame determination unit, a moving target association data determination unit, and a target key video frame determination unit. The preliminary key video frame determination unit is used to determine the preliminary key video frames of the video frame sequence to be processed based on the key frame threshold and the inter-frame difference index. The moving target association data determination unit is used to identify moving targets in the preliminary key video frames according to the target detection model to obtain the moving target association data of the preliminary key video frames. The target key video frame determination unit is used to determine the target key video frames based on the moving target association data and the intersection-union ratio (IU) threshold.
[0121] Optionally, the initial screening key video frame determination unit is specifically used to group the inter-frame difference index according to the continuity relationship of each of the continuous video frames to be processed, to obtain each current inter-frame difference index group; to perform minimum value deletion processing on each current inter-frame difference index group, to obtain the processing result of each group; to compare the processing result of each group with the key frame threshold, and if there is at least one group processing result whose data is less than the key frame threshold, then according to the processing result of each group, update each current inter-frame difference index group, and return to perform the minimum value deletion processing on each current inter-frame difference index group, until the processing result of each group is greater than the key frame threshold.
[0122] Optionally, the target key video frame determination unit is specifically used to, when determining that there is a moving target in the initial screening key video frames based on the moving target association data, remove redundant video frames in the initial screening key video frames according to the non-maximum suppression (NMS) algorithm and the intersection-union ratio (IU) threshold, to obtain the target key video frame.
[0123] The keyframe extraction device provided in this embodiment of the invention can execute the keyframe extraction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0124] Example 4
[0125] Figure 5 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0126] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0127] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0128] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as keyframe extraction methods.
[0129] In some embodiments, the keyframe extraction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the keyframe extraction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the keyframe extraction method by any other suitable means (e.g., by means of firmware).
[0130] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0132] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0135] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0136] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0137] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for extracting keyframes, characterized in that, include: A sequence of video frames to be processed for a moving target is obtained, and motion and color features are extracted from the sequence of video frames to be processed to obtain multi-dimensional feature extraction results. Based on the multidimensional feature extraction results and the number of video shots that match the video frame sequence to be processed, the inter-frame difference index is determined; The keyframe threshold is determined based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed; The target key video frames in the video frame sequence to be processed are determined by the target detection model, the inter-frame difference index, and the key frame threshold. The step of determining the inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots matching the video frame sequence to be processed includes: Based on the multi-dimensional feature extraction results that match the current consecutive video frames to be processed, determine the feature vector between the current consecutive frames; Clustering is performed on the video frame sequence to be processed to obtain the number of video shots that match the video frame sequence to be processed. Based on the number of video shots, a target weight coefficient is calculated, and based on the current consecutive inter-frame feature vector and the target weight coefficient, the inter-frame difference index of the current consecutive video frames to be processed is determined.
2. The method according to claim 1, characterized in that, The process of extracting motion and color features from the video frame sequence to be processed, resulting in multi-dimensional feature extraction results, includes: Based on the Gaussian mixture background model, the first motion feature is extracted from the video frame sequence to be processed to obtain the motion target area of each video frame to be processed. The second motion feature extraction is performed on the continuous video frames to be processed in the video frame sequence using the diamond search method to obtain the motion vector of the continuous video frames to be processed. Color features are extracted from the video frame sequence to be processed to obtain the color entropy of each video frame to be processed. The multidimensional feature extraction result is determined based on the moving target area of each video frame to be processed, the motion vector of the continuous video frames to be processed, and the color entropy of each video frame to be processed.
3. The method according to claim 2, characterized in that, The step of determining the keyframe threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed includes: The inter-frame difference index of all the continuous video frames to be processed is summed to obtain the target sum value; The target number of all consecutive video frames to be processed is determined based on the total number of frames in the video frame sequence to be processed. The keyframe threshold is determined based on the quotient of the target and the value and the number of targets.
4. The method according to claim 2, characterized in that, The step of determining the target key video frame in the video frame sequence to be processed using the target detection model, the inter-frame difference index, and the key frame threshold includes: Based on the key frame threshold and the inter-frame difference index, the initial key video frames of the video frame sequence to be processed are determined. Based on the target detection model, moving targets in the initial screening key video frames are identified to obtain moving target association data of the initial screening key video frames; Based on the moving target association data and the cross-union ratio threshold, the key video frames of the target are determined.
5. The method according to claim 4, characterized in that, The step of determining the initial screening of key video frames in the video frame sequence to be processed based on the key frame threshold and the inter-frame difference index includes: Based on the continuity relationship of each of the continuous video frames to be processed, the inter-frame difference index is grouped to obtain each current inter-frame difference index group. Minimum value deletion is performed on each current inter-frame difference index group to obtain the processing results for each group. The processing results of each group are compared with the keyframe threshold. If the data in at least one group processing result is less than the keyframe threshold, the current inter-frame difference index group is updated according to the processing results of each group, and the operation of deleting the minimum value of each current inter-frame difference index group is returned until the processing results of each group are greater than the keyframe threshold.
6. The method according to claim 4, characterized in that, The step of determining the key video frames of the target based on the moving target association data and the intersection-over-union (IoU) threshold includes: When it is determined that there is a moving target in the initial screening key video frame based on the moving target association data, redundant video frames in the initial screening key video frame are removed according to the non-maximum suppression (NMS) algorithm and the intersection-union ratio (IU) threshold to obtain the target key video frame.
7. A keyframe extraction device, characterized in that, include: The feature extraction module is used to acquire the video frame sequence of the moving target to be processed, and to extract motion features and color features from the video frame sequence to be processed to obtain multi-dimensional feature extraction results. The inter-frame difference index determination module is used to determine the inter-frame difference index based on the multi-dimensional feature extraction results and the number of video shots that match the video frame sequence to be processed. The keyframe threshold determination module is used to determine the keyframe threshold based on the inter-frame difference index and the total number of frames in the video frame sequence to be processed. The target key video frame determination module is used to determine the target key video frame in the video frame sequence to be processed by using the target detection model, the inter-frame difference index and the key frame threshold. The inter-frame difference index determination module includes a feature vector determination unit, a video shot number determination unit, and an inter-frame difference index calculation unit. The feature vector determination unit is used to determine the feature vector between the current consecutive frames based on the multi-dimensional feature extraction results that match the current consecutive video frames to be processed. The video shot number determination unit is used to perform clustering processing on the video frame sequence to be processed to obtain the number of video shots that match the video frame sequence to be processed. The inter-frame difference index calculation unit is used to calculate the target weight coefficient based on the number of video shots, and determine the inter-frame difference index of the current continuous video frame to be processed based on the current continuous inter-frame feature vector and the target weight coefficient.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which enables the at least one processor to perform the keyframe extraction method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the keyframe extraction method according to any one of claims 1-6.