A method for extracting video key frames based on active learning technology
By constructing directed network graphs and user interaction judgments, video keyframe extraction method based on active learning technology, the problem of inaccurate keyframe extraction in video is solved, and more efficient video summary and content retrieval is achieved.
Patent Information
- Application Number
- CN202510585476.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The prior art is difficult to extract keyframes from videos efficiently and accurately, affecting the accuracy of video summary and content retrieval.
Using the video keyframe extraction method based on active learning technology, by constructing a directed network graph, using image feature extraction algorithm and network centrality index, combining user interaction to judge the similarity of frame nodes, and iteratively generate accurate keyframes.
Improves the accuracy and flexibility of video keyframe extraction, can identify multiple scene categories, and generate more accurate video summary and content retrieval results.
Smart Images

Figure CN120107867B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video processing, and particularly to a method for extracting key frames of a video based on active learning technology. Background Art
[0002] With the rapid development of social platforms and the wide application of multimedia technology, as the main carrier of multimedia information, the amount of digital video data has shown an explosive growth. How to efficiently store, manage, and retrieve massive video content has become a key problem in realizing functions such as video semantic analysis, personalized recommendation, and intelligent content review. The video summarization technology is an important way to address this problem, and the core lies in the selection of key frames. A key frame refers to an image frame with the most concentrated information in a video segment. The purpose of key frame extraction is to select several representative frames from a continuous video sequence in order to condense and summarize the main content of the entire video.
[0003] Based on the above situation, there is an urgent need for a method that can accurately extract the key frames of a video. Summary of the Invention
[0004] To solve the above problems, the present application provides a method for extracting key frames of a video based on active learning technology, which can output all the key frames corresponding to various scenes in the video to improve the accuracy of video key frame extraction.
[0005] To achieve the above object, in a first aspect, the present application provides a method for extracting key frames of a video based on active learning technology, including:
[0006] S1. Obtain an image frame sequence corresponding to the video to be processed;
[0007] S2. Use an image feature extraction algorithm to obtain the feature vector of each image frame in the image frame sequence;
[0008] S3. Take the feature vector as the frame node of the image frame, and then obtain a frame node set, and each frame node in the frame node set is used as a candidate frame node;
[0009] S4. Calculate the Euclidean distance between any two candidate frame nodes respectively, connect each candidate frame node with the candidate frame node with the closest Euclidean distance to it to form an edge, and the direction of the edge is from the current candidate frame node to its closest candidate frame node, so as to obtain a plurality of first connected subgraphs; the first connected subgraph contains at least two candidate frame nodes, and any one of the first connected subgraphs only includes a set of mutually nearest neighbor node pairs, the Euclidean distances between the two candidate frame nodes in the mutually nearest neighbor node pair are the closest to each other, and the two candidate frame nodes in the mutually nearest neighbor node pair point to each other;
[0010] S5. Calculate the network centrality indices corresponding to the two candidate frame nodes in the mutually nearest neighbor node pairs, and disconnect the edge connecting the candidate frame node with the larger network centrality index to the candidate frame node with the smaller network centrality index. Among them, the candidate frame node being pointed to among the two candidate frame nodes corresponding to the mutually nearest neighbor nodes is the target frame node, and the other candidate frame node is the source frame node;
[0011] S6. Take the candidate frame node with the larger network centrality index in the mutually nearest neighbor node pairs as the representative frame node, and all the representative frame nodes corresponding to all the current first connected subgraphs form a representative frame node set;
[0012] S7. Based on the representative frame nodes obtained in S6, repeat the above S4 - S6 and update the representative frame node set until there is only one representative frame node in the representative frame node set, and obtain the first directed network graph;
[0013] S8. Calculate the ambiguity value of each edge in the first directed network graph; select the frame node pair corresponding to the edge with the largest ambiguity value as the object to be judged, and let the user determine whether the object to be judged is similar: if the user's judgment result is yes, retain the edge corresponding to the object to be judged; if the user's judgment result is no, disconnect the edge corresponding to the object to be judged, update the first directed network graph, and enter S9;
[0014] S9. Present the source frame node corresponding to the disconnected edge and the representative frame nodes in the representative frame node set to the user one by one to judge whether they are similar: if they are similar, connect the source frame node and the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and set the ambiguity value of the edge connecting the source frame node and the representative frame node to 0, update the first directed network graph again, and terminate this round of judgment; if they are not similar in all cases, take the source frame node as the representative frame node and add it to the representative frame node set;
[0015] S10. Repeat the above S8 and S9 for the first directed network graph until the preset number of repetitions is reached, and obtain the final second directed network graph, where the second directed network graph contains multiple second connected subgraphs;
[0016] S11. Based on any one of the second connected subgraphs, calculate the network centrality indices of each frame node within each second connected subgraph, and select the frame node with the largest centrality index as the key frame of the second connected subgraph. After aggregating the key frames of all the second connected subgraphs, output them to complete the key frame extraction.
[0017] The beneficial effects of the present invention are as follows: The present invention discloses a method for extracting key frames of a video based on active learning. First, a first directed network graph is constructed according to the nearest neighbor rule, and a pair of frame nodes composed of the two frame nodes corresponding to the edge with the largest ambiguity value is selected and submitted to the user for judgment to determine whether the node pair belongs to the same type of scene. If the user determines that they are not similar, taking the source node of this edge as a reference, new pairs of frame nodes are formed in the order of the Euclidean distance from near to far to each frame node in the set of representative frame nodes and submitted to the user for judgment again, and the network structure is adjusted in real time according to the user's results, and the second directed network graph is iteratively generated. Finally, the node with the highest network centrality index in each second connected subgraph is selected as the key frame. The method of the present invention introduces an active learning mechanism through interaction with the user, uses the user's judgment result as a topological guidance, and generates accurate key frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is the flowchart of the embodiment of the present invention;
[0020] Figure 2 is the schematic diagram of the initial frame node distribution map and the first connected subgraph; where Figure 2 (a) is the initial frame node distribution map, Figure 2 (b) is the first connected subgraph;
[0021] Figure 3 is the schematic diagram of the representative frame nodes in the first connected subgraph;
[0022] Figure 4 is the schematic diagram of the first directed network graph of the embodiment of the present invention;
[0023] Figure 5 is the schematic diagram of the second connected subgraph of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, other embodiments obtained by those of ordinary skill in the art without creative efforts all fall within the protection scope of the present application.
[0025] Hereinafter, terms such as "first", "second", etc. are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more such features. In the description of this application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0026] With the rapid development of social platforms and the wide application of multimedia technologies, as the main carrier of multimedia information, the amount of digital video data has shown an explosive growth. How to efficiently store, manage, and retrieve massive video content has become a key problem in realizing functions such as video semantic analysis, personalized recommendation, and intelligent content review. The video summarization technology is an important way to address this problem, and the core lies in the selection of key frames. A key frame refers to the image frame with the most concentrated information in a video segment. The purpose of key frame extraction is to screen out several representative frames from a continuous video sequence in order to condense and summarize the main content of the entire video.
[0027] Exemplarily, in an intelligent driving system, a vehicle's driving recorder will collect many segments of driving records, and the vehicle's memory will store the driving records collected within a certain period of time and update them regularly. In the intelligent driving system, when storing the driving records, summary information can be generated based on the key frames corresponding to the segment of driving record. At this time, the driving records stored in the memory will include video segments and summary information. In this way, if a user wants to extract a certain segment of driving record, they can retrieve it based on the summary information corresponding to the driving record to quickly obtain the desired driving record without having to search through each segment one by one.
[0028] For the above reasons, the accuracy of video key frame extraction will greatly affect the accuracy of the summary information corresponding to the segment of video. Therefore, the embodiments of this application provide a video key frame extraction method based on active learning technology, which can output all the key frames corresponding to various types of scenes in the video to improve the accuracy of video key frame extraction.
[0029] Figure 1 is a flowchart of a video key frame extraction method based on active learning technology provided by the embodiments of this application.
[0030] As Figure 1 shown, the video key frame extraction method based on active learning technology provided by the embodiments of this application includes:
[0031] S1. Obtain an image frame sequence corresponding to the video to be processed, where the image frame sequence is obtained by performing shot boundary segmentation processing on the video to be processed.
[0032] Specifically, in this step, there are many methods to obtain the sequence of video image frames to be processed, such as conventional shot boundary detection, motion analysis extraction, etc. According to the characteristics of each method, in this embodiment, the shot boundary segmentation method is selected as the method to obtain the sequence of image frames of the video to be processed. Particularly, the sequence of image frames refers to the sequence obtained by sorting all the obtained image frames according to time; at the same time, when the acquisition time is long and the number of image frames is too large, the sequence of image frames can also be the sequence obtained by sampling the image frames at a certain time interval and then sorting them according to time.
[0033] Shot Boundary Detection (SBD) is a technique in video processing used to identify the switching points between different shots (or scenes) in a video. A shot refers to a continuous segment of footage in a video, and a shot boundary refers to the transition between two adjacent shots, which is usually accompanied by a significant change in the content of the frame. Shot boundary segmentation can detect the switching of scenes in a video and distinguish when a transition between shots occurs. Shot boundary segmentation also helps in understanding the structure of the video, facilitating subsequent analysis, retrieval, and processing.
[0034] Shot boundaries usually represent significant changes in the scene. Focusing on processing these frames can reduce the consumption of computing resources. The image frames obtained after shot boundary segmentation are often better for subsequent analysis and processing tasks, such as video summarization, content retrieval, etc. These frames usually represent different scenes or plot developments, helping to understand the overall structure of the video. In addition, in a video, some frames may be noise caused by minor changes within a scene. Without shot boundary segmentation, they may interfere with the results of key frame extraction. Shot boundary segmentation can effectively filter out these noise frames. Therefore, the sequence of image frames after shot boundary segmentation can improve the efficiency of key frame extraction.
[0035] S2. Use an image feature extraction algorithm to obtain the feature vector of each image frame in the sequence of image frames.
[0036] Among them, the image feature extraction algorithm is a technology in the fields of computer vision and image processing used to extract useful information and descriptions from images. These features can be used for various tasks, such as classification, recognition, retrieval, and image analysis. The purpose of feature extraction is to convert the original image into a more concise and high-dimensional representation for subsequent processing.
[0037] Common image feature extraction algorithms include edge detection algorithms, corner detection algorithms, feature descriptors, and texture feature algorithms, etc.
[0038] The feature vector of an image frame can compress complex image information into a relatively concise numerical representation. This representation retains the key features of the image (such as color, texture, shape, etc.), making subsequent processing and analysis more efficient. Feature vectors can be used to calculate the similarity between image frames. By comparing the feature vectors of adjacent frames, it is possible to identify which frames have significant content changes, thereby helping to determine shot transition points and key frames.
[0039] In this embodiment, an edge detection algorithm is used to extract the feature vector corresponding to the image frame, which may specifically include the following steps:
[0040] S201. Convert each image frame in the image frame sequence into a grayscale image frame.
[0041] S202. Use the HOG algorithm to divide the grayscale image frame into multiple image blocks and perform feature extraction to obtain an initial feature vector, where the initial feature vector is , is the constant value of the grayscale image frame in the Mth dimension.
[0042] Among them, the Histogram of Oriented Gradient (HOG) is a technique for image feature description and is widely used in object detection, especially in fields such as pedestrian detection, vehicle detection, and face recognition. HOG features describe the shape and structure of an image by capturing the gradient direction and amplitude in local regions of the image.
[0043] S203. Calculate the information entropy weight coefficient of each image block.
[0044] S204. Perform weighted processing on the initial feature vector based on the information entropy weight coefficient obtained in S203 to obtain a weighted feature vector;
[0045] S205. Perform dimensionality reduction processing on the weighted feature vector to obtain the feature vector of this image frame, and use the feature vector as the frame node of this image frame;
[0046] S206. Repeat S201~S205 until the feature vectors of all image frames in the image frame sequence are obtained.
[0047] In the above-mentioned feature vector extraction process, converting the image frame into a grayscale image frame provides a more efficient, more stable and more robust processing method for edge detection. Grayscale images play an important role in edge detection algorithms by simplifying calculations and improving the extractability of edge features. The information entropy weight coefficient of each image block indicates the amount of image information corresponding to the image block. In other words, the larger the information entropy weight coefficient of the image block, the more information the image block contains; the smaller the information entropy weight coefficient of the image block, the less information the image block contains. In this way, each image block can be weighted based on the information entropy weight coefficient to obtain a weighted feature vector. Finally, the weighted feature vector is subjected to dimensionality reduction processing to simplify the dimension of the feature vector through dimensionality reduction processing, reduce the amount of subsequent calculations, and obtain the final feature vector.
[0048] Exemplarily, the information entropy weight coefficient can be calculated based on the following formula: , where ω b (i) is the information entropy weight coefficient of the i-th image block, is the gray value of the i-th image block The probability of pixel occurrence.
[0049] For example, principal component analysis (PCA) can be used for dimensionality reduction. PCA is a commonly used dimensionality reduction technique that converts high-dimensional data into low-dimensional data through linear transformation while retaining the variability of the original data as much as possible. PCA is widely used in many fields such as data preprocessing, feature extraction, and image processing.
[0050] In dynamic scenes, image frames may be affected by noise, lighting changes and other environmental factors. After reducing their dimensionality through methods such as PCA, unnecessary noise can be filtered out and the robustness of target recognition and tracking can be improved.
[0051] S3. Using the feature vector as a frame node of the image frame, and then obtaining a frame node set, wherein each frame node in the frame node set is used as a candidate frame node.
[0052] By using the feature vector corresponding to each image frame as the frame node of the image frame, in the video processing task, converting each frame of the image into a frame node helps to capture the changes and dynamic features in time, thereby analyzing action sequences and events. At the same time, after using the feature vector as the frame node of the image frame, the feature vector is no longer just a data representation, but a node entity in the network diagram.
[0053] S4. Calculate the Euclidean distance between any two candidate frame nodes respectively, connect each of the candidate frame nodes with the candidate frame node with the closest Euclidean distance to it to form an edge, and the direction of the edge is from the current candidate frame node to its closest candidate frame node, so as to obtain a plurality of first connected subgraphs; each of the first connected subgraphs contains at least two candidate frame nodes, and any one of the first connected subgraphs only includes a set of mutually nearest neighbor node pairs, the Euclidean distances between the two candidate frame nodes in the mutually nearest neighbor node pair are the closest to each other, and the two candidate frame nodes in the mutually nearest neighbor node pair point to each other.
[0054] As Figure 2 shown, we call the subgraph formed after multiple pairings in this step the first connected subgraph. Among them, Figure 2 (a) is the initial frame node graph, Figure 2 (b) is the first connected subgraph. In the figure, the direction indicated by the arrow is the connection direction between node pairs, and the mutually nearest neighbor node pairs are connected to each other.
[0055] S5. Calculate the network centrality indices corresponding to the two candidate frame nodes in the mutually nearest neighbor node pair, and disconnect the edge from the candidate frame node with a larger network centrality index to the candidate frame node with a smaller network centrality index. Among them, the candidate frame node pointed to among the two candidate frame nodes corresponding to the mutually nearest neighbor nodes is the target frame node, and the other candidate frame node is the source frame node.
[0056] In this step, the network centrality index is an index used to measure the importance or influence of nodes in the network. For frame nodes, the commonly used centrality indices include local centrality and semi-local centrality. The local centrality may be overly influenced by certain frame nodes (for example, the degree of a frame node is particularly high but it does not have actual influence), making it unable to fully reflect the structural characteristics in the complex network. The semi-local centrality can capture more complex connection patterns, such as the strength of second-order connections, which is crucial for understanding the role of frame nodes in the network. By introducing the semi-local centrality, this influence can be balanced, making the evaluation of the importance of frame nodes more stable and reliable.
[0057] To balance the local centrality and semi-local centrality, in this embodiment, a new network centrality index is proposed, and its calculation formula is as follows: , where φ Total (v) represents the network centrality index of frame node v, α is the first weight coefficient, φ Local (v) represents the local centrality of frame node v, and φ Semilocal (v) represents the semi-local centrality of frame node v;
[0058] The local centrality is calculated by the following formula:
[0059] ;
[0060] In the formula, Degree v represents the degree of the frame node v, and N represents the number of all frame nodes;
[0061] The semi-local centrality is calculated by the following formula:
[0062] ;
[0063] In the formula, φ Semilocal (v) represents the semi-local centrality of the frame node v; |G N2 (v)| represents the total number of first-order and second-order neighborhood nodes of the frame node v; represents the set of frame nodes within the first-order neighborhood of the frame node v, represents the set of frame nodes within the second-order neighborhood of the frame node v; k v represents the degree of the frame node v; k u represents the degree of the frame node u; the frame node u is a frame node within the first-order or second-order neighborhood of the frame node v; d u,v represents the distance between the frame node v and the frame node u; d max1 represents the maximum distance between the frame node v and all frame nodes within its first-order neighborhood, d max2 represents the maximum distance between the frame node v and all frame nodes within its second-order neighborhood.
[0064] The first-order neighborhood referred to here is the set of frame nodes directly connected to the nearest neighbor node pair where the frame node v is located; the second-order neighborhood referred to is the set of frame nodes directly connected to the frame nodes in the first neighborhood.
[0065] The local centrality refers to the importance of a certain frame node among its direct neighbors. The semi-local centrality is an extension of the local centrality, usually considering the influence of the neighbors of the frame node and their neighbors.
[0066] In this embodiment, the local centrality and the semi-local centrality are weighted and combined to obtain a network centrality index, which can comprehensively evaluate the importance of the frame node in the entire network and avoid information loss caused by a single perspective. It can more accurately identify key frames or important events and effectively improve the efficiency of target tracking and event detection.
[0067] S6. Use the candidate frame node with a larger network centrality index in the nearest neighbor node pairs as the representative frame node, and all the representative frame nodes corresponding to all the current first-connected subgraphs form a set of representative frame nodes.
[0068] Such as Figure 3As shown, one of the connecting edges between the mutually nearest neighbor node pairs has been disconnected and a representative frame node has been selected.
[0069] S7. Taking the representative frame node obtained in S6 as a reference, repeat S4 - S6 and update the set of representative frame nodes until there is only one such representative frame node in the set of representative frame nodes, and obtain a first directed network graph.
[0070] In this step, only the representative frame nodes obtained in the previous step are selected for operation, while the remaining frame nodes are temporarily ignored. The finally obtained first directed network graph includes the frame nodes corresponding to all image frames.
[0071] As Figure 4 shown, it is the first directed network graph finally formed.
[0072] S8. Calculate the ambiguity value of each connecting edge in the first directed network graph; select the pair of frame nodes corresponding to the connecting edge with the largest ambiguity value as the object to be judged, and let the user determine whether the object to be judged is similar: if the user's judgment result is yes, retain the connecting edge corresponding to the object to be judged; if the user's judgment result is no, disconnect the connecting edge corresponding to the object to be judged, update the first directed network graph, and enter S9;
[0073] Among them, the ambiguity value of a connecting edge is the ambiguity values of the two frame nodes in the node pair. The ambiguity value is mainly used to measure the ambiguity of the relationship and mutual influence between a node and the nodes connected to it in the network. In the first directed network graph, the ambiguity value usually refers to how to quantify the ambiguity of the similarity, influence or connection strength of these frame nodes considering the relationship between the two frame nodes connected by the connecting edge. The ambiguity value in the network is the uncertainty of the relationship between frame nodes, which may be caused by incomplete information, dynamic changes in node states, etc. The calculation of the ambiguity value can help understand this uncertainty, and thus provide more reliable results for network analysis. This comprehensive value can be calculated in various ways, for example, considering factors such as the similarity of node attributes, connection strength, distance, etc., and using fuzzy logic or fuzzy set theory to evaluate.
[0074] Specifically, in this step, for all the connecting edges in the first directed network graph, calculate their ambiguity values , where frame node p and frame node q are the two frame nodes corresponding to the connecting edge respectively, and the ambiguity value is calculated by the following formula:
[0075] ;
[0076] In the formula, d p,q represents the Euclidean distance between frame node p and frame node q, d maxDenotes the maximum distance of all connected edges in the first directed network graph, k p Denotes the degree of the frame node p, k q Denotes the degree of the frame node q, k max Denotes the maximum value of the degrees of frame nodes in the first directed network graph.
[0077] Meanwhile, in this step, when selecting the pair of frame nodes corresponding to the connected edge with the largest ambiguity value as the object to be judged, if there are multiple connected edges with the largest ambiguity value, randomly select one of the connected edges with the largest ambiguity value.
[0078] In this embodiment, by calculating the ambiguity values of node pairs, the overall uncertainty between them can be effectively quantified. This is very important for analyzing the relationships, information transmission, and propagation capabilities between frame nodes. It can also analyze their interactions more deeply. For example, in video analysis, it may be found that the changes between certain image frames have a high degree of ambiguity, indicating that the content of these image frames needs further attention.
[0079] The algorithm provided in this embodiment can interact with the user. In this way, the user can judge whether the source frame node is similar to the representative frame node. The judgment results include yes and no, where yes represents similarity and no represents dissimilarity.
[0080] When the user compares the similarity of these two frame nodes, what is compared is the similarity of the two image frames corresponding to the two frame nodes. In this way, the user can make a judgment by comparing the images, with a lower difficulty, so that any user can operate, which can improve the universality of this method.
[0081] In addition, comparing two images basically does not require special professional skills and knowledge reserves, so the requirements for the user's professional skills and knowledge reserves can be reduced, and the application flexibility is higher.
[0082] S9. Submitting the source frame node corresponding to the disconnected connected edge and the representative frame nodes in the representative frame node set to the user one by one for judging similarity: If they are similar, connect the source frame node and the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and set the ambiguity value of the connected edge corresponding to the source frame node and the representative frame node to 0, update the first directed network graph again, and terminate this round of judgment; if they are all dissimilar, use the source frame node as the representative frame node and add it to the representative frame node set.
[0083] In this step, first, the Euclidean distances between each representative frame node in the representative frame node set and the source frame node;
[0084] Then, sort each representative frame node in the set of representative frame nodes in ascending order of the Euclidean distance from the source frame nodes to obtain a sequence of representative frame nodes;
[0085] Finally, submit the representative frame nodes and the source frame nodes to the user one by one in the order of the sequence of representative frame nodes to determine whether they are similar: the determination method is as shown above.
[0086] S10. Repeat the above S8 and S9 for the first directed network graph until a preset number of repetitions is reached to obtain a final second directed network graph, where the second directed network graph includes multiple second connected subgraphs:
[0087] During the repetition process, the mutually nearest neighbor node pairs in the first directed network graph can be continuously judged, and the scene categories can be continuously determined to identify all the scenes in the first directed network graph, thereby improving the accuracy of scene recognition for the video to be processed. Among the output multiple second connected subgraphs, each subgraph represents a different scene.
[0088] The condition for ending the repetition is reaching the preset number of repetitions. Those skilled in the art can set a suitable number of repetitions according to the actual situation. As Figure 5 shown, it is a second directed network graph composed of multiple second connected subgraphs, and each second connected subgraph represents a different scene.
[0089] S11. Based on each second connected subgraph in the second directed network graph, calculate the network centrality index of each frame node in each second connected subgraph, and select the frame node with the largest centrality index as the key frame of the second connected subgraph, and output the key frames of all second connected subgraphs after aggregation to complete the key frame extraction. This step includes the following sub-steps:
[0090] S1101. Take the frame node with the largest network centrality index in each of the second connected subgraphs as a key frame;
[0091] S1102. Output the key frame corresponding to the second connected subgraph to obtain the key frames corresponding to different category scenes.
[0092] In summary, the method of this embodiment can extract a key frame based on the scenarios of each category to obtain more accurate key frames, so as to accurately understand the content of the video to be processed. Compared with the key frame extraction method based on clustering, in the method provided by this embodiment of the present application, all scene categories in the video to be processed can be recognized. The scene categories in the clustering method are set in advance and cannot accurately correspond to the actual scene categories of the video to be processed, which has certain limitations. The key frame extraction method provided by this embodiment combines active learning technology, which can achieve a more comprehensive, accurate and flexible key frame extraction process.
[0093] The present invention discloses a video key frame extraction method based on active learning. First, a first directed network graph is constructed according to the nearest neighbor rule, and a pair of frame nodes composed of the two frame nodes corresponding to the edge with the largest ambiguity value is selected and handed over to the user for judgment to determine whether the node pair belongs to the same category of scenes. If the user determines that they are not similar, taking the source node of this edge as the benchmark, new node pairs are sequentially formed in the order of the Euclidean distance from near to far to each node in the set of representative frame nodes and handed over to the user for judgment again, and the network structure is adjusted in real time according to the user's results, and a second directed network graph is iteratively generated. Finally, the node with the highest network centrality index is selected as the key frame in each connected subgraph. In this way, an active learning mechanism can be introduced through interaction with the user. And the user's judgment result is used as a kind of topological guidance to realize the understanding of the scene category corresponding to the first directed network graph, so as to improve the accuracy of the system's understanding of the video content, generate accurate key frames, and further provide efficient and stable technical support for subsequent video summary generation and content retrieval.
[0094] It should be noted that those skilled in the art will easily think of other implementation schemes of the present application after considering the specification and practicing the application disclosed herein. The present application aims to cover any variations, uses or adaptations of the present application, and these variations, uses or adaptations follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0095] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The true scope is indicated by the present application.
Claims
1. A method for extracting key frames of a video based on active learning technology, characterized in that, Including: S1. Obtain the image frame sequence corresponding to the video to be processed; S2. Use an image feature extraction algorithm to obtain the feature vector of each image frame in the image frame sequence; S3. Take the feature vector as the frame node of the image frame, and then obtain a set of frame nodes. Each frame node in the set of frame nodes is used as a candidate frame node; S4. Calculate the Euclidean distance between any two candidate frame nodes respectively, and connect each candidate frame node to the candidate frame node with the closest Euclidean distance to it to form an edge. The direction of the edge points from the current candidate frame node to its closest candidate frame node to obtain a plurality of first connected subgraphs; The first connected subgraph contains at least two candidate frame nodes, and any one of the first connected subgraphs only includes a set of mutually nearest neighbor node pairs. The Euclidean distances between the two candidate frame nodes in the mutually nearest neighbor node pair are the closest to each other, and the two candidate frame nodes in the mutually nearest neighbor node pair point to each other; S5. Calculate the network centrality index of the two candidate frame nodes corresponding to the mutually nearest neighbor node pair, and disconnect the edge where the candidate frame node with the larger network centrality index points to the candidate frame node with the smaller network centrality index. Among them, the candidate frame node being pointed to among the two candidate frame nodes corresponding to the mutually nearest neighbor node is the target frame node, and the other candidate frame node is the source frame node; S6. Take the candidate frame node with the larger network centrality index in the mutually nearest neighbor node pair as the representative frame node, and all the representative frame nodes corresponding to all the current first connected subgraphs form a set of representative frame nodes; S7. Based on the representative frame nodes obtained in S6, repeat the above S4 - S6 and update the set of representative frame nodes until there is only one representative frame node in the set of representative frame nodes, and obtain the first directed network graph; S8. Calculate the ambiguity value of each edge in the first directed network graph; select the frame node pair corresponding to the edge with the largest ambiguity value as the object to be judged, and let the user determine whether the object to be judged is similar: if the user's judgment result is yes, retain the edge corresponding to the object to be judged; if the user's judgment result is no, disconnect the edge corresponding to the object to be judged, update the first directed network graph, and enter S9; S9. Present the source frame node corresponding to the disconnected edge and the representative frame nodes in the set of representative frame nodes to the user one by one to judge whether they are similar: if they are similar, connect the source frame node to the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and set the ambiguity value of the edge corresponding to the source frame node and the representative frame node to 0, and update the first directed network graph again and terminate this round of judgment; if they are not similar, take the source frame node as the representative frame node and add it to the set of representative frame nodes; S10. Repeat the above S8 and S9 for the first directed network graph until a preset number of repetitions is reached to obtain the final second directed network graph, and the second directed network graph contains a plurality of second connected subgraphs; S11. Based on any second connected subgraph, calculate the network centrality index of each frame node in each second connected subgraph, select the frame node with the largest centrality index as the key frame of the second connected subgraph, and output the key frames of all second connected subgraphs after aggregation to complete the key frame extraction.
2. The method for extracting video key frames based on active learning technology according to claim 1, wherein, In S1, an image frame sequence is obtained through shot boundary segmentation.
3. The method for extracting video key frames based on the active learning technology according to claim 1, wherein The S2 includes the following steps: S201. Convert each image frame in the image frame sequence into a grayscale image frame; S202. Use the HOG algorithm to divide the grayscale image frame into multiple image blocks and perform feature extraction to obtain an initial feature vector; S203. Calculate the information entropy weight coefficient of each image block; S204. Perform weighted processing on the initial feature vector based on the information entropy weight coefficient in S203 to obtain a weighted feature vector; S205. Perform dimensionality reduction processing on the weighted feature vector to obtain the feature vector of the image frame, and use the feature vector as the frame node of the image frame; S206. Repeat S201 - S205 until the feature vectors of all image frames in the image frame sequence are obtained.
4. The method for extracting video key frames based on the active learning technique according to claim 3, wherein The information entropy weight coefficient is calculated based on the following formula: ; where ω b (i) is the information entropy weight coefficient of the i-th image block, is the occurrence probability of the pixel with the gray value of in the i-th image block.
5. The video key frame extraction method based on the active learning technology according to claim 1, wherein The network centrality index is calculated by the following formula: , where φ Total (v) represents the network centrality index of frame node v, α is the first weight coefficient, and φ Local (v) represents the local centrality of frame node v, and φ Semilocal (v) represents the semi-local centrality of frame node v; The local centrality is calculated through the following formula: ; where Degree v represents the degree of the frame node v, and N represents the number of all frame nodes; The semi - local centrality is calculated through the following formula: ; where φ Semilocal (v) represents the semi-local centrality of the frame node v; |G N2 (v)| represents the total number of first-order and second-order neighborhood nodes of the frame node v; represents the set of frame nodes within the first-order neighborhood of the frame node v, represents the set of frame nodes within the second-order neighborhood of the frame node v; k v represents the degree of the frame node v; k u represents the degree of the frame node u; the frame node u is a frame node within the first-order or second-order neighborhood of the frame node v; d u,v represents the distance between the frame node v and the frame node u; d max1 represents the maximum distance between the frame node v and all frame nodes within its first-order neighborhood, d max2 represents the maximum distance between the frame node v and all frame nodes within its second-order neighborhood.
6. The method for extracting video key frames based on active learning technology according to claim 1, wherein In step S8, for all the connecting edges in the first directed network graph, calculate their ambiguity values , where the frame node p and the frame node q are respectively the two frame nodes corresponding to the connecting edge, and the ambiguity value is calculated by the following formula: ; where, d p,q represents the Euclidean distance between frame node p and frame node q, d max represents the maximum distance of all edges in the first directed network graph, k p represents the degree of the frame node p, k q represents the degree of the frame node q, k max represents the maximum value of the degrees of frame nodes in the first directed network graph.
7. The video key frame extraction method based on active learning technology according to claim 1, wherein In S8, when selecting the frame node pair corresponding to the edge with the largest ambiguity value as the object to be judged, if there are multiple edges with the largest ambiguity value, randomly select one edge with the largest ambiguity value.
8. The video key frame extraction method based on the active learning technique according to claim 1, characterized in that In S9, the source frame node corresponding to the disconnected edge and the representative frame nodes in the representative frame node set are handed over to the user to judge whether they are similar one by one, including the following sub - steps: Determine the Euclidean distance between each representative frame node in the representative frame node set of S9 and the source frame node; Sort each representative frame node in the representative frame node set in ascending order of the Euclidean distance from the source frame node to obtain a representative frame node sequence; Hand over the representative frame nodes and the source frame node to the user to judge whether they are similar one by one in the order of the representative frame node sequence: If they are similar, connect the source frame node and the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and set the ambiguity value of the edge corresponding to the source frame node and the representative frame node to 0, and update the first directed network graph again, and terminate this round of judgment; If they are not similar, use the source frame node as a representative frame node and add it to the representative frame node set.
9. The method for extracting video key frames based on active learning technology according to claim 1, characterized in that, The S11 includes: S1101. Use the frame node with the largest network centrality index in each second connected subgraph existing in the second directed network graph as a key frame; S1102. Output the key frames corresponding to the second connected subgraph to obtain the key frames corresponding to different category scenarios.
Citation Information
Patent Citations
Video scene detection method based on graph partitioning and instance learning
CN104318208A
Intelligent video splitting method based on graph convolutional neural network
CN111126126A