Video key frame extraction method based on active learning technology

By constructing a directed network graph and using network centrality index and user interaction, the accurate extraction of video keyframes is achieved, the problem of inaccurate extraction of keyframes in the prior art is solved, and the management and understanding of video content is improved.

CN120107867AActive Publication Date: 2025-06-06SOUTHWEST PETROLEUM UNIV

Patent Information

Application Number
CN202510585476.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately extract keyframes in videos, which affects the accuracy of video summary and the efficient management of massive video content.

Method used

The video keyframe extraction method based on active learning technology is adopted, and the keyframe nodes are gradually determined by building a directed network graph, using the network centrality index and user interaction, and the keyframe nodes are gradually determined to achieve accurate keyframe extraction.

Benefits of technology

It improves the accuracy of video keyframe extraction, can identify keyframes of various categories of scenes in the video, and enhances the understanding and management capabilities of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107867A_ABST
    Figure CN120107867A_ABST
Patent Text Reader

Abstract

The invention discloses a video key frame extraction method based on an active learning technology, and relates to the technical field of video processing. And constructing a first directed network graph according to a nearest neighbor rule, and selecting a frame node pair consisting of two frame nodes corresponding to the connecting edge with the maximum fuzziness value to be submitted to a user for judgment, so as to determine whether the node pair belongs to the same class of scenes. If the user judges that the connection edges are not similar, taking the source node of the connection edge as a reference, sequentially forming new frame nodes according to the sequence of the Euclidean distances between the source node of the connection edge and each frame node of the representative frame node set from near to far so as to send the new frame nodes to the user again for judgment, adjusting the network structure according to a user result, and generating a second directed network graph comprising a plurality of second connected sub-graphs, and finally, selecting a node with the highest network centrality index from each second connected sub-graph as a key frame. According to the method, an active learning mechanism is introduced through interaction with the user, the judgment result of the user serves as topological guidance, and the accurate key frame is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a video key frame extraction method based on active learning technology. Background Art

[0002] With the rapid development of social platforms and the widespread application of multimedia technology, as the main carrier of multimedia information, the amount of digital video data has shown explosive growth. How to efficiently store, manage and retrieve massive video content has become a key problem in realizing functions such as video semantic analysis, personalized recommendation and intelligent content review. Video summarization technology is an important way to deal with this problem, and the core lies in the selection of key frames. Key frames refer to image frames with the most concentrated information in video clips. The purpose of key frame extraction is to select several representative frames from a continuous video sequence in order to condense and summarize the main content of the entire video.

[0003] Based on the above situation, there is an urgent need for a method that can accurately extract key frames of a video. Summary of the invention

[0004] In order to solve the above problems, the present application provides a video key frame extraction method based on active learning technology, which can output all key frames corresponding to scenes of various categories in the video to improve the accuracy of video key frame extraction.

[0005] In order to achieve the above objectives, in a first aspect, the present application provides a video key frame extraction method based on active learning technology, comprising: S1. Obtaining an image frame sequence corresponding to the video to be processed; S2. using an image feature extraction algorithm to obtain a feature vector of each image frame in the image frame sequence; S3, taking the feature vector as a frame node of the image frame, and then obtaining a frame node set, wherein each frame node in the frame node set is taken as a candidate frame node; S4, respectively calculating the Euclidean distance between any two candidate frame nodes, connecting each candidate frame node with the candidate frame node with the closest Euclidean distance to form an edge, the direction of the edge is from the current candidate frame node to the nearest candidate frame node, so as to obtain multiple first connected subgraphs; the first connected subgraphs contain at least two candidate frame nodes, and any of the first connected subgraphs only includes a set of mutually nearest neighbor node pairs, the Euclidean distances between the two candidate frame nodes in the mutually nearest neighbor node pairs are the closest to each other, and the two candidate frame nodes of the mutually nearest neighbor node pair point to each other; S5, calculating the network centrality indexes corresponding to the two candidate frame nodes in the mutually nearest neighbor node pair, disconnecting the edge from the candidate frame node with a larger network centrality index to the candidate frame node with a smaller network centrality index, wherein the candidate frame node pointed to by the two candidate frame nodes corresponding to the mutually nearest neighbor nodes is the target frame node, and the other candidate frame node is the source frame node; S6, taking the candidate frame node with a larger network centrality index among the mutual nearest neighbor node pairs as a representative frame node, and forming a representative frame node set with all representative frame nodes corresponding to all current first connected subgraphs; S7, taking the representative frame node obtained in S6 as a reference, repeating the above S4-S6 and updating the representative frame node set until there is only one representative frame node in the representative frame node set, and obtaining a first directed network graph; S8, calculating the fuzziness value of each edge in the first directed network graph; selecting the frame node pair corresponding to the edge with the largest fuzziness value as the object to be judged, and letting the user determine whether the objects to be judged are similar: if the user's judgment result is yes, retaining the edge corresponding to the object to be judged; if the user's judgment result is no, disconnecting the edge corresponding to the object to be judged, updating the first directed network graph, and entering S9; S9, the source frame node corresponding to the disconnected edge and the representative frame node in the representative frame node set are submitted to the user for judgment one by one to see whether they are similar: if they are similar, the source frame node is connected to the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and the fuzziness value of the edge corresponding to the source frame node and the representative frame node is set to 0, the first directed network graph is updated again, and this round of judgment is terminated; if they are not similar, the source frame node is taken as the representative frame node and added to the representative frame node set; S10, repeating the above S8 and S9 for the first directed network graph until a preset number of repetitions is reached, to obtain a final second directed network graph, wherein the second directed network graph includes a plurality of second connected subgraphs; S11. Based on any second connected subgraph, the network centrality index of each frame node in each second connected subgraph is calculated, and the frame node with the largest centrality index is selected as the key frame of the second connected subgraph. The key frames of all the second connected subgraphs are aggregated and output to complete the key frame extraction.

[0006] The beneficial effects of the present invention are as follows: the present invention discloses a method for extracting video key frames based on active learning, firstly constructing a first directed network graph according to the nearest neighbor rule, and selecting a frame node pair consisting of two frame nodes corresponding to the edge with the largest fuzziness value for the user to judge, so as to determine whether the node pair belongs to the same type of scene. If the user determines that they are not similar, then taking the source node of the edge as the reference, a new frame node pair is sequentially formed in the order of the Euclidean distance from each frame node representing the frame node set from near to far, and then submitted to the user for judgment again, and the network structure is adjusted in real time according to the user results, and the second directed network graph is iteratively generated. Finally, the node with the highest network centrality index in each second connected subgraph is selected as the key frame. The method of the present invention introduces an active learning mechanism through interaction with the user, uses the user's judgment result as a topology guide, and generates accurate key frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0008] Figure 1 is a flow chart of an embodiment of the present invention; Figure 2 is a schematic diagram of the initial frame node distribution graph and the first connected subgraph; wherein, Figure 2 (a) is the initial frame node distribution diagram, Figure 2 (b) is the first connected subgraph; Figure 3 is a schematic diagram representing frame nodes in the first connected subgraph; Figure 4 is a schematic diagram of a first directed network graph according to an embodiment of the present invention; Figure 5 It is a schematic diagram of a second connected subgraph according to an embodiment of the present invention. DETAILED DESCRIPTION

[0009] The technical solutions in the embodiments of the present application will be described clearly below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, other embodiments obtained by ordinary technicians in this field without making creative work all belong to the protection scope of the present application.

[0010] In the following, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0011] With the rapid development of social platforms and the widespread application of multimedia technology, as the main carrier of multimedia information, the amount of digital video data has shown explosive growth. How to efficiently store, manage and retrieve massive video content has become a key problem in realizing functions such as video semantic analysis, personalized recommendation and intelligent content review. Video summarization technology is an important way to deal with this problem, and the core lies in the selection of key frames. Key frames refer to image frames with the most concentrated information in video clips. The purpose of key frame extraction is to select several representative frames from a continuous video sequence in order to condense and summarize the main content of the entire video.

[0012] For example, in an intelligent driving system, the vehicle's driving recorder will collect many driving records, and the vehicle's memory will store the driving records collected within a certain period of time and update them regularly. In the intelligent driving system, when storing the driving records, summary information can be generated based on the key frames corresponding to the driving records. At this time, the driving records stored in the memory will include video clips and summary information. In this way, if the user wants to extract a certain driving record, he can search based on the summary information corresponding to the driving record to quickly obtain the desired driving record without having to search for it one by one.

[0013] Based on the above reasons, the accuracy of video key frame extraction will greatly affect the accuracy of the summary information corresponding to the video segment. Therefore, the embodiment of the present application provides a video key frame extraction method based on active learning technology, which can output all the key frames corresponding to the scenes of each category in the video to improve the accuracy of video key frame extraction.

[0014] Figure 1 This is a flowchart of a video key frame extraction method based on active learning technology provided in an embodiment of the present application.

[0015] like Figure 1 As shown, the video key frame extraction method based on active learning technology provided in the embodiment of the present application includes: S1. Obtain an image frame sequence corresponding to a video to be processed, wherein the image frame sequence is obtained by performing shot boundary segmentation processing on the video to be processed.

[0016] Specifically, in this step, there are many methods for obtaining the image frame sequence of the video to be processed, such as conventional shot boundary monitoring, motion analysis extraction, etc. According to the characteristics of each method, in this embodiment, the shot boundary segmentation method is selected as the method for obtaining the image frame sequence of the video to be processed. In particular, the image frame sequence refers to a sequence obtained by sorting all the image frames in time; at the same time, when the acquisition time is long and the number of image frames is too large, the image frame sequence at this time can also be a sequence obtained by sampling the image frames at a certain time interval and sorting them in time.

[0017] Shot Boundary Detection (SBD) is a technique in video processing that is used to identify the switching points between different shots (or scenes) in a video. A shot refers to a continuous segment of a video, while a shot boundary refers to the transition between two adjacent shots, which is usually accompanied by a significant change in the content of the picture. Shot boundary segmentation can detect the switching of scenes in a video and distinguish when the transition between shots occurs. Shot boundary segmentation also helps understand the structure of the video, facilitating subsequent analysis, retrieval, and processing.

[0018] Shot boundaries usually represent major changes in the scene, and centralized processing of these frames can reduce the consumption of computing resources. The image frames obtained after shot boundary segmentation are often better used for subsequent analysis and processing tasks, such as video summarization, content retrieval, etc. These frames usually represent different scenes or plot developments, which help to understand the overall structure of the video. In addition, in the video, some frames may be noise caused by slight changes in the scene, which may interfere with the results of key frame extraction without shot boundary segmentation. Shot boundary segmentation can effectively filter out these noisy frames. Therefore, the image frame sequence after shot boundary segmentation can improve the efficiency of key frame extraction.

[0019] S2. using an image feature extraction algorithm to obtain a feature vector of each image frame in the image frame sequence; Among them, image feature extraction algorithms are technologies used in the fields of computer vision and image processing to extract useful information and descriptions from images. These features can be used for various tasks such as classification, recognition, retrieval, and image analysis. The purpose of feature extraction is to convert the original image into a more concise and high-dimensional representation for subsequent processing.

[0020] Common image feature extraction algorithms include edge detection algorithms, corner detection algorithms, feature descriptors, and texture features.

[0021] The feature vector of an image frame can compress complex image information into a relatively concise numerical representation. This representation retains the key features of the image (such as color, texture, shape, etc.), making subsequent processing and analysis more efficient. The feature vector can be used to calculate the similarity between image frames. By comparing the feature vectors of adjacent frames, it is possible to identify which frames have significant changes in content, thereby helping to determine shot switching points and key frames.

[0022] In this embodiment, an edge detection algorithm is used to extract a feature vector corresponding to an image frame, which may specifically include the following steps: S201. Convert each image frame in the image frame sequence into a grayscale image frame.

[0023] S202, using the HOG algorithm to divide the grayscale image frame into multiple image blocks and perform feature extraction to obtain an initial feature vector, where the initial feature vector is , is the constant value of the grayscale image frame in the Mth dimension.

[0024] Among them, the Histogram of Oriented Gradient (HOG) is a technology used to describe image features and is widely used in object detection, especially in the fields of pedestrian detection, vehicle detection, and face recognition. The HOG feature describes the shape and structure of an image by capturing the gradient direction and magnitude of a local area in the image.

[0025] S203: Calculate the information entropy weight coefficient of each image block.

[0026] S204, re-weighting the initial feature vector based on the information entropy weight coefficient of S203 to obtain a weighted feature vector; S205, performing dimensionality reduction processing on the weighted feature vector to obtain a feature vector of the image frame, and using the feature vector as a frame node of the image frame; S206, repeat S201 to S205 until the feature vectors of all the image frames in the image frame sequence are obtained.

[0027] In the above-mentioned feature vector extraction process, converting the image frame into a grayscale image frame provides a more efficient, more stable and more robust processing method for edge detection. Grayscale images play an important role in edge detection algorithms by simplifying calculations and improving the extractability of edge features. The information entropy weight coefficient of each image block indicates the amount of image information corresponding to the image block. In other words, the larger the information entropy weight coefficient of the image block, the more information the image block contains; the smaller the information entropy weight coefficient of the image block, the less information the image block contains. In this way, each image block can be weighted based on the information entropy weight coefficient to obtain a weighted feature vector. Finally, the weighted feature vector is subjected to dimensionality reduction processing to simplify the dimension of the feature vector through dimensionality reduction processing, reduce the amount of subsequent calculations, and obtain the final feature vector.

[0028] Exemplarily, the information entropy weight coefficient can be calculated based on the following formula: , where ω b (i) is the information entropy weight coefficient of the i-th image block, is the gray value of the i-th image block The probability of pixel occurrence.

[0029] For example, principal component analysis (PCA) can be used for dimensionality reduction. PCA is a commonly used dimensionality reduction technique that converts high-dimensional data into low-dimensional data through linear transformation while retaining the variability of the original data as much as possible. PCA is widely used in many fields such as data preprocessing, feature extraction, and image processing.

[0030] In dynamic scenes, image frames may be affected by noise, lighting changes and other environmental factors. After reducing their dimensionality through methods such as PCA, unnecessary noise can be filtered out and the robustness of target recognition and tracking can be improved.

[0031] S3. Using the feature vector as a frame node of the image frame, and then obtaining a frame node set, wherein each frame node in the frame node set is used as a candidate frame node.

[0032] By using the feature vector corresponding to each image frame as the frame node of the image frame, in the video processing task, converting each frame of the image into a frame node helps to capture the changes and dynamic features in time, thereby analyzing action sequences and events. At the same time, after using the feature vector as the frame node of the image frame, the feature vector is no longer just a data representation, but a node entity in the network diagram.

[0033] S4. Calculate the Euclidean distance between any two candidate frame nodes respectively, connect each candidate frame node with the candidate frame node with the closest Euclidean distance to form an edge, and the direction of the edge is from the current candidate frame node to the nearest candidate frame node to obtain multiple first connected subgraphs; the first connected subgraph contains at least two candidate frame nodes, and any of the first connected subgraphs only includes a set of mutually nearest neighbor node pairs, the Euclidean distances between the two candidate frame nodes in the mutually nearest neighbor node pairs are the closest to each other, and the two candidate frame nodes of the mutually nearest neighbor node pair point to each other.

[0034] like Figure 2 As shown, we call the subgraph formed after multiple pairings in this step the first connected subgraph. Figure 2 (a) is the initial frame node diagram, Figure 2 (b) is the first connected subgraph. In the figure, the direction indicated by the arrow is the connection direction between node pairs, and the nearest neighbor node pairs are connected to each other.

[0035] S5. Calculate the network centrality indexes corresponding to the two candidate frame nodes in the mutually nearest neighbor node pair, and disconnect the edge that points the candidate frame node with a larger network centrality index to the candidate frame node with a smaller network centrality index, wherein the candidate frame node pointed to by the two candidate frame nodes corresponding to the mutually nearest neighbor nodes is the target frame node, and the other candidate frame node is the source frame node.

[0036] In this step, the network centrality index is an indicator used to measure the importance or influence of nodes in the network. For frame nodes, commonly used centrality indices include local centrality and semi-local centrality. Local centrality may be too affected by some frame nodes (for example, the degree of frame nodes is particularly high and has no actual influence), making it unable to fully reflect the structural characteristics in complex networks. Semi-local centrality can capture more complex connection patterns, such as the strength of second-order connections, which is critical for understanding the role of frame nodes in the network. By introducing semi-local centrality, this influence can be balanced, making the importance assessment of frame nodes more stable and reliable.

[0037] In order to balance local centrality and semi-local centrality, in this embodiment, a new network centrality index is proposed, and its calculation formula is as follows: , where φ Total (v) represents the network centrality index of frame node v, α is the first weight coefficient, φ Local (v) represents the local centrality of frame node v, φ Semilocal (v) represents the semi-local centrality of frame node v; The local centrality is calculated by the following formula: ; In the formula, Degree v represents the degree of the frame node v, and N represents the number of all frame nodes; The semi-local centrality is calculated by the following formula: ; In the formula, φ Semilocal (v) represents the semi-local centrality of the frame node v; |G N2 (v)| represents the total number of first-order neighboring nodes and second-order neighboring nodes of the frame node v; represents the set of frame nodes in the first-order neighborhood of the frame node v, represents the set of frame nodes in the second-order neighborhood of the frame node v; k v represents the degree of the frame node v; k u represents the degree of frame node u; the frame node u is a frame node in the first-order neighborhood or second-order neighborhood of frame node v; d u,v represents the distance between the frame node v and the frame node u; d max1 represents the maximum distance between the frame node v and all frame nodes in its first-order neighborhood, d max2 Represents the maximum distance between the frame node v and all frame nodes in its second-order neighborhood.

[0038] The first-order neighborhood referred to here refers to the set of frame nodes directly connected to the nearest neighbor node pair of the frame node v; the second-order neighborhood referred to here refers to the set of frame nodes directly connected to the frame nodes in the first neighborhood.

[0039] Local centrality refers to the importance of a certain frame node among its direct neighbors, and semi-local centrality is an extension of local centrality, which usually takes into account the influence of the frame node's neighbors and its neighbors.

[0040] In this embodiment, the network centrality index is obtained by combining the weighted local centrality and semi-local centrality, which can comprehensively evaluate the importance of frame nodes in the entire network and avoid information loss caused by a single perspective. It can more accurately identify key frames or important events, effectively improving the efficiency of target tracking and event detection.

[0041] S6. The candidate frame node with a larger network centrality index in the mutual nearest neighbor node pair is taken as a representative frame node, and all representative frame nodes corresponding to all current first connected subgraphs form a representative frame node set.

[0042] like Figure 3 As shown, one of the edges between the nearest neighbor node pairs has been disconnected and the representative frame node has been selected.

[0043] S7. Taking the representative frame node obtained in S6 as a reference, repeat S4-S6 and update the representative frame node set until there is only one representative frame node in the representative frame node set, and obtain a first directed network graph.

[0044] In this step, only the representative frame nodes obtained in the previous step are selected for operation, and the remaining frame nodes are temporarily ignored. Finally, the first directed network graph is obtained including the frame nodes corresponding to all image frames.

[0045] like Figure 4 As shown, it is the first directed network graph finally formed.

[0046] S8, calculating the fuzziness value of each edge in the first directed network graph; selecting the frame node pair corresponding to the edge with the largest fuzziness value as the object to be judged, and letting the user determine whether the objects to be judged are similar: if the user's judgment result is yes, retaining the edge corresponding to the object to be judged; if the user's judgment result is no, disconnecting the edge corresponding to the object to be judged, updating the first directed network graph, and entering S9; Among them, the fuzziness value of the edge is the fuzziness value of the two frame nodes in the node pair. The fuzziness value is mainly used to measure the fuzziness of the relationship and mutual influence between a node and its connected nodes in the network. In the first directed network graph, the fuzziness value usually refers to how to quantify the fuzziness of the similarity, influence or connection strength of these frame nodes, considering the relationship between the two frame nodes connected by the edge lock. The fuzziness value in the network is the uncertainty of the relationship between frame nodes, which may be caused by incomplete information, dynamic changes in node status, etc. The calculation of fuzziness values ​​can help understand this uncertainty and provide more reliable results for network analysis. This comprehensive value may be calculated in a variety of ways, for example, considering factors such as the similarity of node attributes, connection strength, distance, etc., and using fuzzy logic or fuzzy set theory to evaluate.

[0047] Specifically, in this step, for all the edges in the first directed network graph, the fuzziness value is calculated. , where frame node p and frame node q are the two frame nodes corresponding to the edge, and the fuzziness value Calculated by the following formula: ; Where, d p,q represents the Euclidean distance between frame node p and frame node q, d max represents the maximum distance of all edges in the first directed network graph, k p represents the degree of the frame node p, k q represents the degree of the frame node q, k max Represents the maximum value of the frame node degree in the first directed network graph.

[0048] Meanwhile, in this step, when the frame node pair corresponding to the edge with the largest fuzziness value is selected as the object to be judged, if there are multiple edges with the largest fuzziness value, one edge with the largest fuzziness value is randomly selected.

[0049] In this embodiment, by calculating the fuzziness value of the node pair, the overall uncertainty between them can be effectively quantified. This is very important for analyzing the relationship, information transfer and communication capabilities between frame nodes. It can also analyze the interaction between them more deeply. For example, in video analysis, it may be found that the changes between certain image frames have high fuzziness, suggesting that further attention needs to be paid to the content of these image frames.

[0050] The algorithm provided in this embodiment can interact with the user, so that the user can judge whether the source frame node is similar to the representative frame node. The judgment result includes yes and no, where yes means similarity and no means dissimilarity.

[0051] When the user compares the similarity of the two frame nodes, what is compared is the similarity of the two image frames corresponding to the two frame nodes. In this way, the user can make a judgment by comparing the images, which is relatively easy, so that any user can perform the operation, which can improve the universality of the method.

[0052] In addition, comparing two images basically does not require special professional skills and knowledge reserves, thereby reducing the requirements for users' professional skills and knowledge reserves and providing higher application flexibility.

[0053] S9. The source frame node corresponding to the disconnected edge and the representative frame nodes in the representative frame node set are submitted to the user for judgment one by one to see whether they are similar: if they are similar, the source frame node is connected to the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and the fuzziness value of the edge corresponding to the source frame node and the representative frame node is set to 0, the first directed network graph is updated again, and this round of judgment is terminated; if they are not similar, the source frame node is taken as the representative frame node and added to the representative frame node set.

[0054] In this step, first, the Euclidean distance between each representative frame node in the representative frame node set and the source frame node is calculated; Then, each representative frame node in the representative frame node set is sorted from near to far according to the Euclidean distance between the representative frame node and the source frame node, so as to obtain a representative frame node sequence; Finally, the representative frame nodes and the source frame nodes are handed over to the user one by one in the order of the representative frame node sequence to determine whether they are similar: the determination method is as shown above.

[0055] S10, repeating the above S8 and S9 for the first directed network graph until a preset number of repetitions is reached, to obtain a final second directed network graph, wherein the second directed network graph includes a plurality of second connected subgraphs: In the repetitive process, the nearest neighbor node pairs in the first directed network graph can be continuously judged, and the scene categories can be continuously determined to identify all scenes in the first directed network graph, thereby improving the accuracy of scene recognition of the processed video. Each of the multiple second connected subgraphs outputted represents a different scene.

[0056] The condition for ending the repetition is to reach the preset number of repetitions. Those skilled in the art can set an appropriate number of repetitions according to actual conditions. Figure 5 As shown, it is a second directed network graph composed of multiple second connected subgraphs, and each second connected subgraph represents a different scene.

[0057] S11, based on each second connected subgraph in the second directed network graph, calculate the network centrality index of each frame node in each second connected subgraph, and select the frame node with the largest centrality index as the key frame of the second connected subgraph, and output the key frames of all the second connected subgraphs after aggregation, and complete the key frame extraction. This step includes the following sub-steps: S1101, taking a frame node with the largest network centrality index in each of the second connected subgraphs as a key frame; S1102: Output the key frame corresponding to the second connected subgraph to obtain the key frames corresponding to different categories of scenes.

[0058] In summary, the method of this embodiment can extract a key frame based on each category of scenes to obtain a more accurate key frame, so that the content of the video to be processed can be accurately understood. Compared with the key frame extraction method of the clustering method, the method provided in the embodiment of the present application can identify all scene categories in the video to be processed. However, the scene categories in the clustering method are set in advance and cannot accurately correspond to the actual scene categories of the video to be processed, which has certain limitations. The key frame extraction method provided in this embodiment is combined with active learning technology to achieve a more comprehensive, accurate and flexible key frame extraction process.

[0059] The invention discloses a video key frame extraction method based on active learning. First, a first directed network graph is constructed according to the nearest neighbor rule, and a frame node pair consisting of two frame nodes corresponding to the edge with the largest fuzziness value is selected and handed over to the user for judgment, so as to determine whether the node pair belongs to the same type of scene. If the user determines that they are not similar, a new node pair is formed in sequence from near to far in the order of the Euclidean distance with each node representing the frame node set based on the source node of the edge, and handed over to the user for judgment again, and the network structure is adjusted in real time according to the user result, and a second directed network graph is iteratively generated. Finally, the node with the highest network centrality index is selected as the key frame in each connected subgraph. In this way, an active learning mechanism can be introduced through interaction with the user. And the user's judgment result is used as a topological guide to realize the understanding of the scene category corresponding to the first directed network graph, so as to improve the accuracy of the system's understanding of the video content, and generate accurate key frames, thereby providing efficient and stable technical support for subsequent video summary generation and content retrieval.

[0060] It should be noted that those skilled in the art will easily think of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variation, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application.

[0061] It should be understood that the present application is not limited to the precise construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof, the true scope being indicated by the present application.

Claims

1. A video key frame extraction method based on active learning technology, characterized in that: include: S1. Obtaining an image frame sequence corresponding to the video to be processed; S2. using an image feature extraction algorithm to obtain a feature vector of each image frame in the image frame sequence; S3, taking the feature vector as a frame node of the image frame, and then obtaining a frame node set, wherein each frame node in the frame node set is taken as a candidate frame node; S4, respectively calculating the Euclidean distance between any two candidate frame nodes, connecting each candidate frame node to the candidate frame node with the closest Euclidean distance to form an edge, and the direction of the edge is from the current candidate frame node to the nearest candidate frame node, so as to obtain multiple first connected subgraphs; The first connected subgraph includes at least two candidate frame nodes, and any of the first connected subgraphs includes only one set of mutually nearest neighbor node pairs, the Euclidean distances between the two candidate frame nodes in the mutually nearest neighbor node pairs are the shortest, and the two candidate frame nodes of the mutually nearest neighbor node pair point to each other; S5, calculating the network centrality indexes corresponding to the two candidate frame nodes in the mutually nearest neighbor node pair, disconnecting the edge from the candidate frame node with a larger network centrality index to the candidate frame node with a smaller network centrality index, wherein the candidate frame node pointed to by the two candidate frame nodes corresponding to the mutually nearest neighbor nodes is the target frame node, and the other candidate frame node is the source frame node; S6, taking the candidate frame node with a larger network centrality index among the mutual nearest neighbor node pairs as a representative frame node, and forming a representative frame node set with all representative frame nodes corresponding to all current first connected subgraphs; S7, taking the representative frame node obtained in S6 as a reference, repeating the above S4-S6 and updating the representative frame node set until there is only one representative frame node in the representative frame node set, and obtaining a first directed network graph; S8, calculating the fuzziness value of each edge in the first directed network graph; selecting the frame node pair corresponding to the edge with the largest fuzziness value as the object to be judged, and letting the user determine whether the objects to be judged are similar: if the user's judgment result is yes, retaining the edge corresponding to the object to be judged; if the user's judgment result is no, disconnecting the edge corresponding to the object to be judged, updating the first directed network graph, and entering S9; S9, the source frame node corresponding to the disconnected edge and the representative frame node in the representative frame node set are submitted to the user for judgment one by one to see whether they are similar: if they are similar, the source frame node is connected to the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and the fuzziness value of the edge corresponding to the source frame node and the representative frame node is set to 0, the first directed network graph is updated again, and this round of judgment is terminated; if they are not similar, the source frame node is taken as the representative frame node and added to the representative frame node set; S10, repeating the above S8 and S9 for the first directed network graph until a preset number of repetitions is reached, to obtain a final second directed network graph, wherein the second directed network graph includes a plurality of second connected subgraphs; S11. Based on any second connected subgraph, the network centrality index of each frame node in each second connected subgraph is calculated, and the frame node with the largest centrality index is selected as the key frame of the second connected subgraph. The key frames of all the second connected subgraphs are aggregated and output to complete the key frame extraction.

2. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: In S1, an image frame sequence is obtained by segmenting the shot boundaries.

3. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: The S2 comprises the following steps: S201, converting each image frame in the image frame sequence into a grayscale image frame; S202, using the HOG algorithm to divide the grayscale image frame into multiple image blocks and perform feature extraction to obtain an initial feature vector; S203, calculating the information entropy weight coefficient of each image block; S204, performing weighted processing on the initial feature vector based on the information entropy weight coefficient of S203 to obtain a weighted feature vector; S205, performing dimensionality reduction processing on the weighted feature vector to obtain a feature vector of the image frame, and using the feature vector as a frame node of the image frame; S206, repeat S201 to S205 until the feature vectors of all the image frames in the image frame sequence are obtained.

4. The video key frame extraction method based on active learning technology according to claim 3 is characterized in that: The information entropy weight coefficient is calculated based on the following formula: ; In the formula, ω b (i) is the information entropy weight coefficient of the i-th image block, is the gray value of the i-th image block The probability of pixel occurrence.

5. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: The network centrality index is calculated by the following formula: , where φ Total (v) represents the network centrality index of frame node v, α is the first weight coefficient, φ Local (v) represents the local centrality of frame node v, φ Semilocal (v) represents the semi-local centrality of frame node v; The local centrality is calculated by the following formula: ; In the formula, Degree v represents the degree of the frame node v, and N represents the number of all frame nodes; The semi-local centrality is calculated by the following formula: ; In the formula, φ Semilocal (v) represents the semi-local centrality of the frame node v; |G N2 (v)| represents the total number of first-order neighboring nodes and second-order neighboring nodes of the frame node v; represents the set of frame nodes in the first-order neighborhood of the frame node v, represents the set of frame nodes in the second-order neighborhood of the frame node v; k v represents the degree of the frame node v; k u represents the degree of frame node u; the frame node u is a frame node in the first-order neighborhood or second-order neighborhood of frame node v; d u,v represents the distance between the frame node v and the frame node u; d max1 represents the maximum distance between the frame node v and all frame nodes in its first-order neighborhood, d max2 Represents the maximum distance between the frame node v and all frame nodes in its second-order neighborhood.

6. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: In step S8, for all the edges in the first directed network graph, the fuzziness values ​​are calculated. , where frame node p and frame node q are the two frame nodes corresponding to the edge, and the fuzziness value Calculated by the following formula: ; Where, d p,q represents the Euclidean distance between frame node p and frame node q, d max represents the maximum distance of all edges in the first directed network graph, k p represents the degree of the frame node p, k q represents the degree of the frame node q, k max Represents the maximum value of the frame node degree in the first directed network graph.

7. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: In S8, when the frame node pair corresponding to the edge with the largest fuzziness value is selected as the object to be judged, if there are multiple edges with the largest fuzziness value, one edge with the largest fuzziness value is randomly selected.

8. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: In S9, the source frame node corresponding to the disconnected edge and the representative frame nodes in the representative frame node set are submitted to the user one by one to determine whether they are similar, including the following sub-steps: Determine the Euclidean distance between each representative frame node in the representative frame node set of S9 and the source frame node; Sort the representative frame nodes in the representative frame node set from near to far according to the Euclidean distance between the representative frame nodes and the source frame node to obtain a representative frame node sequence; According to the order of the representative frame node sequence, the representative frame nodes and the source frame nodes are handed over to the user one by one to judge whether they are similar: if they are similar, the source frame node is connected to the representative frame node to merge the first connected subgraph corresponding to the source frame node and the first connected subgraph corresponding to the representative frame node, and the fuzziness value of the edge corresponding to the source frame node and the representative frame node is set to 0, the first directed network graph is updated again, and this round of judgment is terminated; If they are not similar, the source frame node is taken as a representative frame node and added to the representative frame node set.

9. The video key frame extraction method based on active learning technology according to claim 1 is characterized in that: The S11 includes: S1101, taking a frame node with the largest network centrality index in each second connected subgraph in the second directed network graph as a key frame; S1102: Output the key frame corresponding to the second connected subgraph to obtain the key frames corresponding to different categories of scenes.

Citation Information

Patent Citations

  • Video scene detection method based on graph partitioning and instance learning

    CN104318208A

  • Intelligent video splitting method based on graph convolutional neural network

    CN111126126A

  • Image frame extraction apparatus and image frame extraction method

    US20220207874A1

Cited By

  • Portrait video micro-expression frame extraction method based on active learning technology

    CN121482851A