Video frame extraction method and device, equipment, storage medium and program product
By extracting and segmenting similar frames from the video, the problem of excessive duplicate frames in video frame extraction is solved, achieving effective deduplication of video frames and preservation of key information, thus improving the accuracy of downstream processing.
Patent Information
- Application Number
- CN202511072191.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing video frame extraction methods cannot effectively remove video frames containing duplicate information, resulting in insufficient processing capabilities of downstream processing models.
By extracting time-series video frames from the video to be processed, determining the distance between frames using the feature vectors of the video frames, grouping similar frames into the same cluster, and extracting frames based on the intra-cluster frame distance, the duplicate frames are removed and key frames are retained.
This reduces duplicate video frames in the target frame sequence, preventing excessively long frame sequences from affecting downstream task processing, ensuring that the frame sequence contains key information, and improving the accuracy of downstream task processing.
Smart Images

Figure CN120913128A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present document relates to the technical field of video processing, and particularly relates to a video frame extraction method and device, equipment, storage medium and program product. BACKGROUND
[0002] Video frame extraction refers to a process of extracting a plurality of video frames from a video sequence according to a certain rule. The extracted video frames can be used for various application scenarios such as video analysis, image processing, target detection and tracking. However, if the number of extracted video frames exceeds the processing capacity of a downstream processing model, the downstream processing result will be affected.
[0003] For example, current video conferences are increasingly focusing on document projection, and the analysis of document projection video often relies on uniformly extracted video frame sequences for analysis. For example, if a one-hour video conference is extracted every 3 seconds, a frame sequence containing 1200 video frames will be formed. However, most large models currently used for document projection video analysis do not have the ability to process 1200 video frame sequences, and need to remove as many video frames containing repetitive information as possible, retain video frames containing key information, and then input them to the downstream large model for processing. SUMMARY
[0004] The purpose of the embodiments of the present specification is to provide a video frame extraction method, device, equipment, storage medium and program product to remove as many video frames containing repetitive information as possible, retain video frames containing key information, and supply the downstream processing model for better processing.
[0005] In order to achieve the above-mentioned purpose, the technical solutions adopted by the embodiments of the present specification are as follows: In a first aspect, a video frame extraction method is provided, comprising: extracting a first frame sequence from a to-be-processed video, the first frame sequence comprising a plurality of video frames arranged based on time sequence; determining distances between different video frames in the first frame sequence based on feature vectors of the video frames in the first frame sequence; dividing video frames with similar distances in the first frame sequence into the same cluster to obtain a plurality of first clusters; extracting frames from each first cluster based on distances between video frames in each first cluster to obtain a target frame sequence of the to-be-processed video.
[0006] In a second aspect, a video frame extraction device is provided, comprising: a first extraction module configured to extract a first frame sequence from a to-be-processed video, the first frame sequence comprising a plurality of video frames arranged based on time sequence; determining a distance between different video frames in the first frame sequence based on the feature vectors of the video frames in the first frame sequence; dividing the video frames with similar distances in the first frame sequence into a same cluster to obtain a plurality of first clusters; the second extracting module is configured to extract frames from each first cluster based on the distance between the video frames in each first cluster to obtain a target frame sequence of the video to be processed.
[0007] In a third aspect, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the video frame extraction method according to the first aspect.
[0008] In a fourth aspect, a computer-readable storage medium is provided, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the video frame extraction method according to the first aspect.
[0009] In a fifth aspect, a computer program product is provided, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform some or all of the steps in the video frame extraction method according to the first aspect.
[0010] According to the scheme of the embodiments of the present specification, a plurality of video frames arranged in time sequence are first extracted from a video to be processed as a candidate first frame sequence; then, considering that the feature vectors of the video frames reflect the semantics of the video, the distance between the video frames determined based on the feature vectors can reflect the similarity between the video frames; on this basis, the video frames with similar distances in the first frame sequence are divided into a same cluster to obtain a plurality of first clusters, so that the video frames in a same first cluster are similar; further, based on the distance between the video frames in each first cluster, a part of the video frames in the first cluster are extracted to form a target frame sequence, which can realize fast deduplication of the first cluster, reduce the repeated video frames in the final target frame sequence, avoid the target frame sequence being too long to affect the processing of downstream tasks, and ensure that the video frames in the target frame sequence contain key information, thereby providing reliable data support for improving the accuracy of the processing of downstream tasks. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings, which are included to provide a further understanding of the present specification, constitute a part of the present specification, and the illustrative embodiments of the present specification and their description serve to explain the present specification, and do not constitute improper limitations on the present specification. In the drawings: Figure 1A flowchart of a video frame extraction method provided for an embodiment of the present specification; Figure 2 A flowchart of a video frame extraction method provided for another embodiment of the present specification; Figure 3 A structural diagram of a video frame extraction device provided for an embodiment of the present specification; Figure 4 A structural diagram of an electronic device provided for an embodiment of the present specification. DETAILED DESCRIPTION
[0012] For the purpose, technical solutions and advantages of the present specification to be clearer, the technical solutions of the present specification will be described clearly and completely in combination with the embodiments of the present specification and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present specification, but not all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present document.
[0013] The term "comprising" and its variants thereof used in the present document are open and inclusive, i.e. "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The related definitions of other terms will be given in the following description. The term "in response to" is used to indicate the condition or state on which the operation is performed, when the dependent condition or state is met, the performed operation can be in real time or have a set delay, and there is no limitation on the order of multiple operations performed without special description.
[0014] It should be noted that the "first", "second", and the like concepts mentioned in the present document are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0015] It should be noted that the "one", "multiple" and the like modifiers in the present document are illustrative and not restrictive, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0016] The names of the messages or information exchanged between the devices in the embodiments of the present document are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0017] The embodiment of the present specification provides a video frame extraction method, first extracts a plurality of video frames arranged based on time sequence from a to-be-processed video as a candidate first frame sequence; then, considering that the feature vector of the video frame reflects the semantics of the video, the distance between the video frames determined therefrom can reflect the similarity between the video frames; on this basis, the adjacent video frames in the first frame sequence are divided into the same cluster to obtain a plurality of first clusters, so that the video frames in the same first cluster are similar; further, based on the distance between the video frames in each first cluster, a part of the video frames in the first cluster are extracted to form a target frame sequence, which can realize fast deduplication of the first cluster, reduce the repeated video frames in the finally obtained target frame sequence, avoid the influence of the long target frame sequence on the downstream task processing, and can ensure that the video frames in the target frame sequence contain key information, and provide reliable data support for improving the accuracy of downstream task processing.
[0018] The technical solutions provided by the embodiments of the present specification are described in detail below with reference to the drawings.
[0019] Please refer to Figure 1 A flowchart of a video frame extraction method provided by an embodiment of the present specification is shown, which can include the following steps: S102, extracting a first frame sequence from a to-be-processed video.
[0020] The to-be-processed video can be any video that needs to be processed for frame extraction. In an implementation, if a document projection video needs to be analyzed, the to-be-processed video is obtained by recording the document projection process in a meeting.
[0021] The first frame sequence includes a plurality of video frames arranged based on time sequence in the to-be-processed video.
[0022] In an implementation, the above S102 includes the following steps: sampling the to-be-processed video according to a preset interval duration to obtain the first frame sequence.
[0023] The preset interval duration can be set according to actual needs, such as being set according to the duration of the to-be-processed video, which is not limited in the embodiments of the present specification. The number of video frames contained in the first frame sequence obtained in this way is obtained according to the ratio between the duration of the to-be-processed video and the preset interval duration.
[0024] For example, the duration of the to-be-processed video is y seconds (such as 3600 seconds), and the to-be-processed video is sampled every x seconds (such as 3 seconds), and a first frame sequence with a length of y / x (such as 1200) is obtained, that is, the first frame sequence contains 1200 video frames.
[0025] In another implementation, the above S102 includes the following steps: S1022, sample the video to be processed according to a preset interval length to obtain a second frame sequence.
[0026] S1024, convert a video frame in the second frame sequence into a grayscale image.
[0027] The video frame is usually a multi-channel image (such as including R, G, and B three channels), and the grayscale image of the video frame is a single-channel image, which retains the key visual information of the video frame, but has less calculation amount compared with the video frame, and can accelerate the inter-frame difference analysis.
[0028] S1026, perform deduplication processing on the second frame sequence based on difference information between grayscale images corresponding to different video frames in the second frame sequence to obtain a first frame sequence.
[0029] The grayscale image of the video frame reflects the grayscale value of each pixel in the video frame. For any two video frames, the difference information between the grayscale images corresponding to the two video frames can include the absolute difference value between the grayscale values of the pixels at the same position in the two video frames.
[0030] The smaller the difference between the grayscale images corresponding to any two video frames in the second frame sequence, the more similar the two video frames are, and the two video frames are repeated video frames, and thus one of the two video frames can be deleted and the other video frame can be retained; if the difference between the grayscale images corresponding to the two video frames is large, the two video frames contain different contents and are not repeated video frames, and thus both of the two video frames can be retained.
[0031] As an example, two-by-two adjacent video frames in the second frame sequence are combined to obtain a plurality of first combinations; for each first combination, a difference degree of the first combination is determined based on the difference information between the grayscale images corresponding to the video frames in the first combination; the video frame that is later in time in the first combination and has a difference degree less than a first threshold is deleted from the second frame sequence to obtain the first frame sequence.
[0032] The difference degree of the first combination represents the difference size between the grayscale images corresponding to the video frames in the first combination, which can be represented by the mean value of the absolute difference values of the grayscale values of each pixel. For example, the grayscale images corresponding to the video frames include 64 pixels, and the difference degree of the first combination is the sum of the absolute difference values of the grayscale values of the 64 pixels divided by 64.
[0033] The first threshold can be set according to actual needs, such as 0.84, and the embodiments of the present specification do not limit the same.
[0034] Specifically, from the first video frame in the second frame sequence, the first video frame is combined with the next video frame to obtain a first combination, and a difference degree of the first combination is determined based on difference information between the gray images corresponding to the video frames in the first combination; if the difference degree of the first combination is less than a first threshold, the video frame later in time sequence in the first combination is deleted; if the difference degree of the first combination is greater than or equal to the first threshold, all video frames in the first combination are retained; further, the last video frame retained in the first combination is combined with the next video frame to obtain a new first combination, and the above operation is repeatedly executed until the combination and deduplication of all video frames in the second frame sequence are completed.
[0035] For example, Figure 2 The second frame sequence shown includes video frames 1 to video frames n arranged based on time sequence. Video frame 1 and video frame 2 are combined to obtain a first combination: (video frame 1, video frame 2).
[0036] If the difference degree corresponding to the first combination is less than the first threshold, video frame 2 is deleted, and video frame 1 and video frame 3 are combined to obtain a new first combination (video frame 1, video frame 3). If the difference degree corresponding to the first combination is greater than or equal to the first threshold, video frame 1 and video frame 2 are retained, and video frame 2 and video frame 3 are combined to obtain a new first combination (video frame 2, video frame 3).
[0037] The above operation is repeatedly executed on the new first combination until the combination and deduplication of video frame n are completed.
[0038] The first frame sequence obtained in this way is: {video frame 1, video frame 3, …, video frame n-1}.
[0039] Since the repeated video frames are usually continuous, the above method can quickly filter out the almost identical continuous video frames in the second frame sequence by frame difference method, greatly reducing the number of repeated video frames in the first frame sequence, thereby reducing the subsequent calculation amount, improving the calculation efficiency, and saving the calculation resources.
[0040] The above shows part of the implementation mode of S102. Of course, it should be understood that S102 can also be implemented in other ways, and the embodiments of the present application do not limit this.
[0041] S104, based on the feature vectors of the video frames in the first frame sequence, determining the distance between different video frames in the first frame sequence.
[0042] The feature vector of the video frame can be extracted by a pre-trained model. The feature vector of the video frame is a high-dimensional feature vector, which contains rich semantic information. The distance between different video frames, i.e., the distance between the feature vectors of different video frames, reflects the difference in content between different video frames. For any two video frames in the first frame sequence, the smaller the distance between them, the more similar they are in content.
[0043] In actual applications, the distance between different video frames can be the Euclidean distance, Manhattan distance, cosine similarity, etc., which are not limited by the embodiments of the present application.
[0044] S106, video frames with similar distances in the first frame sequence are divided into the same cluster to obtain a plurality of first clusters.
[0045] Each first cluster contains a plurality of video frames with similar distances, i.e., a plurality of similar video frames.
[0046] In an implementation, S106 includes the following steps: combining every two adjacent video frames in the first frame sequence to obtain a plurality of second combinations; for each second combination, if the distance between the video frames in the second combination is less than a second threshold, the second combination is divided into the same cluster to obtain the division result of the second combination; and based on the division result of each second combination, a plurality of first clusters are obtained.
[0047] The second threshold can be set according to actual needs, which is not limited by the embodiments of the present application.
[0048] Taking the first frame sequence as an example, the distance between each two adjacent video frames is calculated to obtain the following distances: distance (video frame 1, video frame 2), distance (video frame 2, video frame 3), …, distance (video frame n-2, video frame n-1). Figure 2
[0049] For the first second combination (video frame 1, video frame 3), assuming that the distance between video frame 1 and video frame 3 is less than the second threshold, video frame 1 and video frame 3 are divided into the same cluster.
[0050] For the second second combination (video frame 3, video frame 5), assuming that the distance between video frame 3 and video frame 5 is less than the second threshold, video frame 3 and video frame 5 are divided into the same cluster.
[0051] By analogy, the division results of all second combinations are obtained, and then the division results of all second combinations are integrated to obtain a plurality of first clusters.
[0052] In the above implementation manner, the continuous similar video frames can be quickly divided together, and reliable data support is provided for subsequent extraction of video frames containing key information and without repetition.
[0053] In another implementation manner, a clustering algorithm commonly used in the art can be used to cluster the video frames in the first frame sequence to obtain a plurality of first clusters. For example, through a k-means clustering algorithm, each video frame is assigned to the nearest preset initial cluster center based on the distance between different video frames in the first frame sequence to obtain at least one candidate cluster, and then the position of the cluster center is updated according to the distance between the video frames in each candidate cluster until convergence. In this way, the clustering of the video frames in the first frame sequence is completed, and a plurality of first clusters are obtained.
[0054] The above shows part of the implementation manner of S106. Of course, it should be understood that S106 can also be implemented in other manners, which are not limited by the embodiments of the present specification.
[0055] S108, frame extraction is performed on each first cluster based on the distance between the video frames in the first cluster to obtain a target frame sequence of the video to be processed.
[0056] For each first cluster, the most representative video frame can be extracted from the first cluster based on the distance between the video frames in the first cluster, and then the most representative video frames in all clusters are sorted in time sequence, and the target frame sequence of the video to be processed is obtained.
[0057] In an implementation manner, S108 includes the following steps: S1082, for each first cluster, the video frames in the first cluster with a distance less than a third threshold value are divided into the same cluster to obtain a plurality of second clusters in the first cluster, and at least one video frame is extracted from the second cluster containing the largest number of video frames in the first cluster.
[0058] As an example, the division of the video frames in the first cluster can include: combining the video frames in the first cluster that are time-sequentially adjacent to each other to obtain a plurality of third combinations; for each third combination, if the distance between the video frames in the third combination is less than a third threshold value, the third combination is divided into the same cluster to obtain a division result of the third combination; and based on the division result of each third combination in the first cluster, a plurality of second clusters in the first cluster are obtained. The third threshold value can be set according to actual needs, for example, the third threshold value can be greater than the second threshold value described above, which is not limited by the embodiments of the present specification.
[0059] For example, in the case of Figure 2The first cluster 1 shown as an example contains video frame 1, video frame 3, video frame 5, video frame 7, video frame 9 and video frame 11. The video frames in the first cluster are combined in pairs of time-sequentially adjacent video frames to obtain the following third combinations: (video frame 1, video frame 3), (video frame 3, video frame 5), (video frame 5, video frame 7), (video frame 7, video frame 9), (video frame 9, video frame 11). The third video frames in the third combinations are divided based on the distance between the video frames in each third combination to obtain two second clusters, wherein the second cluster 1 contains video frame 1 and video frame 3, and the second cluster 2 contains video frame 5, video frame 7, video frame 9 and video frame 11.
[0060] For each first cluster, the most representative video frame in the first cluster can be extracted in various appropriate manners after the first cluster is divided into multiple second clusters, which are not limited in the embodiments of the present specification.
[0061] As an example, any video frame in the second cluster containing the largest number of video frames is extracted.
[0062] As another example, the video frame in the first position in the second cluster containing the largest number of video frames is extracted.
[0063] For example, as shown in Figure 2 The second cluster 2 contains the largest number of video frames, and the video frame 5 in the second cluster 2 is in the first position, so the video frame 5 in the second cluster 2 is extracted as the most representative video frame in the first cluster 1.
[0064] It can be understood that, since the video frames in the same second cluster are similar in content, if a second cluster contains the largest number of video frames, it means that the content presented by the video frames in the second cluster contains the largest amount of information, is the most important or has the longest dwell time, and is the most important for the entire video to be processed, especially in the document projection scenario, the video to be processed is obtained by recording the process of projecting a document in a meeting, according to the characteristic of the document sliding downward, the content presented by the video frames in such a second cluster is often the key content of the document, and the video frame in the first position in such a second cluster usually contains the most critical information, therefore, by extracting the video frame in the first position in the second cluster containing the largest number of video frames, it can be ensured that the extracted video frame can cover the key information of the video frames in the first cluster, and is the most representative video frame in the first cluster.
[0065] S1084, based on the video frame extracted from each first cluster, a target frame sequence of the video to be processed is obtained.
[0066] The video frames extracted from each first cluster are sorted in time sequence to obtain a target frame sequence of the video to be processed.
[0067] Through the above implementation manner, the video frames with close distances in the first frame sequence are first divided into the same cluster to obtain a plurality of first clusters, each first cluster is further divided to obtain a plurality of second clusters, and at least one video frame is extracted from the second cluster with the largest number of contained video frames, so that the repeated video frames can be removed as much as possible while the video frame with the largest amount of information is retained.
[0068] In another implementation manner, the S108 includes the following steps: for each first cluster, based on the distance between each two video frames in the first cluster, the similar video frames in the first cluster are de-duplicated to obtain a third cluster; for each third cluster, based on the frame information of the video frames in the third cluster, a target video frame in the third cluster is determined; and the target video frame in each third cluster is extracted to obtain a target frame sequence of the video to be processed.
[0069] Specifically, the video frames in the first cluster are combined in pairs to obtain a plurality of fourth combinations; for each fourth combination, if the distance between the video frames in the fourth combination is less than a third threshold value, the video frame with a later time sequence in the fourth combination is deleted from the first cluster to obtain a third cluster. The third threshold value can be set according to actual requirements, and the embodiments of the present specification do not limit this. Exemplarily, the third threshold value is greater than the second threshold value described above, which helps to select the most representative video frame from each first cluster.
[0070] For example, assuming that a first cluster includes a video frame 1, a video frame 3, and a video frame 5, these video frames are combined in pairs to obtain a plurality of fourth combinations as follows: (video frame 1, video frame 3), (video frame 1, video frame 5), (video frame 3, video frame 5). In the first fourth combination, the distance between the video frame 1 and the video frame 3 is less than the third threshold value, so the video frame 3 is deleted from the first cluster. In the second fourth combination, the distance between the video frame 1 and the video frame 5 is greater than the second threshold value, so the video frame 1 and the video frame 5 are retained. In the third fourth combination, the distance between the video frame 3 and the video frame 5 is equal to the third threshold value, so the video frame 5 is retained. The third cluster obtained thereby includes the video frame 1 and the video frame 5.
[0071] Further, considering that generally the longer the duration of a video frame, the more information it contains and the more important it is for the entire video to be processed, especially in the document projection scenario, the video to be processed is a video obtained by recording the process of projecting a document in a meeting, and due to the characteristic of the document sliding downward, the video frame extracted at a fixed position or randomly from the third cluster is often not the video frame with the largest amount of information and the most important, and the longer the duration of the video frame, the longer the duration of the content contained in the video frame in the meeting, which means that the content is more critical. Based on this, for each third cluster, the video frame with a duration greater than or equal to the duration threshold in the third cluster is determined as the target video frame in the third cluster. In this way, the video frame with the largest amount of information can be retained while removing as many duplicate video frames as possible. For example, continuing with the third cluster in the example above, assuming that the duration of video frame 1 is greater than the duration of video frame 5, video frame 1 is determined as the target video frame in the third cluster.
[0072] Finally, the target video frames in all third clusters are sorted in time sequence, and the target frame sequence of the video to be processed is obtained.
[0073] In the above implementation manner, the video frames in the first cluster are first de-duplicated, and then the target video frame containing the key information is extracted, so that the video frame with the largest amount of information can be retained while removing as many duplicate video frames as possible, and fine frame processing of the video to be processed is realized.
[0074] The above shows part of the implementation manner of S108. Of course, it should be understood that S108 can also be implemented in other manners, which are not limited in the embodiments of the present specification.
[0075] In another embodiment, after S108, it further includes: identifying, by an image processing model, content data of each video frame in the target frame sequence; and generating a content summary of the video to be processed based on the time sequence information of the video frames in the target frame sequence and the content data of each video frame.
[0076] The image processing model can be any model with image processing capability, such as a large model, a traditional deep learning model, etc. The image processing model identifies each video frame in the target frame sequence to output the content data of each video frame, and then splices the content data of each video frame in time sequence, and takes the splicing result as the content summary of the video to be processed, or further processes the splicing result (such as de-duplication, error correction, summary, etc.) to obtain the content summary of the video to be processed.
[0077] The content summary of the video to be processed describes the main content of the video to be processed, and can be applied to various business scenarios such as video classification, information retrieval, knowledge question and answer, etc., which are not limited in the embodiments of the present specification.
[0078] For example, in a video classification scenario, a multimedia platform classifies a to-be-processed video based on a content summary of the to-be-processed video, and determines a category of the to-be-processed video. In an information retrieval scenario, an information platform adds the to-be-processed video and a content summary thereof to a video library, so as to search for a video containing specific content based on content summaries of videos in the video library. In a knowledge question and answer scenario, a question and answer system takes a content summary of a to-be-processed video as a corpus of knowledge question and answer, and uses the content summary to quickly retrieve related knowledge in a knowledge question and answer process.
[0079] The video frame extraction method provided in the embodiments of the present specification first extracts a plurality of video frames arranged based on time sequence from a to-be-processed video as a candidate first frame sequence. Then, considering that a feature vector of a video frame reflects the semantics of the video, the distance between the video frames determined based on this can reflect the similarity between the video frames. On this basis, the video frames with adjacent distances in the first frame sequence are divided into the same cluster to obtain a plurality of first clusters, so that the video frames in the same first cluster are similar. Further, based on the distance between the video frames in each first cluster, a part of the video frames in the first cluster are extracted to form a target frame sequence, which can realize fast deduplication of the first cluster, reduce the repeated video frames in the finally obtained target frame sequence, avoid the target frame sequence being too long to affect the processing of downstream tasks, and can ensure that the video frames in the target frame sequence contain key information, thereby providing reliable data support for improving the accuracy of the processing of downstream tasks.
[0080] In addition, corresponding to the video frame extraction method shown in the above Figure 1 The present specification also provides a video frame extraction device. Figure 3 FIG. 3 is a structural schematic diagram of a video frame extraction device 300 provided by the present specification, which comprises a first extraction module 310, a determination module 320, a division module 330, and a second extraction module 340.
[0081] The first extraction module 310 is configured to extract a first frame sequence from a to-be-processed video, wherein the first frame sequence comprises a plurality of video frames arranged based on time sequence.
[0082] The determination module 320 is configured to determine the distance between different video frames in the first frame sequence based on the feature vectors of the video frames in the first frame sequence.
[0083] The division module 330 is configured to divide the video frames with adjacent distances in the first frame sequence into the same cluster to obtain a plurality of first clusters.
[0084] The second extraction module 340 is configured to extract frames from each first cluster based on the distance between the video frames in each first cluster to obtain a target frame sequence of the to-be-processed video.
[0085] The video frame extraction device provided by the embodiment of the present specification first extracts a plurality of video frames arranged based on time sequence from a to-be-processed video as a candidate first frame sequence; then, considering that the feature vectors of the video frames reflect the semantics of the video, the distance between the video frames determined therefrom can reflect the similarity between the video frames; on this basis, the video frames adjacent in distance in the first frame sequence are divided into the same cluster to obtain a plurality of first clusters, so that the video frames in the same first cluster are similar; further, based on the distance between the video frames in each first cluster, a part of the video frames in the first cluster are extracted to form a target frame sequence, which can realize fast deduplication of the first cluster, reduce the repeated video frames in the finally obtained target frame sequence, avoid the target frame sequence being too long to affect the downstream task processing, and can ensure that the video frames in the target frame sequence contain key information, thereby providing reliable data support for improving the accuracy of downstream task processing.
[0086] In another embodiment, the first extraction module comprises: a sampling sub-module configured to sample the to-be-processed video according to a preset interval duration to obtain a second frame sequence; a conversion sub-module configured to convert the video frames in the second frame sequence into grayscale images; a first deduplication sub-module configured to perform deduplication processing on the second frame sequence based on the difference information between the grayscale images corresponding to different video frames in the second frame sequence to obtain the first frame sequence.
[0087] In another embodiment, the first deduplication sub-module is configured to: combine the video frames adjacent in time sequence in pairs in the second frame sequence to obtain a plurality of first combinations; for each first combination, determine a difference degree of the first combination based on the difference information between the grayscale images corresponding to the video frames in the first combination; delete the video frame later in time sequence in the first combination whose difference degree is less than a first threshold from the second frame sequence to obtain the first frame sequence.
[0088] In another embodiment, the division module comprises: a first combination sub-module configured to combine the video frames adjacent in time sequence in pairs in the first frame sequence to obtain a plurality of second combinations; a first division sub-module configured to, for each second combination, if the distance between the video frames in the second combination is less than a second threshold, divide the second combination into the same cluster to obtain a division result of the second combination; a first determination sub-module configured to obtain a plurality of first clusters based on the division result of each second combination.
[0089] In another embodiment, the second extraction module comprises: The second partitioning submodule is used to partition video frames within the first cluster whose distance is less than the third threshold into the same cluster for each first cluster, thereby obtaining multiple second clusters within the first cluster, and to extract at least one video frame from the second cluster containing the most video frames within the first cluster. The second determining submodule is used to obtain the target frame sequence of the video to be processed based on the video frames extracted from each first cluster.
[0090] In another embodiment, the second partitioning submodule is used for: The video frames that are temporally adjacent to each other in the first cluster are combined to obtain multiple third combinations; For each third group, if the distance between video frames in the third group is less than a third threshold, then the third group is divided into the same cluster to obtain the division result of the third group; Based on the partitioning results of each third combination within the first cluster, multiple second clusters within the first cluster are obtained.
[0091] In another embodiment, the first extraction submodule is used to: Extract the video frame that is first in time from the second cluster, which contains the most video frames.
[0092] In another embodiment, the video to be processed is obtained by recording the document projection process during a meeting.
[0093] In another embodiment, the video frame extraction device further includes: The recognition module is used to identify the content data of each video frame in the target frame sequence through an image processing model; The generation module is used to generate a summary of the video to be processed based on the timing information of the video frames in the target frame sequence and the content data of each video frame.
[0094] Obviously, the video frame extraction device in the embodiments of this specification can be used as described above. Figure 1 The execution body of the video frame extraction method shown is therefore capable of implementing the video frame extraction method in... Figure 1 The functions implemented are the same, so they will not be described in detail here.
[0095] Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this specification. Please refer to it. Figure 4At the hardware level, the electronic device includes a processor, and optionally further includes an internal bus, a network interface, and a memory. The memory can include a memory such as a random-access memory (RAM), and can further include a non-volatile memory such as at least one disk memory. Of course, the electronic device can further include other hardware required by a business.
[0096] The processor, the network interface, and the memory can be connected to each other through the internal bus, which can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, and a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0097] The memory is used to store a program. Specifically, the program can include program code including computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0098] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs, and forms a video frame extraction device at the logical level. The processor executes the program stored in the memory, and is specifically used to perform the following operations: extract a first frame sequence from the video to be processed, the first frame sequence including a plurality of video frames arranged based on time sequence; determine distances between different video frames in the first frame sequence based on feature vectors of the video frames in the first frame sequence; divide video frames with similar distances in the first frame sequence into the same cluster to obtain a plurality of first clusters; extract frames from each first cluster based on distances between video frames in each first cluster to obtain a target frame sequence of the video to be processed.
[0099] The above as described in the specification Figure 1The method performed by the video frame extracting device disclosed in the embodiment can be applied in a processor or implemented by the processor. The processor can be an integrated circuit chip with signal processing capability. In the implementation, each step of the above method can be completed by integrated logic circuits of hardware in the processor or instructions in the form of software. The processor can be a general processor, including a central processing unit (CPU), a network processor (NP), etc. It can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiment of the present specification can be implemented or executed. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiment of the present specification can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method.
[0100] It should be understood that the electronic device of the embodiment of the present specification can implement the functions of the video frame extracting device in the embodiment. Figure 1 The same principles apply, and the embodiment of the present specification will not be repeated here.
[0101] Of course, in addition to the software implementation, the electronic device of the present specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0102] The embodiment of the present specification also proposes a computer readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by an electronic device including a plurality of applications, can cause the electronic device to perform the method of the embodiment. Figure 1 The method of the embodiment, and specifically for performing the following operations: extracting a first frame sequence from a video to be processed, the first frame sequence including a plurality of video frames arranged based on time sequence; determine distances between different video frames in the first frame sequence based on the feature vectors of the video frames in the first frame sequence; divide the video frames with similar distances in the first frame sequence into a same cluster to obtain a plurality of first clusters; perform frame extraction on each first cluster based on the distances between the video frames in the first cluster to obtain a target frame sequence of the video to be processed.
[0103] The embodiments of the present specification also provide a computer program product, which comprises a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform part or all of the steps in the video frame extraction method according to the first aspect.
[0104] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0105] In conclusion, the above only describes preferred embodiments of the present specification, and is not intended to limit the protection scope of the present specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the protection scope of the present specification.
[0106] The systems, devices, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may, for example, be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0107] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0108] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not explicitly listed, or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0109] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
Claims
1. A method of video decimation, characterized by, The method comprises the following steps: extracting a first frame sequence from a to-be-processed video, the first frame sequence comprising a plurality of video frames arranged based on time sequence; determining distances between different video frames in the first frame sequence based on feature vectors of the video frames in the first frame sequence; dividing video frames with close distances in the first frame sequence into the same cluster to obtain a plurality of first clusters; extracting frames from each first cluster based on distances between video frames in the first cluster to obtain a target frame sequence of the to-be-processed video.
2. The method of claim 1, wherein, The method comprises the following steps: sampling the to-be-processed video according to a preset interval duration to obtain a second frame sequence; converting video frames in the second frame sequence into grayscale images; performing deduplication processing on the second frame sequence based on difference information between grayscale images corresponding to different video frames in the second frame sequence to obtain the first frame sequence.
3. The method of claim 2, wherein, The method comprises the following steps: combining two-by-two time-sequentially adjacent video frames in the second frame sequence to obtain a plurality of first combinations; for each first combination, determining a difference degree of the first combination based on difference information between grayscale images corresponding to video frames in the first combination; deleting a time-sequentially later video frame in the first combination with a difference degree less than a first threshold from the second frame sequence to obtain the first frame sequence.
4. The method of claim 1, wherein, The method comprises the following steps: combining two-by-two time-sequentially adjacent video frames in the first frame sequence to obtain a plurality of second combinations; for each second combination, if distances between video frames in the second combination are less than a second threshold, dividing the second combination into the same cluster to obtain a division result of the second combination; obtaining a plurality of first clusters based on the division result of each second combination.
5. The method of claim 1, wherein, The method comprises the following steps: for each first cluster, dividing video frames with distances less than a third threshold in the first cluster into the same cluster to obtain a plurality of second clusters in the first cluster, and extracting at least one video frame from a second cluster with the largest number of contained video frames in the first cluster; obtaining a target frame sequence of the to-be-processed video based on the extracted video frames from each first cluster.
6. The method of claim 5, wherein, The method comprises the following steps: combining two-by-two time-sequentially adjacent video frames in the first cluster to obtain a plurality of third combinations; for each third combination, if distances between video frames in the third combination are less than the third threshold, dividing the third combination into the same cluster to obtain a division result of the third combination; obtaining a plurality of second clusters in the first cluster based on the division result of each third combination in the first cluster.
7. The method of claim 5, wherein, The method comprises the following steps: extract a video frame with the first time sequence from the second cluster with the largest number of included video frames.
8. The method of claim 1, wherein, The to-be-processed video is obtained by recording a document screen projection process in a conference.
9. The method of claim 1, wherein, After the target frame sequence of the to-be-processed video is obtained by frame extraction on each first cluster based on the distance between video frames in each first cluster, the method further includes: content data of each video frame in the target frame sequence is identified through an image processing model; a content summary of the to-be-processed video is generated based on the time sequence information of the video frames in the target frame sequence and the content data of each video frame.
10. A video decimation device, characterized by, The method includes: a first extraction module configured to extract a first frame sequence from a to-be-processed video, the first frame sequence including a plurality of video frames arranged based on time sequence; a determination module configured to determine the distance between different video frames in the first frame sequence based on feature vectors of the video frames in the first frame sequence; a division module configured to divide video frames with similar distances in the first frame sequence into the same cluster to obtain a plurality of first clusters; a second extraction module configured to perform frame extraction on each first cluster based on the distance between video frames in each first cluster to obtain a target frame sequence of the to-be-processed video.
11. An electronic device, comprising: The method includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the video frame extraction method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can perform the video frame extraction method according to any one of claims 1 to 9.
13. A computer program product, characterised in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform some or all of the steps of the video frame extraction method according to any one of claims 1 to 9.