Video frame selection method and device, electronic equipment, storage medium and program product
By using information entropy and edge feature values to select frames, the method optimizes both quality and efficiency in video frame extraction, addressing the balance issue in existing technologies.
Patent Information
- Application Number
- CN202510572009.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-15
AI Technical Summary
The existing video frame selection technology is difficult to balance the quality of frame selection and processing efficiency. Isometric frame extraction misses important information, keyframe extraction has strong dependence on the encoding format, model-based frame selection technology has high computational complexity and is affected in real time.
By extracting frames on the video, the information entropy and edge information feature values of the candidate video frames are determined, and the frame selection quality and efficiency of the frame selection are optimized based on the key.
The quality of the selected video frames is improved, the calculation amount is reduced, and the real-time frame selection requirements are met, achieving a balance of quality and efficiency.
Smart Images

Figure CN120321431A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, a storage medium, and a program product for video frame selection. Background Art
[0002] Video frame selection generally refers to extracting video frames from a continuous video stream to reduce the data volume without losing important information and improve the processing efficiency of downstream processes. In most current video application scenarios, to reduce the data volume involved in video processing, the data is usually first converted from the video modality to the image modality through video frame selection techniques. However, it is difficult for existing video frame selection techniques to achieve a balance between the quality of frame selection and the processing efficiency of frame selection. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a method, an apparatus, an electronic device, a storage medium, and a program product for video frame selection, which can solve or partially solve the above problems to a certain extent.
[0004] In some embodiments of the present disclosure, the method for video frame selection according to the embodiments of the present disclosure may include: extracting frames from the video to obtain at least one candidate video frame; respectively determining the information entropy of each candidate video frame; respectively determining the edge information eigenvalue of each candidate video frame; grouping the at least one candidate video frame to determine a plurality of video frame selection windows; respectively using each video frame selection window as a target video frame selection window, and performing the following operations on the candidate video frames within the target video frame selection window: determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame; and determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame; and determining the frame selection result of the video based on the target video frame corresponding to the target video frame selection window.
[0005] In some embodiments of the present disclosure, extracting frames from the video to obtain at least one candidate video frame includes: determining the number of candidate video frames based on a preset frame selection number index and a preset candidate frame number index; and performing equidistant frame extraction on the video based on the number of candidate video frames to obtain the at least one candidate video frame.
[0006] In some embodiments of the present disclosure, extracting frames from the video to obtain at least one candidate video frame includes: dividing the video into at least one video segment based on key frames in the video; wherein each of the video segments includes one of the key frames; determining the number of candidate video frames corresponding to each of the video segments respectively based on the number of the video segments, a preset frame selection index, and a preset candidate frame number index; respectively for each of the video segments, performing equidistant frame extraction on the video segment based on the number of candidate video frames corresponding to the video segment to obtain alternative video frames corresponding to the video segment; wherein the alternative video frames corresponding to the video segment include the key frame included in the video segment; and using the alternative video frames corresponding to at least one of the video segments as the at least one candidate video frame.
[0007] In some embodiments of the present disclosure, extracting frames from the video to obtain at least one candidate video frame includes: dividing the video into at least one video shot through shot segmentation; determining the number of candidate video frames corresponding to each of the video shots respectively based on the number of the video shots, a preset frame selection index, and a preset candidate frame number index; respectively for each of the video shots, performing equidistant frame extraction on the video shot based on the number of candidate video frames corresponding to the video shot to obtain alternative video frames corresponding to the video shot; and using the alternative video frames corresponding to at least one of the video shots as the at least one candidate video frame.
[0008] In some embodiments of the present disclosure, respectively determining the information entropy of each of the candidate video frames includes: for each of the candidate video frames, performing the following steps: determining one or a combined information of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame; and using one or a combined information of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame as the information entropy of the candidate video frame.
[0009] In some embodiments of the present disclosure, respectively determining the edge information eigenvalue of each of the candidate video frames includes: for each of the candidate video frames, performing the following steps: determining one or a combined information of the variance of the Laplacian operator of the candidate video frame and the edge ratio of the candidate video frame; and using one or a combined information of the variance of the Laplacian operator of the candidate video frame and the edge ratio of the candidate video frame as the edge information eigenvalue of the candidate video frame.
[0010] In some embodiments of the present disclosure, grouping the at least one candidate video frame and determining multiple video frame selection windows includes: dividing the at least one candidate video frame into multiple candidate video frame groups based on the candidate frame number metric; wherein the number of candidate video frames included in each candidate video frame group is less than or equal to the candidate frame number metric; and respectively taking each candidate video frame group as one of the video frame selection windows.
[0011] In some embodiments of the present disclosure, determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame includes: performing the following steps for each candidate video frame within the target video frame selection window: determining the maximum value and the minimum value of the information entropy of the candidate video frame; determining the normalized information entropy of the candidate video frame based on the maximum value and the minimum value of the information entropy of the candidate video frame; determining the maximum value and the minimum value of the edge information eigenvalue of the candidate video frame; determining the normalized edge information eigenvalue of the candidate video frame based on the maximum value and the minimum value of the edge information eigenvalue of the candidate video frame; and performing a weighted sum of the normalized information entropy of the candidate video frame and the normalized edge information eigenvalue of the candidate video frame to obtain the key degree of the candidate video frame.
[0012] In some embodiments of the present disclosure, determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame includes: taking the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window.
[0013] In some embodiments of the present disclosure, determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frames includes: determining the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window; in response to determining that the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window is the first candidate video frame within the target video frame selection window and the target video frame corresponding to the previous video frame selection window of the target video frame selection window is the last candidate video frame within the previous video frame selection window, using the candidate video frame corresponding to the second largest value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window; in response to determining that the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window is not the first candidate video frame within the target video frame selection window or the target video frame corresponding to the previous video frame selection window of the target video frame selection window is not the last candidate video frame within the previous video frame selection window, using the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window.
[0014] In some embodiments of the present disclosure, determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frames includes: determining the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window; determining the similarity between the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window and the target video frame corresponding to the previous video frame selection window of the target video frame selection window; in response to determining that the similarity is greater than a preset similarity threshold, using the candidate video frame corresponding to the second largest value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window; in response to determining that the similarity is less than or equal to the similarity threshold, using the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window.
[0015] In some embodiments of the present disclosure, after obtaining the at least one candidate video frame, the above method further includes: scaling each of the candidate video frames to a preset image scale.
[0016] Corresponding to the above video frame selection method, an embodiment of the present disclosure also discloses a video frame selection device, including:
[0017] A frame extraction module, configured to extract frames from the video to obtain at least one candidate video frame;
[0018] An information entropy determination module, configured to determine the information entropy of each of the candidate video frames respectively;
[0019] An edge feature determination module, configured to respectively determine the edge information feature values of each of the candidate video frames;
[0020] A grouping module, configured to group the at least one candidate video frame to determine a plurality of video selection frame windows;
[0021] A frame selection module, configured to respectively use each of the video selection frame windows as a target video selection frame window, and perform the following operations on the candidate video frames within the target video selection frame window: determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information feature value of the candidate video frame; and determining the target video frame corresponding to the target video selection frame window based on the key degree of the candidate video frame; and
[0022] An output module, determining the frame selection result of the video based on the target video frame corresponding to the target video selection frame window.
[0023] In addition, an embodiment of the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the above-mentioned frame selection method for a video is implemented.
[0024] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the above-mentioned frame selection method for a video.
[0025] An embodiment of the present disclosure further provides a computer program product, including computer program instructions, where when the computer program instructions run on a computer, the computer is caused to execute the above-mentioned frame selection method for a video.
[0026] It can be seen from this that in the above-mentioned frame selection method, device, electronic device, storage medium, and program product for a video, the information entropy and edge features of the video frames included in a video can be comprehensively considered to select target video frames with rich information content and rich edge features from the candidate video frames. Therefore, the quality of the selected video frames can be greatly improved. On the other hand, the above-mentioned frame selection scheme for a video is relatively simple to implement and consumes little computing power, and can meet the requirements of real-time frame selection. That is to say, the frame selection method, device, electronic device, storage medium, and program product for a video described in the embodiments of the present disclosure can optimize the quality of frame selection and the processing efficiency of frame selection, so as to achieve a balance between the quality of frame selection and the processing efficiency of frame selection. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only the embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 provided by an embodiment of the present disclosure.
[0029] Figure 2 FIG. shows the implementation process of a method for selecting video frames according to some embodiments of the present disclosure.
[0030] Figure 3 FIG. shows an example of a candidate video frame according to some embodiments of the present disclosure.
[0031] Figure 4 FIG. shows the implementation process of a specific method for determining the key degree of a candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame according to some embodiments of the present disclosure.
[0032] Figure 5A FIG. shows an example of selecting a target video frame from candidate video frames according to some embodiments of the present disclosure.
[0033] Figure 5B FIG. shows an example of selecting a target video frame from candidate video frames according to some other embodiments of the present disclosure.
[0034] Figure 6 FIG. shows the internal structure of a video frame selection device according to some embodiments of the present disclosure.
[0035] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device according to some embodiments of the present disclosure. Detailed implementation manners
[0036] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further elaborates on the present disclosure in detail in combination with specific embodiments and with reference to the drawings.
[0037] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprise" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connect" or "be connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0038] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0039] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0040] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0041] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that comply with relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0042] As mentioned above, frame selection of a video refers to extracting video frames from a continuous video stream to reduce the amount of data without losing important information. Through high-quality frame extraction technology, users can quickly obtain the essence of the video content and improve the processing efficiency downstream.
[0043] Currently, common video frame selection methods can include equidistant frame extraction, key frame extraction, model-based frame selection, and so on. Among them, equidistant frame extraction is a simple method that reduces the data volume by uniformly selecting video frames in the video. The advantage of equidistant frame extraction is its simplicity in implementation and low computational overhead, making it suitable for real-time applications. However, this method usually misses important information, resulting in the loss of key details. In addition, in some scenarios of film and television works and video editing user-generated content (UGC), there are a large number of video transitions or fade-in / fade-out shots, etc. Equidistant frame extraction is likely to select video frames of these transitions or fade-in / fade-out frames, thus affecting the quality of video frame extraction. Key frame extraction relies on predefined key frames in the video, such as I frames in the video. This method can retain important scene changes, ensure the integrity of information, be efficient and provide good picture quality. However, this method is highly dependent on the video coding format, lacks flexibility in processing, and is difficult to meet the requirements of frame selection applications with fixed frame number requirements. Model-based frame selection technology uses machine learning or deep learning models to analyze video content and dynamically select the most informative frames. The advantage of model-based frame selection technology is that it can accurately capture important scenes and has strong adaptability. However, this method usually has a high computational complexity, requires a large amount of labeled data for training, and its real-time performance may be affected. In addition, the computing power consumption of this method is usually proportional to the number of video frames to be processed. For example, for a frame selection scenario that requires 8 video frames, if each video frame needs to select the best from 10 similar video frames, 80 video frames need to be input into the model, resulting in huge computing power consumption. It can be seen that it is difficult for existing various video frame selection technologies to achieve a balance between the quality of frame selection and the processing efficiency of frame selection.
[0044] In view of this, embodiments of the present disclosure provide a video frame selection method, apparatus, electronic device, storage medium, and program product, which can solve or partially solve the above problems to a certain extent.
[0045] For the sake of clarity in description, before describing the specific technical solutions of the embodiments of the present disclosure, the technical terms involved in the embodiments of the present disclosure are first explained.
[0046] The shot boundary detection algorithm is one of the key technologies in video processing and can be used to identify the switching points between different shots in the video (such as hard cuts, fades, wipes, etc.). The core principle of the shot boundary detection algorithm is to achieve segmentation by analyzing the visual feature differences between consecutive video frames. Currently, common shot boundary detection algorithms can include: pixel / histogram-based threshold methods, edge / feature-based matching methods, and machine learning-based adaptive detection methods, etc.
[0047] The information entropy of an image can be an index to measure the information content and complexity of the image, which comes from information theory. The information entropy of an image can reflect the degree of uniformity of the pixel value distribution in the image. Generally, the higher the information entropy value of an image, the greater the amount of information in the image, the more complex the image, and the richer the details; while a lower information entropy value of the image means that the image content is relatively single or simple. The information entropy of an image has important applications in image processing, compression and analysis, and can help identify the diversity and changes of image content.
[0048] The Laplacian operator can be a second-order differential operator in an n-dimensional Euclidean space, and its definition can be the divergence of the gradient. Geometrically, the Laplacian operator reflects the "curvature" of the surrounding area of a certain point. Among them, positive curvature represents a local maximum (such as a wave peak), and negative curvature represents a local minimum (such as a wave trough). The Laplacian operator can be used to detect the gray mutation regions in an image, and its output is a second-order derivative response matrix with the same size as the input image. In this second-order derivative response matrix, the value corresponding to each pixel point of the input image represents the second-order derivative value of the gray level of this pixel point.
[0049] Canny edge detection is a classic algorithm proposed by John Canny in 1986, and can achieve high-precision and low-noise interference edge extraction through multi-step processing. The core steps of Canny edge detection can include: convolving the original image with a Gaussian kernel to smooth the noise and reduce the interference of isolated noise points; performing gradient calculation and direction determination; comparing the gradient magnitudes of adjacent pixels along the gradient direction, and only retaining the local maximum to refine the edge width; obtaining edge pixels through double-threshold hysteresis processing; finally, connecting the broken edges into a continuous contour through edge tracking and connection to obtain a closed edge. The output result of Canny edge detection is usually the Canny edge map corresponding to the detected image, which is specifically manifested as a binary image with the same size as the input image (usually a black and white image, where white pixels represent the detected edges and black pixels represent the background).
[0050] Variance can be a quantity to measure the magnitude of data fluctuation, which represents the average of the squares of the differences between each data and the average.
[0051] Candidate video frames can be multiple video frames preselected from all the video frames included in the video, and are used as the range of video frame selection in the method described in the embodiments of the present disclosure. It can be understood that using candidate video frames instead of all the video frames included in the video as the range of video frame selection can reduce the amount of data involved in video frame selection without affecting the frame selection quality, thereby reducing the computational amount of video frame selection and improving the efficiency of video frame selection.
[0052] The target video frame can be a video frame selected from the candidate video frames of the video as the output by the video frame selection method according to the embodiments of the present disclosure.
[0053] Figure 1 FIG. 4 shows a schematic diagram of an exemplary system 100 provided by the embodiments of the present disclosure.
[0054] As Figure 1 shown, the system 100 may include a terminal device 102, a terminal device 104, and a server 106. A medium (e.g., a network) for providing a communication link may be included between the terminal device 102, the terminal device 104, and the server 106. The above network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0055] Exemplarily, an application program (APP) or software that can implement video frame selection may be installed on the terminal device 102 and the terminal device 104. Here, the terminal device 102 and the terminal device 104 may be hardware or software. When the terminal device 102 and the terminal device 104 are hardware, they may be various electronic devices with a display screen, including but not limited to smart phones, tablet computers, laptop portable computers (Laptops), and desktop computers (PCs), etc. When the terminal device 102 and the terminal device 104 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (e.g., for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0056] The server 106 may be a server that provides video frame selection services, such as a background server that supports the application program or software displayed on the terminal device 102 and the terminal device 104. Here, the server 106 may also be hardware or software. When the server 106 is hardware, it may be implemented as a distributed server cluster composed of multiple servers, or may be implemented as a single server. When the server 106 is software, it may be implemented as multiple software or software modules (e.g., for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0057] It should be understood that Figure 1 the numbers of terminal devices, users, and servers in
[0058] As an exemplary scenario, server 106 can provide video frame selection service. User 112 can use the video frame selection application on terminal device 102 to send a video to server 106. Server 106 can select frames from the video uploaded by user 112 through terminal device 102 to obtain multiple target video frames in the video. Then, server 106 can feedback the obtained multiple target video frames to user 112 through terminal device 102.
[0059] In another exemplary scenario, the video frame selection application downloaded and installed by the above terminal device 104 from server 106 can support video frame selection in an offline manner. In this case, user 114 can directly use the video frame selection application on terminal device 104 to select multiple target video frames from the submitted video. That is to say, the specific process of the above video frame selection can be independently completed offline by terminal device 104 without the real-time participation of server 106.
[0060] Based on the above system 100, in order to solve various problems existing in video frame selection in the related art, embodiments of the present disclosure provide a method for selecting frames of a video. The method for selecting frames of a video provided by embodiments of the present disclosure will be described below in conjunction with specific embodiments and the accompanying drawings.
[0061] Figure 2 Shows the implementation process of the method for selecting frames of a video described in embodiments of the present disclosure. As Figure 2 shown, the method for selecting frames of a video described in embodiments of the present disclosure may specifically include the following multiple steps.
[0062] In step 210, extract frames from the video to obtain at least one candidate video frame.
[0063] In step 220, determine the information entropy of each candidate video frame respectively.
[0064] In step 230, determine the edge information feature value of each candidate video frame respectively.
[0065] In step 240, group at least one candidate video frame to determine multiple video frame selection windows.
[0066] In step 250, take each video frame selection window as the target video frame selection window respectively, and perform the following operations on the candidate video frames within the target video frame selection window: determine the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information feature value of the candidate video frame; and determine the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame.
[0067] In step 260, determine the frame selection result of the video based on the target video frame corresponding to the target video frame selection window.
[0068] As can be seen, in the above video frame selection method, the information entropy and edge features of the video frames included in a video can be comprehensively considered to select target video frames with rich information and rich edge features from the candidate video frames. Therefore, the quality of the selected video frames can be greatly improved. On the other hand, the implementation of the above video frame selection method is relatively simple and consumes little computing power, which can meet the requirements of real-time frame selection. That is to say, the video frame selection method described in the embodiments of the present disclosure can optimize the quality of frame selection and the processing efficiency of frame selection, so as to achieve a balance between the quality of frame selection and the processing efficiency of frame selection.
[0069] The following will specifically describe the specific implementation manners of the above video frame selection method in combination with the accompanying drawings and specific examples.
[0070] For the above step 210, the embodiments of the present disclosure can specifically adopt a variety of frame extraction methods to determine candidate video frames.
[0071] In some embodiments of the present disclosure, the above step 210 may include: First, based on a preset frame selection number index and a preset candidate frame number index, determine the number of candidate video frames; then, perform equidistant frame extraction on the video based on the number of candidate video frames to obtain at least one candidate video frame.
[0072] In the above embodiments, the above-mentioned selected frame number index may be the number of target video frames to be extracted from the video. That is to say, the above-mentioned selected frame number index determines the number of selected video frames (i.e., target video frames) finally output by the frame selection method of the above-mentioned video. The above-mentioned candidate frame number index may be defined as how many candidate video frames are required to select each target video frame. In some embodiments, the above-mentioned candidate frame number index also corresponds to the size of the video frame selection window. Generally, a single frame selection operation can be performed within a video frame selection window. For example, assuming that the candidate frame number index is preset to 10, the frame selection method of the video described in the embodiments of the present disclosure will usually divide every 10 candidate video frames into a video frame selection window (each video frame selection window includes 10 candidate video frames), and select one video frame as the target video frame within each video frame selection window. Based on the above content, in the embodiments of the present disclosure, determining the number of candidate video frames based on the preset selected frame number index and the preset candidate frame number index may specifically include: calculating the product of the selected frame number index and the candidate frame number index, and using the above product as the number of candidate video frames. For example, assuming that the selected frame number index is preset to 10 and the candidate frame number index is also preset to 10, the number of candidate video frames determined by the above method may be 10×10 = 100. Thus, after determining the number of candidate video frames, equal-interval frame extraction can be further performed according to the number of all video frames included in the video, so as to obtain candidate video frames matching the above-mentioned number of candidate video frames. For example, if the number of candidate video frames determined by the above method is 100 and the number of all video frames included in the video is 300, then the frame extraction interval of equal-interval frame extraction can be determined to be 3 (300÷100 = 3). Then, by means of equal-interval frame extraction of extracting one video frame every three video frames, 100 candidate video frames can be obtained. It can be understood that the above-mentioned equal-interval frame extraction method can adapt to any selected frame number index and candidate frame number index, and has the advantages of high flexibility, simple implementation and small computing power consumption.
[0073] In addition to the above-mentioned method of evenly sampling frames at equal intervals, in some other embodiments of the present disclosure, when the video includes key frames, the above-mentioned candidate video frames may also be determined in combination with the key frame sampling method. Specifically, in these embodiments, step 210 may include: dividing the video into at least one video segment based on the key frames in the video; wherein each video segment includes one key frame; determining the number of candidate video frames corresponding to each video segment respectively based on the number of video segments, a preset frame selection number index, and a preset candidate frame number index; respectively for each video segment, performing equal-interval frame sampling on the video segment based on the number of candidate video frames corresponding to the video segment to obtain the alternative video frames corresponding to the video segment, wherein the alternative video frames corresponding to the video segment include the key frames included in the video segment; finally, using the alternative video frames corresponding to at least one video segment as the above-mentioned at least one candidate video frame.
[0074] In the above embodiments, the method of dividing a video into at least one video segment based on the key frames in the video may specifically be to use each key frame in the video as the first video frame of a video segment, so as to divide the video into at least one video segment and ensure that each video segment includes one key frame. The method of respectively determining the number of candidate video frames corresponding to each video segment based on the number of video segments, a preset selected frame number index, and a preset candidate frame number index may include: First, calculate the product of the selected frame number index and the candidate frame number index, and use the above product as the number of candidate video frames; then calculate the quotient of the number of candidate video frames and the number of video segments, and use the above quotient as the number of candidate video frames corresponding to each video segment. Based on the previous example, assuming that the selected frame number index is preset to 10 and the candidate frame number index is also preset to 10, the number of candidate video frames determined by the above method can be 100. Then, assuming that the number of video segments is 10, the number of candidate video frames corresponding to each video segment can be 100÷10 = 10. Next, in each video segment, equidistant frame extraction can be performed according to the number of video frames included in each video segment and the number of candidate video frames corresponding to each video segment, so as to obtain the alternative video frames corresponding to each video segment. For example, for a video segment including 40 video frames, it can be determined that the frame extraction interval for equidistant frame extraction is 4 (40÷10 = 4). Then, through the equidistant frame extraction method of extracting one video frame every four video frames, 10 alternative video frames can be obtained. Finally, the alternative video frames corresponding to each video segment are combined together as all the candidate video frames. It can be understood that the above method of determining candidate video frames can also adapt to any selected frame number index and candidate frame number index, and also has the advantages of high flexibility, simple implementation, and low computing power consumption. Further, the candidate video frames determined by the above method may include all the key frames of the video. In addition, it can be understood that this method is generally applicable to the case where the video includes key frames and the number of key frames in the video is less than or equal to the calculated number of candidate video frames. For the case where the number of key frames in the video is greater than the calculated number of candidate video frames, it may not be applicable.
[0075] In still other embodiments of the present disclosure, when the video includes video scenes, the above-mentioned candidate video frames can also be determined in the following manner. Specifically, in these embodiments, step 210 may include: dividing the video into at least one video scene through shot segmentation; determining the number of candidate video frames corresponding to each video scene respectively based on the number of video scenes, a preset frame selection index, and a preset candidate frame index; respectively for each video scene, performing equidistant frame extraction on the video scene based on the number of candidate video frames corresponding to the video scene to obtain alternative video frames corresponding to the video scene; using the alternative video frames corresponding to at least one video scene as the above-mentioned at least one candidate video frame.
[0076] In the above embodiments, the method of dividing a video into at least one video scene by shot segmentation can be implemented by any existing shot segmentation method. Specifically, for example, a pixel-based threshold method, a histogram-based threshold method, an edge detection-based method, a feature matching-based method, an adaptive detection method based on machine learning, etc. can be used. The embodiments of the present disclosure do not limit the shot segmentation method adopted. The method of respectively determining the number of candidate video frames corresponding to each video scene based on the number of video scenes, a preset selected frame number index, and a preset candidate frame number index may include: First, calculate the product of the selected frame number index and the candidate frame number index, and use the above product as the number of candidate video frames; then calculate the quotient of the number of candidate video frames and the number of video scenes, and use the above quotient as the number of candidate video frames corresponding to each video scene. Based on the previous example, assuming that the selected frame number index is preset to 10 and the candidate frame number index is also preset to 10, the number of candidate video frames determined by the above method can be 100. Then assume that the number of video scenes is 2, then the number of candidate video frames corresponding to each video scene can be 100÷2 = 50. Next, in each video scene, equidistant frame extraction can be performed according to the number of video frames included in each video scene and the number of candidate video frames corresponding to each video scene, so as to obtain the alternative video frames corresponding to each video scene. For example, for a video scene including 150 video frames, it can be determined that the frame extraction interval for equidistant frame extraction is 3 (150÷50 = 3), and then, by the equidistant frame extraction method of extracting one video frame every three video frames, 50 alternative video frames can be obtained. Finally, the alternative video frames corresponding to each video scene are combined together as all the candidate video frames. It can be understood that the above method of determining candidate video frames can also adapt to any selected frame number index and candidate frame number index, and also has the advantages of high flexibility, simple implementation, and low computing power consumption. Further, it can be ensured that each scene in the video has basically an equal number of candidate video frames among the candidate video frames determined by the above method. It can also be understood that this method is generally applicable to the case where the video includes scenes and the number of video scenes is less than or equal to the calculated number of candidate video frames. For the case where the number of video scenes is greater than the calculated number of candidate video frames, it may not be applicable.
[0077] For the above step 220, in the embodiments of the present disclosure, the information entropy of the candidate video frames may include: one of the original image information entropy of the candidate video frames and the grayscale image information entropy of the candidate video frames, or a combined information thereof.
[0078] Specifically, in the embodiments of the present disclosure, the original image information entropy of the above candidate video frames can be determined based on the calculation method of image information entropy, and specifically may include the following multiple steps. First, the single-channel histograms of the candidate video frames on each color channel are respectively determined (for example, for each color channel among the R, G, and B channels of a color RGB image). Then, the histograms of the candidate video frames on each color channel are concatenated in the pixel value dimension to obtain the multi-channel histogram corresponding to the candidate video frames. Next, the multi-channel histogram corresponding to the candidate video frames is normalized. Further, based on the normalized multi-channel histogram, the occurrence probabilities of each pixel value are statistically obtained to obtain the probability distribution of each pixel value. Finally, the original image information entropy of the candidate video frames is determined based on the probability distribution of each pixel value. Specifically, the original image information entropy of the candidate video frames can be calculated through the following expression:
[0079] H1 = -∑(p(x)log2(p(x)))
[0080] Where x represents each pixel value on each color channel of the candidate video frame; p(x) represents the probability of each pixel value appearing in the candidate video frame; and H1 represents the original image information entropy of the candidate video frame.
[0081] In a specific example, assuming that the candidate video frame is a color RGB image, the number of pixel points corresponding to each pixel value on the R, G, and B color channels of the candidate video frame can be respectively counted to obtain the single-channel histogram of each channel. For example, for the R channel, the number of pixel points with pixel values of 0, 1, 2, ……, 255 on the R channel of the candidate video frame is counted to obtain the histogram of the candidate video frame corresponding to the R channel. Among them, the horizontal axis of the single-channel histogram is the pixel value of the R channel (the value range is 0 to 255, a total of 256 points); the vertical axis is the number of pixel points. After obtaining the histograms of the R, G, and B channels, the single-channel histograms of the three channels can be concatenated in the pixel value dimension to obtain the multi-channel histogram of the candidate video frame. Among them, the horizontal axis of the multi-channel histogram is the pixel values of the three channels (the value range can be [0,0,0] to [255,255,255], a total of 256*3 points); the vertical axis is the number of pixel points. Next, the above multi-channel histogram is normalized. For example, the number of pixel points counted for each pixel value on the horizontal axis of the multi-channel histogram is divided by the total number of pixel points on the candidate video frame and then divided by 3 to obtain the probability distribution of each pixel value (the sum of the probabilities corresponding to each pixel value is 1). Finally, the original image information entropy of the candidate video frame is determined based on the probability distribution of each pixel value using the above expression.
[0082] In addition, in the embodiments of the present disclosure, the grayscale image information entropy of the above candidate video frames can be determined based on the calculation method of the grayscale image information entropy, which specifically may include the following multiple steps. First, convert the color candidate video frame into a grayscale image (for example, the grayscale value range of each pixel is 0 - 255). Then, determine the histogram of the grayscale image. Next, perform normalization processing on the histogram of the grayscale image. Further, based on the normalized histogram, statistically obtain the occurrence probability of each grayscale value to obtain the probability distribution of each grayscale value. Finally, determine the grayscale image information entropy of the candidate video frame based on the probability distribution of each grayscale value. Among them, the grayscale image information entropy of the candidate video frame can be calculated through the following expression:
[0083] H2 = -∑(p(y)log2(p(y)))
[0084] Where y represents each grayscale value on the grayscale image corresponding to the candidate video frame; p(y) represents the occurrence probability of each grayscale value on the grayscale image corresponding to the candidate video frame; and H2 represents the grayscale image information entropy of the candidate video frame.
[0085] In a specific example, after converting the candidate video frame into a grayscale image, the number of pixel points of the grayscale image at each grayscale value can be statistically obtained to obtain the histogram of the grayscale image. Among them, the horizontal axis of the histogram is the grayscale value (the value range is 0 - 255, a total of 256 points); the vertical axis is the number of pixel points. Next, perform normalization processing on the histogram of the above grayscale image. For example, divide the number of pixel points statistically obtained for each grayscale value on the horizontal axis of the histogram of the grayscale image by the total number of pixel points on the candidate video frame, so as to obtain the probability distribution of each grayscale value (the sum of the probabilities corresponding to each grayscale value is 1). Finally, use the above expression to determine the grayscale image information entropy of the candidate video frame based on the probability distribution of each grayscale value.
[0086] Based on this, determining the information entropy of the candidate video frame in step 220 above may include: using one or a combination of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame as the information entropy of the candidate video frame. In these embodiments, the above combination information may represent two types of information included by using the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame together as the information entropy of the candidate video frame, and may be further fused in subsequent steps.
[0087] In addition, for the above step 230, in the embodiments of the present disclosure, the edge information feature value of the candidate video frame may include: one or a combination of the variance of the Laplacian operator of the candidate video frame and the edge ratio of the candidate video frame.
[0088] As described above, the Laplace operator can be a second-order differential operator in an n-dimensional Euclidean space, and its definition can be the divergence of the gradient. In the field of image edge detection, the Laplace operator of an image can be a second-order derivative response matrix with the same size as the input image, where the pixel value of each pixel represents the second-order gray-scale derivative value of that pixel. Based on this, in the embodiments of the present disclosure, the variance of the Laplace operator of the candidate video frame described in step 230 above can be determined by the following method: First, determine the Laplace operator of the candidate video frame; then further calculate the variance of the Laplace operator of the candidate video frame as the variance of the Laplace operator of the candidate video frame. It can be understood that generally, the larger the variance of the Laplace operator of an image, the more Laplace peaks there are in the image, that is, the more edges and details there are, and the greater the complexity and information content of the image. On the contrary, if the variance of the Laplace operator of an image is small, it means that the image is relatively flat or simple, or the picture is blurred. It should be noted that the specific calculation method of the Laplace operator of the candidate video frame is not limited in the embodiments of the present disclosure.
[0089] In addition, in some embodiments of the present disclosure, the edge ratio of the candidate video frame described in step 230 above can specifically be the ratio of edge pixels in the Canny edge map. As described above, the Canny edge map is an image that shows the edges of an image detected by the Canny edge detection method, and it is specifically a binary image with the same size as the input image. Based on this, the ratio of the pixel points representing edges in the Canny edge map to all pixel points can be used as the ratio of edge pixels in the Canny edge map. For example, assume that the Canny edge map obtained by Canny edge detection is a black-and-white image, where white pixels represent the detected edges and their pixel values are usually 255, and black pixels represent the background and their pixel values are usually 0. Then, in the above steps, the ratio of the part of the Canny edge map with pixel values greater than 0 in the image can be counted to obtain the edge ratio of the candidate video frame. It can be understood that generally, the higher the edge ratio of an image, the more contour content there is in the image, and it also means that the information content of the picture is relatively high. On the contrary, the lower the edge ratio of an image, the less contour content there is in the image, and it also means that the information content of the picture is relatively low.
[0090] Based on this, the determination of the edge information eigenvalue of the candidate video frame in step 230 above can include: using one or a combination of the variance of the Laplace operator of the candidate video frame and the edge ratio of the candidate video frame as the edge information eigenvalue of the candidate video frame. In these embodiments, the above combination information can only represent the two types of information included when using the variance of the Laplace operator of the candidate video frame and the edge ratio of the candidate video frame together as the edge information eigenvalue of the candidate video frame, and can be fused in subsequent steps.
[0091] In particular, in some specific embodiments, through the above steps 220 and 230, the following four metrics corresponding to each candidate video frame can be obtained respectively: the information entropy of the original image, the information entropy of the grayscale image, the variance of the Laplacian operator, and the edge ratio.
[0092] For the above step 240, the following method can be used to group the candidate video frames and determine multiple video frame selection windows: First, divide the at least one candidate video frame into multiple candidate video frame groups based on the candidate frame number metric; wherein, the number of candidate video frames included in each candidate video frame group is equal to or less than the candidate frame number metric; then, each candidate video frame group is respectively used as a video frame selection window. Specifically, in the above method, the candidate video frames corresponding to the number of candidate frame number metrics can be sequentially divided into a candidate video frame group in the order of the candidate video frames, as a video frame selection window. Figure 3 Shows an example of the candidate video frames described in some embodiments of the present disclosure. Figure 3 A total of 18 candidate video frames are shown, where each candidate video frame is represented by a small rectangular box. Assuming that the candidate frame number metric is 6, the candidate video frames can be divided into three candidate video frame groups and three video frame selection windows can be determined by the above method. Among them, each video frame selection window is represented by six small rectangular boxes with a bold border. The goal of frame selection for the video described in the embodiments of the present disclosure is to select one target video frame from each video frame selection window. It should be noted that Figure 3 The number of candidate video frames shown, the number of video selection windows, and the number of candidate video frames within each video selection window are all examples, and the embodiments of the present disclosure do not limit the above numbers. In particular, when determining the video selection window, if the candidate frame number metric cannot divide the number of candidate video frames evenly, N + 1 candidate video frame groups can be obtained; where the number of candidate video frames in N candidate video frame groups is equal to the value of the candidate frame number metric, and N is the quotient obtained by dividing the number of candidate video frames by the candidate frame number metric; and the number of candidate video frames in 1 candidate video frame group is equal to the remainder obtained by dividing the number of candidate video frames by the candidate frame number metric.
[0093] For the above step 250, first, each video frame selection window is sequentially used as the target video frame selection window, and then the following operations are respectively performed on the multiple candidate video frames within the target video frame selection window: determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame; and determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame.
[0094] Specifically, in some embodiments of the present disclosure, the specific method for determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame may refer to Figure 4 , and is used to calculate the key degree of each candidate video frame within the target video frame selection window, and specifically may include the following multiple steps:
[0095] In step 410, determine the maximum value and the minimum value of the information entropy of the candidate video frame.
[0096] Specifically, as described above, the information entropy of the candidate video frame may include: one or a combination of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame. Assuming that the information entropy of the candidate video frame is the original image information entropy of the candidate video frame or the grayscale image information entropy of the candidate video frame, the above steps may determine the maximum value and the minimum value of the original image information entropy of the candidate video frame included in the target video frame selection window or the grayscale image information entropy of the candidate video frame. Assuming that the information entropy of the candidate video frame is the combined information of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame, the above steps may respectively determine the maximum value and the minimum value of the original image information entropy of the candidate video frame included in the target video frame selection window and the maximum value and the minimum value of the grayscale image information entropy of the candidate video frame.
[0097] In step 420, determine the normalized information entropy of the candidate video frame based on the maximum value and the minimum value of the information entropy of the candidate video frame.
[0098] In the embodiments of the present disclosure, for the original image information entropy of the candidate video frame, the normalized original image information entropy of the candidate video frame may be determined based on the maximum value and the minimum value of the original image information entropy of the candidate video frame, and specifically, the normalized original image information entropy may be determined through the following expression:
[0099]
[0100] wherein, H1 represents the original image information entropy of the candidate video frame; H1 min represents the minimum value of the original image information entropy of the candidate video frame; H1 max represents the maximum value of the original image information entropy of the candidate video frame; and H1 norm represents the normalized original image information entropy of the candidate video frame.
[0101] In the embodiments of the present disclosure, for the grayscale image information entropy of the candidate video frame, the normalized grayscale image information entropy of the candidate video frame may be determined based on the maximum value and the minimum value of the grayscale image information entropy of the candidate video frame, and specifically, the normalized grayscale image information entropy may be determined through the following expression:
[0102]
[0103] where H2 represents the grayscale image entropy of the candidate video frame; H2 min represents the minimum value of the grayscale image entropy of the candidate video frame; H2 max represents the maximum value of the grayscale image entropy of the candidate video frame; and H2 norm represents the normalized grayscale image entropy of the candidate video frame.
[0104] In step 430, determine the maximum and minimum values of the edge information eigenvalue of the candidate video frame.
[0105] Specifically, as described above, the edge information eigenvalue of the candidate video frame may include: one or a combination of the variance of the Laplacian operator of the candidate video frame and the edge ratio of the candidate video frame. Assuming that the edge information eigenvalue of the candidate video frame is the variance of the Laplacian operator of the candidate video frame or the edge ratio of the candidate video frame, the above steps can determine the maximum and minimum values of the variance of the Laplacian operator or the edge ratio of the candidate video frames included in the target video frame selection window. Assuming that the edge information eigenvalue of the candidate video frame is the combined information of the variance of the Laplacian operator of the candidate video frame and the edge ratio of the candidate video frame, the above steps can respectively determine the maximum and minimum values of the variance of the Laplacian operator of the candidate video frames included in the target video frame selection window and the maximum and minimum values of the edge ratio of the candidate video frames.
[0106] In step 440, determine the normalized edge information eigenvalue of the candidate video frame based on the maximum and minimum values of the edge information eigenvalue of the candidate video frame.
[0107] In an embodiment of the present disclosure, for the variance of the Laplacian operator of the candidate video frame, the normalized variance of the Laplacian operator of the candidate video frame can be determined based on the maximum and minimum values of the variance of the Laplacian operator of the candidate video frame. Specifically, the normalized variance of the Laplacian operator can be determined through the following expression:
[0108]
[0109] where LD represents the variance of the Laplacian operator of the candidate video frame; LD min represents the minimum value of the variance of the Laplacian operator of the candidate video frame; LD max represents the maximum value of the variance of the Laplacian operator of the candidate video frame; and LD norm represents the normalized variance of the Laplacian operator of the candidate video frame.
[0110] In an embodiment of the present disclosure, for the edge ratio of the candidate video frame, the normalized edge ratio of the candidate video frame can be determined based on the maximum and minimum values of the edge ratio of the candidate video frame. Specifically, the normalized edge ratio can be determined through the following expression:
[0111]
[0112] where C represents the edge ratio of the candidate video frame; C min represents the minimum value of the edge ratio of the candidate video frame; C max represents the maximum value of the edge ratio of the candidate video frame; and H2 norm represents the normalized edge ratio of the candidate video frame.
[0113] In step 450, the normalized information entropy of the candidate video frame and the normalized edge information eigenvalue of the candidate video frame are weighted and summed to obtain the key degree of the candidate video frame.
[0114] In the embodiments of the present disclosure, the normalized information entropy of the candidate video frame and the normalized edge information eigenvalue of the candidate video frame can be weighted and summed according to the following expression to determine the key degree of the candidate video frame:
[0115] K = α1·H1 norm + α2·H2 norm + α3·LD norm + α4·C norm
[0116] where K represents the key degree of the candidate video frame; α1, α2, α3, and α4 respectively represent the weight coefficients corresponding to the four parameters of the above-mentioned normalized original image information entropy, normalized grayscale image information entropy, normalized Laplacian operator variance, and normalized edge ratio. It should be noted that when a certain parameter among the above four parameters is not included in the normalized information entropy of the candidate video frame or the normalized edge information eigenvalue of the candidate video frame, its corresponding weight coefficient can be set to 0, without affecting the accuracy of the key degree described in the embodiments of the present disclosure. In the embodiments of the present disclosure, the weight coefficients corresponding to the four parameters of the above-mentioned normalized original image information entropy, normalized grayscale image information entropy, normalized Laplacian operator variance, and normalized edge ratio can be flexibly set according to the actual situation, and the specific values of the above weight coefficients are not limited in the embodiments of the present disclosure.
[0117] Regarding the above step 260, in some embodiments of the present disclosure, the specific method for determining the frame selection result of the video based on the target video frames corresponding to the target video frame selection window may include: taking the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window. That is to say, in the embodiments of the present disclosure, the candidate video frame corresponding to the maximum value of the key degree can be found in each video selection window respectively, and the found candidate video frame is taken as the target video frame corresponding to each video selection window. It can be seen from this that through the above method, the information amount carried by the candidate video frames within each video selection window and the complexity of the contour and other aspects can be evaluated respectively from the above two to four dimensions, and the candidate video frame with a large information amount and high complexity is found as the target video frame corresponding to each video selection window, so as to ensure the quality of frame selection.
[0118] Figure 5A Shows an example of selecting target video frames from candidate video frames according to some embodiments of the present disclosure. Figure 5A A total of 18 candidate video frames are shown, where each candidate video frame is represented by a small rectangular box. Assuming that the candidate frame number index is 6, then through the above method, the candidate video frames can be divided into three groups of candidate video frames, and three video frame selection windows are determined. Among them, each video frame selection window is represented by six small rectangular boxes with thickened borders. As described above, through the above method, a target video frame can be selected from each video frame selection window respectively. In Figure 5A the above target video frame can be represented by a small rectangular box with slash shading. It should be noted that Figure 5A the number of candidate video frames, the number of video selection windows, the number of candidate video frames within each video selection window, and the number of target video frames shown in the figure are all examples, and the embodiments of the present disclosure do not limit the above numbers.
[0119] In some embodiments of the present disclosure, the target video frames selected by the above method may be adjacent. For example, Figure 5A in the figure, the target video frame corresponding to the second video frame selection window is the last candidate video frame within this video frame selection window; while the target video frame corresponding to the third video frame selection window is the first candidate video frame within this video frame selection window. Since the target video frames corresponding to the two video frame selection windows are adjacent candidate video frames, there may be a situation where the similarity of the two target video frames is relatively high, thus affecting the final frame selection quality.
[0120] In order to break up the selected target video frames and avoid selecting target video frames with high similarity, as an alternative to the above solution, the specific method for determining the video frame selection result based on the target video frames corresponding to the target video frame selection window in step 260 described above may include: determining a candidate video frame with the maximum key degree within the target video frame selection window; in response to determining that the candidate video frame with the maximum key degree within the target video frame selection window is the first candidate video frame within the target video frame selection window and the target video frame corresponding to the previous video frame selection window of the target video frame selection window is the last candidate video frame within the previous video frame selection window, using the candidate video frame with the second largest key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window; in response to determining that the candidate video frame with the maximum key degree within the target video frame selection window is not the first candidate video frame within the target video frame selection window or the target video frame corresponding to the previous video frame selection window of the target video frame selection window is not the last candidate video frame within the previous video frame selection window, using the candidate video frame with the maximum key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window. It can be seen that through the above method, the selected target video frames can be broken up to avoid the situation of selecting adjacent candidate video frames.
[0121] Figure 5B Shows an example of selecting target video frames from candidate video frames according to some other embodiments of the present disclosure. Figure 5B A total of 18 candidate video frames are shown, where each candidate video frame is represented by a small rectangular box. Assuming the candidate frame number index is 6, the candidate video frames can be divided into three candidate video frame groups through the above method, and three video frame selection windows are determined. Among them, each video frame selection window is represented by six small rectangular boxes with thickened borders. As described above, one target video frame can be selected from each video frame selection window through the above method. And assuming that Figure 5B in, the target video frame corresponding to the first video frame selection window is the third candidate video frame within this video frame selection window; the target video frame corresponding to the second video frame selection window is the last candidate video frame within this video frame selection window (represented by a small rectangular box with diagonal shading); and the candidate video frame with the maximum key degree in the third video frame selection window is the first candidate video frame within this video frame selection window (represented by a black small rectangular box). Then, according to the above method, the candidate video frame with the second largest key degree within the video frame selection window (the second candidate video frame within the third video frame selection window) is used as the target video frame corresponding to the target video frame selection window (represented by a small rectangular box with diagonal shading). It can be seen that through the above method, the selected target video frames can be broken up to avoid the situation of selecting adjacent candidate video frames. Similarly, Figure 5BThe number of candidate video frames shown, the number of video selection windows, the number of candidate video frames within each video selection window, and the number of target video frames are also all examples, and the embodiments of the present disclosure do not limit the above numbers.
[0122] As an alternative to the above solution, also in order to scatter the selected target video frames and avoid selecting target video frames with high similarity, the specific method of determining the video frame selection result based on the target video frames corresponding to the target video frame selection window described in step 260 above may alternatively include: determining the candidate video frame with the maximum key degree within the target video frame selection window; determining the similarity between the candidate video frame with the maximum key degree within the target video frame selection window and the target video frame corresponding to the previous video frame selection window of the target video frame selection window; in response to determining that the similarity is greater than a pre-set similarity threshold, using the candidate video frame with the second largest key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window; in response to determining that the similarity is less than or equal to the pre-set similarity threshold, using the candidate video frame with the maximum key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window. It can be seen that by the above method of comparing the similarity with the target video frame corresponding to the previous video selection window, the purpose of scattering the selected target video frames and avoiding the situation of selecting candidate video frames with too high similarity can also be achieved.
[0123] Furthermore, as an additional solution, in Figure 2 the embodiment shown, after obtaining at least one candidate video frame in step 210, it may further include: scaling each candidate video frame to a pre-set image scale. Then, the above step 220 may be continued. It can be understood that by scaling the candidate video frames, the computational amount in subsequent operations can be adjusted, thereby achieving the goal of saving computing power. In a specific application, the above-mentioned pre-set image scale can be set according to the actual situation. For example, in some examples, the above-mentioned pre-set image scale can be the resolution of the short side of the pre-set video frame. For example, the resolution of the short side of the video frame can be pre-set to 540 pixels. In this way, in the above step, before executing the above step 220, the candidate video frames can be scaled to a short side resolution of 540 pixels.
[0124] It can be seen from this that through the above method, it is possible to evaluate aspects such as the amount of information carried by each candidate video frame in each video selection window and the complexity of the contour, and find the candidate video frames with a large amount of information and high complexity as the target video frames corresponding to each video selection window, thereby effectively ensuring the quality of frame selection. In addition, the various parameters for determining the key degree of the candidate video frames are not complex, the calculation amount is small, and the computing power consumption is not large. Therefore, the above method can optimize the quality of frame selection and the processing efficiency of frame selection, so as to achieve a balance between the quality of frame selection and the processing efficiency of frame selection.
[0125] Corresponding to the above video frame selection method, some embodiments of the present disclosure also disclose a video frame selection device. Figure 6 shows the internal structure of the video frame selection device described in the embodiments of the present disclosure. As Figure 6 shown, the above video frame selection device may include the following multiple modules:
[0126] The frame extraction module 610 is used to extract frames from the video to obtain at least one candidate video frame;
[0127] The information entropy determination module 620 is used to respectively determine the information entropy of each candidate video frame;
[0128] The edge feature determination module 630 is used to respectively determine the edge information feature value of each candidate video frame;
[0129] The grouping module 640 is used to group at least one candidate video frame to determine multiple video selection windows;
[0130] The frame selection module 650 is used to respectively use each video selection window in the multiple video selection windows as the target video selection window, and perform the following operations on the candidate video frames in the target video selection window: determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information feature value of the candidate video frame; and determining the target video frame corresponding to the target video selection window based on the key degree of the candidate video frame;
[0131] The output module 660 determines the frame selection result of the video based on the target video frame corresponding to the target video selection window.
[0132] It should be noted that the implementation methods and the specific technical effects that can be achieved by each module in the above video frame selection device can refer to the implementation methods of each step in the foregoing embodiments, and will not be repeated here.
[0133] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for selecting frames of a video described in any of the above embodiments is implemented.
[0134] Figure 7 FIG. shows a schematic diagram of the hardware structure of a more specific electronic device provided in this embodiment. The device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. Among them, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are communicatively connected to each other inside the device through the bus 2050.
[0135] The processor 2010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0136] The memory 2020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 2020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 2020 and are called and executed by the processor 2010.
[0137] The input / output interface 2030 is used to connect to input / output devices to achieve information input and output. Among them, the input / output devices may be configured as components in the device or externally connected to the device to provide corresponding functions. Among them, the input devices may include microphones, various sensors, etc., and the output devices may include displays, speakers, vibrators, indicator lights, etc.
[0138] The communication interface 2040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module may communicate through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0139] The bus 2050 includes a path for transmitting information among various components of the device, such as the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040.
[0140] It should be noted that although only the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050 are shown in the above device, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0141] The electronic device of the above embodiment is used to implement the corresponding video frame selection method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0142] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the video frame selection method as described in any of the foregoing embodiments.
[0143] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0144] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the video frame selection method as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0145] Based on the same inventive concept, corresponding to the video frame selection method of any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, when the computer program instructions run on a computer, the computer is caused to execute the steps in each of the embodiments of the video frame selection method. Corresponding to the execution subject of each step in each of the embodiments of the video frame selection method, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0146] The computer program product of the above embodiment is used to cause a processor to execute the video frame selection method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0147] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; Under the idea of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.
[0148] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0149] Although the present disclosure has been described in connection with specific embodiments of the present disclosure, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0150] Embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for selecting frames of a video, comprising: Performing frame extraction on the video to obtain at least one candidate video frame; Respectively determining the information entropy of each of the candidate video frames; Respectively determining the edge information eigenvalue of each of the candidate video frames; Grouping the at least one candidate video frame to determine a plurality of video frame selection windows; Respectively taking each of the video frame selection windows as a target video frame selection window, and performing the following operations on the candidate video frames within the target video frame selection window: determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame; and determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame; And Determining the frame selection result of the video based on the target video frame corresponding to the target video frame selection window.
2. The method according to claim 1, wherein, The performing frame extraction on the video to obtain at least one candidate video frame includes: Determining the number of the candidate video frames based on a preset frame selection number index and a preset candidate frame number index; and Performing equidistant frame extraction on the video based on the number of the candidate video frames to obtain the at least one candidate video frame.
3. The method according to claim 1, wherein The performing frame extraction on the video to obtain at least one candidate video frame includes: Dividing the video into at least one video segment based on the key frames in the video; wherein each video segment includes one of the key frames; Respectively determining the number of candidate video frames corresponding to each video segment based on the number of the video segments, a preset frame selection number index, and a preset candidate frame number index; Respectively for each video segment, performing equidistant frame extraction on the video segment based on the number of candidate video frames corresponding to the video segment to obtain the alternative video frames corresponding to the video segment; wherein the alternative video frames corresponding to the video segment include the key frame included in the video segment; and Taking the alternative video frames corresponding to at least one of the video segments as the at least one candidate video frame.
4. The method according to claim 1, wherein, The performing frame extraction on the video to obtain at least one candidate video frame includes: Dividing the video into at least one video shot through shot segmentation; Respectively determining the number of candidate video frames corresponding to each video shot based on the number of the video shots, a preset frame selection number index, and a preset candidate frame number index; Respectively for each video shot, performing equidistant frame extraction on the video shot based on the number of candidate video frames corresponding to the video shot to obtain the alternative video frames corresponding to the video shot; and Taking the alternative video frames corresponding to at least one of the video shots as the at least one candidate video frame.
5. The method according to claim 1, wherein, The respectively determining the information entropy of each of the candidate video frames includes: Performing the following steps for each of the candidate video frames: Determining one or a combination of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame; and Taking one or a combination of the original image information entropy of the candidate video frame and the grayscale image information entropy of the candidate video frame as the information entropy of the candidate video frame.
6. The method according to claim 1, wherein, The separately determining the edge information eigenvalue of each of the candidate video frames includes: Performing the following steps for each of the candidate video frames: Determining one or a combination of the variance of the Laplacian operator of the candidate video frame and the edge proportion of the candidate video frame; and Taking one or a combination of the variance of the Laplacian operator of the candidate video frame and the edge proportion of the candidate video frame as the edge information eigenvalue of the candidate video frame.
7. The method according to any one of claims 2 to 4, wherein The grouping the at least one candidate video frame and determining multiple video frame selection windows includes: Dividing the at least one candidate video frame into multiple candidate video frame groups based on the candidate frame number index; wherein, the number of candidate video frames included in each candidate video frame group is less than or equal to the candidate frame number index; and Respectively taking each candidate video frame group as one of the video frame selection windows.
8. The method according to claim 1, wherein The determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information eigenvalue of the candidate video frame includes: Performing the following steps for each of the candidate video frames within the target video frame selection window: Determining the maximum value and the minimum value of the information entropy of the candidate video frame; Determining the normalized information entropy of the candidate video frame based on the maximum value and the minimum value of the information entropy of the candidate video frame; Determining the maximum value and the minimum value of the edge information eigenvalue of the candidate video frame; Determining the normalized edge information eigenvalue of the candidate video frame based on the maximum value and the minimum value of the edge information eigenvalue of the candidate video frame; and Performing weighted summation on the normalized information entropy of the candidate video frame and the normalized edge information eigenvalue of the candidate video frame to obtain the key degree of the candidate video frame.
9. The method according to claim 1, wherein The determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame includes: Taking the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window.
10. The method according to claim 1, wherein, The determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame includes: Determining the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window; In response to determining that the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window is the first candidate video frame within the target video frame selection window and the target video frame corresponding to the previous video frame selection window of the target video frame selection window is the last candidate video frame within the previous video frame selection window, taking the candidate video frame corresponding to the second largest value of the key degree within the target video frame selection window as the target video frame corresponding to the target video frame selection window; In response to determining that the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window is not the first candidate video frame within the target video frame selection window or the target video frame corresponding to the previous video frame selection window of the target video frame selection window is not the last candidate video frame within the previous video frame selection window, the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window is used as the target video frame corresponding to the target video frame selection window.
11. The method according to claim 1, wherein The determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame includes: determining the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window; determining the similarity between the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window and the target video frame corresponding to the previous video frame selection window of the target video frame selection window; In response to determining that the similarity is greater than a preset similarity threshold, the candidate video frame corresponding to the second largest value of the key degree within the target video frame selection window is used as the target video frame corresponding to the target video frame selection window; In response to determining that the similarity is less than or equal to the similarity threshold, the candidate video frame corresponding to the maximum value of the key degree within the target video frame selection window is used as the target video frame corresponding to the target video frame selection window.
12. The method according to claim 1, wherein After obtaining the at least one candidate video frame, the method further includes: scaling each of the candidate video frames to a preset image scale.
13. A video frame selection device, comprising: a frame extraction module, configured to extract frames from the video to obtain at least one candidate video frame; an information entropy determination module, configured to determine the information entropy of each of the candidate video frames respectively; an edge feature determination module, configured to determine the edge information feature value of each of the candidate video frames respectively; a grouping module, configured to group the at least one candidate video frame to determine a plurality of video frame selection windows; a frame selection module, configured to use each of the video frame selection windows as a target video frame selection window, and perform the following operations on the candidate video frames within the target video frame selection window: determining the key degree of the candidate video frame based on the information entropy of the candidate video frame and the edge information feature value of the candidate video frame; and determining the target video frame corresponding to the target video frame selection window based on the key degree of the candidate video frame; and an output module, configured to determine the frame selection result of the video based on the target video frame corresponding to the target video frame selection window.
14. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the video frame selection method according to any one of claims 1-12 is implemented.
15. A non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the video frame selection method according to any one of claims 1-12.
16. A computer program product comprising computer program instructions which, when run on a computer, cause the computer to execute the method for frame selection of a video according to any one of claims 1-12.