Video generation method, electronic device, and readable medium
By employing interval frame extraction and adaptive configuration in video generation, the problem of incomplete keyframe extraction in existing technologies is solved, resulting in better video display effects and analysis performance.
Patent Information
- Application Number
- CN202410038433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-01-10
AI Technical Summary
Existing technologies that extract keyframes from videos lead to biased video understanding and incomplete video content parsing, resulting in poor display quality of the generated one-click blockbuster videos.
Multiple first keyframes are extracted at a first preset time interval for large-scale extraction, and continuous extraction is performed in key areas by combining small-span frame extraction. By adaptively configuring the number and duration of extraction frames, the accuracy and comprehensiveness of keyframes are ensured.
It enables accurate understanding of video content within a limited time, generating one-click blockbuster and short videos with better display effects, thus improving the efficiency and accuracy of video analysis.
Smart Images

Figure CN119254906B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a video generation method, electronic device, computer program product, and computer-readable storage medium. Background Technology
[0002] Electronic devices involve video editing in the process of performing various functions. For example, when an electronic device generates a short video with a one-click feature, it needs to extract keyframes from the user-selected video. The electronic device uses keyframes to represent a segment of the video in order to understand the content of that segment.
[0003] The common method for electronic devices to extract keyframes from video is to extract video frames at fixed intervals by scrolling from front to back. However, using keyframes extracted in this way to understand video content can lead to inaccuracies in video interpretation and inadequate summarization of video content, resulting in poor display quality of one-click large-scale or small-scale videos generated from keyframes. Summary of the Invention
[0004] This application provides a video generation method, electronic device, computer program product, and computer-readable storage medium, with the aim of extracting key frames that can comprehensively and accurately represent the meaning of the video, so as to ensure that the video generated based on the key frames has a better display effect.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] Firstly, this application provides a video generation method, comprising: displaying a first interface showing a thumbnail of a video and a first button; responding to a user selecting at least one video on the first interface and clicking the first button, extracting multiple first keyframes from the video at first preset time intervals, wherein the multiple first keyframes are discontinuous; obtaining highlight segments of the video based on the multiple first keyframes; and synthesizing the highlight segments and special effects of the video to obtain a short video. In the scenario of a one-click blockbuster function, the first interface can refer to the interface of the one-click blockbuster function.
[0007] As can be seen from the above, multiple first keyframes are extracted from the video at intervals of a first preset duration. Since the images of the multiple first keyframes are not continuous, it indicates that the multiple first keyframes are extracted using a large-span method. This allows for the extraction of keyframes over a large or full range of the video, thereby enabling an accurate and comprehensive understanding of the video content based on the extracted video frames, and ensuring that the video generated based on the keyframes has a good display effect.
[0008] In one possible implementation, obtaining a highlight segment of the video based on a first keyframe includes: obtaining a key region of the video based on the first keyframe; extracting multiple second keyframes from the key region using a frame-skipping method with a second preset duration, wherein the second preset duration is less than the first preset duration, and the multiple second keyframes are continuous; and obtaining a highlight segment in the key region based on the second keyframes.
[0009] In the above possible implementations, the continuous images of multiple second keyframes indicate that the multiple second keyframes are extracted in a small span manner. After the first stage of large-scale frame extraction for rapid overview, and the second stage of selecting a small number of high-value regions for continuous video segment analysis based on the analysis results of the first stage, the efficiency is improved within the limited video analysis time and hardware performance limitations, and the precision and recall rate of highlight segment recognition are improved.
[0010] In one possible implementation, before extracting multiple first keyframes from the video at a frame-skipping interval of a first preset duration, the method further includes: configuring the number of first keyframes to be skimmed and the frame-skipping duration; extracting multiple first keyframes from the video at a frame-skipping interval of a first preset duration includes: extracting the number of first keyframes from the video within the frame-skipping duration at a frame-skipping interval of a first preset duration.
[0011] In one possible implementation, configuring the number of frames extracted and the extraction duration of the first keyframe includes: determining a first number of frames extracted and a second number of frames extracted based on the duration of the video, wherein the second number of frames extracted is greater than the first number of frames extracted; using a portion of the total video analysis duration as the extraction duration; and configuring the number of frames extracted as the first number of frames extracted and the extraction duration as a portion of the total video analysis duration if it is determined that there are enough keyframes to be extracted from the video within the extraction duration, and the number of keyframes extracted does not exceed the second number of frames extracted.
[0012] In the above possible implementations, determining the first and second number of extracted frames for the first keyframe based on the video duration means that the first and second number of extracted frames are positively correlated with the video duration. If the extraction time is sufficient to extract the first number of keyframes from the video, and the extracted keyframes do not exceed the second number of extracted frames, configuring the extraction number of the first keyframe as the first number of extracted frames and the extraction time as a portion of the total video analysis time ensures that a reasonable number of first keyframes are extracted from the video, and that the extraction time of the first keyframe only occupies a portion of the total video analysis time. Thus, the remaining portion of the total video analysis time can be used for detailed analysis of the extracted first keyframes to obtain highlight segments of the video.
[0013] In one possible implementation of the first and second frame extraction counts for extracting the first keyframe, when the user selects the first video and the second video on the first interface, the first and second frame extraction counts for extracting the first keyframe are determined based on the video duration, with the second frame extraction count being greater than the first frame extraction count. This includes: determining the third and fourth frame extraction counts for extracting the first keyframe based on the duration of the first video, and determining the fifth and sixth frame extraction counts for extracting the first keyframe based on the duration of the second video, with the fourth frame extraction count being greater than the third frame extraction count, and the sixth frame extraction count being greater than the fifth frame extraction count; determining the first frame extraction count for extracting the first keyframe as the sum of the third and fifth frame extraction counts, and determining the second frame extraction count for extracting the first keyframe as the sum of the fourth and sixth frame extraction counts.
[0014] In one possible implementation, the method further includes: if it is determined that the number of keyframes to be extracted from the first video and the second video within the frame extraction duration is insufficient, configuring the number of keyframes to be extracted as the fifth number of keyframes, and the frame extraction duration as a portion of the total video analysis duration; wherein, extracting the first number of keyframes to be extracted from the video within the frame extraction duration at a frame extraction interval of the first preset duration includes: extracting the fifth number of first keyframes to be extracted from the second video within the frame extraction duration at a frame extraction interval of the first preset duration; obtaining highlight segments of the video based on multiple first keyframes includes: obtaining highlight segments of the second video based on the fifth number of first keyframes extracted from the second video, and obtaining highlight segments of the first video based on the first keyframes in the historical frame extraction results of the first video.
[0015] In the above possible implementations, if it is determined that the number of keyframes required to extract the first number of keyframes from the first and second videos is insufficient within the extraction time, the extraction number of the first keyframes is configured to be the fifth number of keyframes, and the extraction time is a portion of the total video analysis time. This reduces the number of keyframes required to complete the extraction within the extraction time, ensuring that the required number of keyframes can be extracted within the extraction time even when the electronic device performance is poor. Furthermore, since the first video uses the first keyframe from historical extraction results to obtain the highlight segment of the first video, it also ensures that the keyframes in the first video participate in the generation of the short video, guaranteeing that the obtained short video is based on both the first and second videos.
[0016] In one possible implementation, after configuring the number of frames extracted from the first keyframe as the fifth number of frames extracted and the frame extraction duration as a portion of the total video analysis duration, the method further includes: if it is determined that the frame extraction duration is insufficient to extract the fifth number of keyframes from the second video, increasing the frame extraction duration to a first frame extraction duration, wherein the first frame extraction duration is sufficient to extract the fifth number of keyframes from the second video; and configuring the frame extraction duration as the first frame extraction duration.
[0017] In one possible implementation, after configuring the number of frames extracted from the first keyframe to be the number of frames extracted from the fifth keyframe, and the frame extraction duration to be a portion of the total video analysis duration, the method further includes: if it is determined that there are not enough keyframes to be extracted from the second video within the frame extraction duration, configuring the frame extraction duration to be the total video analysis duration, where there are either enough keyframes to be extracted from the second video within the total video analysis duration, or not enough keyframes to be extracted from the second video within the total video analysis duration.
[0018] In one possible implementation, the method further includes: if it is determined that there is sufficient time within the frame extraction duration to extract a first number of keyframes from the video, and the number of keyframes extracted exceeds the number of second frame extractions, reducing the frame extraction duration to a second frame extraction duration, wherein there is sufficient time within the second frame extraction duration to extract a first number of keyframes from the video, and the number of keyframes extracted is substantially the same as the number of second frame extractions; configuring the number of keyframes extracted for the first keyframes as the first frame extraction duration and the frame extraction duration as the second frame extraction duration.
[0019] In one possible implementation, based on the duration of the video, a first frame extraction number and a second frame extraction number are determined to extract the first keyframe. After the second frame extraction number is greater than the first frame extraction number, the method further includes: when it is determined that the electronic device stores historical frame extraction results of the video and the video analysis algorithm capability of the electronic device has not been upgraded, updating the first frame extraction number based on the seventh frame extraction number in the historical frame extraction results of the video, and updating the second frame extraction number based on the eighth frame extraction number in the historical frame extraction results of the video; wherein the updated first frame extraction number is less than the original first frame extraction number, and the updated second frame extraction number is less than the original second frame extraction number.
[0020] In one possible implementation, the first frame count is updated based on the seventh frame count in the historical frame-sampling results of the video, and the second frame count is updated based on the eighth frame count in the historical frame-sampling results of the video, including: subtracting the product of the seventh frame count and the first weight from the first frame count, and subtracting the product of the eighth frame count and the second weight from the second frame count, wherein both the first weight and the second weight are less than 1.
[0021] In a second aspect, this application provides an electronic device, including: one or more processors, a memory, and a display screen; the memory and the display screen are coupled to one or more processors, the memory is used to store a computer program, the computer program including computer instructions, and when one or more processors execute the computer instructions, the electronic device performs a video generation method as provided in any of the first aspect and possible embodiments.
[0022] Thirdly, this application provides a computer-readable storage medium for storing a computer program, which, when executed, specifically implements the video generation method provided in any of the first aspects and possible embodiments.
[0023] Fourthly, this application provides a computer program product that, when run on a computer, causes the computer to perform a video generation method as provided in any of the first aspects and possible embodiments. Attached Figure Description
[0024] Figure 1 A hardware structure diagram of the electronic device provided in the embodiments of this application;
[0025] Figure 2 A software structure diagram of the electronic device provided in the embodiments of this application;
[0026] Figure 3 An illustration of obtaining the highlight segment of video 1 in the video generation method provided in the embodiments of this application;
[0027] Figure 4 A flowchart of the method for adjusting the frame extraction duration of keyframes in the video generation method provided in the embodiments of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise.
[0029] References to "some embodiments" and the like in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, phrases such as "in some embodiments," "in other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiments, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including, but not limited to," unless otherwise specifically emphasized.
[0030] The "multiple" mentioned in the embodiments of this application refers to two or more. It should be noted that in the description of the embodiments of this application, terms such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order.
[0031] The "One-Click Blockbuster" or "One-Click Video Creation" function refers to a user selecting photos or videos and clicking the "One-Click Blockbuster" button (or "One-Click Video Creation" button). The electronic device automatically analyzes the user's input material, selects highlight clips, and combines these clips with special effects and background music configured by the electronic device to create a one-click blockbuster music video (or simply a one-click blockbuster short video). Special effects refer to unique effects that can be added to video frames, supported by the source material, such as snowflakes, fireworks, and other animation effects, as well as filters, stickers, borders, etc. In some embodiments, special effects may also be referred to as styles or style themes.
[0032] When electronic devices analyze user-input materials and select highlight segments, they need to edit the video. This editing may include extracting keyframes (also called representative frames, I-frames, etc.) to represent a segment of the video to understand its content. Currently, the common method for extracting keyframes from video is for the electronic device to extract video frames at fixed intervals by scrolling from front to back. This frame extraction method involves very short intervals between frames.
[0033] After studying the aforementioned method of extracting keyframes from videos, the inventors discovered that there are issues with understanding video content based on extracted video frames, including biases in video comprehension and inadequate analysis and summarization of video content. Furthermore, the keyframes extracted by this method cannot accurately reflect the meaning of the video, resulting in poor display quality of the one-click generation of large and small videos based on keyframes.
[0034] The reason for this was found to be that, to avoid users waiting a long time for electronic devices to generate one-click short videos, the devices are configured with a time limit for generation. Extracting keyframes from the video is one step in the one-click short video generation process, and it too is subject to time constraints. Because the electronic device extracts keyframes scrolling from beginning to end, the time limit limits the extraction of keyframes, sometimes resulting in only keyframes being extracted from a portion of the content. This is especially true for longer videos, where the device can only extract keyframes from a limited range. Keyframes extracted from this limited range are incomplete or even biased in reflecting the overall meaning of the video, failing to provide a comprehensive analysis and accurate understanding of its overall message. Of course, keyframes extracted from only a portion of the video also cannot accurately represent its meaning.
[0035] To balance user waiting time and comprehensive video comprehension, this application provides a video generation method. This method can be applied to electronic devices such as mobile phones, tablets, personal digital assistants (PDAs), desktop, laptop, and notebook computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, and wearable devices.
[0036] Taking mobile phones as an example, Figure 1 This is an example of the composition of an electronic device provided in an embodiment of this application. For example... Figure 1 As shown, the electronic device 100 may include a processor 110, an internal memory 120, a camera 130, and a display screen 140, etc.
[0037] It should be understood that, for the purpose of facilitating the understanding of the embodiments of this application, Figure 1 The electronic device 100 shown only includes some components related to the video frame extraction method provided in this application embodiment. The electronic device provided in this application embodiment may include more than Figure 1 The electronic device shown has 100 or more or fewer components. That is to say, Figure 1 The electronic device 100 shown does not constitute a specific limitation on the electronic device provided in the embodiments of this application.
[0038] Processor 110 may include one or more processing units, such as an application processor (AP), a graphics processing unit (GPU), an image signal processor (ISP), or a video codec. Processor 110 may also include memory for storing instructions and data.
[0039] Internal memory 120 can be used to store computer executable program code, including instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 120. In some embodiments, internal memory 120 stores instructions for video generation methods. Processor 110 can execute the instructions stored in internal memory 120 to enable electronic device to extract keyframes that can comprehensively and accurately represent the meaning of the video, thereby ensuring that the video generated based on the keyframes has a better display effect.
[0040] The electronic device 100 can perform shooting functions through an ISP, camera 130, video codec, GPU, display screen 140, and application processor. The electronic device 100 can also perform display functions through a GPU, display screen 140, and application processor.
[0041] The following section introduces the software architecture of electronic devices, which can be understood as the layered architecture of the electronic device's operating system. The operating system of an electronic device runs on top of its hardware components and can be, for example, iOS, the Android open-source operating system, or Windows.
[0042] This application uses the layered architecture of the Android system as an example to illustrate the software structure of an electronic device.
[0043] Figure 2 This is a software structure block diagram of an electronic device according to an embodiment of this application.
[0044] A layered architecture divides software into several layers, which communicate with each other through software interfaces. In some embodiments, the Android system is divided into two layers: an application layer and an application framework layer, from top to bottom. Below this two-layer architecture, other layers and a hardware layer may also be included. The hardware layer includes the hardware components of the electronic device, such as… Figure 1 The hardware components on display.
[0045] The application layer can include multiple applications. For example, Figure 2 It showcases two applications: a gallery and a video editing application.
[0046] In some embodiments, the gallery app is used to display videos and images saved on the electronic device. The gallery app can invoke a video editing app based on user actions; during the operation of the video editing app, the display screen switches from showing the gallery app's interface to showing the video editing app's interface.
[0047] In some embodiments, the electronic device implements a one-click blockbuster function through a video editing application. For example, the gallery application displays images and videos saved on the electronic device. The user long-presses on an image or video, and the screen displays the interface for the one-click blockbuster function, including a one-click blockbuster button. The user selects the video and image materials for which a one-click blockbuster video needs to be generated, clicks the one-click blockbuster button, and the electronic device invokes the video editing application. The application automatically analyzes the user-input materials, selects highlight clips, and combines these highlight clips with effects and background music configured by the electronic device to create a one-click blockbuster video.
[0048] In some embodiments, such as Figure 2 As shown, the video editing application includes three functional modules: material analysis, highlight clip editing and splicing, and cinematic final product compositing. Material analysis analyzes the video and image materials selected by the user to generate one-click high-quality short videos; the results are provided to the lower-level modules. Highlight clip editing and splicing uses the timestamps of the highlight clips provided by the lower-level modules to edit and splice them. The lower-level modules mentioned here may refer to... Figure 2 The media processing platform showcased below demonstrates its functions. The cinematic final effect compositing feature combines highlight clips with special effects and background music from electronic devices to create stunning short videos with a single click.
[0049] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example, Figure 2 The application framework layer shown may include a media processing platform that provides application programming interfaces and programming frameworks for video editing applications.
[0050] In some embodiments, the media processing platform first performs a large-scale I-frame extraction operation on the video footage, enabling the extraction of I-frames over a wide or even full range. Then, it uses the extracted I-frames to quickly overview the video and identify the approximate areas (or key areas) where highlight segments of the video footage are located. The media processing platform also performs detailed analysis on these key areas to determine the precise location, range, and highlight characteristics (such as highlight actions, postures, and expressions) of the highlight segments in the video footage. Finally, it reports the determined precise location, range, and highlight characteristics (such as highlight actions, postures, and expressions) of the highlight segments in the video footage to the video editing application.
[0051] The following combination Figure 3This section describes how an electronic device obtains highlight clips from the user-selected video footage during the process of generating one-click short videos. It explains that the media processing platform within the electronic device's application framework layer obtains these highlight clips.
[0052] For example, such as Figure 3 As shown, the video materials selected by the user include Video 1 and Video 2. First, it should be noted that when the user selects multiple video materials, the electronic device stitches these materials together into a single long video for I-frame extraction analysis. After I-frame extraction analysis, the electronic device obtains the number of I-frames extracted from each video material. In some embodiments, the electronic device uniformly plans and allocates the number of I-frames extracted to each video material based on the total duration of video analysis and the specifications and length of the input video materials.
[0053] After the electronic device performs I-frame extraction analysis to obtain the number of I-frames extracted for each video clip, it extracts keyframes corresponding to the number of I-frames extracted from each video clip. Based on the extracted keyframes, the electronic device can infer the theme of the video clip through a theme analysis algorithm. In this way, the electronic device can obtain special effects templates based on the theme of the video clip. The electronic device can also compare the image similarity between keyframes based on the extracted keyframes to initially obtain similar image groups, and then obtain the shot position of each video clip, that is, the timestamp of the image indicating the shot.
[0054] The following uses video 1 as an example to describe in detail the process by which an electronic device obtains the highlight segment of video 1 based on the I-frame extracted from video 1. The electronic device obtains the highlight segment of video 2 in the same way, but will not be described in detail.
[0055] After obtaining the number of I-frames extracted from Video 1, the electronic device performs the first round of rapid preview operation, specifically the frame extraction phase. This involves extracting a large number of I-frames (keyframes) from Video 1 using a wide-span extraction method. Wide-span extraction means the electronic device extracts video frames at intervals of several seconds, resulting in non-continuous frames. This allows the device to extract keyframes across a large or even full range of Video 1, enabling an accurate and comprehensive understanding of the video content based on these extracted frames.
[0056] Furthermore, based on the keyframes extracted from a wide or full range of video 1, the electronic device accurately and comprehensively understands the video content. As a result, the theme of video 1 inferred by the electronic device is also accurate, thus enabling the acquisition of special effects templates that better match the theme of video 1, ensuring a better display effect for one-click blockbuster and short video formats.
[0057] In some embodiments, the electronic device further groups the extracted keyframes based on image similarity to obtain similar image groups, such as... Figure 3 As shown, the first two keyframes extracted are grouped as similar scene group 1, the third to fifth keyframes extracted are grouped as similar scene group 2, and the sixth and seventh keyframes extracted are grouped as similar scene group 3.
[0058] In some embodiments, a group of similar frames obtained by the electronic device is a storyboard obtained from a rough analysis. For example, such as... Figure 3 As shown, the electronic device obtains three segments of video 1 based on similar scene group 1, similar scene group 2 and similar scene group 3.
[0059] In other embodiments, the electronic device obtains key regions corresponding to the video footage based on each group of similar frames, and on the content and image quality evaluation of keyframes. These key regions then serve as candidate regions for a second round of fine-tuning analysis to identify highlight segments.
[0060] In some embodiments, the electronic device can determine which of the multiple similar scene groups belongs to the "high-quality" similar scene groups based on the content of keyframes in each similar scene group. The electronic device then selects a video segment from the first image's timestamp to the last image's timestamp within the "high-quality" similar scene group as the key region corresponding to that video material. Alternatively, the electronic device can also obtain the key region corresponding to the video material based on high-rated images within the "high-quality" similar scene group. In some embodiments, the electronic device uses a video analysis algorithm to determine a high-rated image within a similar scene group. The electronic device evaluates the image quality of the images in the similar scene group using the video analysis algorithm, obtains a score for each image, and designates the image with the highest overall score as the high-rated image. Multiple video analysis algorithms can be used; the electronic device comprehensively evaluates the images based on the scores of each algorithm and designates the image with the highest overall score as the high-rated image.
[0061] The electronic device will also conduct a second round of detailed analysis on key areas. This involves extracting frames at small intervals, such as 10 frames per second. Because the key frames extracted in this small-interval extraction are very close together, it can be essentially understood as extracting continuous video segments. The electronic device then analyzes the extracted video frames to obtain the actual location of the highlight segments in the key areas.
[0062] In some embodiments, during the process of small-span frame extraction in key areas, the electronic device may also base its actions on the following principles: extract fewer frames from areas with similar content and extract more frames from areas with changing content; extract fewer frames from low-quality areas and extract more frames from high-quality areas.
[0063] The video generation method provided in this application does not primarily focus on analyzing continuous video segments using related technologies, but rather revolves around extracting keyframes from video footage. The main reason for this is:
[0064] 1. The inventors' research found that the speed of video analysis is a key performance indicator and a reflection of competitiveness.
[0065] 2. The video decoding and video analysis capabilities of electronic devices are limited, making it impossible to support the analysis of large-scale continuous video segments;
[0066] 3. Reserves support for multi-round progressive analysis capabilities for subsequent image library application analysis scenarios.
[0067] Furthermore, through the first stage of large-scale frame sampling for rapid overview, and the second stage of selecting a small number of high-value regions for continuous video segment analysis based on the analysis results of the first stage, it is possible to improve efficiency and enhance the precision and recall rate of highlight segment recognition within the limited analysis time and hardware performance constraints.
[0068] The following describes the process of adaptively configuring the frame extraction duration of keyframes in the video generation method provided in the embodiments of this application.
[0069] like Figure 4 As shown, the method for adaptively configuring the frame extraction duration of keyframes includes:
[0070] S401. Calculate the basic number of keyframe extractions based on the duration of the video footage.
[0071] The basic number of keyframes extracted from video footage can be understood as the minimum number of keyframes extracted from that video footage that can participate in the generation of one-click blockbuster short videos.
[0072] In some embodiments, when a user selects only one video clip in the "One-Click Movie" feature interface displayed in the gallery application, the electronic device can calculate the basic frame extraction count of keyframes based on the duration of that video clip. In other embodiments, when a user selects multiple video clips in the "One-Click Movie" feature interface displayed in the gallery application, the electronic device needs to calculate the basic frame extraction count of keyframes for each selected video clip based on its duration. The sum of the basic frame extraction counts of keyframes from multiple video clips is then used as the basic frame extraction count of the long video composed of the multiple video clips.
[0073] In some embodiments, the basic frame extraction number for calculating keyframes based on the duration of video footage must meet the following rules:
[0074] 1. Extract at least one keyframe from a video clip; if the video clip is 3 seconds or longer, extract 2 keyframes.
[0075] 2. For every 5 seconds of video footage, add one more keyframe, up to 30 seconds, for a total of 8 keyframes.
[0076] 3. If the video footage is longer than 30 seconds, for every 10 seconds of video content after 30 seconds, add one more keyframe, up to 90 seconds, which means a total of 14 keyframes will be extracted.
[0077] It should be noted that, to ensure a relatively accurate video theme, the video footage must have at least 5 keyframes. Specifically, if the user selects only one video clip, and the base number of keyframes calculated according to the above three rules is less than 5, then the base number of keyframes will be adjusted to 5.
[0078] When the user selects any one of the video clips from 2 to 4, if the total number of basic frame extractions of the keyframes of the multiple video clips calculated according to the above three rules is less than 5, for video clips with less than 2 keyframes, the basic frame extraction count will be adjusted to 2. If the total number of basic frame extractions after adjustment is still less than 5, the longest video clip among the video clips with a basic frame extraction count of 2 will be adjusted to a basic frame extraction count of 3.
[0079] S402. Calculate the maximum number of keyframes to be extracted based on the duration of the video footage.
[0080] The maximum number of keyframes extracted from a video clip can be understood as the maximum number of keyframes extracted from that video clip that can be used in the one-click generation of large-scale and short videos.
[0081] In some embodiments, when a user selects only one video clip in the "One-Click Movie" feature interface displayed in the gallery application, the electronic device can calculate the maximum number of keyframes to be extracted based on the duration of that video clip. In other embodiments, when a user selects multiple video clips in the "One-Click Movie" feature interface displayed in the gallery application, the electronic device needs to calculate the maximum number of keyframes to be extracted for each selected video clip based on its duration. The sum of the maximum number of keyframes to be extracted from the multiple video clips is then used as the maximum number of keyframes to be extracted from the combined long video.
[0082] In some embodiments, the maximum number of frames extracted from keyframes based on the duration of the video footage must meet the following rules:
[0083] 1. Extract at least one keyframe from a video clip; if the video clip is longer than or equal to 2 seconds, extract 2 keyframes.
[0084] 2. For every 3 seconds of video footage, add one more keyframe, up to 30 seconds, for a total of 12 keyframes.
[0085] 3. If the video footage is longer than 30 seconds, for every 10 seconds of video content after 30 seconds, add 1 keyframe, up to 120 seconds, which is a total of 21 keyframes.
[0086] It should be noted that the maximum number of keyframes extracted from the video footage must also meet the requirement of obtaining a relatively accurate video theme as mentioned above.
[0087] S403. Determine whether historical frame extraction results of video material are stored, and whether the historical video analysis algorithm capability is equal to the current video analysis algorithm capability.
[0088] The historical frame extraction results of video footage include: the historical base number of keyframes extracted from the video footage and the historical maximum number of keyframes extracted; it may also include the analysis results of keyframes extracted from the video footage, such as the content of the keyframes and the picture quality evaluation results.
[0089] Before a user selects one or more video clips to generate a one-click blockbuster short video, they may have already performed the one-click blockbuster short video generation operation based on the currently selected video clips. For the user's historical one-click blockbuster short video generation operations, the electronic device can record the basic and maximum number of keyframe extractions from the selected video clips, and can also record the analysis results of the extracted keyframes. The electronic device can determine whether historical frame extraction results for the video clips are stored by reading the history table.
[0090] Electronic devices use video analysis algorithms to perform content and image quality analysis on keyframes extracted from video footage. The analysis results are then used to determine the keyframes needed to generate one-click large-scale short videos.
[0091] As electronic devices' operating systems and video editing applications are updated, the video analysis algorithms configured in those devices also need to be updated. Typically, after an update, the video analysis algorithm's processing power or stability is enhanced. After an update, the previous video analysis algorithm is called the historical video analysis algorithm, and the updated video analysis algorithm is called the current video analysis algorithm.
[0092] In some embodiments, the electronic device compares historical video analysis algorithms with the current video analysis algorithm. If it determines that the number of current video analysis algorithms is greater than the number of historical video analysis algorithms, meaning the electronic device can perform more functions based on the current video analysis algorithm, then the electronic device determines that the capabilities of the historical video analysis algorithms are lower than those of the current video analysis algorithms. For example, if there are three historical video analysis algorithms, the electronic device can analyze extracted keyframes based on these three algorithms in terms of clarity, aesthetics, and face recognition. The analysis result determines whether the keyframe is used to generate a short, high-quality video. The current video analysis algorithm, however, adds the function of recognizing pet faces. The current video analysis algorithm includes four algorithms, meaning the electronic device can use these four historical video analysis algorithms to determine whether the extracted keyframe is used to generate a short, high-quality video based on clarity, aesthetics, face recognition, and pet face recognition.
[0093] In other embodiments, if the electronic device determines that the current video analysis algorithm and the historical video analysis algorithm perform the same function, but the accuracy of the current video analysis algorithm is higher than that of the historical video analysis algorithm, the electronic device will determine that the capability of the historical video analysis algorithm is not equal to the capability of the current video analysis algorithm, that is, lower than the capability of the current video analysis algorithm.
[0094] If the electronic device determines that it has stored historical frame extraction results of video material and that the historical video analysis algorithm capability is equal to the current video analysis algorithm capability, then it executes step S404, followed by step S405 and subsequent steps; otherwise, it directly executes step S405 and subsequent steps.
[0095] S404. Update the base frame rate to base frame rate - historical base frame rate × 0.7, and update the maximum frame rate to maximum frame rate - historical maximum frame rate × 0.7.
[0096] For the video material selected by the user, the electronic device stores historical frame extraction results for that video material. Since the historical video analysis algorithm's capability is equal to the current video analysis algorithm's capability, it indicates that the video analysis algorithm has not been upgraded to update its capabilities. Furthermore, the historical frame extraction results obtained from the historical video analysis algorithm can be used in the current process of generating one-click large-scale short videos. Therefore, the electronic device can reduce the basic number of keyframes extracted from the video material obtained in step S401, and the maximum number of keyframes extracted from the video material obtained in step S402, to save frame extraction time and the analysis time for the extracted keyframes.
[0097] The electronic device can subtract the historical base frame rate from the base frame rate to obtain the updated base frame rate. Similarly, it can subtract the historical maximum frame rate from the maximum frame rate to obtain the updated maximum frame rate. Considering that the historical base frame rate and historical maximum frame rate in the historical frame rate results may be small, the electronic device multiplies both the historical base frame rate and historical maximum frame rate by a constant less than 1; 0.7 is an example of this constant. The electronic device subtracts the historical base frame rate multiplied by the constant from the base frame rate to obtain the updated base frame rate. Similarly, the electronic device subtracts the historical maximum frame rate multiplied by the constant from the maximum frame rate to obtain the updated maximum frame rate.
[0098] For the video footage selected by the user, since the electronic device does not store historical frame extraction results, it uses the base frame extraction count and maximum frame extraction count obtained in steps S401 and S402 to perform the following steps. Furthermore, after completing this one-click movie function, the electronic device can store the frame extraction results of the video footage selected by the user.
[0099] Alternatively, for the video material selected by the user, the electronic device stores historical frame extraction results. However, if the capability of the historical video analysis algorithm is not equal to that of the current video analysis algorithm, it indicates that the video analysis algorithm has been upgraded and its capabilities have been improved. Therefore, the historical frame extraction results are unsuitable for this "One-Click Movie" function. Thus, the electronic device uses the base frame extraction number and maximum frame extraction number obtained in steps S401 and S402 to perform the following steps. Furthermore, after the electronic device completes this "One-Click Movie" function, it can update the historical frame extraction results for the video material selected by the user based on the frame extraction results of this "One-Click Movie" function execution.
[0100] It should be noted that the use of historical frame extraction results obtained from historical video analysis algorithms in the generation of one-click blockbuster short videos in this process means that: the electronic device, based on the frame extraction duration configuration method provided in this embodiment, obtains the first round of frame extraction duration and number of frames for the video material selected by the user. Based on this frame extraction duration, the electronic device extracts key frames from the user-selected video material that meet the required number of frames. Then, the electronic device performs content analysis and image quality analysis on the extracted key frames using video analysis algorithms to obtain the analysis results. The electronic device comprehensively evaluates the analysis results of key frames from historical frame extraction results and the analysis results of the currently extracted key frames to determine the theme corresponding to the video material and the highlight segments of the video material (by first determining the key areas of the video material, and then determining the highlight segments within the key areas).
[0101] In some embodiments, steps S403 and S404 may be omitted.
[0102] S405, the frame extraction duration of keyframes is configured to be 30% of the total video analysis duration.
[0103] The keyframe extraction time refers to the time taken to extract the basic number of keyframes from the video footage. In some embodiments, when a user selects a single video clip, the keyframe extraction time refers to the time taken to extract the basic number of keyframes from the video clip. In other embodiments, when a user selects multiple video clips, the keyframe extraction time refers to the time taken to extract the basic total number of keyframes from the long video composed of the multiple video clips, where the basic total number of keyframes is the sum of the basic number of keyframes from the multiple video clips.
[0104] Typically, the frame extraction time from video footage accounts for only a small fraction of the total time it takes for an electronic device to generate a one-click blockbuster video. Ignoring the time spent downloading effects templates, compositing highlights, images, effects, and background music from the video footage, the total time it takes for an electronic device to generate a one-click blockbuster video can be understood as the total video analysis time. It can also refer to the time it takes for a user to select video or image footage, click "one-click blockbuster video," and wait for the electronic device to generate and display the video.
[0105] Since the frame extraction time for keyframes from video footage accounts for only a small fraction of the total time spent by an electronic device generating a short, high-quality video with a single click, electronic devices are typically configured to allocate less than 50% of the total video analysis time for keyframe extraction. For example, the frame extraction time for keyframes from video footage is configured to be 30% of the total video analysis time. This allows the electronic device to reserve a larger portion of the total video analysis time for subsequent, more detailed analysis.
[0106] In some embodiments, the electronic device can determine the total video analysis time based on the number of video clips selected by the user and the duration of the video clips. The determination rule is that the longer the total duration of the video clips selected by the user, the longer the total video analysis time will be. Of course, in order to prevent the user from waiting too long, the total video analysis time is configured with a maximum value, and the total video analysis time determined by the electronic device must not exceed the maximum value.
[0107] S406. Determine whether 30% of the total video analysis time is insufficient to extract the keyframes for the basic number of frame extractions.
[0108] As described in step S405, 30% of the total video analysis time is designated as the keyframe extraction time. Developers expect the electronic device to extract the basic number of keyframes within this timeframe. However, in reality, when the electronic device's performance is insufficient, it may be unable to extract the basic number of keyframes within 30% of the total video analysis time. Therefore, step S406 determines whether the electronic device will be unable to extract the basic number of keyframes within 30% of the total video analysis time.
[0109] In some embodiments, the electronic device determines whether 30% of the total video analysis time is insufficient to extract the basic number of keyframes in the following way:
[0110] Based on its own performance, the electronic device estimates the number of keyframes to be extracted within 30% of the total video analysis time. The electronic device compares this estimated number with the base number of extracted frames. If the estimated number is greater than the base number of extracted frames, it is determined that the electronic device is sufficient to extract the base number of keyframes within 30% of the total video analysis time; otherwise, it is determined that the electronic device is insufficient to extract the base number of keyframes within 30% of the total video analysis time.
[0111] When the user selects multiple video clips, the base frame extraction count in this step refers to the sum of the base frame extraction counts of the keyframes of the multiple video clips.
[0112] If the electronic device determines that 30% of the total video analysis time is insufficient to extract the key frames required for the basic number of frame extractions, it will proceed to step S407; otherwise, it will proceed to step S410.
[0113] S407. When it is determined that the historical video analysis algorithm capability is lower than the current video analysis algorithm capability, configure not to analyze the first video material, and the electronic device stores the historical frame extraction results of the first video material.
[0114] The electronic device determines that 30% of the total video analysis time is insufficient to extract the basic number of keyframes, indicating that the device's performance is inadequate to support extracting the required number of keyframes within that 30% timeframe. Therefore, the number of keyframes extracted by the electronic device within that 30% timeframe needs to be reduced. Based on this, the electronic device retrieves video footage containing stored historical frame extraction results from the user's currently selected video footage, designated as the first video footage. If the electronic device determines that the historical video analysis algorithm for the first video footage is less powerful than the current video analysis algorithm, it will not analyze the first video footage.
[0115] In some embodiments, for a first video clip configured not to be analyzed, the electronic device may reuse the historical frame extraction results of the first video clip without performing keyframe extraction and subsequent keyframe analysis on the first video clip.
[0116] In other embodiments, the electronic device should also configure the first video material to not be analyzed according to the following rules.
[0117] 1. For video footage configured not to be analyzed, the historical frame extraction results stored in the electronic device can indicate that the electronic device has completed the key frame analysis process for the key frames of the basic frame extraction number obtained in step S401 for the video footage.
[0118] For example, when an electronic device previously executed the "One-Click Large-Scale Short Video" function on a video clip, it extracted 10 keyframes from the video clip and performed subsequent keyframe analysis on these 10 keyframes. When the electronic device executes the "One-Click Large-Scale Short Video" function again based on the same video clip, the basic frame extraction count is 8, obtained through step S401. Therefore, based on the historical frame extraction results of the video clip, the electronic device determines that 8 keyframes have previously undergone keyframe analysis. Thus, this video clip can be configured not to be analyzed.
[0119] Conversely, when the electronic device executes the one-click blockbuster video function, it extracts three keyframes from the video footage and performs subsequent keyframe analysis on these three keyframes. In this instance, the electronic device obtains a basic frame extraction count of 8 through step S401. Based on the historical frame extraction results of this video footage, the electronic device determines that these 8 keyframes have not undergone keyframe analysis previously; therefore, this video footage cannot be configured to not be analyzed.
[0120] 2. In the first video material that satisfies 1, the electronic device randomly selects the first video material and configures it to not be analyzed.
[0121] The purpose of using a random selection method is to avoid situations where a specific selection method, such as prioritizing those with a large number of analyzed frames, results in some video materials being consistently configured not to be analyzed, and their historical frame extraction results not being updated for a long time.
[0122] 3. The electronic device selects one or more first video clips and configures them to not be analyzed until any one of the following conditions is met:
[0123] (1) 30% of the total video analysis time is sufficient to complete the basic frame extraction of the remaining video material; the remaining video material refers to the remaining video material in the user-selected video material excluding the first video material configured not to be analyzed.
[0124] (2) The total number of remaining video clips has been reduced to 50% of the original number (rounded up), or the first video clip satisfying 1 has been reduced to only 1; the original number refers to the number of video clips selected by the user.
[0125] The purpose of setting condition (2) is to avoid the situation where all or too many video materials are not analyzed, which would result in the historical frame extraction results of the video materials not being updated.
[0126] The method by which electronic devices determine whether the capabilities of historical video analysis algorithms are lower than those of current video analysis algorithms can be found in step S403 above, and will not be repeated here.
[0127] S408. Determine whether 30% of the total video analysis time is insufficient to extract the keyframes of the basic frame extraction number.
[0128] After executing step S407, the electronic device executes step S408 again to determine whether 30% of the total video analysis time is insufficient to extract the basic number of keyframes. In some embodiments, the electronic device compares the estimated number of keyframes extracted during 30% of the total video analysis time with the basic number of keyframes. If the estimated number of keyframes extracted during 30% of the total video analysis time is less than the basic number of keyframes, it indicates that 30% of the total video analysis time is insufficient to extract the basic number of keyframes. The basic number of keyframes in this step refers to the sum of the basic number of keyframes of the remaining video material. The remaining video material refers to the remaining video material selected by the user, excluding the first video material configured not to be analyzed.
[0129] If the electronic device determines that 30% of the total video analysis time is insufficient to extract the key frames required for the basic number of frame extractions, it will proceed to step S409; otherwise, it will proceed to step S410.
[0130] S409. Increase the keyframe extraction time until the time requirement for extracting the basic number of keyframes is met or until the total video analysis time is reached.
[0131] After the electronic device reduces the number of keyframes extracted within 30% of the total video analysis time in step S407, it determines in step S408 that 30% of the total video analysis time is insufficient to extract the basic number of keyframes. The electronic device can then increase the extraction time of the keyframes to test whether it can achieve sufficient extraction of the basic number of keyframes within 30% of the total video analysis time. The basic number of keyframes in this step also refers to the sum of the basic number of keyframes in the remaining video material. The remaining video material refers to the video material selected by the user, excluding the first video material configured not to be analyzed.
[0132] As shown in step S404, the initial frame extraction duration of the electronic device is configured to be 30% of the total video analysis time. Based on this, the electronic device gradually increases the frame extraction duration and determines whether it can extract the basic number of key frames based on the increased extraction duration. If the electronic device determines that it can extract the basic number of key frames, it stops increasing the frame extraction duration; otherwise, it continues to increase the extraction duration until the increased extraction duration meets the time requirement for extracting the basic number of key frames.
[0133] The electronic device continuously increases the keyframe extraction time until the keyframe extraction time reaches the total video analysis time. If the total video analysis time is still insufficient to extract the basic number of keyframes, the electronic device will no longer increase the keyframe extraction time and will configure the keyframe extraction time to the total video analysis time.
[0134] It should be noted that the keyframe extraction time configured on the electronic device is equal to the total video analysis time. This means that after the electronic device extracts keyframes from the video footage to obtain the key areas, it doesn't have enough time to analyze the highlight segments from those key areas. Instead, it directly obtains the highlight segments based on the keyframes. A potential problem with this is:
[0135] 1. Electronic devices directly obtain highlight segments based on keyframes of video footage. Because the process of analyzing highlight segments from key areas is missing, it may be impossible to extract highlight motion images from video footage to form highlight segments. For example, in video footage of a person jumping, a highlight motion image indicates the image of the highest point of the jump, or in video footage of someone playing, an image indicates the image of the most radiant smile.
[0136] 2. When electronic devices directly obtain highlight clips from keyframes of video footage, there may be some cases of poor image quality. For example, the keyframe extracted from the video footage may have good image quality, but the image quality of some video frames within a certain period before and after the keyframe may be very poor. Thus, the highlight clips obtained based on the keyframe and the video frames before and after it may have poor image quality. Another example is that the keyframe extracted from the video footage may have very poor image quality, but the image quality of the intermediate segment between the keyframes may be very good. The image of the intermediate segment may not be used as the image in the highlight clip, which also leads to poor image quality in the highlight clip.
[0137] The inventors discovered that problem 1 in the one-click video streaming function of electronic devices does not fundamentally affect the one-click video streaming feature. Compared to related technologies that only extract keyframes from a small portion of the video, the impact of this problem is negligible. The keyframes extracted by the electronic device using the embodiments of this application can basically indicate the overall meaning of the video material, balancing the user's waiting time and comprehensive understanding of the video. Furthermore, the inventors also found that problem 2 occurs relatively rarely. The reason for this is that with the development of shooting technology, the video captured by electronic devices can generally guarantee good image quality for each frame, thus reducing the probability of problem 2 occurring.
[0138] It should also be noted that when the frame extraction time of the electronic device is increased to the total video analysis time, or when it is increased to meet the time requirement for extracting the basic number of key frames, the electronic device performs the operation of extracting key frames from the video material using the increased frame extraction time of the key frames. In other words, the time taken for the electronic device to extract key frames from the video material is the increased frame extraction time of the key frames.
[0139] During the process of extracting keyframes from multiple video clips based on the total video analysis time, the electronic device obtains the actual number of keyframes extracted for each video clip based on the ratio between the number of keyframes extracted from the total video analysis time and the basic number of extracted frames for the multiple video clips obtained in step S401. The electronic device then extracts keyframes from the video clips according to the actual number of extracted frames for each video clip.
[0140] S410. Determine whether the maximum number of keyframes to be extracted is exceeded within 30% of the total video analysis time.
[0141] In practice, the electronic equipment performs well enough to extract the basic number of keyframes within 30% of the total video analysis time, with even some spare capacity. Based on this, the electronic equipment executes step S410.
[0142] Alternatively, after the electronic device reduces the number of keyframes extracted within 30% of the total video analysis time through step S407, the electronic device then determines through step S408 that there are enough keyframes to extract the basic number of keyframes within 30% of the total video analysis time. Based on this, the electronic device also executes step S410.
[0143] In some embodiments, the electronic device compares the estimated number of keyframes extracted during 30% of the total video analysis time with the maximum number of keyframes extracted. If the estimated number is greater than the maximum number of keyframes extracted, it is determined that the electronic device has extracted more keyframes than the maximum number of keyframes extracted during 30% of the total video analysis time; otherwise, it has not.
[0144] If the electronic device determines that the number of keyframes extracted exceeds the maximum number of keyframes for 30% of the total video analysis time, then step S411 is executed. If the electronic device determines that the number of keyframes extracted does not exceed the maximum number of keyframes for 30% of the total video analysis time, then the electronic device performs the operation of extracting keyframes from the video material, and the time taken for extracting keyframes is 30% of the total video analysis time.
[0145] S411. Reduce the keyframe extraction time until it just meets the time requirement for extracting the maximum number of keyframes.
[0146] If the electronic device determines that the number of keyframes extracted during 30% of the total video analysis time exceeds the maximum number of keyframes extracted, it means that the time taken by the electronic device to extract the basic number of keyframes from the video footage is less than 30% of the total video analysis time. In this way, the electronic device can reduce the keyframe extraction time and avoid wasting time.
[0147] In some embodiments, the electronic device gradually reduces the keyframe extraction time based on the initial keyframe extraction time, and determines whether the maximum number of keyframes to be extracted has been exceeded based on the reduced keyframe extraction time. If the electronic device determines that the maximum number of keyframes to be extracted has been exceeded, the electronic device continues to reduce the keyframe extraction time until the reduced keyframe extraction time exactly meets the time requirement for extracting the maximum number of keyframes.
[0148] It should be noted that the electronic device performs the operation of extracting keyframes from the video footage using the reduced keyframe extraction time. In other words, the time taken for the electronic device to extract keyframes from the video footage is the reduced keyframe extraction time.
[0149] Another embodiment of this application provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods.
[0150] Computer-readable storage media can be non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices.
[0151] Another embodiment of this application provides a computer program product containing instructions. When the computer program product is run on a computer or processor, it causes the computer or processor to perform one or more steps of any of the methods described above.
Claims
1. A method for generating video, characterized in that, include: The first interface is displayed, which shows a thumbnail of the video and a first button; In response to a user selecting at least one video on the first interface and clicking the first button, a first number of keyframes are extracted from the video within the extraction time of a first keyframe using a frame extraction method at first preset intervals. The extracted first keyframes are not continuous. The number of frames extracted and the extraction time of the first keyframes are configured before extracting the multiple first keyframes. The method of configuring the number of frames extracted and the extraction time of the first keyframes includes: determining a first number of frames extracted and a second number of frames extracted based on the duration of the video, wherein the second number of frames extracted is greater than the first number of frames extracted; using a portion of the total duration of the video analysis as the extraction time; and configuring the number of frames extracted as the first number of frames extracted and the extraction time as a portion of the total duration of the video analysis when it is determined that the extraction time is sufficient to extract the first number of keyframes from the video, and the extracted keyframes do not exceed the second number of frames extracted. The highlight segments of the video are obtained based on multiple first keyframes; The highlight segments and special effects of the video are combined to obtain a short video.
2. The video generation method according to claim 1, characterized in that, The highlight segment of the video is obtained based on the first keyframe, including: Based on the first keyframe, the key areas of the video are obtained; Multiple second key frames are extracted from the key area using a frame-skipping method with a second preset time interval. The second preset time interval is less than the first preset time interval, and the images of the multiple second key frames are continuous. Based on the second keyframe, the highlight fragment in the key area is obtained.
3. The video generation method according to claim 1, characterized in that, When the user selects a first video and a second video on the first interface, determining the first and second number of frames to extract the first keyframe based on the video duration, wherein the second number of frames is greater than the first number of frames, includes: Based on the duration of the first video, the third and fourth number of frames to be extracted from the first keyframe of the first video are determined, and based on the duration of the second video, the fifth and sixth number of frames to be extracted from the first keyframe of the second video are determined, wherein the fourth number of frames is greater than the third number of frames, and the sixth number of frames is greater than the fifth number of frames. The first number of frames extracted from the first keyframe is determined to be the sum of the third and fifth number of frames extracted, and the second number of frames extracted from the first keyframe is determined to be the sum of the fourth and sixth number of frames extracted.
4. The video generation method according to claim 3, characterized in that, The method further includes: If it is determined that there are not enough keyframes to be extracted from the first video and the second video within the frame extraction time, the frame extraction number of the first keyframe is configured as the fifth frame extraction number, and the frame extraction time is a portion of the total video analysis time. The step of extracting the first keyframe of the specified number of frames from the video within the specified frame extraction time using a frame extraction method with a first preset time interval includes: extracting the first keyframe of the specified number of frames from the second video within the specified frame extraction time using a frame extraction method with a first preset time interval. The step of obtaining the highlight segment of the video based on multiple first keyframes includes: obtaining the highlight segment of the second video based on the first keyframe of the fifth number of extracted frames from the second video, and obtaining the highlight segment of the first video based on the first keyframe in the historical frame extraction results of the first video.
5. The video generation method according to claim 4, characterized in that, After configuring the number of frames extracted from the first keyframe to be the fifth number of frames extracted, and the frame extraction duration to be a portion of the total video analysis duration, the method further includes: If it is determined that the frame extraction duration is insufficient to extract the fifth number of keyframes from the second video, the frame extraction duration is increased to a first frame extraction duration, wherein the first frame extraction duration is sufficient to extract the fifth number of keyframes from the second video. The frame extraction duration is configured to be the first frame extraction duration.
6. The video generation method according to claim 4, characterized in that, After configuring the number of frames extracted from the first keyframe to be the fifth number of frames extracted, and the frame extraction duration to be a portion of the total video analysis duration, the method further includes: If it is determined that the frame extraction duration is insufficient to extract the fifth number of keyframes from the second video, the frame extraction duration is configured as the total video analysis duration, where the total video analysis duration is either sufficient to extract the fifth number of keyframes from the second video, or insufficient to extract the fifth number of keyframes from the second video.
7. The video generation method according to claim 1, characterized in that, Also includes: If it is determined that the first number of keyframes can be extracted from the video within the frame extraction duration, and the number of keyframes extracted exceeds the second number of keyframes extracted, the frame extraction duration is reduced to a second frame extraction duration, wherein the first number of keyframes can be extracted from the video within the second frame extraction duration, and the number of keyframes extracted is substantially the same as the second number of keyframes extracted. The number of frames extracted from the first keyframe is set to the first number of frames extracted, and the duration of the frame extraction is set to the second duration of the frame extraction.
8. The video generation method according to claim 1, characterized in that, Based on the duration of the video, the method of determining the first and second number of extracted frames for extracting the first keyframe, wherein the second number of extracted frames is greater than the first number of extracted frames, further includes: When it is determined that the electronic device stores the historical frame extraction results of the video, and the video analysis algorithm capability of the electronic device has not been upgraded, the first frame extraction number is updated based on the seventh frame extraction number in the historical frame extraction results of the video, and the second frame extraction number is updated based on the eighth frame extraction number in the historical frame extraction results of the video; wherein, the updated first frame extraction number is less than the original first frame extraction number, and the updated second frame extraction number is less than the original second frame extraction number.
9. The video generation method according to claim 8, characterized in that, The step of updating the first frame count based on the seventh frame count in the historical frame-sampling results of the video, and updating the second frame count based on the eighth frame count in the historical frame-sampling results of the video, includes: Subtract the product of the seventh frame extraction number and the first weight from the first frame extraction number, and subtract the product of the eighth frame extraction number and the second weight from the second frame extraction number, wherein both the first weight and the second weight are less than 1.
10. An electronic device, characterized in that, include: One or more processors, memory, and a display screen; The memory and the display screen are coupled to the one or more processors. The memory is used to store a computer program, the computer program including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the video generation method as described in any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed, is specifically used to implement the video generation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Video processing method and device, equipment and storage medium
CN115766977A