Highlight segment extraction method and electronic equipment

By using a two-round frame extraction method to process video clips, the problem of large computational load and long processing time caused by frame-by-frame analysis is solved, the efficiency of human-computer interaction and the richness of highlight clips are improved, and the extraction of highlight clips is achieved efficiently.

CN121837992APending Publication Date: 2026-04-10HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing electronic devices require frame-by-frame video analysis to extract highlight segments, which is computationally intensive, time-consuming, and has low human-computer interaction efficiency. Furthermore, some analysis methods may result in concentrated highlight segments and monotonous content.

Method used

A two-round frame extraction method is used to process the video into scenes. First, N frames are extracted, and then a second round of frame extraction is performed between adjacent frames with similarity below a threshold. Scene information is calculated and highlight segments are extracted from multiple scenes.

Benefits of technology

It reduces computational load, lowers analysis time, improves human-computer interaction efficiency, and enhances the richness of highlight clips, preventing highlight clips from being concentrated in a single shot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837992A_ABST
    Figure CN121837992A_ABST
Patent Text Reader

Abstract

The invention discloses a highlight segment extraction method and electronic equipment, and relates to the technical field of terminals. The method comprises the following steps: executing a first round of frame extraction on a to-be-analyzed video, extracting N frames, N being greater than or equal to 2, and N being an integer; executing a second round of frame extraction between adjacent frames of which the similarity is lower than a first similarity threshold value in the N frames, and extracting at least one frame; based on the frame extraction results of the first round of frame extraction and the second round of frame extraction, calculating the split mirror information of the to-be-analyzed video, the split mirror information comprising a plurality of split mirrors; a highlight segment is extracted from at least two of the plurality of sub-mirrors. In this way, the operand can be reduced, the man-machine interaction efficiency can be improved, and the richness of the highlight segment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of terminal, and in particular, to a highlight segment extraction method and an electronic device. BACKGROUND

[0002] Currently, some electronic devices provide a function of extracting a highlight segment from a video, and the extracted highlight segment can be used for video secondary creation, sharing, etc.

[0003] However, when extracting a highlight segment, the electronic device usually needs to analyze a large number of video frames included in the video one by one, and then extract the highlight segment based on the analysis result. This way of analyzing frame by frame has a large amount of calculation, a long time-consuming analysis process, and low human-computer interaction efficiency. SUMMARY

[0004] The present application provides a highlight segment extraction method and an electronic device, which can perform split-screen processing on a video based on two rounds of frame extraction, and then extract a highlight segment, thereby reducing the amount of calculation, improving human-computer interaction efficiency, and improving the richness of the highlight segment.

[0005] To achieve the above object, embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, the present application provides a highlight segment extraction method, comprising: performing a first round of frame extraction on a video to be analyzed, extracting N frames, N≥2, N being an integer. Performing a second round of frame extraction between adjacent frames with a similarity lower than a first similarity threshold in the N frames, extracting at least one frame. Calculating split-screen information of the video to be analyzed based on the frame extraction results of the first and second rounds of frame extraction, the split-screen information including a plurality of split screens. Extracting a highlight segment from at least two split screens in the plurality of split screens.

[0007] To sum up, on the one hand, the extraction of the highlight segment can be completed based on the analysis of the two rounds of frame extraction, without the need for frame-by-frame analysis, thereby reducing the amount of calculation, reducing the time-consuming analysis, and improving the human-computer interaction efficiency. On the other hand, the two rounds of frame extraction can approximate the boundaries of the scene split screens, and the highlight segment can be extracted from at least two split screens subsequently, thereby improving the richness of the highlight segment.

[0008] In a possible implementation manner of the first aspect, before performing the first round of frame extraction on the video to be analyzed to extract N frames, the method further comprises: calculating a minimum frame extraction number and an actual frame extraction number of the video to be analyzed, the minimum frame extraction number representing a theoretically minimum frame extraction number of the video to be analyzed, and the actual frame extraction number representing an actual frame extraction number of the video to be analyzed. And in the case that the actual frame extraction number exceeds the minimum frame extraction number, such as case B in the following, N is the minimum frame extraction number, and the total number of frames of the second round of frame extraction does not exceed the difference between the actual frame extraction number and the minimum frame extraction number.

[0009] That is, in the case that the actual frame number of the video to be analyzed is more than the minimum frame number, the first round of frame extraction can extract the minimum frame number, thereby meeting the requirement of the minimum frame number and ensuring the effect of extracting the highlight segment, and the second round of frame extraction can extract the remaining frames of the actual frame number (i.e., the difference between the actual frame number and the minimum frame number), thereby avoiding too many frames from affecting the efficiency and ensuring that the time consumption of frame extraction is within a certain range.

[0010] In a possible implementation of the first aspect, the method further includes: in the case that the actual frame number does not exceed the minimum frame number, as Case A below, N is the actual frame number, and the second round of frame extraction is not performed on the video to be analyzed.

[0011] That is, in the case that the actual frame number of the video to be analyzed is less than the minimum frame number, the first round of frame extraction can extract the actual frame number, thereby extracting as many frames as possible when the minimum frame number cannot be met. Correspondingly, there is no remaining frame for the second round of frame extraction, and thus the second round of frame extraction is not performed.

[0012] In a possible implementation of the first aspect, the minimum frame number is related to the length of the video to be analyzed.

[0013] Generally, the longer the length of the video to be analyzed is, the more the minimum frame number is. In particular, after the length of the video to be analyzed exceeds a certain length, the minimum frame number can remain a relatively large value. In this way, the longer the length of the video to be analyzed is, the more frame extraction can be used to calculate the shot information.

[0014] In a possible implementation of the first aspect, before calculating the minimum frame number and the actual frame number of the video to be analyzed, the method further includes: receiving an extraction requirement of the highlight segment, the extraction requirement including a time consumption requirement and all videos to be analyzed. For example, in response to a triggering operation of "one-key big video" 1045 in the interface 104 shown in FIG. 1 or "one-key big video" 2031 in the interface 203 shown in FIG. 2, the extraction requirement of the highlight segment can be generated, and the extraction requirement can include the time consumption requirement specified by the gallery application and the video to be analyzed selected by the user. Figure 1 Figure 2 In a possible implementation of the first aspect, the method further includes: calculating a maximum supported frame number (which can be referred to as a maximum supported frame number for short in the following embodiments) of the electronic device based on the time consumption requirement. The maximum supported frame number is allocated to each video to be analyzed to obtain the actual frame number of each video to be analyzed.

[0015] In this way, each video to be analyzed is frame-extracted according to the actual frame number, and the frame extraction of all videos to be analyzed can meet the time consumption requirement.

[0016] ​In a possible implementation manner of the first aspect, the maximum supported number of frames is allocated to each video to be analyzed to obtain an actual number of frames of each video to be analyzed, including: in a case where the maximum supported number of frames is less than a sum of minimum numbers of frames of all videos to be analyzed, as case 1 below, a preset number of frames is allocated to each video to be analyzed first, and a remaining number of frames in the maximum supported number of frames is allocated according to a time length proportion of all videos to be analyzed to obtain the actual number of frames of each video to be analyzed, so that the actual number of frames allocated to each video to be analyzed matches a time length of each video to be analyzed.

[0017] In a case where the maximum supported number of frames is greater than or equal to a sum of minimum numbers of frames of all videos to be analyzed and less than a sum of maximum numbers of frames of all videos to be analyzed, as case 2 below, a minimum number of frames is allocated to each video to be analyzed first, and a remaining number of frames in the maximum supported number of frames is allocated according to a time length proportion of all videos to be analyzed to obtain the actual number of frames of each video to be analyzed, so that the actual number of frames allocated to each video to be analyzed matches a time length of each video to be analyzed on the premise of meeting the minimum number of frames.

[0018] In a case where the maximum supported number of frames is greater than or equal to a sum of maximum numbers of frames of all videos to be analyzed, as case 3 below, a maximum number of frames is allocated to each video to be analyzed to obtain the actual number of frames of each video to be analyzed.

[0019] In a possible implementation manner of the first aspect, the shot information of the video to be analyzed is calculated based on the frame extraction results of the first round of frame extraction and the second round of frame extraction, including: merging continuous frames with a similarity higher than a second similarity threshold (similarity threshold 3 below) in the first round of frame extraction and the second round of frame extraction to obtain a consistent region. Wherein, the frames not merged are recorded as isolated frames. The adjacent consistent regions are merged, and the consistent regions and the isolated frames are merged to obtain a noisy consistent region. Wherein, the continuous isolated frames not merged constitute a non-stability region. The shot information is obtained based on the noisy consistent region and the non-stability region. In this way, the shot information of the video to be analyzed can be obtained through the analysis of the two rounds of frame extraction, thereby facilitating the subsequent extraction of highlight segments.

[0020] In a possible implementation manner of the first aspect, the adjacent consistent regions are merged, and the consistent regions and the isolated frames are merged to obtain the noisy consistent region, including:

[0021] If the similarity between adjacent consistent regions is higher than the third similarity threshold (similarity threshold 4 below), or if the similarity between adjacent consistent regions is higher than the third similarity threshold and the time interval between adjacent consistent regions is less than the first duration (duration 3 below), the adjacent consistent regions are merged to obtain a noisy consistent region. In other words, two consistent regions with high similarity or a small time interval can be merged to obtain a noisy consistent region composed of the two consistent regions and the region between them.

[0022] If the similarity between a consistent region and a lonely frame is higher than the fourth similarity threshold, or if the similarity between a consistent region and a lonely frame is higher than the fourth similarity threshold (similarity threshold 5 below) and the time interval between the consistent region and the lonely frame is less than the second duration (duration 4 below), the consistent region and the lonely frame are merged to obtain a noisy consistent region. In other words, consistent regions with high similarity, or with a further small time interval, can be merged with lonely frames to obtain a noisy consistent region composed of the consistent region, the lonely frame, and the regions between them.

[0023] In one possible implementation of the first aspect, the first duration includes the sum of the region lengths of adjacent consistent regions.

[0024] In one possible implementation of the first aspect, extracting highlight segments from at least two of a plurality of shots includes: extracting highlight segments, including an optimal frame, from at least two shots according to the priority of the plurality of shots. The optimal frame includes preset content.

[0025] In other words, the higher the priority of a storyboard, the earlier it will be considered when extracting highlight clips; the lower the priority, the later it will be considered. This allows for the extraction of highlight clips, including the optimal frame, from higher-priority storyboards, ensuring the rationality of the extracted highlight clips. Furthermore, extracting highlight clips from at least two storyboards avoids the concentration of highlight clips in a single storyboard, thus ensuring the richness of the highlight clips.

[0026] In one possible implementation of the first aspect, the multiple scenes include at least two of the following: persistently stable scenes, non-persistently stable scenes, abruptly changing scenes, and alternative scenes. The priority of persistently stable scenes, non-persistently stable scenes, and alternative scenes decreases sequentially.

[0027] In other words, highlight clips can be extracted first from continuously stable storyboards, then from non-continuously stable storyboards, and finally from alternative storyboards. This ensures that highlight clips are extracted first from storyboards with long-term, stable shooting of highly similar content, then from storyboards with short-term, continuous shooting of highly similar content, and finally from storyboards with continuous shooting of low-similarity content. This guarantees the rationality of highlight clip extraction.

[0028] In one possible implementation of the first aspect, the preset content includes at least one of the following: a face, a body, and a voice. This allows for the extraction of highlight fragments from image frames related to the person.

[0029] In one possible implementation of the first aspect, the method further includes: performing aesthetic scoring on frames other than the optimal frame. After extracting highlight segments including the optimal frame from at least two storyboards according to the priority of multiple storyboards, the method further includes: extracting highlight segments in descending order of aesthetic score.

[0030] Understandably, if none of the shots contain the optimal frame, highlight fragments cannot be extracted from the optimal frame. Alternatively, if only one shot contains the optimal frame, the highlight fragments extracted from the optimal frame will be concentrated in that single shot. In these cases, further extracting highlight fragments based on aesthetic scores can increase the richness of the extracted highlight fragments, avoid having too few highlight fragments, and reduce the possibility of highlight fragments being concentrated in a single shot.

[0031] In a second aspect, an electronic device is provided, comprising a memory and one or more processors. The memory is coupled to the processors. The memory stores computer program code, which includes instructions. When the instructions are executed by the processor, the electronic device performs the method as described in any one of the first aspects above.

[0032] Thirdly, a computer-readable storage medium is provided that stores instructions which, when executed on an electronic device, cause the electronic device to perform any of the methods described in the first aspect.

[0033] Fourthly, a computer program product containing instructions is provided, which, when run on an electronic device, enables the electronic device to perform the method described in any one of the first aspects above.

[0034] Fifthly, embodiments of this application provide a chip, the chip including a processor, the processor being configured to invoke a computer program in memory to execute the method as described in the first aspect.

[0035] It is understood that the beneficial effects of the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect can be referred to the beneficial effects of the first aspect and any of its possible design embodiments, which will not be repeated here. Attached Figure Description

[0036] Figure 1 One of the interface interaction diagrams for one-click movie streaming provided in the embodiments of this application;

[0037] Figure 2 The second interface interaction diagram for the one-click blockbuster movie provided in the embodiments of this application;

[0038] Figure 3 This is a schematic diagram illustrating the effect of extracting highlight fragments;

[0039] Figure 4 This is a schematic diagram illustrating the effect of extracting highlight fragments provided in an embodiment of this application;

[0040] Figure 5 A schematic diagram of the two-round frame extraction process provided in an embodiment of this application;

[0041] Figure 6 A schematic diagram of the obtained consistent region provided in the embodiments of this application;

[0042] Figure 7 A schematic diagram illustrating the merging of adjacent consistent regions to obtain a noisy consistent region, as provided in an embodiment of this application.

[0043] Figure 8 A schematic diagram illustrating the noisy consistency region obtained by merging the consistency region and isolated frames, as provided in an embodiment of this application.

[0044] Figure 9 A complete example diagram of the segmentation provided for embodiments of this application;

[0045] Figure 10 A flowchart illustrating the segmentation of scenes provided in this application embodiment;

[0046] Figure 11 A hardware structure diagram of the electronic device provided in the embodiments of this application;

[0047] Figure 12 A software architecture diagram of the electronic device provided in the embodiments of this application;

[0048] Figure 13 A temporal interaction diagram of the highlight fragment extraction method provided in the embodiments of this application;

[0049] Figure 14 This is a structural diagram of the chip system provided in an embodiment of this application. Detailed Implementation

[0050] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0051] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.

[0052] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0053] The highlight segment extraction method provided in this application embodiment can be applied to scenarios where highlight segments are extracted from videos.

[0054] In some scenarios, electronic devices can extract highlight segments from one or more video clips and use them for secondary video creation, such as creating a video composed of spliced ​​highlight segments.

[0055] In some scenarios, electronic devices can extract highlight clips from a video and store them in a gallery, which can then be shared with friends later.

[0056] This article primarily uses the extraction of highlight clips for video editing as an example. This function of extracting highlight clips for video editing can also be referred to as "one-click blockbuster" (or "one-click video creation") on some electronic devices. For simplicity, the following explanation will use "one-click blockbuster" as an example.

[0057] In the "One-Click Blockbuster" feature, the electronic device first decodes the user-selected video, then analyzes it, and extracts highlight segments based on the analysis results. Finally, the electronic device stitches together all the extracted highlight segments into a single, finished video.

[0058] In one specific implementation, the electronic device has image processing applications installed, such as gallery apps and video editing apps. These applications can provide the aforementioned ability to create blockbuster movies with a single click.

[0059] The following example uses a mobile phone as an electronic device and a photo gallery application that provides the ability to take one-click photos to great pictures to illustrate two types of interface interaction processes for implementing one-click photos in a photo gallery application:

[0060] First type of interface interaction process

[0061] See Figure 1 The phone's desktop 101 includes an application icon 1011 for the Gallery app. In response to a user's click on the application icon 1011, the phone can run the Gallery app in the foreground. The Gallery app can provide menu pages including, but not limited to, "Photos," "Albums," "Memories," and "Creative." After the Gallery app is running in the foreground, the phone can display interface 102, which is the menu page for the "Creative" menu. Interface 102 includes "One-Click Blockbuster" 1021, used to trigger the selection of videos and / or images for One-Click Blockbuster.

[0062] See also Figure 1 In response to the user's click on "One-Click Movies" 1021 in interface 102, the phone can display interface 103, which includes video and image options in the gallery application. In response to the user's selection of one or more videos or images in interface 103, the phone can select the corresponding video and / or image.

[0063] See also Figure 1After selecting two videos, the phone displays interface 104. In interface 104, the selection boxes in the lower right corners of videos 1041, 1042, and 1043 are all selected, indicating that videos 1041, 1042, and 1043 are selected. Interface 104 also includes a pop-up window 1044. Pop-up window 1044 includes the selected videos, such as videos 1041, 1042, and 1043. Pop-up window 1044 also includes "One-Click Blockbuster" 1045, which triggers the phone to perform one-click blockbuster material (videos and / or images) analysis and highlight extraction. In response to the user clicking on "One-Click Blockbuster" 1045 in pop-up window 1044, the phone displays interface 105, which includes pop-up window 1051. Pop-up window 1051 includes the text "Intelligently analyzing materials," indicating that the selected videos and / or images are being analyzed. After completing the analysis and generating a new video, the phone can display interface 106, which includes a new video 1061. Video 1061 is composed of highlight clips from selected videos (such as video 1041, video 1042, and video 1043) and selected images.

[0064] The second type of interface interaction process

[0065] In response to a user's long-press selection of any image or video in the "Photos" menu or "Album" menu of the Gallery app, the phone can pop up a "One-Click Blockbuster" control to trigger the phone to perform one-click blockbuster material (video and / or image) analysis, highlight extraction and other processing.

[0066] See Figure 2 The phone can display interface 201, which is the menu page for the "Album" menu. Interface 201 includes multiple albums such as "Camera" 2011, "All Photos" 2012, and "Videos" 2013. In response to a user's action on "Videos" 2013, the phone can display interface 202. Interface 202 is the album page for "Videos" 2013, and includes multiple videos from "Videos" 2013. In response to a user's selection and long-press action on video 2021 in interface 202, the phone can display interface 203. In interface 203, video 2021 is selected, and interface 203 includes multiple operation items such as "Share," "Favorite," "Edit," and "Delete," as well as "One-Click Movie" 2031. "One-Click Movie" 2031 is a control for creating one-click movies. In response to the user's click on "One-Click Movie" 2031 in interface 203, the mobile phone can also complete the one-click movie by referring to the response process of similar interfaces 105 and 106.

[0067] However, during the analysis of "One-Click Movies," whether triggered by a user clicking the "One-Click Movies" button 1045 on interface 104 or by a user clicking the "One-Click Movies" button 2031 on interface 203, electronic devices typically analyze the selected video frame by frame. Frame-by-frame analysis involves analyzing each frame to identify highlight segments. This method is computationally intensive, time-consuming, and inefficient in human-computer interaction.

[0068] Taking the one-click movie feature as an example, if the user selects a 30-second video with a frame rate of 30 frames per second, then the video has a total of 900 frames. If the electronic device analyzes each of the 900 frames one by one, the user will have to wait for a long time. If the user stays on the interface 105 mentioned above for a long time, the human-computer interaction efficiency will be low.

[0069] Furthermore, when a user selects a long video or multiple videos, using frame-by-frame analysis may significantly increase the power consumption of electronic devices and affect their performance, such as causing overheating and slow operation.

[0070] In some improved approaches, electronic devices can analyze only a few segments or frames of a video to balance the human-computer interaction effect with the extraction of highlight segments, such as achieving the final effect of a one-click blockbuster video. However, this partial analysis method is highly likely to result in a concentrated distribution of highlight segments extracted from a video, leading to a lack of diversity in the content of these highlight segments.

[0071] See Figure 3 For a video segment that transitions from outdoors (e.g., frames 1-150) to indoors (e.g., frames 151-200) and back to outdoors (e.g., frames 201-300), using the partial analysis method described above, the electronic device might only analyze a few frames from the initial outdoor scene sequence. Thus, the highlight fragments extracted by the electronic device based on the analysis results could be highlight fragment 1, highlight fragment 2, and highlight fragment 3, all concentrated within frames 1-150, i.e., concentrated within the initial outdoor scene sequence.

[0072] To address the aforementioned issues, this application provides a method for extracting highlight segments. An electronic device can extract frames from a video, analyze the extracted frames, and determine the video's scene composition information based on the analysis results. Finally, highlight segments are extracted from at least two scenes. This approach eliminates the need for frame-by-frame analysis, reducing computational load and analysis time, thus improving human-computer interaction efficiency. Furthermore, extracting highlight segments from at least two scenes enhances the richness of the highlight segments.

[0073] SeeFigure 4 A video that transitions from outdoors (e.g., frames 1-150) to indoors (e.g., frames 151-200) and back to outdoors (e.g., frames 201-300) can be processed using the method provided in this application. By extracting and analyzing frames, highlight fragment a can be extracted from the outdoor scene shot at the beginning of the shoot, highlight fragment b can be extracted from the indoor scene shot, and highlight fragment c can be extracted from the outdoor scene shot at the end of the shoot.

[0074] The specific implementation of the highlight fragment extraction method provided in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0075] The highlight segment extraction method provided in this application mainly includes the following steps: Step 1, Calculation of the expected number of frames extracted. Step 2, Calculation of the actual number of frames extracted. Step 3, Frame extraction. Step 4, Segmentation calculation. Step 5, Highlight segment extraction.

[0076] For example, in the One-Click Movie feature, steps 1-5 can all be completed during the process of the electronic device displaying the interface 105.

[0077] The following is a detailed explanation of each of the steps 1-5 above.

[0078] Step 1: Calculate the expected number of frames to be drawn.

[0079] The expected frame rate refers to the theoretically expected number of frames to be extracted from the video being analyzed. The expected frame rate is primarily related to the duration of the video. Generally, the longer the video, the higher the expected frame rate, which is necessary for boundary analysis.

[0080] The video to be analyzed can be any video from which the electronic device needs to extract highlight segments. For example, the video to be analyzed can be any video selected by the user in the "One-Click Blockbuster" app, such as video 1041, video 1042, and video 1043 in the interface 104 above.

[0081] The expected frame rate includes the maximum expected frame rate and the minimum expected frame rate. The minimum expected frame rate is greater than or equal to 2, and the maximum expected frame rate is greater than or equal to the minimum expected frame rate. Both the maximum and minimum expected frame rates are integers.

[0082] For any video to be analyzed, if too many frames are extracted, the computational load of the frame extraction process (as in step 3 below) and the subsequent analysis process (as in step 4) will be higher. By controlling the maximum expected number of frames extracted, the maximum number of frames can be controlled, thereby controlling the computational load and reducing the computation time.

[0083] If the number of frames extracted is too small for any video to be analyzed, the results of subsequent analysis processes may be inaccurate, such as the inaccurate storyboard results, which will affect the accuracy of highlight segment extraction. By using the minimum expected number of frames extracted, the minimum number of frames extracted can be controlled, thereby ensuring the accuracy of highlight segment extraction.

[0084] For example, the correspondence between the duration of the video to be analyzed (which can be denoted as t) and the maximum expected number of frames extracted (which can be denoted as Nmax) and the minimum expected number of frames extracted (which can be denoted as Nmin) includes the following three cases:

[0085] Case a) t≤2 seconds.

[0086] The maximum and minimum number of frame drops are both 2. That is, Nmax = Nmin = 2. For example, if t = 1 second, Nmax = Nmin = 2.

[0087] Case b) 2 seconds < t ≤ 30 seconds.

[0088] The maximum number of frames to be extracted is one frame every 2 seconds, i.e., Nmax = (t / 2 + 2) frames, where "2" in "t / 2" represents a 2-second interval and "2" in "+2" represents adding one frame at the beginning and end.

[0089] The minimum number of frames to be extracted is one frame every 3 seconds, i.e., Nmin = (t / 3 + 2) frames, where "3" in "t / 3" represents a 3-second interval and "2" in "+2" represents adding one frame at the beginning and end.

[0090] Case c), 30 seconds < t.

[0091] Regarding the minimum number of frames drawn:

[0092] The minimum number of frames to be sampled is 5 seconds for the portion of the video to be analyzed that is between 30 seconds (excluding) and 90 seconds (including), i.e., Nmin = (t – 30) / 5 + 12. Here, "30" in "(t – 30) / 5" refers to the portion of the video to be analyzed that is less than or equal to 30 seconds, "5" refers to the 5-second interval, and "12" in "+12" refers to the minimum number of frames to be sampled for the portion of the video to be analyzed that is less than or equal to 30 seconds.

[0093] Alternatively, if 90 seconds < t, the minimum number of frames drawn is equal to 24 frames.

[0094] Regarding the maximum number of frame drops:

[0095] The maximum number of frames to be sampled is between 30 seconds (excluding) and 120 seconds (including), with one frame sampled every 5 seconds, i.e., Nmax = (t – 30) / 5 + 17. Here, "30" in "(t – 30) / 5" refers to the portion of the video to be analyzed that is less than or equal to 30 seconds, "5" refers to the 5-second interval, and "17" in "+17" refers to the maximum number of frames to be sampled within the portion of the video to be analyzed that is less than or equal to 30 seconds.

[0096] Alternatively, if 120 seconds < t, the maximum number of frame drops is equal to 35 frames.

[0097] Taking a video of 30 seconds as an example, the electronic device can detect that it belongs to the above situation b). Therefore, the Nmin = (30 / 3+2) = 12 frames and Nmax = (t / 2+2) = 17 frames of the video to be analyzed are calculated.

[0098] The above relationship between the duration of the video to be analyzed and the maximum and minimum number of frames extracted is merely illustrative and is used as an example in this article. In actual implementation, it is not a limitation. Generally, satisfying the following principles while balancing computational load and frame extraction effect is sufficient: the maximum number of frames extracted from the same video to be analyzed is greater than or equal to the minimum number of frames extracted; among different videos to be analyzed, the maximum number of frames extracted from the longer video to be analyzed is greater than or equal to the maximum number of frames extracted from the shorter video to be analyzed, and the minimum number of frames extracted from the longer video to be analyzed is greater than or equal to the minimum number of frames extracted from the shorter video to be analyzed.

[0099] Therefore, it should be noted that the above description of step 1 only illustrates the calculation process for the expected number of frames extracted from one video to be analyzed. In actual implementation, if there are multiple videos to be analyzed, such as multiple videos selected by the user for one-click video generation, the electronic device can calculate the expected number of frames extracted for each video to be analyzed according to step 1 above, which will not be elaborated upon in this article.

[0100] Step 2: Calculate the actual number of frames extracted.

[0101] The actual frame extraction count refers to the number of frames actually extracted from the video being analyzed. The actual frame extraction count is primarily related to the time requirements in the actual application scenario. Generally, the shorter the time requirement, such as needing to extract highlight segments and generate video within a very short time, the smaller the actual frame extraction count can be allocated.

[0102] The time requirement can be specified by the application that initiates the highlight segment extraction request (such as providing one-click blockbuster).

[0103] Taking the "One-Click Movie" feature as an example, after a user selects a video and triggers the analysis of the footage, within the time interval of the aforementioned interface 105 on the phone, the gallery / video editing app can determine the required time based on information such as the number and duration of the videos selected by the user, such as a time limit of no more than 15 seconds. If the user selects a large number of videos or a long duration, a lower time requirement can be determined, meaning the duration can be longer. Of course, the application can limit the upper limit of efficiency requirements, meaning the time requirement will not increase indefinitely with the number or duration of videos.

[0104] Of course, the time requirement can also be specified by the user. For example, after selecting a video but before triggering the analysis of the "One-Click Blockbuster" footage, such as before clicking the "One-Click Blockbuster" 1045 in the aforementioned interface 104, the user can input their time requirement. In other words, the electronic device can provide an entry point for inputting efficiency requirements. In response to the user's input of the time requirement, the electronic device can obtain the user's specified time requirement. For example, if the user wants to extract highlight segments faster, such as wanting to generate a "One-Click Blockbuster" video faster, they can input a shorter time duration. In this way, the electronic device can flexibly decide the actual number of frames extracted based on the user's time requirement.

[0105] Electronic devices can calculate the maximum number of frame extractions they can support based on time consumption requirements (hereinafter referred to as the maximum supported frame extraction number), and allocate an actual number of frame extractions to the video to be analyzed based on this maximum supported frame extraction number, so that extracting frames with this actual number of frame extractions can meet the time consumption requirements.

[0106] In some embodiments, the electronic device may first allocate the time required for frame extraction based on the time consumption requirement, and then calculate the actual number of frames extracted from the video to be analyzed based on the time required for frame extraction.

[0107] Understandably, after frame extraction, the electronic device still needs to perform subsequent processing, such as shot calculation and highlight extraction. The time consumed by these processing processes is included in the time consumption requirement. Therefore, the electronic device can allocate the time consumption of frame extraction based on the time consumption requirement, that is, determine the time consumed by the frame extraction process, so that the frame extraction and subsequent processing as a whole can meet the time consumption requirement.

[0108] For example, an electronic device can allocate the frame extraction time according to a preset ratio. For instance, if the time requirement is 15 seconds and the preset ratio is 70%, then the frame extraction time will be approximately 10 seconds.

[0109] Furthermore, the maximum supported frame rate can be calculated based on the performance of the electronic device. This performance includes the video decoding speed and the frame rate extraction algorithm speed. In this way, the maximum supported frame rate aligns with the actual performance of the electronic device, resulting in a more accurate calculated actual frame rate.

[0110] Taking a frame extraction time of 10 seconds as an example, based on the decoding speed and algorithm speed, the electronic device can calculate that the maximum number of frames it supports within 10 seconds is 120. Subsequently, the electronic device can allocate the actual number of frames to be extracted for the video to be analyzed based on 120 frames.

[0111] In some embodiments, the electronic device can combine the maximum number of frame extractions and the minimum number of frame extractions calculated in step 1 above to allocate the maximum supported number of frame extractions to each video to be analyzed, thereby obtaining the actual number of frame extractions for each video to be analyzed. This ensures that the actual number of frame extractions can meet the minimum number of frame extractions as much as possible, thus guaranteeing the effect of frame extraction and ensuring that the actual number of frame extractions does not exceed the maximum number of frame extractions, thereby avoiding excessive computation.

[0112] Understandably, if there are multiple videos to be analyzed, such as multiple videos included in the one-click video creation tool selected by the user, the maximum supported frame extraction number must be allocated to multiple videos to be analyzed. Taking a total of 5 videos to be analyzed and a maximum supported frame extraction number of 120 frames as an example, the electronic device must allocate 120 frames to the 5 videos to be analyzed. For example, the actual frame extraction numbers allocated to the 5 videos to be analyzed would be 30 frames, 10 frames, 50 frames, 20 frames, and 10 frames respectively.

[0113] For example, the electronic device allocates the actual number of frames to be extracted for each video to be analyzed as follows:

[0114] Case 1: The maximum number of frames that can be extracted is less than the sum of the minimum number of frames extracted from all the videos to be analyzed.

[0115] In this case, the maximum supported frame rate cannot support each video to be analyzed to have at least the minimum frame rate. The electronic device can first assign a preset number of frames to each video to be analyzed, such as 2 frames.

[0116] For example, if the maximum supported frame extraction number is 60 frames, and there are 5 videos to be analyzed, with minimum frame extraction numbers of 10, 15, 18, 20, and 20 frames respectively, then the maximum supported frame extraction number of 60 frames is less than the sum of the minimum frame extraction numbers of all the videos to be analyzed, which is 83 frames. The electronic device can first allocate 2 frames to each of the 5 videos to be analyzed.

[0117] Furthermore, after allocating a preset number of frames to each video to be analyzed, if there are still frames remaining after the maximum supported number of frame extractions, the electronic device can allocate the remaining frames according to the duration ratio of all videos to be analyzed, so that the actual number of frame extractions allocated to each video to be analyzed matches its respective duration.

[0118] Continuing with the example in Case 1 above, the maximum number of frames supported is 60. After allocating 2 frames to each of the 5 videos to be analyzed, there are 60 - (5 * 2) = 50 frames remaining. The electronic device will allocate the 50 frames according to the duration ratio of the 5 videos to be analyzed.

[0119] Understandably, when allocating frames according to duration, the longer the duration, the more frames are allocated, and the shorter the duration, the fewer frames are allocated.

[0120] After allocating frames according to the duration ratio, the electronic device can obtain the actual number of frames extracted from each video to be analyzed as the sum of the preset number of frames and the number of frames allocated according to the duration ratio.

[0121] Furthermore, if the duration of a video to be analyzed is too high, the sum of the number of frames allocated to that video according to its duration and the preset number of frames may exceed the minimum number of frames to be analyzed. In this case, the electronic device can reduce the number of frames to be analyzed to the minimum number. This ensures a balance in the number of frames to be analyzed across multiple videos. On the other hand, the actual number of frames to be analyzed that exceeds the minimum number needs to be extracted in the second round (see step 3 below for details), which takes a long time. Therefore, controlling the number of frames to be extracted to not exceed the minimum number can control the time spent on frame extraction.

[0122] For example, if the preset number of frames is 2, and the number of frames allocated to the video to be analyzed according to the duration ratio is 17, and the minimum number of frames to be extracted from the video to be analyzed is 18, then obviously (2+17) frames exceeds 18 frames. The electronic device can adjust the actual number of frames extracted from the video to be analyzed from (2+17) frames to 18 frames.

[0123] Understandably, after allocating a preset number of frames to each video to be analyzed, if there are no frames left to be extracted from the maximum supported number of frames, the electronic device does not need to continue allocating frames. Accordingly, the actual number of frames extracted from each video to be analyzed is the preset number of frames.

[0124] Case 2: The maximum supported frame extraction number is greater than or equal to the sum of the minimum frame extraction numbers of all videos to be analyzed, but less than the sum of the maximum frame extraction numbers of all videos to be analyzed.

[0125] In this case, the maximum supported frame extraction number allows each video to be analyzed to extract at least the minimum frame extraction number, but it cannot support each video to be analyzed to extract the maximum frame extraction number.

[0126] For example, if the maximum supported frame rate is 90 frames, and there are 5 videos to be analyzed, with minimum frame rate values ​​of 10, 15, 18, 20, and 20 frames respectively, the sum of the minimum frame rate values ​​is 83 frames. The maximum frame rate values ​​are 12, 17, 20, 25, and 26 frames respectively, resulting in a maximum sum of 100 frames. Clearly, if the maximum supported frame rate is greater than the sum of the minimum frame rate values ​​(83 frames), then each video to be analyzed will only have its minimum frame rate extracted. However, if the maximum supported frame rate is less than the sum of the maximum frame rate values ​​(100 frames), then it is not possible to extract the maximum frame rate from each video to be analyzed.

[0127] The electronic device can first assign a minimum number of frames to each video to be analyzed. Continuing with the example in Case 2 above, the electronic device can assign 10 frames, 15 frames, 18 frames, 20 frames, and 20 frames to the 5 videos to be analyzed in sequence.

[0128] Furthermore, after allocating the minimum number of frames to be extracted for each video to be analyzed, if there are still frames remaining from the maximum supported number of frames, the electronic device can distribute the remaining frames according to the duration ratio of all videos to be analyzed, so that the actual number of frames allocated to each video to be analyzed not only meets the minimum number of frames but also matches its respective duration. The specific implementation principle is similar to the allocation according to the duration ratio in Case 1 above, and will not be repeated here.

[0129] After allocating frames according to the duration ratio, the electronic device can obtain the actual number of frames extracted from each video to be analyzed as the sum of the minimum number of frames extracted and the number of frames allocated according to the duration ratio.

[0130] Similarly, if the duration of a video to be analyzed is too high, the sum of the number of frames allocated to that video according to its duration and the minimum number of frames to be extracted may exceed the maximum number of frames to be extracted. In this case, the electronic device can reduce the number of frames extracted to the maximum number of frames to be extracted. This avoids excessive frame extraction that leads to long processing times. Furthermore, since video content usually changes slowly, too many frames extracted can result in significant information redundancy during analysis. Therefore, controlling the number of frames to not exceed the maximum number of frames can control the redundancy in frame extraction analysis. The specific implementation principle is similar to the allocation based on duration proportion in case 1 above, and will not be elaborated here.

[0131] Understandably, after allocating the minimum number of frames to be extracted for each video to be analyzed, if there is no remaining maximum supported number of frames to be extracted, the electronic device does not need to continue allocating them. Accordingly, the actual number of frames extracted for each video to be analyzed is the minimum number of frames to be extracted.

[0132] Case 3: The maximum supported frame extraction number is greater than or equal to the sum of the maximum frame extraction numbers of all videos to be analyzed.

[0133] In this case, the maximum supported frame extraction number allows each video to be analyzed to extract the maximum number of frames.

[0134] For example, if the maximum supported frame extraction is 120 frames, and there are 5 videos to be analyzed, with maximum frame extraction numbers of 12, 17, 20, 25, and 26 frames respectively, then the sum of the maximum frame extraction numbers is 100 frames. Clearly, if the maximum supported frame extraction number is greater than the sum of the maximum frame extraction numbers (100 frames), then each video to be analyzed can be extracted to the maximum frame extraction number.

[0135] The electronic device can allocate a maximum number of frames to each video to be analyzed. Continuing with the example in Case 3 above, the electronic device can allocate 12, 17, 20, 25, and 26 frames to the 5 videos to be analyzed, respectively.

[0136] It should be noted that, unlike cases 1 and 2 above, after allocating the maximum number of frames to be extracted for each video to be analyzed, the electronic device will not allocate any more frames, even if there is still a remaining maximum number of frames to be extracted. In other words, in case 3, the actual number of frames extracted for each video to be analyzed is the maximum number of frames to be extracted, and will not exceed the maximum number of frames to be extracted, thus controlling the computational load of frame extraction.

[0137] Step 3, frame extraction.

[0138] In this embodiment of the application, the electronic device can approximate the shot position of the video to be analyzed through two rounds of frame extraction processing, so as to facilitate subsequent shot segmentation.

[0139] In the first round of frame extraction, electronic devices can use a uniform frame extraction method.

[0140] For example, an electronic device can use the first second of the video to be analyzed as the starting point of the first round of frame extraction and the end point of the video to be analyzed as the ending point of frame extraction, and extract frames evenly.

[0141] All videos to be analyzed will undergo a first round of frame extraction, which can be divided into the following cases:

[0142] Case A: The actual number of frames extracted from the video to be analyzed does not exceed the minimum number of frames extracted.

[0143] For example, in step 2 above, the actual number of frames allocated in case 1 did not exceed the minimum number of frames.

[0144] For example, in step 2 above, under case 2, the actual number of frames allocated may partially or entirely not exceed the minimum number of frames. For instance, after allocating the minimum number of frames to each video to be analyzed, if there is no remaining maximum supported number of frames, then the actual number of frames allocated to all videos to be analyzed is equal to the minimum number of frames, satisfying the condition of not exceeding the minimum number of frames. Another example is that after allocating the minimum number of frames to each video to be analyzed, if there is still a remaining maximum supported number of frames, but some videos to be analyzed have too short a duration and are not allocated further frames according to their duration proportions, then the actual number of frames allocated to these videos to be analyzed is equal to the minimum number of frames, also falling under the condition of not exceeding the minimum number of frames.

[0145] In this case, the electronic device can extract frames from the video to be analyzed according to the actual number of frames extracted. For example, if the minimum number of frames to be extracted from the video to be analyzed is 20 frames, and the actual number of frames extracted is 18 frames, the electronic device can start from the first second of the video to be analyzed, calculate the time interval for extracting 18 frames according to the duration of the video, and complete the frame extraction.

[0146] Case B: The actual number of frames extracted from the video to be analyzed exceeds the minimum number of frames extracted.

[0147] For example, in step 2 above, case 2, the actual number of frames allocated may partially or completely exceed the minimum number of frames to be extracted. For instance, after allocating the minimum number of frames to each video to be analyzed, if there is still a remaining maximum supported number of frames to be extracted, and at least one frame is allocated to each video to be analyzed according to its duration, then the actual number of frames extracted for all videos to be analyzed exceeds the minimum number of frames to be extracted. As another example, after allocating the minimum number of frames to each video to be analyzed, if there is still a remaining maximum supported number of frames to be extracted, but some videos to be analyzed have a duration that is too short to be allocated further frames according to their duration, then the actual number of frames extracted for these videos to be analyzed is equal to the minimum number of frames to be extracted, while the actual number of frames extracted for the remaining videos to be analyzed exceeds the minimum number of frames to be extracted.

[0148] For example, in case 3 of step 2 above, the actual number of frames allocated is the maximum number of frames to be drawn, all of which exceed the minimum number of frames to be drawn.

[0149] In this scenario, the electronic device can extract frames from the video to be analyzed according to the minimum number of frames extracted. The portion of the actual number of frames extracted that exceeds the minimum number can be used for a second round of frame extraction. For example, if the minimum number of frames to be extracted from the video to be analyzed is 20, and the actual number of frames extracted is 22, the electronic device can start from the first second of the video to be analyzed, calculate the time interval for extracting 20 frames according to the duration of the video, and complete the frame extraction.

[0150] See Figure 5 In (a), the duration of the video X to be analyzed is 30 seconds (different fill colors represent different scene shots), the minimum number of frames extracted is 12, the actual number of frames extracted is 17, and the electronic device can use 1 second as the starting point. The calculated frame extraction interval for 12 frames is (30-1) / 12 = 2.42 seconds. See also... Figure 5 In (b), the 12 extracted frames include Z1, Z2, Z3...Z12.

[0151] In the actual implementation of step 3, electronic preparation can detect whether each video to be analyzed belongs to either situation A or situation B, and extract the corresponding number of frames based on the detection results.

[0152] The second round of frame extraction involves further frame extraction based on the similarity between adjacent frames from the first round.

[0153] See also Figure 5 In (b), the electronic device can perform a second round of frame dropping based on the phase velocities of Z1 and Z2, Z2 and Z3, Z3 and Z4...Z11 and Z12.

[0154] The lower the similarity between adjacent frames, the greater the possibility of a scene switch between adjacent frames. The electronic device can further extract frames between these adjacent frames, thereby making the extracted frames approach the position where the scene switch will occur.

[0155] The electronic device can perform a second round of frame extraction for parts of the video to be analyzed that still have remaining frame extraction counts. For example, for the video to be analyzed that belongs to case B in the first round of frame extraction, after the electronic device extracts the minimum number of frames in the first round, it can extract (actual number of frames extracted - minimum number of frames extracted) in the second round. In this way, the electronic device can ensure that the total number of frames extracted in the two rounds does not exceed the actual number of frames extracted.

[0156] This application does not specifically limit the strategy for similarity-based frame extraction in the second round of frame extraction. The following examples illustrate only a few typical strategies.

[0157] Strategy 1: Electronic devices can select target adjacent frames with similarity below the similarity threshold of 1, and further extract frames between the target adjacent frames.

[0158] See Figure 5 In step (c), the electronic device calculates that the similarity between Z5 and Z6 is below the similarity threshold of 1, and can extract frames between Z5 and Z6, such as extracting Z13 at the center position between Z5 and Z6. Similarly, the electronic device calculates that the similarity between Z9 and Z10 is below the similarity threshold of 1, and can extract frames between Z9 and Z10, such as extracting Z14 at the center position between Z9 and Z10. Thus, Z13 extracted in the second round is closer to the scene transition position q1, and Z14 is closer to the scene transition position q2.

[0159] Strategy 2: The electronic device can select the target adjacent frames with the lowest similarity (number 1) for further frame extraction, and then extract frames between the target adjacent frames. Here, the number 1 can be less than or equal to (actual number of extracted frames - minimum number of extracted frames).

[0160] For example, if the actual number of frames extracted minus the minimum number of frames extracted equals 5 frames, the electronic device can select the 5 groups of targets with the lowest similarity to extract frames from each other, and then extract more frames.

[0161] Furthermore, the electronic device can filter out adjacent frames with a time interval less than 1 second, thus allowing for further frame extraction for adjacent frames with a time interval greater than or equal to 1 second. For example, if 1 second is 1 second, the electronic device can calculate the similarity of adjacent frames with a time interval greater than or equal to 1 second and extract further frames accordingly.

[0162] Furthermore, electronic devices can extract one or more frames between adjacent frames.

[0163] For example, an electronic device can decide the number of frames to be extracted based on the time interval between adjacent frames; the longer the time interval, the more frames are extracted. For instance, if the time interval is between 1 and 2 seconds, one frame can be extracted; if the time interval is greater than 2 seconds, one frame can be extracted every 2 seconds.

[0164] As another example, the electronic device can iteratively extract frames. For instance, after extracting a frame between adjacent frames, the electronic device can further calculate the similarity between the extracted frame and the preceding and following frames, and then determine whether to extract another frame based on the similarity. This process can be repeated to extract one or more frames between adjacent frames.

[0165] At this point, it should be noted that:

[0166] First, in step 3 above, two rounds of frame extraction are performed using the minimum number of frames extracted determined in steps 1 and 2, and the actual number of frames extracted. This allows frame extraction to balance time consumption, computational load, and the accuracy of highlight segment extraction. However, in actual implementation, this is not the limitation.

[0167] For example, regarding the first round of frame extraction: the electronic device can set a fixed number of frames to be extracted in the first round, such as 10 frames; or, the electronic device can also directly determine the number of frames to be extracted in the first round based on the duration of the video to be analyzed, the longer the duration, the more frames are extracted.

[0168] For example, regarding the second round of frame skipping: the electronic device can be set to a fixed number of frames skipped in the second round, such as 3 frames.

[0169] Second, in step 3 above, the minimum number of frames extracted and the actual number of frames extracted are used to determine whether to perform two rounds of frame extraction for each video to be analyzed. In actual implementation, this is not a limitation.

[0170] For example, an electronic device may perform a second round of frame extraction for videos to be analyzed that have a duration exceeding 2 seconds, while for videos to be analyzed that have a duration of less than 2 seconds, there is no need to perform a second round of frame extraction.

[0171] For example, the electronic device can decide whether to perform a second round of frame extraction based on the results of the first round of frame extraction. For instance, if the similarity between adjacent frames after the first round of frame extraction is higher than the similarity threshold of 2, the electronic device can decide not to perform a second round of frame extraction on the video to be analyzed.

[0172] Step 4, storyboard calculation.

[0173] Electronic devices can perform scene segmentation calculations to complete the segmentation of the video to be analyzed.

[0174] In some embodiments, the storyboard includes one or more of the following: a continuously stable storyboard, an alternative storyboard, a non-continuously stable storyboard, and a sudden change shot. The calculation process for these types of shots is described below.

[0175] Step 41: Merge to obtain consistent regions and isolated frames.

[0176] The electronic device can calculate the similarity between frames extracted in two rounds, and merge the frames with similarity greater than a similarity threshold of 3 and the regions between them into a consistent region. After the calculation in step 41, the electronic device can obtain the consistent regions and isolated frames of the video to be analyzed.

[0177] Among them, a lonely frame refers to a frame in the extracted frame whose similarity with the adjacent extracted frames before and after it is lower than the similarity threshold of 3.

[0178] Among them, the similarity threshold 3 is higher than the similarity threshold 1 and similarity threshold 2 mentioned above.

[0179] See Figure 6 A total of 8 frames were extracted in the two rounds. Note that the numbers 1, 2, ... represent the first frame, second frame, ... extracted sequentially from the beginning of the video to be analyzed, not the first frame, second frame, ... of the entire video to be analyzed.

[0180] If the similarity between 1 and 2 is higher than the similarity threshold 3, the electronic device can merge 1 to 2 into a consistent region, such as [1, 2], which represents the video region in the video to be analyzed, starting from the first extracted frame and ending at the second extracted frame.

[0181] The similarity between 2 and 3, 3 and 4, and 4 and 5 is all below the similarity threshold of 3. The electronic device will not merge 3 and 4 with the frames extracted before and after, nor will it merge 3 and 4. Therefore, 3 and 4 are lonely frames.

[0182] The similarity between 5 and 6, and between 6 and 7 are both higher than the similarity threshold of 3. The electronic device can merge 5 to 7 into a consistent region, which can be denoted as [5, 7], representing the video region in the video to be analyzed, starting from the extracted 5th frame and ending at the extracted 7th frame.

[0183] Since the similarity between 7 and 8 is below the similarity threshold of 3, the electronic device will not merge the consistent regions 8 and 7, and 8 is a lonely frame.

[0184] Step 42: Merge the consistent regions and / or isolated frames to obtain the noisy consistent regions.

[0185] Step 42 mainly includes merging two adjacent consistent regions, and merging adjacent consistent regions and isolated frames, which are explained below:

[0186] First, merge two adjacent consistent regions.

[0187] Adjacent consistent regions refer to two consistent regions that are not separated by other consistent regions. For example, the above... Figure 6 The consistent regions [1,2] and [5,7] are separated only by lonely frames 3 and 4, without any other consistent regions. Therefore, [1,2] and [5,7] are two adjacent consistent regions.

[0188] When two adjacent consistent regions meet merging condition 1, the electronic device can merge the two adjacent consistent regions. Merging condition 1 includes the similarity between the two consistent regions being higher than the similarity threshold 4. It should be noted that the embodiments of this application do not specifically limit the representation of the similarity between two consistent regions; only a few typical representations are listed below.

[0189] Form 1: The similarity between two consistent regions can be represented by the similarity between the intermediate frames of the two consistent regions, as described above. Figure 6 The similarity between the consistent regions [1,2] and [5,7] can be represented by the similarity between the intermediate frames of [1,2] and [5,7].

[0190] In Form 1, if the similarity of the intermediate frames is higher than the similarity threshold 4, it indicates that the similarity between the two consistent regions is higher than the similarity threshold 4.

[0191] Form 2: The similarity between two consistent regions can be represented by the similarity of multiple sets of image frames. The multiple sets of image frames include the k-th frame from front to back of the preceding consistent region and the k-th frame from back to front of the following consistent region.

[0192] For example, the above Figure 6 The similarity between the consistent regions [1,2] and [5,7] can be represented by the similarity of the following sets of image frames: the first frame of [1,2] and the last frame of [5,7], the second frame of [1,2] and the penultimate frame of [5,7], the third frame of [1,2] and the penultimate frame of [5,7], and so on...

[0193] In Form 2, if the similarity of each group of image frames in multiple groups of image frames is higher than the similarity threshold 4, or if the average similarity of multiple groups of image frames is higher than the similarity threshold 4, it indicates that the similarity of two consistent regions is higher than the similarity threshold 4.

[0194] In this way, electronic devices can further merge highly similar and consistent regions, see [link to relevant documentation]. Figure 7In (a), the electronic device can merge the consistent regions [1,2] and [4,5] to obtain a noisy consistent region consisting of [1,2], [4,5], and the region between them (including lonely frame 3), which can be denoted as [1,5]. Note that in this paper, for the purpose of distinction, the consistent region is represented by "[]", and the noisy consistent region is represented by "【】".

[0195] Furthermore, merging condition 1 also includes: the time interval between two adjacent consistent regions is less than duration 3. In other words, the time interval between the two consistent regions being merged is relatively small. This time interval includes the time interval between the end frame of the foreground consistent region and the start frame of the subsequent consistent region.

[0196] In one specific implementation, the duration 3 is a fixed value, such as 10 seconds.

[0197] In another specific implementation, the duration 3 is related to the region lengths of the two consistent regions. The region length of a consistent region is the time interval between the start and end frames of that region. Specifically, the duration 3 can be the sum of the region lengths of the two consistent regions.

[0198] Of course, in practice, electronic devices can also combine the two implementation methods described above. The electronic device will only merge two consistent regions if the time interval between them is less than the duration of 3 in either of the two implementation methods. Conversely, if the time interval between two consistent regions is greater than or equal to the duration of 3 in either of the two implementation methods, they will not be merged.

[0199] Electronic devices can improve the plausibility of merging two consistent regions with short time intervals. For example, a segment of video to be analyzed may be identical or similar to another segment that is far apart from it. This could be because the first segment was filmed to form one object, then another object was filmed, and then the camera switched back to film the first object to form another segment. In this case, another object was filmed between the two segments, so they do not belong to the same shot, and merging the two segments would be unreasonable.

[0200] See Figure 7 In (b), although there are lonely frames 3 and 4 between the consistency regions [1,2] and [5,6], the time interval between [1,2] and [5,6] is less than the duration 3. Therefore, if the similarity between [1,2] and [5,6] is higher than the similarity threshold 4, the electronic device can merge [1,2] and [5,6] to obtain a noisy consistency region consisting of [1,2], [5,6] and the region between them (including lonely frames 3 and 4), which can be denoted as [1,6].

[0201] Second, merge consistent regions and isolated frames.

[0202] When a consistent region and a lonely frame meet merging condition 2, the electronic device can merge the consistent region and the lonely frame. Merging condition 2 includes: the similarity between the consistent region and the lonely frame is higher than a similarity threshold 5. It should be noted that the embodiments of this application do not specifically limit the representation of the similarity between the consistent region and the lonely frame; only a few typical representations are listed below.

[0203] Form 3: The similarity between consistent regions and lonely frames can be represented by the similarity between intermediate frames of the consistent region and lonely frames, such as... Figure 8 The similarity between the consistent region [1, 2] and the lonely frame 4 in (a) can be represented by the similarity between the intermediate frames of the consistent region [1, 2] and the lonely frame 4.

[0204] In Form 3, if the similarity between the intermediate frames and the lonely frames in the consistent region is higher than the similarity threshold of 5, it indicates that the similarity between the consistent region and the lonely frames is higher than the similarity threshold of 5.

[0205] Form 4: The similarity between a consistent region and a lonely frame can be represented by the similarity between multiple image frames within the consistent region and the lonely frame. These multiple image frames can be some or all of the image frames within the consistent region.

[0206] For example, Figure 8 The similarity between the consistent region [1, 2] and the lonely frame 4 in (a) can be represented by the similarity between the first frame, the second frame, the third frame, ... in the consistent region [1, 2] and the lonely frame 4, respectively.

[0207] In Form 4, if the similarity between each image frame and the lonely frame is higher than the similarity threshold 5, or if the average similarity between each image frame and the lonely frame is higher than the similarity threshold 5, it indicates that the similarity between the consistent region and the lonely frame is higher than the similarity threshold 5.

[0208] In this way, electronic devices can further merge highly similar consistent regions and lonely frames, see [link to relevant documentation]. Figure 8 In (a), the electronic device can merge the consistent region [1,2] and the lonely frame 4 to obtain a noisy consistent region consisting of the consistent region [1,2], the lonely frame 4 and the region between them (including the lonely frame 3), which can be denoted as [1,4].

[0209] Furthermore, similar to merging condition 1 mentioned above, merging condition 2 also includes: the time interval between the consistent region and the lonely frame is less than 4 seconds. In other words, the time interval between the merged consistent region and the lonely frame is relatively small.

[0210] The time interval between a consistent region and a lonely frame includes the time interval between the start frame, middle frame, or end frame of the consistent region and the lonely frame.

[0211] In one specific implementation, the duration of 4 is a fixed value, such as 5 seconds or 10 seconds.

[0212] In another specific implementation, the duration of 4 is related to the region length of the consistency region. The region length of a consistency region is the time interval between the start and end frames of that consistency region. Specifically, the duration of 4 can be the region length of the consistency region.

[0213] Of course, in actual implementation, electronic devices can also combine the two implementation methods described above. The electronic device will only merge the consistent region and the lonely frame if the time interval between them is less than the duration of 4 in either of the two implementation methods. Conversely, if the time interval between the consistent region and the lonely frame is greater than or equal to the duration of 4 in either of the above implementation methods, they will not be merged.

[0214] Electronic devices can also improve the rationality of merging by combining consistent regions and isolated frames with short time intervals. See also Figure 8 In (b), although there are lonely frames 3 and 4 between the consistent region [1,2] and the lonely frame 5, the time interval between the consistent region [1,2] and the lonely frame 5 is less than the duration 4. Therefore, if the similarity between the consistent region [1,2] and the lonely frame 5 is higher than the similarity threshold 5, the electronic device can merge the consistent region [1,2] and the lonely frame to obtain a noisy consistent region composed of the consistent region [1,2], the lonely frame 5 and the region between them (including lonely frames 3 and 4), which can be denoted as [1,5].

[0215] In addition, if a consistent region is not merged, then the consistent region constitutes a noisy consistent region on its own.

[0216] It is important to note that isolated frames are not merged. If a region contains multiple isolated frames, it can be named an unstable region.

[0217] The following will continue to use... Figure 5 The example further illustrates the effects of steps 41 and 42: After obtaining... Figure 5 After the frame extraction result shown in (c) above, the electronic device executes step 41 to obtain... Figure 9 The consistent regions Y1, Y2, Y3, and Y4 are shown in (a), along with isolated frames Z3, Z8, Z9, and Z14. Then, the electronic device performs step 42, which merges the consistent regions Y1 and Y2 to obtain... Figure 9(b) includes the noisy consistency region M1 of Y1, isolated frame Z3, and Y2. After step 42, Y3 and Y4 are not merged, but instead form separate noisy consistency regions, such as... Figure 9 As shown in (b) M2 and M3. Furthermore, after step 42, the isolated frames Z8, Z9, and Z14 do not synthesize with the preceding and following consistent regions, thus forming unstable regions, such as... Figure 9 As shown in T1 in (b) of the diagram.

[0218] Step 43, divide the shots.

[0219] Based on the merging results of step 42 above, the electronic device can divide the video to be analyzed into continuously stable scenes, alternative scenes, non-continuously stable scenes, and / or abrupt scenes. These are described below:

[0220] The first type is a consistently stable storyboard.

[0221] A sustained stable shot refers to a shot that consistently captures the same scene (i.e., content with high similarity) over a relatively long period. Electronic devices can determine a sustained stable shot based on a noisy, consistent region 1 with a length greater than 5 seconds of duration. The duration 5 can be 5 seconds, 8 seconds, etc.

[0222] For example, Figure 9 In (b), the duration of regions M1 and M3 is greater than 5, so M1 and M3 are noisy and consistent regions 1.

[0223] Furthermore, the electronic device can extend the noisy consistency region 1 by a certain area before and after it to obtain a continuously stable storyboard. Specifically, the electronic device can extend the starting frame of the noisy consistency region 1 forward to the point between the starting frame and a previous frame, and approximately one-third of the way from the starting frame. And / or, the electronic device can extend the ending frame of the noisy consistency region 1 backward to the point between the ending frame and a subsequent frame, and approximately one-third of the way from the ending frame.

[0224] For example, electronic devices can be based on Figure 9 M1 in (b) is determined Figure 9 In the continuous stable segment M1' shown in (c), the ending frame V1 of M1' is extended backward compared to the ending frame Z13 of M1. V1 is located between Z13 and Z6, and is close to one-third of Z13.

[0225] For example, electronic devices can be based on Figure 9 M3 in (b) is determined Figure 9The continuously stable segment M3' is shown in (c). The starting frame V2 of M3' is extended forward compared to the starting frame Z10 of M3, and V2 is located between Z14 and Z10, about one-third of the way from Z10.

[0226] Furthermore, if the starting frame of the noisy consistency region 1 is a sampled frame from the first frame of the video to be analyzed, the electronic device may not need to extend the starting frame forward when determining a continuously stable sequence. For example, Figure 9 In (c) of the middle, the starting frame of M1' is maintained with Figure 9 In (b) of the above, the starting frame of M1 is the same, which is the first frame extracted from frame Z1.

[0227] Additionally, if the ending frame of noisy consistency region 1 is the last frame extracted from the video to be analyzed, the electronic device can extend the ending frame of noisy consistency region 1 backward, but not beyond the ending frame of the video to be analyzed. For example, it can extend backward between the ending frame of noisy consistency region 1 and the ending frame of the video to be analyzed, approximately one-third of the way down from the ending frame of noisy consistency region 1. Figure 9 In (c), the end frame V3 of M3' is compared to Figure 9 In (b), the end frame Z12 of M3 extends backward, and V3 is located about one-third of the way from Z12 between Z12 and the end frame of the video to be analyzed.

[0228] The second category is non-continuously stabilized lenses.

[0229] Non-continuously stable shots refer to shots that depict the same scene within a relatively short period of time. Electronic devices can determine continuously stable shots based on a noisy, consistent region 2 with a length greater than duration 6 and less than or equal to duration 5. Duration 6 is less than duration 5, and duration 6 can be 2 seconds, 3 seconds, etc.

[0230] For example, Figure 9 If the duration of region M2 in (b) is greater than the duration of 6 and less than the duration of 5, then M2 is a noisy and consistent region 2.

[0231] Furthermore, the electronic device can extend the noisy consistency region 2 by a certain area before and after it to obtain a non-constantly stable storyboard. The principle of extension is the same as the principle of extension to obtain a continuously stable storyboard mentioned above, and will not be repeated here.

[0232] For example, electronic devices can be based on Figure 9 M2 in (b) is determined Figure 9 The non-steady-state segment M2' is shown in (c). The starting frame of M2' is extended forward compared to the starting frame of M2, and the ending frame of M2' is extended backward compared to the ending frame of M2.

[0233] The third type is mutation storyboarding.

[0234] A sudden change in scene refers to a sequence of shots depicting the same scene within a very short period of time. Electronic devices can determine sudden change in scene based on a noisy, consistent region 3 with a region length less than or equal to a duration of 6.

[0235] Furthermore, the electronic device can extend the noisy consistency region 3 by a certain area before and after it to obtain a sudden change in the storyboard. The principle of extension is the same as the principle of extension to obtain a continuous and stable storyboard mentioned above, and will not be repeated here.

[0236] The fourth category is alternative storyboards.

[0237] The electronic device can determine candidate storyboards based on a target instability region whose length is greater than 7 seconds in duration. The duration 7 can be 5 seconds, 8 seconds, etc.

[0238] For example, Figure 9 If the duration of region T1 in (b) is greater than 6, then T1 is the target unstable region.

[0239] Furthermore, electronic devices can extend the unstable region of the target by a certain area before and after to obtain alternative storyboards. The principle of extension is the same as that of extension to obtain a continuously stable storyboard mentioned earlier, and will not be repeated here.

[0240] For example, electronic devices can be based on Figure 9 In (b), T1 is determined Figure 5 The alternative storyboard T1' is shown in (c). The starting frame of T1' is extended forward compared to the starting frame of T1, and the ending frame of T1' is extended backward compared to the ending frame of T1.

[0241] Additionally, if the length of an unstable region is less than or equal to the duration of 7, the electronic device can discard that unstable region.

[0242] After step 4 above, the electronic device divides the video to be analyzed into segments based on the results of two rounds of frame extraction, and can obtain various types of segments. The segmentation is basically consistent with the actual scene segments of the video to be analyzed.

[0243] For example, the above Figure 9 In the video to be analyzed shown in (a), different fill colors represent different scene sequences, resulting in a total of 3 scene sequences. After frame extraction in step 3 and scene sequence calculation in step 4, we can obtain... Figure 10 The storyboard results are shown in (c). Among them, the continuously stable storyboard M1' basically covers the first scene storyboard, the continuously stable storyboard M3' basically covers the third scene storyboard, and the second scene storyboard between the first and third scene storyboards is divided into the non-continuously stable storyboard M2' and the alternative storyboard T1'.

[0244] The following further combines Figure 5 The following example illustrates the complete process of step 4:

[0245] S1001. Perform consistency identification. See step 41 above for details.

[0246] After S1001, the electronic equipment can merge to obtain a consistent region. In addition, frames that are not merged can be recorded as isolated frames.

[0247] S1002, Check if it is a consistent region.

[0248] That is, to detect consistent regions in the video to be analyzed.

[0249] S1103. After obtaining the consistent region, perform segment expansion to obtain the noisy consistent region.

[0250] The execution segment expansion refers to merging a consistent region with an adjacent consistent region or with a lonely frame. After merging, a noisy consistent region can be obtained.

[0251] S1004. Check whether the time for detecting the noisy uniformity region exceeds 5 seconds.

[0252] The 5 seconds in S1004 is the same as the duration 5 in step 43 above.

[0253] The electronic device can detect the length of each noisy, consistent region. If the region length exceeds 5 seconds, a continuous and stable segment can be obtained, as shown in S1005 below. If the region length does not exceed 5 seconds, S1006 can be further executed to divide the non-continuous and stable segment and the abrupt segment.

[0254] S1005, obtain a continuously stable storyboard.

[0255] S1006. Check whether the time for detecting the noisy uniform region exceeds 3 seconds.

[0256] The 3 seconds in S1004 is the same as the duration 6 in step 43 above.

[0257] If the length of the noisy, consistent region does not exceed 5 seconds but exceeds 3 seconds, a non-stable segment can be obtained, as shown in S1007 below. If the length of the noisy, consistent region does not exceed 3 seconds, a sudden change segment can be obtained, as shown in S1008 below.

[0258] S1007, Obtain the non-constantly stable storyboard.

[0259] S1008, obtain the mutation storyboard.

[0260] S1009. Detect whether the duration of the unstable region formed by lonely frames exceeds 3 seconds.

[0261] The 3 seconds in S1009 is the same as the duration 7 in step 43.

[0262] The electronic device can classify regions consisting of consecutive isolated frames that cannot be merged into a noisy, consistent region as unstable regions. If the length of an unstable region exceeds 3 seconds, it is divided into candidate storyboards, as shown in S1010 below. If the length of an unstable region does not exceed 3 seconds, it can be discarded.

[0263] S1010, Obtain alternative storyboards.

[0264] In the above-described scene breakdown, when extracting highlight clips, continuously stable scenes have the highest priority, followed by non-continuously stable scenes, with alternative scenes being the least important. This allows the electronic device to prioritize extracting highlight clips from long-duration, continuous shooting content, then from short-duration, continuous shooting content, and finally from scenes with significant changes in visual quality. This ensures that the optimal highlight clips are extracted.

[0265] It should be noted that, in actual implementation, the method of storyboard calculation is not limited to the one specified in step 4 above.

[0266] For example, the electronic device may also calculate only a portion of the above-mentioned storyboards, such as continuously stable storyboards and non-continuously stable storyboards.

[0267] For example, electronic devices can also obtain various storyboards directly based on noisy, consistent regions and unstable regions without extending forward or backward.

[0268] As another example, electronic devices can also directly calculate various scene sequences based on the similarity of two rounds of frame sampling. For instance, an electronic device can merge consecutive frames with a similarity higher than a similarity threshold of 6 and divide the scene sequence based on the length of the merged region. (Continuing from the previous text...) Figure 11 Taking the frame extraction result shown in (c) as an example, the electronic device calculates that the similarity between Z1 and Z2, Z2 and Z3, Z3 and Z4, Z4 and Z5, and Z5 and Z13 are all higher than the similarity threshold of 6. The electronic device can merge Z1 as the starting frame and Z13 as the ending frame, and then divide the merged region into a continuous and stable segment based on the length of the region between Z1 and Z13.

[0269] Step 5, extract highlight fragments.

[0270] In step 5, the electronic device can extract highlight segments from at least two split shots, thereby ensuring the richness of the extracted highlight segments.

[0271] The electronic device can perform two rounds of frame extraction to obtain the optimal frame based on the content contained in the image frames. The optimal frame is defined as an image frame containing preset content. The preset content includes at least one of the following: a face, a body, and / or a voice.

[0272] Subsequently, the electronic device can extract highlight segments based on the best frame in each shot, according to the priority of the shot.

[0273] The electronic device can first select the best frame 1 from the continuously stable shot and extract the highlight fragment a that includes the best frame 1, then select the best frame 2 from the non-continuously stable shot and extract the highlight fragment b that includes the best frame 2, and finally select the best frame 3 from the candidate shot and extract the highlight fragment c that includes the best frame 3.

[0274] When extracting highlight segments including the optimal frame, the electronic device can expand forward and backward from the optimal frame to form a video segment of a preset duration (such as 3 seconds), thereby obtaining the highlight segment.

[0275] Of course, if the optimal frame is close to the start or end frame of the storyboard, the electronic device can also extend towards the center of the storyboard to form a video clip of a preset duration, thus obtaining a highlight clip. This ensures that the highlight clip does not exceed the scope of the storyboard.

[0276] It should be noted that if the length of the segment is less than the preset duration, the electronic device can extend beyond the segment to form a video clip of the preset duration, thereby ensuring the duration of the highlight clip. This application does not impose specific limitations on this.

[0277] Therefore, the following points need to be explained:

[0278] First, a storyboard may include one or more optimal frames. The electronic device may extract highlight segments based on each optimal frame in the storyboard, or it may extract highlight segments based only on a preset number of optimal frames (such as 1 frame).

[0279] Second, a storyboard may not include the optimal frame. If a storyboard does not include the optimal frame, then the highlight fragment cannot be extracted from that storyboard. For example, if the optimal frame 1 is not included in a continuously stable storyboard, then the highlight fragment 1 cannot be extracted from the continuously stable storyboard.

[0280] When extracting highlight segments based on the optimal frame, electronic devices can extract highlight segments from at least two shots, thereby avoiding the concentration of highlight segments in one shot and improving the richness of highlight segments.

[0281] In practice, if only one shot contains the optimal frame, the electronic device may not be able to extract highlight fragments from at least two shots. Alternatively, if none of the shots contain the optimal frame, the electronic device may not be able to extract highlight fragments based on the optimal frame. In these cases, the electronic device can extract highlight fragments based on the aesthetic scores of image frames, thus enabling the extraction of more highlight fragments. Specifically, the electronic device can evaluate and rank the aesthetic scores of frames other than the optimal frame in two rounds of frame extraction, and then extract highlight fragments from high to low aesthetic scores, ensuring that highlight fragments can be extracted from at least two shots.

[0282] It should be noted that, in actual implementation, the method of storyboard calculation is not limited to the one specified in step 5 above.

[0283] For example, the electronic device can also comprehensively score the image frames in each scene based on their content and aesthetic ratings, and then extract the highlight fragment containing the highest comprehensive score from at least two scenes according to the priority of the scenes. For instance, the electronic device first extracts highlight fragment 'a' containing the image frame with the highest comprehensive score from a continuously stable scene, and then extracts highlight fragment 'b' containing the image frame with the highest comprehensive score from a non-continuously stable scene, and so on.

[0284] For example, the electronic devices involved in the embodiments of this application may also be referred to as terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Electronic devices may include mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technologies or device forms used in the electronic devices.

[0285] Figure 11 A schematic diagram of the hardware structure of an electronic device according to an embodiment of this application is shown. The following description, in conjunction with...Figure 11 The hardware structure of electronic devices will be introduced.

[0286] Take a mobile phone as an example. Figure 1 As shown, the electronic device 200 may include: a processor 210, a memory 220, a universal serial bus (USB) interface 230, a power management module 240, an antenna, a communication module 250, a display screen 260, an audio module 270, a camera 280, a sensor module 290, etc.

[0287] Processor 210 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, memory, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). Different processing units may be independent devices or integrated into one or more processors. The controller may serve as the central nervous system and command center of the electronic device 200. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0288] Memory 220 can be used to store executable program code, including instructions. Processor 210 executes various functional applications and data processing of the electronic device by running the instructions stored in memory 220. Memory 220 may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, interface display, etc.). The data storage area may store data created during the use of the electronic device (such as notification messages). Furthermore, memory 220 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0289] The power management module 240 is used to connect the battery and the processor 210. The power management module 240 receives battery and / or power input to power the processor 210, memory 220, communication module 250, display 260, audio module 270, and camera 280, etc. The power management module 240 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 240 may also be located within the processor 210.

[0290] The communication module 250 can provide solutions for wireless communication applications on the electronic device 200, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR). The communication module 250 can be one or more devices integrating at least one communication processing module. The communication module 250 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 210. The communication module 250 can also receive signals to be transmitted from the processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via the antenna.

[0291] In some embodiments, the antenna of the electronic device 200 is coupled to the communication module 250, enabling the electronic device 200 to communicate with networks and other devices via wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BitTorrent, Global Navigation Satellite System (GNSS), WLAN, NFC, FM, and / or IR technologies. The GNSS may include Global Positioning System (GPS), BeiDou Navigation Satellite System (BDS), GLONASS, and / or Galileo.

[0292] Electronic device 200 implements display functions through GPU, display screen 260, and application processor, such as displaying the above-mentioned... Figure 2 , Figure 11 The interface shown.

[0293] Electronic device 200 can perform shooting functions through ISP, camera 280, video codec, GPU, display 260, and application processor. Audio module 270 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 270 can also be used for encoding and decoding audio signals.

[0294] The sensor module 290 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors, etc.

[0295] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 200. In other embodiments, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0296] Generally, the implementation of one-click movie streaming functionality in electronic devices requires not only hardware support but also software cooperation. The software system of electronic devices can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application's embodiment uses a layered architecture of Android... Taking the operating system as an example, combined with Figure 12 The software architecture of the electronic device involved in the embodiments of this application will be described.

[0297] Figure 12 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application is shown.

[0298] like Figure 12 As shown, the layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, Android... The operating system is divided into four layers, from top to bottom: application (APP) layer, media middleware framework layer, application framework (FWK) layer, and hardware abstraction layer (HAL).

[0299] The application layer can include a series of application packages, such as camera, gallery, third-party video editing software, calendar, map, and navigation. When these application packages are run, they can access the various service modules provided by the media platform framework layer and the application framework layer through the application programming interface (API) and execute corresponding intelligent business logic.

[0300] In one embodiment, the camera is used to capture photos, videos, slow-motion images, and panoramic images in response to user actions. After these images are captured by the camera, or after the user triggers a screenshot, or after the user triggers screen recording, or after the terminal device downloads images from other devices, the terminal device can save these images in a gallery, allowing the user to perform video editing operations on the images in the gallery, such as one-click blockbuster editing.

[0301] In this embodiment of the application, the image library is divided into two layers from top to bottom: the business layer and the application function layer.

[0302] The business layer, also known as the video editing business layer, includes several modules such as multi-camera recording with automatic video editing, AI-powered music videos, one-click blockbuster videos, and highlight moments, each providing corresponding services. These services are presented as controls in the gallery's user interface (UI). Users can trigger corresponding video processing actions by operating a control. For example, after a user selects one or more clips (including video clips and / or images) and clicks a control like "one-click blockbuster video" in the gallery, the electronic device decodes the clips, automatically analyzes and extracts highlight segments from the video clips and highlight images from the images, and then combines all highlight segments and images into a pre-edited video. The clips to be analyzed can be all or part of the clips selected by the user.

[0303] The one-click blockbuster service is based on an automatic editing framework. The automatic editing framework includes basic functions such as clip selection, storyline organization, layout splicing, and special effects enhancement. After generating the service, it calls the strategy control module in the lower layer (such as the media middle platform framework) to issue parameters such as the available allocation time of video analysis, the minimum interval time of highlight clips, and the longest highlight clip time.

[0304] The application functionality layer includes an automatic editing framework. Each business function in the business layer can call the automatic editing framework to provide automatic editing services for the materials. For example, the automatic editing framework may include functional modules such as segment selection, storyline organization, layout splicing, and special effects enhancement. Segment selection is used to call the channel interface to decode, reduce resolution, and convert the format of the materials to be analyzed, and to store the processed materials. Segment selection is also used to call the image highlight segment analysis interface, strategy monitoring module, etc., to extract highlight segments and / or highlight images from the materials to be analyzed. Storyline organization is used to sequentially splice multiple materials to be analyzed in the form of a storyline based on the content of the materials to be analyzed. Layout splicing is used to adjust the interface layout of the materials to be analyzed. Special effects enhancement is used to adjust the enhancement effects of the materials to be analyzed, such as adjusting the screen brightness and beautifying facial features.

[0305] The basic functionality layer is used to perform basic processing on the edited video clips after the automatic editing framework has edited multiple source materials. For example, the basic functionality layer may include modules such as video splicing, compositing and saving, video speed adjustment, and audio / video effects processing. Specifically, video splicing is used to stitch together all extracted highlight segments and / or highlight images to obtain a final video. Compositing and saving is used to store the spliced ​​final video. Video speed adjustment is used to speed up the final video. For example, video speed adjustment can add a slow-motion effect to the parts of the spliced ​​final video that include highlight actions. The audio / video effects processing module is used to add video effects and sound effects to the spliced ​​final video. For example, it can add style filters and themes, and add background music to the final video.

[0306] The media middle platform framework layer is a software layer positioned between the application layer and the application framework. This layer primarily handles the basic data processing for the one-click video production service, including interfaces for image highlight segment analysis, theme summarization, and analysis performance queries, as well as a strategy monitoring module. The image highlight segment analysis interface and the strategy monitoring module mainly control the analysis location and duration of actual highlight segments in the source video, decoding and analyzing them according to the allocated duration and location based on the strategy algorithm.

[0307] The strategy scheduling in the strategy monitoring module has multiple selection strategies. When there are a large number of selected video materials and the video materials are relatively long, performing algorithm analysis on all videos would cause resource shortages and excessively long analysis times. In this case, the strategy scheduling layer will select intensive and sparse key analysis strategies to allocate and select analysis segments.

[0308] The media middleware framework layer will send the processed video segment data to the hardware abstraction layer through the application framework layer. The hardware abstraction layer will process the data into data that can be recognized and calculated by the chip algorithm, perform algorithm analysis, and return the analysis results.

[0309] Specifically, the media middle platform framework layer may include an analysis performance query interface, a strategy monitoring (solution) module, a channel (pipeline) interface, a topic summary interface, an integration capability interface, and an image highlight segment analysis interface, etc.

[0310] The performance analysis query interface is used to obtain the estimated analysis time for each video to be analyzed.

[0311] The strategy monitoring module is used to obtain the material to be analyzed from all materials, determine the image frames that need to be analyzed in each video, determine the target area of ​​each video, and determine the highlight segments of each video.

[0312] The channel interface is used to decode, convert formats, and reduce resolution of the material to be analyzed. It forwards the frame data address of the processed material to the hardware abstraction layer through the application framework layer, and then reports the analysis results of the analyzed image frames returned by the hardware abstraction layer to the application function layer.

[0313] The channel interface is also used to perform frame reduction processing on the decoded video material to be analyzed, retaining a portion of the image frames from each video material. This reduces the time spent by electronic devices on format conversion and resolution reduction, thus improving efficiency.

[0314] The theme summary interface is used to invoke the theme algorithm to obtain theme types that match all highlight fragments and / or highlight images, and then sends these theme types to the application functional layer. The application functional layer then retrieves the theme template corresponding to the theme type.

[0315] The integration capability interface provides a public interface for querying and calling underlying algorithms, and can be connected to underlying algorithm modules.

[0316] It should be noted that this application uses the one-click blockbuster function provided by the image library as an example for illustration, and it does not limit the embodiments of this application. In actual implementation, third-party video editing software can use the video segment processing method provided in the embodiments of this application to combine multiple images and videos selected by the user into a single video with one click.

[0317] The application framework layer, or simply the framework layer, supports the operation of various modules within the media middleware framework layer. For example, the framework layer may include service interfaces such as a one-click blockbuster interface, parameter management interface, image data transmission interface, theme analysis interface, resource management module, and performance analysis interface.

[0318] The hardware abstraction layer can include a capability query interface and algorithm modules. The capability query interface is used to retrieve the algorithm capabilities supported by the hardware abstraction layer. The algorithm modules can include specular highlight algorithms, face detection algorithms, video acceleration algorithms, image super-resolution algorithms, etc.

[0319] The highlight segment algorithm is used to analyze image frames in the video footage to determine whether an image frame belongs to a highlight segment. The highlight segment algorithm is also used to analyze images in the pictures to determine whether an image is a highlight image.

[0320] Among them, the face detection algorithm is used to detect faces in image frames (or pictures).

[0321] Among them, the video acceleration algorithm is used to accelerate the processing of videos (such as the first batch of videos).

[0322] Among them, image super-resolution algorithms are used to increase image resolution.

[0323] It should be noted that, Figure 13 The layers and components within each layer of the illustrated software architecture do not constitute a specific limitation on the terminal device. In other embodiments, the terminal device may include more layers than illustrated, such as a system library (FWK LIB) layer and a kernel layer. Each layer may include more or fewer components than illustrated. Furthermore, the aforementioned functional modules may be combined into a single functional module, and the layers may be combined into a single layer; for example, the media middleware framework layer may be located within the application framework layer.

[0324] It is understood that, in order to implement the methods in the embodiments of this application, electronic devices include hardware and / or software modules that perform various functions. Based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in a hardware-driven or software-driven manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments.

[0325] The following section further details the highlight extraction method provided in this application, taking the one-click blockbuster video as an example, in conjunction with the hardware and software components of the aforementioned electronic device. See also... Figure 1 Highlight fragment extraction methods include:

[0326] 1301. Load images and videos at the business layer.

[0327] For example, in response to user input Figure 1 The trigger operation of "One-Click Blockbuster" 1021 in the interface 102 shown can trigger the loading of images and videos.

[0328] The business layer allows users to select from images and videos by loading them.

[0329] 1302. The business layer transmits the file descriptor table (FD) and the expected time to the policy monitoring module of the media middle platform framework layer.

[0330] After the user selects materials (videos and / or images), the system responds by performing operations such as one-click material analysis and highlight extraction, for example, for... Figure 13In the pop-up window 1044 of the interface shown in Figure 104, clicking the "One-Click Blockbuster" button 1045 allows the business layer to transmit file descriptors to the strategy monitoring module of the media platform framework layer. A file descriptor uniquely identifies a media clip. In other words, by transmitting file descriptors to the strategy monitoring module, the module can obtain the media clip selected by the user. The business layer can also transmit the expected time to the strategy monitoring module. The expected time can be understood as the time consumption requirement mentioned earlier; this time consumption requirement is used by the strategy monitoring module to allocate the actual number of frame extractions.

[0331] It should be noted that, in practice, the timing of the business layer sending the expected time to the policy monitoring module is not based on... Figure 14 The above is for illustrative purposes only. For example, the business layer may also send the desired time to the policy monitoring module after the subsequent algorithm initialization is successful, as shown in step 1305 below.

[0332] 1303. The strategy monitoring module sends initialization parameters to the channel interface of the media middleware framework layer.

[0333] The initialization parameters are used to initialize the algorithms required by the business logic in the hardware abstraction layer. In the one-click blockbuster service, the required algorithms include highlight segment algorithms, face detection algorithms, video acceleration algorithms, and image super-resolution algorithms.

[0334] 1304. Channel interface initializes the algorithm in the algorithm module of the hardware abstraction layer.

[0335] For example, the channel interface can initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, image super-resolution algorithm, etc. required for the one-click blockbuster service based on the initialization parameters.

[0336] It should be noted that data can be passed between the media middleware framework layer and the hardware abstraction layer through the application framework layer, as not shown in the figure.

[0337] 1305. The algorithm module returns an initialization success message to the business layer through the channel interface and the policy monitoring module.

[0338] After receiving the initialization success message, the policy monitoring module can determine the analysis strategy for videos and / or images, and complete the analysis through the channel interface, as shown in 1309-1314 below:

[0339] 1309. The strategy monitoring module allocates the duration of videos and / or images according to the desired time.

[0340] Typically, the analysis time for a single image is fixed. The policy monitoring module can determine the total analysis time for each image based on its analysis time and the number of images. Furthermore, the policy monitoring module can determine the total video analysis time based on the expected time and the total image analysis time.

[0341] Furthermore, in step 3 above, the electronic device determines the maximum number of frame extractions based on the time consumption requirement, which can be understood as the maximum number of supported frame extractions determined based on the total duration of video analysis, and thus used to determine the actual number of frame extractions for each video.

[0342] 1310. The strategy monitoring module dynamically sets the analysis strategy for the current video and / or image.

[0343] The analysis strategies include frame-by-frame analysis and frame-by-frame analysis. Furthermore, frame-by-frame analysis includes one round of frame-by-frame analysis or two rounds of frame-by-frame analysis.

[0344] The strategy monitoring module can sequentially set analysis strategies for each video and / or image. After analyzing each video and / or image, the module can then set the analysis strategy for the next video and / or image.

[0345] For example, the policy monitoring module can calculate the remaining duration of video analysis based on the time already consumed and the total duration of video analysis, and dynamically adjust the analysis strategy for the remaining videos based on the remaining duration and the remaining videos. It should be noted that, typically, after completing the first round of frame extraction for all videos, the policy monitoring module can determine whether to perform a second round of frame extraction based on the remaining duration.

[0346] In one specific implementation, the strategy monitoring module can perform the following process through steps 1309-1310: calculation of the expected number of frames extracted in step 1; calculation of the actual number of frames extracted in step 2; determination of the number of rounds of frame extraction and the determination of the frame extraction time point in step 3. For example, the strategy monitoring module can calculate the frame extraction time point based on the video duration and the number of frames extracted in the first round, or calculate the frame extraction time point based on the similarity between adjacent frames after the first round of frame extraction.

[0347] 1311. The policy monitoring module sends policy parameters to the channel interface.

[0348] For example, if a frame-by-frame analysis strategy is used, the strategy parameters include instructions for frame-by-frame analysis. After receiving the strategy parameters, the channel module can decode and process the video frame by frame, as shown in sections 1312-1313 below. If a frame-by-frame analysis strategy is used, the strategy parameters include the frame-by-frame extraction time point. After receiving the strategy parameters, the channel module can extract the corresponding video frames, decode and process them, again as shown in sections 1312-1313 below.

[0349] Furthermore, the strategy parameters can also include file descriptors, which are used to obtain the corresponding materials through the subsequent channel interface.

[0350] 1312. Channel interface decodes image frames.

[0351] The channel interface can obtain and decode the corresponding materials based on the file descriptors of video and / or images. For example, the channel interface can decode image frames based on frame-by-frame extraction, or it can decode image frames frame by frame.

[0352] In this context, for video, an image frame refers to the individual video frames that make up the video. For images, an image frame refers to the image itself.

[0353] 1313. The channel interface performs format conversion and resolution reduction.

[0354] Typically, to save analysis time, the channel interface can perform down-resolution processing on image frames.

[0355] Furthermore, since the down-resolution algorithm only supports preset formats, the channel interface can first convert image frames from other formats to the preset format before performing down-resolution processing.

[0356] It is understandable that after completing the processes 1312 and 1313 above, the channel interface can store the processed image frames.

[0357] 1314. The channel interface sends the frame buffer address to the algorithm module.

[0358] The frame buffer address can be used by the algorithm module to obtain the image frame after the channel interface has been processed (such as format conversion or resolution adjustment).

[0359] It should be noted that if the current data is a video, the above steps 1312-1314 can be executed repeatedly to continue decoding, format conversion, and resolution reduction of the undecoded parts of the video, and the frame buffer address can be sent to the algorithm module.

[0360] 1315. The algorithm module executes the algorithm processing.

[0361] For example, the algorithm module can obtain each processed image frame from the frame buffer address and analyze the image frames, such as detecting faces, bodies, and voices in the image frames, and performing aesthetic scoring on the image frames.

[0362] 1316. The algorithm module returns the processing results to the policy monitoring module through the channel interface.

[0363] 1317. The strategy monitoring module analyzes highlight segments.

[0364] After obtaining all the processing results of a video, such as the results of two rounds of frame extraction, the strategy monitoring module can analyze the highlight segments of the video.

[0365] For example, the strategy monitoring module can execute step 4 above to complete the scene segmentation, and execute step 5 to extract highlight segments. This yields the highlight segments of the video, such as determining the start and end frames of the highlight segments.

[0366] It should be noted that after processing an image or video, you can repeat steps 1310-1317 to continue analyzing unanalyzed videos and / or images.

[0367] 1318. The strategy monitoring module returns the analysis results to the application function layer.

[0368] After completing the analysis of all videos and images, the policy monitoring module can summarize the analysis results and send them to the application function layer.

[0369] 1319. The application function layer performs clipping and filtering.

[0370] The application functionality layer can edit and filter videos and / or images based on the analysis results to obtain all highlight images and highlight clips. For example, the application functionality layer can edit videos to obtain highlight clips and filter images to obtain highlight images.

[0371] 1320. The application function layer calls the basic function layer to generate the finished product.

[0372] The application function layer can send highlight clips and highlight images to the basic function layer, and then the basic function layer stitches the highlight clips and highlight images together to generate a finished image.

[0373] 1321. The basic functional layer sends the complete video to the business layer.

[0374] Afterwards, the business layer can display the finished product to the user, allowing the user to view the final effect of the one-click blockbuster.

[0375] This application also provides an electronic device, which may include a display screen, a memory, and one or more processors (such as a CPU, GPU, NPU, etc.). The display screen, memory, and processor are coupled. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the device in the above method embodiments.

[0376] like ​As shown in the illustration, this application also provides a chip system. The chip system 1400 includes at least one processor 1401 and at least one interface circuit 1402. The at least one processor 1401 and the at least one interface circuit 1402 are interconnected via lines. The processor 1401 is used to support an electronic device in implementing the various steps in the above method embodiments, and the at least one interface circuit 1402 can be used to receive signals from other devices (e.g., memory) or to send signals to other devices (e.g., a communication interface). The chip system may include a chip and may also include other discrete devices.

[0377] This application also provides a computer storage medium including instructions that, when executed on the electronic device, cause the electronic device to perform the steps in the above method embodiments.

[0378] This application also provides a computer program product including instructions that, when executed on the electronic device, cause the electronic device to perform the steps in the method embodiments described above.

[0379] The technical effects of the chip system, computer storage medium, and computer program product are similar to those in the preceding method embodiments.

[0380] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0381] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0382] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0383] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms.

[0384] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located on one device or distributed across multiple devices. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0385] In addition, the functional modules in the various embodiments of this application can be integrated into one device, or each module can exist physically separately, or two or more modules can be integrated into one device.

[0386] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted from one computer storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium accessible to a computer or a data storage device including one or more servers, data centers, etc., that can be integrated with the medium. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0387] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for extracting highlight fragments, characterized in that, include: The first round of frame extraction is performed on the video to be analyzed, extracting N frames, where N≥2 and N is an integer; A second round of frame extraction is performed between adjacent frames in the N frames whose similarity is lower than the first similarity threshold, and at least one frame is extracted. Based on the frame extraction results of the first round and the second round, the scene information of the video to be analyzed is calculated, and the scene information includes multiple scenes. Extract highlight segments from at least two of the multiple storyboards.

2. The method according to claim 1, characterized in that, Before performing the first round of frame extraction on the video to be analyzed, extracting N frames, the method further includes: Calculate the minimum number of frames extracted and the actual number of frames extracted in the video to be analyzed. The minimum number of frames extracted represents the theoretical minimum number of frames to be extracted in the video to be analyzed, and the actual number of frames extracted represents the actual number of frames extracted in the video to be analyzed. Wherein, when the actual number of frames drawn exceeds the minimum number of frames drawn, N is the minimum number of frames drawn, and the total number of frames drawn in the second round does not exceed the difference between the actual number of frames drawn and the minimum number of frames drawn.

3. The method according to claim 2, characterized in that, The method further includes: If the actual number of frames extracted does not exceed the minimum number of frames extracted, N is the actual number of frames extracted, and a second round of frame extraction is not performed on the video to be analyzed.

4. The method according to claim 2 or 3, characterized in that, The minimum number of frames to be extracted is related to the duration of the video to be analyzed.

5. The method according to any one of claims 2-4, characterized in that, Before calculating the minimum number of frames extracted and the actual number of frames extracted from the video to be analyzed, the method further includes: Receive the request to extract highlight segments, which includes time requirements and all videos to be analyzed; Calculate the maximum number of frame draws supported by the electronic device based on the time consumption requirement; The maximum supported number of frame extractions is allocated to each video to be analyzed, thus obtaining the actual number of frame extractions for each video to be analyzed.

6. The method according to claim 5, characterized in that, The step of allocating the maximum supported frame extraction number to each video to be analyzed, to obtain the actual frame extraction number of each video to be analyzed, includes: If the maximum supported number of frame extractions is less than the sum of the minimum number of frame extractions of all videos to be analyzed, a preset number of frames are first allocated to each video to be analyzed, and the remaining frames in the maximum supported number of frame extractions are allocated according to the duration ratio of all videos to be analyzed to obtain the actual number of frame extractions for each video to be analyzed. If the maximum supported frame extraction number is greater than or equal to the sum of the minimum frame extraction numbers of all videos to be analyzed, but less than the sum of the maximum frame extraction numbers of all videos to be analyzed, first allocate the minimum frame extraction number to each video to be analyzed, and then allocate the remaining frames in the maximum supported frame extraction number according to the duration ratio of all videos to be analyzed to obtain the actual frame extraction number of each video to be analyzed. If the maximum supported frame extraction number is greater than or equal to the sum of the maximum frame extraction numbers of all videos to be analyzed, the maximum frame extraction number is assigned to each video to be analyzed, and the actual frame extraction number of each video to be analyzed is obtained.

7. The method according to any one of claims 1-6, characterized in that, The step of calculating the scene information of the video to be analyzed based on the frame extraction results of the first and second rounds of frame extraction includes: Frames with similarity higher than the second similarity threshold in the first and second rounds of frame sampling are merged to obtain a consistent region; among them, frames that are not merged are recorded as lonely frames. Adjacent consistent regions are merged, and consistent regions and lonely frames are merged to obtain noisy consistent regions; among them, consecutive lonely frames that are not merged constitute unstable regions. The storyboard information is obtained based on the noisy, consistent region and the unstable region.

8. The method according to claim 7, characterized in that, The process of merging adjacent consistent regions and merging consistent regions with isolated frames to obtain a noisy consistent region includes: If the similarity between adjacent consistent regions is higher than a third similarity threshold, or if the similarity between adjacent consistent regions is higher than a third similarity threshold and the time interval between adjacent consistent regions is less than a first duration, the adjacent consistent regions are merged to obtain a noisy consistent region; and / or, If the similarity between the consistent region and the lonely frame is higher than the fourth similarity threshold, or if the similarity between the consistent region and the lonely frame is higher than the fourth similarity threshold and the time interval between the consistent region and the lonely frame is less than the second duration, the consistent region and the lonely frame are merged to obtain a noisy consistent region.

9. The method according to claim 8, characterized in that, The first duration includes the sum of the lengths of the adjacent consistent regions.

10. The method according to any one of claims 1-9, characterized in that, Extracting highlight segments from at least two of the plurality of shots includes: According to the priority of the multiple scenes, extract the highlight segments including the optimal frame from at least two scenes; The optimal frame includes preset content.

11. The method according to claim 10, characterized in that, The multiple segments include at least two of the following: persistent stability segments, non-persistent stability segments, mutation segments, and alternative segments; The priority of the continuously stable segment, the non-continuously stable segment, and the alternative segment decreases in that order.

12. The method according to claim 10 or 11, characterized in that, The preset content includes at least one of the following: face, body, and voice.

13. The method according to any one of claims 10-12, characterized in that, The method further includes: Aesthetic scoring is performed on frames other than the optimal frame; After extracting the highlight fragment including the optimal frame from at least two scenes according to the priority of the plurality of scenes, the method further includes: Highlight segments were extracted based on aesthetic scores, from highest to lowest.

14. An electronic device, characterized in that, The device includes one or more processors and a memory coupled to the processors; the memory stores computer program code, which includes instructions; when the instructions are executed by the processor, the electronic device performs the method of any one of claims 1-13.

15. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1-13.

16. A computer program product, characterized in that, Includes instructions that, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1-13.