Video processing method, electronic equipment and storage medium

By analyzing and by-selecting highlight clips in video processing, the problem of missing important clips in the prior art is solved, ensuring the integrity of video quality and duration.

CN120343183AActive Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202410042635.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-18
Estimated Expiration
2044-01-10

AI Technical Summary

Technical Problem

The prior art is prone to missing important clips when extracting video highlights, resulting in a decline in video quality.

Method used

By analyzing multiple image materials selected by the user, obtain the highlight clips of each video, and perform by-selecting when the total duration of the highlight clip is insufficient until the preset total duration requirement is met, ensuring the duration and quality of the highlight clips.

Benefits of technology

It effectively avoids the omission of highlight clips, ensures the duration and quality of the target video set after splicing, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343183A_ABST
    Figure CN120343183A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method, electronic equipment and a storage medium, and relates to the technical field of video data, and the method comprises the steps: carrying out the highlight fragment analysis of a plurality of image materials selected by a user through the electronic equipment, and obtaining a highlight fragment of each video in the plurality of image materials, if the sum of the durations of the highlight clips is smaller than the preset suggested total duration of the highlight clips, the electronic equipment can select the highlight clips from all videos in a supplementary mode until the sum of the durations of the highlight clips is equal to or larger than the suggested total duration of the highlight clips, or no supplementary image frames exist in the videos. Wherein the highlight segments of the videos in the plurality of image materials are used for splicing to obtain a target video set. According to the scheme, the duration of the output highlight clip meets the duration requirement, and the problem of missing of the highlight clip is avoided, so that the output target video set has relatively high video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of video data, and in particular, to a video processing method, an electronic device, and a storage medium. Background Art

[0002] With the development of image and video processing technologies, users can trigger an electronic device to further process photos and videos in an album. For example, the electronic device can splice multiple image materials (including multiple pictures and / or multiple videos) to obtain a complete spliced video.

[0003] For example, the electronic device or a third-party video processing software in the electronic device may have a function of generating a video with one click (or referred to as a service of generating a video with one click). The function of generating a video with one click can automatically analyze and extract highlight segments of videos in the selected multiple image materials by an algorithm; then, automatically generate a clipped video based on the extracted highlight segments. Among them, the above-mentioned highlight segments are also called wonderful segments, which refer to video segments composed of single-frame images or continuous multiple-frame images extracted from the above materials for recording wonderful moments. The wonderful moments can be moments when wonderful actions corresponding to a person's smiling face, a moment of winning a championship, an airplane landing, etc. occur.

[0004] However, when extracting highlight segments of a video in the prior art, it is easy to cause the problem of missing other highlight segments in the video. Summary of the Invention

[0005] Embodiments of the present application provide a video processing method, an electronic device, and a storage medium. When the total duration of the highlight segments of a video is less than a preset total duration of the highlight segments, the highlight segments can be supplemented and selected to avoid missing highlight segments in the video. At the same time, the duration and quality of the target video set spliced from the highlight segments are ensured.

[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions.

[0007] In a first aspect, a video processing method is provided. The method includes:

[0008] An electronic device receives a selection operation of a user for multiple image materials in a gallery. Among them, the multiple image materials include videos; or, the multiple image materials may further include videos and pictures.

[0009] The electronic device analyzes each video in response to the selection operation to obtain a highlight segment of each video.

[0010] If the sum of the durations of the highlight segments of all videos in the multiple image materials is less than the recommended total duration of the highlight segments, supplement and select highlight segments from the image frames not covered by the non-highlight segments of all videos until the output condition of the highlight segments is met.

[0011] Among them, the output conditions for satisfying the highlight segments include that the sum of the durations of all highlight segments is equal to or greater than the recommended total duration of the highlight segments, or there are no selectable image frames in the video.

[0012] The highlight segments of each video in multiple image materials are used to splice to obtain a target video set; the recommended total duration of the highlight segments is the recommended value of the sum of the durations of the highlight segments of all videos in multiple image materials.

[0013] In this application, after the electronic device analyzes the highlight segments of multiple image materials selected by the user and obtains the highlight segments of each video in the multiple image materials, if the sum of the durations of the highlight segments is less than the recommended total duration of the preset highlight segments, the electronic device can select additional highlight segments from all videos to satisfy that the sum of the durations of the highlight segments is equal to or greater than the recommended total duration of the highlight segments, or there are no selectable image frames in the video. Among them, the highlight segments of each video in multiple image materials are used to splice to obtain a target video set. In this solution, the problem of missing other highlight segments in the video is effectively avoided, and at the same time, the duration and quality of the target video set spliced from the highlight segments are guaranteed.

[0014] In a possible implementation manner of the first aspect, analyzing each video to obtain the highlight segment of each video includes:

[0015] Performing an overview analysis of a preset number of image frames for each video to obtain the scores of each image frame. Among them, the score is an aesthetic score for the corresponding image frame.

[0016] Based on the scores of the image frames in each video, determine the highlight segment of the corresponding video.

[0017] Among them, the highlight segment includes the first image frame, and the image frames within the first preset duration before and after the first image frame; the first image frame is the image frame with the highest score in the overview analysis.

[0018] In this application, the highlight segments in each video are determined by the first image frame with the highest score, which can ensure that at least one highlight segment in each video has the highest value. In this way, the effect and quality of the video set including the highlight segments output are relatively good.

[0019] In another possible implementation manner of the first aspect, selecting additional highlight segments from the image frames not covered by the non-highlight segments of all videos until the output conditions for the highlight segments are satisfied, including:

[0020] Selecting additional highlight segments from the second image frames of all videos until the output conditions for the highlight segments are satisfied. Among them, the score of the second image frame is greater than the first threshold, and the second image frame is not covered by the selected highlight segments.

[0021] In this application, highlight segments are selected from the second image frames with higher scores, and the selected highlight segments have better effects. After selecting highlight segments based on the second image frames, when the highlight segments meet the output conditions, the duration and quality of the output highlight segments are ensured, and the problem of omission of highlight segments is also avoided.

[0022] In another possible implementation manner of the first aspect, highlight segments are selected from the second image frames of all videos until the output conditions of the highlight segments are met, including:

[0023] According to the order of the scores of each second image frame from high to low, the highlight segments corresponding to each second image frame are obtained in sequence; the highlight segments corresponding to the second image frame include the second image frame, and the image frames within the second preset duration before and after the second image frame.

[0024] If the sum of the durations of the highlight segments of all videos after being selected from the second image frames is greater than or equal to the recommended total duration of the highlight segments, the selection operation of the highlight segments ends.

[0025] If the sum of the durations of the highlight segments of all videos after being selected from the second image frames is less than the recommended total duration of the highlight segments, and if the sum of the durations of the highlight segments is greater than or equal to the recommended total duration of the highlight segments at the first magnification, and there are no second image frames available for selection of highlight segments in the video, the selection operation of the highlight segments ends.

[0026] If the sum of the durations of the highlight segments of all videos after being selected from the second image frames is less than the recommended total duration of the highlight segments at the first magnification, highlight segments are selected from the third image frames of all videos until the output conditions of the highlight segments are met.

[0027] Wherein, the third image frame is an image frame with a score greater than or equal to the second threshold and not covered by the highlight segments; the second threshold is less than the first threshold.

[0028] In this application, there are no second image frames available for selection in the video. In the case where the sum of the durations of the highlight segments is less than the recommended total duration of the highlight segments at the first magnification, highlight segments are selected from the third image frames, so that highlight segments that meet the output conditions can be obtained, ensuring the duration and quality of the output highlight segments.

[0029] In another possible implementation manner of the first aspect, the method further includes:

[0030] If the duration of the highlight segment corresponding to the second image frame is greater than the maximum duration of the preset highlight segment, starting from the left and right boundary positions of the highlight segment, a preset distance is respectively reduced towards the midpoint position of the highlight segment until the duration of the reduced highlight segment is equal to or less than the maximum duration of the highlight segment.

[0031] Among them, the maximum duration of the preset highlight segment is the maximum duration allowed for a highlight segment to ensure the complete highlight effect of the video.

[0032] In this application, the duration of the highlight segment is verified. If the duration is greater than the maximum duration of the highlight segment, the duration of the highlight segment is reduced to control the duration of the highlight segment so that the duration of the highlight segment is less than the maximum duration of the highlight segment, avoiding an increase in the analysis volume due to the long duration of the highlight segment.

[0033] In another possible implementation manner of the first aspect, the method further includes:

[0034] If the duration of the highlight segment corresponding to the second image frame is less than the minimum duration of the preset highlight segment, the second image frame is discarded.

[0035] Among them, the minimum duration of the preset highlight segment is the minimum duration required for a highlight segment to ensure the complete highlight effect of the video.

[0036] In this application, the duration of the highlight segment is verified. If the duration is less than the minimum duration of the highlight segment, it is considered that the highlight segment cannot represent the highlight moment well and the value of the highlight segment is relatively low. To avoid ineffective analysis of such a highlight segment, the image frame corresponding to the highlight segment can be discarded, and subsequently, no supplementary selection of the highlight segment is performed based on this image frame.

[0037] In another possible implementation manner of the first aspect, highlight segments are supplemented and selected from the third image frames of all videos until the output conditions of the highlight segments are met, including:

[0038] The highlight segments corresponding to each third image frame are sequentially obtained in the order of the scores of the third image frames from high to low; the highlight segment corresponding to the third image frame includes the third image frame and the image frames within the third preset duration before and after the third image frame.

[0039] If the sum of the durations of the highlight segments of all videos after supplementing and selecting from the third image frame is greater than or equal to the recommended total duration of the highlight segment, the operation of supplementing and selecting the highlight segment ends.

[0040] If the sum of the durations of the highlight segments of all videos after supplementing and selecting from the third image frame is less than the recommended total duration of the highlight segment, and if the sum of the durations of the highlight segments is greater than or equal to twice the recommended total duration of the highlight segment, there is no third image frame in the video for supplementing and selecting the highlight segment, and the operation of supplementing and selecting the highlight segment ends.

[0041] If the sum of the durations of the highlight segments of all videos after supplementing and selecting from the third image frame is less than the recommended total duration of the highlight segments at the second magnification factor, extend the duration of the highlight segments until the output conditions for the highlight segments are met.

[0042] In this application, there are no second image frames or third image frames available for supplementing and selecting in the video. When the sum of the durations of the highlight segments is less than the recommended total duration of the highlight segments at the second magnification factor, extend the duration of the first target highlight segment, so that the obtained highlight segments meet the output conditions, ensuring the duration and quality of the output highlight segments, and also avoiding the problem of omission of highlight segments.

[0043] In another possible implementation of the first aspect, the method further includes:

[0044] If the duration of the highlight segment corresponding to the third image frame is greater than the preset maximum duration of the highlight segment, reduce the second preset duration from the left and right boundary positions of the highlight segment to the midpoint position of the highlight segment until the duration of the highlight segment after reducing the duration is equal to or less than the maximum duration of the highlight segment.

[0045] Wherein, the preset maximum duration of the highlight segment is the maximum duration allowed for a highlight segment to ensure the complete highlight effect of the video.

[0046] In this application, check the duration of the highlight segment. If the duration is greater than the maximum duration of the highlight segment, reduce the duration of the highlight segment to control the duration of the highlight segment so that the duration of the highlight segment is less than the maximum duration of the highlight segment, avoiding the increase in the analysis volume caused by the long duration of the highlight segment.

[0047] In another possible implementation of the first aspect, the method further includes:

[0048] If the duration of the highlight segment corresponding to the third image frame is less than the preset minimum duration of the highlight segment, discard the third image frame.

[0049] Wherein, the preset minimum duration of the highlight segment is the minimum duration required for a highlight segment to ensure the complete highlight effect of the video.

[0050] In this application, check the duration of the highlight segment. If the duration is less than the minimum duration of the highlight segment, it is considered that the highlight segment cannot represent the highlight moment well and the value of the highlight segment is relatively low. To avoid ineffective analysis of such highlight segments, the image frame corresponding to the highlight segment can be discarded, and subsequent supplementing and selecting of highlight segments are not based on this image frame either.

[0051] In another possible implementation of the first aspect, it includes:

[0052] Select a first target highlight segment from the highlight segments, move the left and right boundary positions of the first target highlight segment to both sides respectively to obtain a second target highlight segment with an extended duration; the second magnification factor is smaller than the first magnification factor.

[0053] Among them, the first target highlight segment is a highlight segment in the highlight segments whose duration is less than or equal to the maximum duration of the preset highlight segments, and there are expandable regions before and after the target highlight segment, and the distances between the left and right boundary positions of the target highlight segment and the boundary positions of adjacent segments are greater than the preset distance.

[0054] The expandable region is a connected region in the video except for the highlight segments and the regions with scores less than the third threshold. The connected region is a region with the same score and continuous in time, and there is no overlapping region between the connected regions.

[0055] The maximum duration of the preset highlight segments is the maximum duration allowed for a highlight segment to ensure the complete highlight effect of the video.

[0056] In this application, there are no second image frames and third image frames that can be selected in the video. When the sum of the durations of the highlight segments is less than the recommended total duration of the highlight segments at the second magnification factor, the duration of the first target highlight segment is extended, so that the highlight segments that meet the output conditions can be obtained, ensuring the duration and quality of the output highlight segments, and also avoiding the problem of omission of highlight segments.

[0057] In another possible implementation manner of the first aspect, moving the left and right boundary positions of the first target highlight segment to both sides respectively to obtain a second target highlight segment with an extended duration includes:

[0058] Take the difference between the recommended total duration of the highlight segments at the second magnification factor and the sum of the durations of all the highlight segments of the video after supplementing and selecting from the third image frame as the required extended duration of the highlight segments.

[0059] According to the required extended duration of the highlight segments, allocate an extended quota duration for each first target highlight segment; the extended quota duration is the duration allowed for the first target highlight segment to be extended in the corresponding expandable region.

[0060] Within the expandable region of the target highlight segment, move the left and right boundary positions of the first target highlight segment to both sides by 1 / 2 of the extended quota duration to obtain the second target highlight segment.

[0061] In this application, extending the duration of the first target highlight segment based on the extended quota duration can ensure that the duration of the extended second target highlight segment will not be too long, and the highlight segments that meet the output conditions can be obtained, ensuring the duration and quality of the output highlight segments, and also avoiding the problem of omission of highlight segments.

[0062] In another possible implementation of the first aspect, allocating an extended quota duration for each first target highlight segment according to the required extended duration of the highlight segment includes:

[0063] Determining the extended quota duration of each target highlight segment according to the required extended duration of the highlight segment and the duration of each first target highlight segment.

[0064] If the sum of the durations of all highlight segments and the cumulative sum of the extended quota durations are greater than the preset maximum total duration of the highlight segments, update the value of the extended quota duration to the difference between the maximum total duration of the highlight segments and the sum of the durations of all highlight segments.

[0065] If the sum of the durations of all highlight segments and the cumulative sum of the extended quota durations are less than or equal to the preset maximum total duration of the highlight segments, the value of the extended quota duration remains unchanged.

[0066] Wherein, the preset maximum total duration of the highlight segments is the maximum duration allowed for the highlight segments of all videos.

[0067] In this application, determining the extended quota duration of the first target highlight segment based on the sum of the durations of all highlight segments and the cumulative sum of the extended quota durations can ensure that the duration of the second target highlight segment after extension does not exceed the preset maximum total duration of the highlight segments, and can ensure the duration and quality of the output highlight segments under the limited time consumption and limited performance of the electronic device.

[0068] In another possible implementation of the first aspect, the method further includes:

[0069] After extending the duration of the highlight segment, if the sum of the durations of all highlight segments is greater than or equal to the recommended total duration of the highlight segments at the second ratio, there is no third image frame available for supplementary selection in the video, and the duration of the highlight segment in the video cannot be extended any further, end the supplementary selection operation of the highlight segment.

[0070] In this application, after there is no second image frame or third image frame available for supplementary selection in the video and after extending the duration of the first target highlight segment, end the supplementary selection operation of the highlight segment. The obtained highlight segment meets the output conditions, ensuring the duration and quality of the output highlight segment and avoiding the problem of omission of highlight segments.

[0071] In a second aspect, an electronic device is provided, which includes a memory, a display screen, and one or more processors; the memory, the display screen are coupled to the processor; computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method according to any one of the above first aspects.

[0072] In a third aspect, a computer-readable storage medium is provided, in which instructions are stored. When the instructions run on an electronic device, the electronic device can execute the method described in any one of the above first aspects.

[0073] In a fourth aspect, a computer program product containing instructions is provided. When the computer program product runs on an electronic device, the electronic device can execute the method described in any one of the above first aspects.

[0074] In a fifth aspect, an embodiment of the present application provides a chip. The chip includes a processor, and the processor is used to call a computer program in a memory to execute the method as described in the first aspect.

[0075] It can be understood that for the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect, reference can be made to the beneficial effects in the first aspect and any of its possible design manners, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 FIG. is a schematic diagram of an application scenario of a video processing method provided by an embodiment of the present application;

[0077] Figure 2 FIG. is a schematic diagram of another application scenario of a video processing method provided by an embodiment of the present application;

[0078] Figure 3 FIG. is a schematic diagram of another application scenario of a video processing method provided by an embodiment of the present application;

[0079] Figure 4 FIG. is a schematic diagram of another application scenario of a video processing method provided by an embodiment of the present application;

[0080] Figure 5 FIG. is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;

[0081] Figure 6 FIG. is a schematic diagram of the software structure of an electronic device provided by an embodiment of the present application;

[0082] Figure 7 FIG. is a schematic diagram of the flowchart of a video processing method provided by an embodiment of the present application;

[0083] Figure 8 FIG. is a schematic diagram of the flowchart of analyzing a video to obtain a highlight segment in a video processing method provided by an embodiment of the present application;

[0084] Figure 9A schematic diagram of a target area provided by an embodiment of the present application;

[0085] Figure 10 Another schematic diagram of a target area provided by an embodiment of the present application;

[0086] Figure 11 A schematic diagram of a high - light segment provided by an embodiment of the present application;

[0087] Figure 12 Another schematic diagram of a high - light segment provided by an embodiment of the present application;

[0088] Figure 13 A schematic flowchart of supplementing a high - light segment in a video processing method provided by an embodiment of the present application;

[0089] Figure 14 A schematic diagram of supplementing a high - light segment in Video 1 provided by an embodiment of the present application;

[0090] Figure 15 Another schematic diagram of supplementing a high - light segment in Video 1 provided by an embodiment of the present application;

[0091] Figure 16 Another schematic diagram of supplementing a high - light segment in Video 1 provided by an embodiment of the present application;

[0092] Figure 17 A schematic diagram of an expandable area 1 of high - light segment 1 of Video 1 provided by an embodiment of the present application;

[0093] Figure 18 A schematic diagram of the allocation of required expansion duration provided by an embodiment of the present application;

[0094] Figure 19 Another schematic diagram of the allocation of required expansion duration provided by an embodiment of the present application;

[0095] Figure 20 A schematic diagram of expanding high - light segment 2 provided by an embodiment of the present application. Detailed implementation manners

[0096] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "the", "above-mentioned", "this" and "this one" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist; for example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0097] Reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.

[0098] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0099] First, some nouns or terms involved in the present application are explained.

[0100] A highlight segment, also known as a wonderful segment, refers to a single-frame image or a video segment composed of consecutive multiple-frame images extracted from a video or a picture for recording wonderful moments. The wonderful moments can be the moments when a person smiles, wins a championship in a competition, takes off in a sport, lands an airplane, or scores a goal in a ball game corresponding to the wonderful actions.

[0101] The one - click video generation function means that the electronic device, in response to the user's selection operation on one or more image materials, automatically analyzes the highlight segments in the image materials through algorithms and combines the highlight segments into a clipped video set. That is to say, the electronic device can extract multiple highlight segments from one or more image materials and synthesize the multiple highlight segments into a video set. Among them, the image materials selected by the user can be pictures or videos; or the image materials can include pictures and videos.

[0102] Among them, the process of the electronic device selecting highlight segments from videos or images can include: obtaining aesthetic scoring parameters such as the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate value of each frame of the image, and performing aesthetic scoring on each frame of the image according to these aesthetic scoring parameters. The electronic device can use the single frame image or consecutive multiple frame images with the highest score as a highlight segment.

[0103] Currently, the one - click video generation function supports the editing of materials such as pictures and videos. In the process of implementing the one - click video generation function, the user may select a relatively large number of pictures or a video with a relatively long duration from the image materials. In this case, if the electronic device randomly selects some image materials from the selected image materials by the user for highlight segment extraction, it is likely to cause the problem of missing other highlight segments in the video.

[0104] In view of the above problems, the embodiments of the present application provide a video processing method. After the electronic device analyzes the highlight segments of multiple image materials selected by the user and obtains the highlight segments of each video in the multiple image materials, if the sum of the durations of the highlight segments is less than the recommended total duration of the preset highlight segments, the electronic device can supplement and select highlight segments from all the videos until the sum of the durations of the highlight segments is equal to or greater than the recommended total duration of the highlight segments, or there are no image frames that can be supplemented and selected in the video. Among them, the highlight segments of each video in the multiple image materials are used to splice and obtain the target video set. This solution effectively avoids the problem of missing other highlight segments in the video. At the same time, it ensures the duration and quality of the target video set spliced from the highlight segments.

[0105] The video processing method provided by the embodiments of the present application can be applied to an electronic device with an image processing function. It should be noted that the image materials for one - click video generation in the embodiments of the present application can include pictures and videos. The user can use the one - click video generation function to generate a video set from the highlight segments of multiple pictures, or use the one - click video generation function to generate a video set from the highlight segments of multiple videos, or use the one - click video generation function to generate a video set from the highlight segments of multiple pictures and videos.

[0106] The above-mentioned electronic device can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city or wireless terminal in smart home, etc. The embodiments of this application do not limit the specific technologies and specific device forms adopted by the electronic device.

[0107] The following will introduce Figures 1 - 4 , taking the electronic device as a mobile phone as an example, the application scenarios and interface implementation for the electronic device to achieve the function of generating a video compilation with one click.

[0108] In an application scenario, the user can use the mobile phone to pre-capture multiple image materials, and the mobile phone stores the multiple image materials in the picture library. The mobile phone can also pre-obtain the image materials transmitted from other devices. As Figure 1 shown in (a) of Figure 1 , an icon of the picture library is displayed on the desktop of the mobile phone. When the user wants to use the mobile phone to generate a clipped video compilation based on multiple image materials, the user can click on the icon of the picture library on the mobile phone desktop. In response to the user's click operation on the icon, the mobile phone displays the picture library interface 101 as shown in (b) of Figure 1 . The picture library interface 101 includes an option of "One-Click Blockbuster". After that, in response to the user's click operation on the "One-Click Blockbuster" option shown in (b) of Figure 1 , the mobile phone can display the picture library interface 102 as shown in (c) of Figure 1 . The picture library interface 102 can include multiple recently captured image materials. The user can select Figure 1 any one or more of the multiple image materials shown in (c) of Figure 1 as candidate image materials for generating a video compilation with one click. For example, in response to the user's selection operation on some of the image materials in (c) of Figure 1 , the mobile phone can display as shown in Figure 1The gallery interface 103 shown in (d) therein. The gallery interface 103 includes all the image materials in the gallery. The gallery interface 103 may also include video generation options, such as the "√" tick option.

[0109] The mobile phone responds to the user's Figure 1 click operation on the "√" tick option shown in (d) therein, and can analyze the 5 image materials selected by the user in the gallery interface 103, select the highlight segments from each image material, and generate a video set based on the selected highlight segments. During this process, the mobile phone can display Figure 2 the gallery interface 201 shown therein. The gallery interface 201 includes the analysis material progress, so that the user can intuitively view the analysis progress.

[0110] In one example, after the mobile phone generates a video set, it can display Figure 3 the gallery interface 301 shown in (a) therein. The generated video set can be displayed in the gallery interface 301. The mobile phone can automatically play the video set in the gallery interface 301. In addition, as Figure 3 shown in (a) therein, the gallery interface 301 may also include a video export option 302 for supporting the export of the generated video set. In response to the user's click operation on the video export option 302, the mobile phone can save the video set in the gallery, so that the user can view the video set from the gallery. In response to the user's click operation on the video export option 302, the mobile phone can also display Figure 3 the video export interface 303 shown in (b) therein.

[0111] In one example, during the process of the user selecting image materials, the mobile phone can prompt the user, so that the user can know how many image materials are more appropriate to select. For example, as Figure 1 shown in (d) therein, the prompt message "Better effect with more than 6 image materials" is displayed in the gallery interface 103, so that the user can know at least how many image materials need to be selected to generate a video set with a better effect.

[0112] In one example, during the process of the user selecting image materials, the mobile phone can prompt the user, so that the user can know the maximum number of image materials that can be selected. For example, as Figure 4 shown therein, the prompt message "Up to 30 image materials can be selected" is displayed in the gallery interface 401, so that the user can know how many image materials can be selected.

[0113] In one example, after generating a video set, the mobile phone can also display other function options on the interface of the video set, so that the user can perform operations such as editing, adding special effects, and analyzing on the generated video set based on these function options. For example, as Figure 3As shown by 301 in [the figure], other function options may include, but are not limited to, options such as templates, music, clips, sharing, etc.

[0114] Taking the electronic device as a mobile phone as an example below, in combination with Figure 5 the hardware structure of the electronic device will be introduced.

[0115] Figure 5 The schematic diagram of the hardware structure of the electronic device 100 provided by the embodiment of the present application is shown. As Figure 5 shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a camera 193, a display screen 194, etc.

[0116] The processor 110 may include one or more processing units. For example, the processor 110 may include a controller, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, the controller may be the nerve center and command center of the mobile phone 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions. A memory may also be set in the processor 110 for storing instructions and data.

[0117] A memory may also be set in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0118] The wireless communication function of the electronic device 100 can be implemented by antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor, baseband processor, etc.

[0119] The electronic device 100 implements the display function through the GPU, display screen 194, application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering, such as rendering Figures 1 - 4 the schematic diagram of the operation interface shown, etc.

[0120] The display screen 194 is used to display the operation interface of the screen mirroring APP, screen mirroring images, screen mirroring videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), and a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0121] In this embodiment, the display screen 194 can be used to display Figure 1 the gallery interfaces 101, 102, 103 in; the display screen 194 can be used to display Figure 2 the gallery interface 201 in; the display screen 194 can be used to display Figure 3 the gallery interface 301 and the video export interface 303 in; the display screen 194 can be used to display Figure 4 the gallery interface 401 in, etc.

[0122] The electronic device 100 can implement the shooting function through the ISP, camera 193, video codec, GPU, display screen 194, application processor, etc.

[0123] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.

[0124] The camera 193 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0125] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0126] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0127] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0128] The external memory interface 120 can be used to connect to an external memory card to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function.

[0129] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs (APPs) required for at least one function (such as a camera APP, a gallery APP, and third-party video editing software, etc.). The data storage area can store data created during the use of the mobile phone 100 (such as photos or videos taken, mobile phone screenshots, mobile phone screen recording content, images downloaded from other devices, and video sets generated using the one-click video creation function, etc.). In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, and a universal flash storage (UFS), etc.

[0130] The electronic device 100 can implement audio functions through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor, etc.

[0131] For example, after a video set is generated using the one-click video creation function, the audio module 170 decodes the audio signal of the video set, and then the speaker 170A, also called the "loudspeaker", converts the audio electrical signal into a sound signal. In this way, the user can hear the background sound synchronized with the video in the highlight segment and the added video background music, etc.

[0132] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0133] Next, in combination with Figure 6 the software structure of the electronic device will be introduced.

[0134] Figure 6 is a schematic diagram of the software structure of the electronic device provided by the embodiments of the present application.

[0135] As Figure 6As shown, the electronic device can adopt a layered architecture, dividing the software into several layers, each layer having a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software layers of the software structure are sequentially divided from top to bottom into: application (APP) layer, media middle platform framework layer, application framework (FWK) layer, and hardware abstract layer (HAL).

[0136] The APP layer, simply referred to as the application layer, can include a series of application packages, such as cameras, galleries, third-party video editing software, calendars, maps, and navigation. When these application packages are run, they can access various service modules provided by the media middle platform framework layer and the application framework layer through the application programming interface (API), and execute corresponding intelligent services.

[0137] In some embodiments, the camera is used to capture photos, videos, slow-motion images, panoramic images, etc. in response to user operations. After these images are captured by the camera, or after the user triggers a phone screenshot, or after the user triggers a phone screen recording, or after the electronic device downloads images from other devices, the electronic device can save these images in the gallery, so that the user can perform video editing operations on the images in the gallery, such as one-click video creation operations.

[0138] In the embodiments of the present application, the gallery is sequentially divided from top to bottom into: a service layer, an application function layer, and a basic function layer.

[0139] Among them, the service layer, also known as the video editing service layer, provides multiple services (which can also be called functions) such as multi-shot video automatic video creation, one-shot multi-gain AI music short film, one-click video creation, and wonderful moments. These services are presented in the form of controls in the user interface (UI) of the gallery. By operating these controls, the user can trigger the gallery to perform corresponding video processing actions. For example, after the user selects image materials (the materials include pictures and / or videos), in response to the user's click operation on the one-click video creation control in the gallery, the gallery can call the underlying module to automatically analyze and extract the highlight segments in the pictures and / or videos through algorithms, and then combine the highlight segments into a clipped video set.

[0140] The application function layer includes an automatic editing framework. Each service in the service layer can call the automatic editing framework to provide automatic editing services for pictures and videos. Exemplarily, the automatic editing framework may include function modules such as segment optimization, storyline organization, layout splicing, and special effect beautification. The segment optimization is used to call the highlight segment analysis interface and the policy monitoring interface in the high media middleware framework layer to extract highlight segments from pictures and / or videos. The storyline organization is used to sequentially splice multiple pictures and / or videos in the form of a storyline based on the content of the pictures and / or videos. The layout splicing is used to adjust the interface layout of the pictures and / or videos. The special effect beautification is used to adjust the beautification effect of the pictures and / or videos, such as adjusting the picture brightness and beautifying the human face, etc.

[0141] The basic function layer is used to perform basic function processing on the clipped picture and / or video segments after the automatic editing framework clips multiple pictures and / or videos. Exemplarily, the basic function layer may include basic function modules such as video splicing, synthesis and saving, video effect rendering, and audio effect processing. Among them, the video splicing is used to splice multiple extracted highlight segments (where the highlight segments include pictures and / or videos). The synthesis and saving is used to store the video set obtained after splicing. The video effect rendering is used to add video effects to the video set obtained after splicing, such as adding style filters and themes to the video set. The audio effect processing is used to add sound effects to the video set obtained after splicing, such as adding background music to the video set.

[0142] The media middleware framework layer is a software layer set between the application layer and the application framework. The media middleware framework layer may include an analysis performance query interface, a highlight segment analysis interface, a policy monitoring interface, a pipeline interface, a theme summary interface, and an initialization interface. Among them, the analysis performance query interface is used to calculate the total duration of all videos according to the video analysis speed. The highlight segment analysis interface is used to call the policy monitoring interface to extract highlight segments. The theme summary interface is used to call the underlying algorithm to analyze the picture content of the highlight segments to determine the theme corresponding to the content of the highlight segments. The initialization interface is used to initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, and image super-resolution algorithm, etc. in the HAL layer.

[0143] The policy monitoring interface is used to configure the expected analysis duration for each video respectively according to the total duration of all videos, and set the analysis policy for each video of the source video based on parameters such as the expected analysis duration and the duration of each video. The channel interface is used to reduce the resolution of the video file according to the file descriptor and analysis policy of each video issued by the policy monitoring interface, and forward the data address of the downscaled video file to the hardware abstraction layer through the application framework layer, and then report the analysis result of the image frame returned by the hardware abstraction layer to the policy monitoring interface.

[0144] The theme summary interface is used to call the underlying algorithm to analyze the content of the high - light segment to determine the theme corresponding to the content of the high - light segment. The initialization interface is used to initialize the high - light segment algorithm, face detection algorithm, video acceleration algorithm, and image super - resolution algorithm, etc. in the HAL layer.

[0145] It should be noted that this application is described by taking the one - click video creation function provided by the gallery as an example, which does not limit the embodiments of this application. In actual implementation, a third - party video editing software can adopt the video processing method provided by the embodiments of this application to synthesize multiple pictures and videos selected by the user into a video set with one click.

[0146] The FWK layer, simply referred to as the framework layer, can be used to support the operation of each module in the media middle - platform framework layer. For example, the framework layer can include the one - click video creation interface, parameter management interface, image data transmission interface, theme analysis interface, and performance analysis interface, etc.

[0147] The HAL layer is a wrapper for the Linux kernel driver, providing interfaces upward. It hides the details of the hardware interfaces of specific platforms, provides a virtual hardware platform for the operating system, making it hardware - independent and portable across multiple platforms. For example, the hardware abstraction layer can include the chip analysis speed interface, high - light segment algorithm, face detection algorithm, video acceleration algorithm, and image super - resolution algorithm. Among them, the high - light segment algorithm is an image - processing algorithm provided by the image signal processor. This algorithm can perform an aesthetic score on each frame of the image according to the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate value, etc. The aesthetic score can be used as a basis for evaluating whether a frame of the image is a high - light segment.

[0148] In this embodiment, the high - light segment algorithm can analyze each image frame in the video to obtain the score of the image frame. For example, if the score of the image frame is greater than the preset score threshold, then the image frame and the image frames within the adjacent preset time period can be used as the high - light segment of the source video. Exemplarily, the preset score threshold can be 90. If the score of the image frame of the video is greater than 90, then the image frame and the image frames within the adjacent preset time period can be used as a high - light segment of the video. Or, for a video including the scores of N image frames, the image frame with the highest score and the image frames within the preset time period before and after this image frame are used as the high - light segment of the video.

[0149] In some other feasible embodiments, assuming that the value range of the score of an image frame is between 0 and 100, the score can also be divided into different levels of score results according to different value ranges of the score. For example, the first threshold is 80, and if the score of the image frame is greater than or equal to the first threshold (the value range of the score is between 80 and 100), it means that the score result of the image frame is "high". The second threshold can be 50, and if the score of the image frame is greater than or equal to the second threshold (the value range of the score is between 50 and 79), it means that the score result of the image frame is "relatively high". The third threshold can be 20, and if the score of the image frame is greater than or equal to the third threshold (the value range of the score is between 20 and 49), it means that the score result of the image frame is "medium". If the score of the image frame is less than the third threshold (the value range of the score is between 0 and 19), it means that the score result of the image frame is "low". Exemplarily, in one implementation manner, the image frames with scores greater than or equal to the first threshold / second threshold / third threshold can be used as optional image frames to extract the highlight segments of the video. In some other implementable manners, the image frame with the highest score can be used as the optional image frame to extract the highlight segments of the video.

[0150] It should be noted that Figure 6 The layers shown in the software structure and the components included in each layer do not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more layers than shown, such as the system library (FWKLIB) layer and the kernel layer. Each layer may include more or fewer components than shown. In addition, the above-mentioned various functional modules may also be combined into one functional module, and each layer may also be combined into one layer. For example, the highlight segment analysis may include policy monitoring. Another example is that the media middle platform framework layer may be set in the application framework layer.

[0151] It can be understood that in order for the electronic device to implement the video processing method in the embodiments of the present application, it includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application in combination with the embodiments.

[0152] In the embodiments of the present application, after the policy monitoring module in the media middle platform framework layer extracts the first highlight segment of each video in the image material, when the sum of the durations of the highlight segments of all videos does not reach the total recommended duration of the highlight segments, it can perform a supplementary selection operation on the analysis results of the analyzed image frames of all videos, so that the duration of the output highlight segment meets the duration requirements and has a high video quality, improving the functional experience of users using the one-click blockbuster function.

[0153] The policy monitoring module in the media middle platform framework layer can analyze the videos in the image material to obtain a highlight segment of each video in the image material. The policy monitoring module can send the image frames in each video to the algorithm module in the HAL layer for analysis and obtain the scores of the image frames from the algorithm in the HAL layer. Further, based on the image frame with the highest score and the corresponding regions before and after the image frame, a highlight segment of each video is determined. When the sum of the durations of the highlight segments of all videos does not reach the total recommended duration of the highlight segments, the policy monitoring module can supplement and select highlight segments from the analyzed image frames of all videos, so that the finally output highlight segment meets the output conditions of the highlight segment and has a high video quality, improving the functional experience of users using the one-click blockbuster function.

[0154] The following takes each module in the software structure diagram as shown in Figure 6 as an example of the execution subject of the audio-visual processing method to exemplarily illustrate the video processing method provided in the embodiments of the present application.

[0155] Figure 7 A method flow diagram before the electronic device policy monitoring module obtains analysis parameters in the video processing method is provided. This method can be applied to a one-click video creation scenario as shown in Figures 1 - 4 Taking the image material including pictures and videos as an example as shown in Figure 7 this method may include the following S01-S16.

[0156] S01. The service layer receives an operation enabling the one-click video creation function input by the user.

[0157] In this embodiment, the service layer refers to the one-click video creation module in the service layer. That is, the one-click video creation module in the service layer receives the operation enabling the one-click video creation function input by the user. For example, this operation may specifically be a click operation on the "one-click blockbuster" card as shown in (b) in Figure 1 .

[0158] S02. The service layer loads and displays candidate pictures and candidate videos.

[0159] S03. The service layer receives the operation of the user selecting multiple pictures and videos and receives the operation of the user inputting to determine to execute the one-click video creation function.

[0160] For example, the operation of a user selecting multiple pictures and videos can be a click operation on photos and videos as shown in (c) of Figure 1 , and the operation of the user inputting to determine the execution of the one - click video compilation function can be a click operation on the checkmark option as shown in (d) of Figure 1 .

[0161] S04. The business layer, through the application function layer, calls the initialization interface of the media middle - platform framework layer to initialize the relevant algorithms of the HAL layer.

[0162] In this embodiment, the relevant algorithms refer to the algorithms for the functions to be implemented by the business layer. Here, the business function is the one - click video compilation function, so the relevant algorithms are the relevant algorithms involved in the one - click video compilation function. For example, the relevant algorithms include the highlight segment algorithm, the face detection algorithm, the video acceleration algorithm, the image super - resolution algorithm, and so on.

[0163] S05. The initialization interface of the media middle - platform framework layer sequentially sends the initialization parameters to the algorithm modules of the HAL layer through the channel interface of the media middle - platform framework layer and the service interface of the FWK layer.

[0164] Among them, in the HAL layer, one algorithm corresponds to one algorithm interface, the FWK layer is provided with multiple service interfaces, and one service interface of the FWK layer corresponds to one algorithm interface of the HAL layer. Each service interface of the FWK layer plays a role in data passthrough between the algorithm interface of the HAL layer and the channel interface of the media middle - platform framework layer.

[0165] Since the initialization parameters involved in different algorithms are different, the algorithm initialization parameters sent to the algorithm interfaces through each service interface may be different.

[0166] S06. The algorithm modules of the HAL layer are initialized according to the initialization parameters.

[0167] S07. The algorithm modules of the HAL layer return an initialization success message to the channel interface of the media middle - platform framework layer through the service interface of the FWK layer.

[0168] S08. The channel interface of the media middle - platform framework layer calls the analysis speed interface of the HAL layer through the service interface of the FWK layer to obtain the chip analysis speed.

[0169] Among them, the service interface of the FWK layer can be a performance analysis interface.

[0170] S09. The analysis speed interface of the HAL layer returns the chip analysis speed to the channel interface of the media middle - platform framework layer through the service interface of the FWK layer.

[0171] S10, The channel interface of the media middleware framework layer returns the chip analysis speed to the initialization interface of the media middleware framework layer.

[0172] Among them, the chip analysis speed can represent the number of image frames analyzed by the image signal processor in a picture / video per unit time; or, the chip analysis speed can also represent the duration (single-frame analysis duration) of the image signal processor analyzing a picture / analyzing an image frame in a video. Therefore, the single-frame analysis duration of the image signal processor and the processing duration of the image signal processor for a picture can be calculated based on the chip analysis speed. The duration of the image signal processor for a picture and the single-frame analysis duration in a video may be different. Exemplarily, the processing duration of a picture can be 400 ms, and the single-frame analysis duration can be 200 ms.

[0173] It should be understood that since the performances of different image signal processors are different, the chip analysis speeds corresponding to different image signal processors may be different. For an electronic device put on the market, the image signal processor is fixed, so the chip analysis speed corresponding to this image signal processor is also fixed.

[0174] In some embodiments, the channel interface of the media middleware framework layer can also return other relevant performance parameters of various algorithms to the initialization interface of the media middleware framework layer.

[0175] S11, The initialization interface of the media middleware framework layer returns an initialization success message to the service layer through the application function layer.

[0176] Among them, the initialization success message can carry performance parameters of various algorithms, such as the chip analysis speed.

[0177] After obtaining the chip analysis speed, for each video in the image materials selected by the user, the following S12 - S15 can be executed to obtain the analysis parameters for the video and all image materials.

[0178] S12, The service layer sends a query message to the analysis performance query interface of the media middleware framework layer through the application function layer.

[0179] Among them, the query message includes the file descriptor fd1 of video 1 selected by the user and the chip analysis speed. Among them, the file descriptor can be used as the unique identifier of the video.

[0180] S13, The analysis performance query interface of the media middleware framework layer obtains the estimated analysis duration of video 1 according to the file descriptor fd1 of video 1.

[0181] Among them, the estimated analysis duration can be the duration required for the image signal processor to analyze a video. Analyzing a video can refer to analyzing a specified number of image frames in a video. For example, the specified number can be 1, or the specified number can also be greater than 1 and less than or equal to the number of image frames contained in the video.

[0182] Among them, exemplarily, when the specified number is 1, that is, the estimated analysis duration represents the duration required for the image signal processor to analyze 1 image frame in a video, then the estimated analysis duration can also represent the analysis duration corresponding to the minimum analysis of the image frame. For example, if the single-frame analysis duration of the image signal processor is 200 ms, then the estimated analysis duration of each video (including Video 1) is 200 ms.

[0183] In some embodiments, the specified number can also be the number of image frames contained in the video. Then, the estimated analysis duration represents the duration corresponding to the image signal processor analyzing all the image frames in a video. For example, if the single-frame analysis duration of the image signal processor is 200 ms and Video 1 includes 10 image frames, then the estimated analysis duration of Video 1 is 10 * 200 ms = 2000 ms. For example, if Video 2 includes 12 image frames, then the estimated analysis duration of Video 2 is 12 * 200 ms = 2400 ms.

[0184] In some embodiments, the specified number can also be a, where a is greater than 1 and less than the number of image frames contained in the video, and a is a natural number. Then, the estimated analysis duration represents the duration corresponding to the image signal processor analyzing a image frames in the video. For example, if the single-frame analysis duration of the image signal processor is 200 ms, Video 1 includes 10 image frames, and a is 5. Then the estimated analysis duration of Video 1 is 5 * 200 ms = 1000 ms.

[0185] S14. The analysis performance query interface of the media middle platform framework layer returns the estimated analysis duration of Video 1 to the service layer through the application function layer.

[0186] After the one-click video creation module in the service layer obtains the estimated analysis duration of Video 1, it can continue to execute S12 - S15 to obtain the estimated analysis duration of the next video among multiple image materials until the estimated analysis durations of all the videos in the multiple image materials are obtained.

[0187] S15. The service layer obtains analysis parameters according to the estimated analysis duration in the image materials.

[0188] Among them, the analysis parameters may include the total analysis duration, the recommended upper limit value of the total analysis duration, the maximum duration of the highlight segment, the minimum duration of the highlight segment, the total recommended duration of the highlight segment, the recommended duration of the highlight segment, whether to force each video to output a highlight segment, whether to enable audio analysis, the selected highlight segments, and so on.

[0189] Among them, whether to force each video to output a highlight segment is defaulted to yes, that is, in this embodiment, each video needs to output a highlight segment.

[0190] Among them, the total analysis duration represents the total duration required to complete the analysis of all the image materials selected by the user (including all the pictures and all the videos selected by the user). The total analysis duration includes the sum of the estimated analysis durations of all the pictures and the sum of the estimated analysis durations of all the videos.

[0191] For the pictures in the image materials, the processing duration of a picture can be directly determined according to the chip speed of the image processor. Therefore, the sum of the estimated analysis durations of all the pictures can be directly determined according to the number of pictures in the image materials.

[0192] For the videos in the image materials, the estimated analysis duration of each video can be obtained according to S12 - S15, and the sum of the estimated analysis durations of all the videos in the image materials can be obtained by adding up the estimated analysis durations of each video.

[0193] Based on the sum of the estimated analysis durations of all the pictures and the sum of the estimated analysis durations of all the videos, the total analysis duration of the image materials can be obtained.

[0194] For example, assume that the materials selected by the user include 10 pictures, Video 1 and Video 2, the estimated analysis duration of a picture is 400ms, the estimated analysis duration of Video 1 is 200ms, and the estimated analysis duration of Video 2 is 300ms. Then the above total analysis duration is 10 * 400ms + 200ms + 300ms = 4500ms.

[0195] The recommended upper limit value of the total analysis duration represents the recommended value of the maximum duration to complete the analysis of all the pictures and all the videos. The recommended upper limit value of the total analysis duration can be determined according to the estimated analysis duration of each video and the estimated analysis duration of the pictures. Generally, the recommended upper limit value of the total analysis duration is greater than the total analysis duration. For example, calculated by analyzing at least 1 image frame for each video, the total analysis duration is 4400ms. Considering that each video may need to analyze multiple image frames, then the recommended upper limit value of the total analysis duration can be much greater than the total analysis duration. For example, the recommended upper limit value of the total analysis duration can be preset to 10000ms.

[0196] In some scenarios where the input of analysis parameters is abnormal, the recommended value of the upper limit of the total analysis duration may also be set to be less than the total analysis duration. When the recommended value of the upper limit of the total analysis duration is less than the total analysis duration, that is, when the recommended value of the upper limit of the total analysis duration is not sufficient to analyze all pictures and all video (one image frame), some materials can be selected from the selected image materials for analysis. This part is the method implemented by the policy monitoring module of the media middleware framework layer and will be introduced in detail in the following embodiments and will not be elaborated here.

[0197] The maximum duration of a highlight segment represents the maximum allowed duration of a highlight segment in a video. The minimum duration of a highlight segment represents the minimum allowed duration of a highlight segment in a video. The maximum duration of a highlight segment and the minimum duration of a highlight segment can be preset values. Exemplarily, the maximum duration of a highlight segment can be 3000ms, and the minimum duration of a highlight segment can be 1000ms.

[0198] The total recommended duration of highlight segments represents the recommended value of the sum of the recommended durations of the highlight segments of all videos in multiple image materials. The total recommended duration of highlight segments can be determined according to the number of videos, the maximum duration of highlight segments, and the minimum duration of highlight segments. Exemplarily, the maximum duration of a highlight segment can be 3000ms, the minimum duration of a highlight segment can be 1000ms, and when the number of videos is 5, the value range of the total recommended duration of highlight segments can be 5000ms - 15000ms. For example, the total recommended duration of highlight segments can be 8000ms.

[0199] The recommended duration of a highlight segment represents the recommended value of the duration of a highlight segment of a video. The highlight segment with this recommended duration value can effectively display the highlight effect. The recommended duration of a highlight segment can be determined according to the maximum duration of highlight segments and the minimum duration of highlight segments. Exemplarily, the maximum duration of a highlight segment can be 3000ms, the minimum duration of a highlight segment can be 1000ms, and the value range of the recommended duration of a highlight segment can be 1000ms - 3000ms. For example, the recommended duration of a highlight segment can be 2000ms.

[0200] It should be understood that the maximum duration of highlight segments, the minimum duration of highlight segments, the total recommended duration of highlight segments, and the recommended duration of highlight segments can all be set according to the actual situation.

[0201] In some embodiments, the analysis parameters may further include the actual duration of each video.

[0202] S16. The service layer, through the application function layer, sends the file descriptor fd and analysis parameters of all materials to be analyzed to the image highlight segment analysis interface of the media middleware framework layer.

[0203] After receiving the analysis parameters, the policy monitoring module in the media middle platform framework layer can determine the material analysis policy based on the recommended value of the upper limit of the total analysis duration and the total analysis duration in the analysis parameters. The material analysis policy refers to analyzing all the image materials selected by the user, or selecting some of the image materials selected by the user for analysis. In the following embodiments, the image materials include all the videos and all the pictures selected by the user.

[0204] After obtaining the total analysis duration of the image materials, the policy monitoring module can determine the number of pictures and videos to be actually analyzed based on the total analysis duration of the image materials and the recommended value of the upper limit of the total analysis duration in the analysis parameters. After Figure 7 S16, execute:

[0205] S17, the policy monitoring module in the media middle platform framework layer determines the material analysis policy according to the total analysis duration and the recommended value of the upper limit of the total analysis duration in the analysis parameters.

[0206] Among them, if the total analysis duration is less than or equal to the recommended value of the upper limit of the total analysis duration, the material analysis policy can be to analyze all pictures and all videos. For example, if the total analysis duration is 4400ms and the recommended value of the upper limit of the total analysis duration is 10000ms, the material analysis policy can be to analyze all pictures and all videos in the image materials selected by the user.

[0207] If the total analysis duration is greater than the recommended value of the upper limit of the total analysis duration, the material analysis policy can be a random sampling policy. For example, if the total analysis duration is 4400ms and the recommended value of the upper limit of the total analysis duration is 3000ms, the material analysis policy can be a random sampling policy. Among them, the random sampling policy refers to randomly selecting some of the image materials selected by the user for analysis.

[0208] Assume that the analysis value of one picture is greater than the analysis value of one image frame of a video. Then, the random sampling policy can be: for every N pictures selected, M videos are allowed to be selected until the total analysis duration of the selected image materials reaches the recommended value of the analysis duration limit or all the pictures in the selected image materials have been taken. Among them, M < N. For example, N can be natural numbers such as 3, 4, 5, etc., and M can be natural numbers less than N such as 1, 2, 3, etc. The specific values of M and N can be determined according to the number of image materials selected by the user. Exemplarily, it can be that for every 4 pictures selected, 1 video is allowed to be selected until the total analysis duration meets 3000ms; or all the pictures in the selected image materials have been taken.

[0209] For example, the total analysis duration of 10 pictures and 2 videos is 4400 ms, which is greater than the recommended upper limit of the total analysis duration of 3000 ms. According to the random selection strategy that allows selecting 1 video for every 4 pictures selected, 4 pictures and 1 video are selected. The total analysis duration of 4 pictures and 1 video is 1800 ms, which does not reach the recommended upper limit of the total analysis duration of 3000 ms. Continue to select image materials. When the 3rd picture is selected in this round, the total analysis duration reaches 3000 ms, and at this time, the selection of image materials stops. Then, the selected image materials are 7 pictures and 1 video. The remaining 3 pictures are not analyzed.

[0210] In another embodiment, assume that the analysis value of one picture is less than the analysis value of one image frame of a video. Then, the random selection strategy can be that for every P videos selected, Q pictures are allowed to be selected until the total analysis duration of the selected image materials reaches the recommended value of the analysis duration upper limit or the videos in the selected image materials have been exhausted. Where Q < P. For example, P can be natural numbers such as 3, 4, 5, etc., and Q can be natural numbers less than P such as 1, 2, 3, etc. The specific values of P and Q can be determined according to the number of image materials selected by the user. Exemplarily, it can be that for every 3 videos selected, 1 picture is allowed to be selected until the total analysis duration meets 3000 ms; or the videos in the selected image materials have been exhausted.

[0211] For example, the total analysis duration of 10 videos and 3 pictures is 3200 ms, which is greater than the recommended upper limit of the total analysis duration of 3000 ms. According to the random selection strategy that allows selecting 1 picture for every 3 videos selected, 3 videos and 1 picture are selected. The total analysis duration of 3 videos and 1 picture is 1000 ms, which does not reach the recommended upper limit of the total analysis duration of 3000 ms. Continue to select image materials. After three rounds of selecting 3 videos and 1 picture, a total of 9 videos and 3 pictures are selected, and the total analysis duration is 3000 ms. At this time, the selection of image materials stops. Then, the selected image materials are 9 videos and 3 pictures. The remaining 1 video is not analyzed.

[0212] In some embodiments, the electronic device determines the analyzable image materials from the materials selected by the user according to the above random selection strategy (such as a mobile phone). Among them, pictures and videos can be randomly selected from the image materials selected by the user in the order of the image materials selected by the user. Or, random numbers less than the number of materials can also be generated through the Random class in Java to select the corresponding videos or pictures.

[0213] In some embodiments, the image material only includes pictures, and the total analysis duration is the sum of the estimated analysis durations of all pictures. When the total analysis duration is greater than the recommended value of the analysis duration upper limit, the number of pictures that can be analyzed is calculated based on the recommended value of the analysis duration upper limit and the estimated analysis duration of analyzing one picture, and the corresponding number of pictures is randomly selected from all pictures for analysis.

[0214] In some embodiments, the image material only includes videos, and the total analysis duration is the sum of the estimated analysis durations of all videos. When the total analysis duration is greater than the recommended value of the analysis duration upper limit, the number of videos that can be analyzed is calculated based on the recommended value of the analysis duration upper limit, and the corresponding number of videos is randomly selected from all videos for analysis. Among them, the duration required for the number of videos that can be analyzed is less than or equal to the total analysis duration.

[0215] After selecting the pictures and / or videos that can be analyzed within the recommended value of the analysis duration upper limit from the image material, the remaining pictures and videos in the selected image material by the user are not analyzed.

[0216] The image material includes pictures. After the electronic device selects the pictures that can be analyzed within the recommended value of the analysis duration upper limit, it executes Figure 7 the S18 - S22 shown to perform high - light segment analysis on the pictures that can be analyzed.

[0217] The following is an example of the process of performing high - light segment analysis on Picture 1 in combination with S18 - S22. Among them, Picture 1 is one of the pictures that can be analyzed within the recommended value of the analysis duration upper limit.

[0218] S18, The policy monitoring module of the media middle - tier framework layer sends an indication message to the channel interface of the media middle - tier framework layer. The indication message includes the file descriptor of Picture 1.

[0219] S19, The channel interface of the media middle - tier framework layer performs processing operations such as decoding, resolution reduction, and format conversion on Picture 1 according to the file descriptor of Picture 1, and stores the processed Picture 1.

[0220] S20, The channel interface of the media middle - tier framework layer sends the frame data address of Picture 1 to the high - light segment algorithm interface of the HAL layer through the service interface of the FWK layer.

[0221] Among them, the service interface can be a one - click video - making interface.

[0222] S21, The high - light segment algorithm interface of the HAL layer obtains Picture 1 according to the frame data address of Picture 1, and then performs aesthetic scoring on Picture 1 based on a preset high - light segment algorithm to obtain an analysis result. Among them, the analysis result can represent the aesthetic score.

[0223] S22. The high-light segment algorithm interface of the HAL layer returns the analysis result of Picture 1 to the policy monitoring module of the media middle platform framework layer through the service interface of the FWK layer and the channel interface of the media middle platform framework layer in sequence.

[0224] After the policy monitoring module of the media middle platform framework layer obtains the analysis result of Picture 1, if there are other pictures, such as Picture 2, the electronic device can continue to execute the above S18 - S22 to obtain the analysis results of other pictures.

[0225] Among them, the policy monitoring module of the media middle platform framework layer can determine the pictures with scores greater than or equal to the first threshold / second threshold as the high-light segments of multiple image materials according to the analysis results of each picture.

[0226] After obtaining the analysis results of all pictures, if the image materials include videos, the electronic device can use the following S23 - S38 to obtain the analysis results of each video. If the image materials do not include videos, the electronic device outputs the target video set composed of the pictures of the high-light segments.

[0227] If the image materials include multiple pictures and videos, since the calculation amount of pictures is small and the time consumption is less, the electronic device usually first analyzes the high-light segments of pictures one by one. After completing the high-light segment analysis of all pictures, it then analyzes the high-light segments of videos one by one. That is, after the electronic device executes S18 - S22 as Figure 7 shown, it continues to execute S23 - S38 as Figure 8 shown.

[0228] If the image materials only include multiple videos, then after the electronic device executes S17, it executes S23 - S38 as Figure 8 shown, and does not need to execute S18 - S22 as Figure 7 shown.

[0229] In this embodiment, an example of the situation where the image materials include pictures and videos is described. The following videos all refer to the videos that can be analyzed within the recommended upper limit value of the total analysis duration.

[0230] Refer to Figure 8 , Figure 8 which gives a schematic flowchart of analyzing a video to obtain high-light segments in a video processing method. It includes:

[0231] S23. The policy monitoring module determines the analysis policy of the video according to the remaining analysis duration.

[0232] Among them, the analysis policy can include an image frame analysis policy, a key frame analysis policy, an analysis policy combining overview analysis and frame-by-frame analysis, and so on.

[0233] Among them, the remaining analysis duration refers to the remaining available duration for analyzing the video after the above S01 - S22.

[0234] In some embodiments, the image frame analysis strategy refers to extracting all the image frames of the video or a preset number of image frames for analysis to obtain the analysis results of the image frames of the video; and determining the highlight segment of the video based on the image frame with the highest score. The key frame analysis strategy refers to extracting all the key frames of the video or a preset number of key frames for analysis to obtain the analysis results of the key frames of the video; and determining the highlight segment of the video based on the key frame with the highest score. The analysis strategy combining overview analysis and frame - by - frame analysis includes two analysis stages. In the first analysis stage, a preset number of image frames in the video are first extracted for analysis to obtain the analysis results of the image frames, and the target area of the video is determined based on the image frame with the highest score. In the second analysis stage, all the image frames in the target area, or a preset number of image frames are analyzed to obtain the analysis results of the image frames in the target area, and the highlight segment of the video is determined based on the image frame with the highest score. Among them, the image frames can be key frames, ordinary image frames, etc.

[0235] If the remaining analysis duration is sufficient to analyze all the image frames or a preset number of image frames of all the videos, the image frame analysis strategy is adopted; if the remaining analysis duration is sufficient to analyze all the key frames or a preset number of key frames of all the videos, the key frame analysis strategy is adopted; if the remaining analysis duration is sufficient to perform the combined overview analysis and frame - by - frame analysis, the analysis strategy combining overview analysis and frame - by - frame analysis is adopted.

[0236] Based on the analysis strategy determined by the policy monitoring module, the following steps are executed to obtain the first highlight segment of each video.

[0237] Taking Video 1 as an example, taking the analysis strategy as the combined overview analysis and frame - by - frame analysis strategy (analyzing the image frames of the video) as an example, the process of analyzing the highlight segment of the video is illustrated in combination with S24 - S38. Among them, S24 - S30 are the first analysis stage, and S31 - S38 are the second analysis stage.

[0238] S24, the policy monitoring module sends the file descriptor fd1 of Video 1 and the first position of the image frames of Video 1 to the channel interface of the media middle - layer framework.

[0239] Among them, the positions of the first number of image frames in the video can be evenly distributed. Then the first position can be the position of the first image frame starting from the start time of the video.

[0240] In some embodiments, the positions of the first number of image frames in the video can also be distributed according to the storyboard points of the video and the similar picture areas of the video. The first position can be the position of any one image frame in Video 1. For example, the first position can be the position of the first image frame of Video 1; or, the first position can also be other specified positions in Video 1.

[0241] S25. The channel interface performs processing operations such as decoding, reducing the resolution, and converting the format on Video 1 according to the file descriptor of Video 1, and stores the processed Video 1.

[0242] S26. The channel interface of the media middle platform framework layer sends the frame data address of Video 1 and the image frame at the corresponding first position in Video 1 to the highlight segment algorithm interface of the HAL layer through the service interface (such as the one-click video compilation interface) of the FWK layer.

[0243] S27. The highlight segment algorithm interface of the HAL layer obtains the image frame at the first position of Video 1 according to the frame data address of Video 1 and the corresponding first position in Video 1, and then analyzes the image frame at the first position based on the preset highlight segment algorithm to obtain the analysis result of the image frame at the first position.

[0244] S28. The highlight segment algorithm interface of the HAL layer returns the analysis result of the image frame at the first position to the policy monitoring module in sequence through the service interface of the FWK layer and the channel interface of the media middle platform framework layer.

[0245] Among them, the analysis result can represent the aesthetic score of the image frame at the first position.

[0246] S29. The policy monitoring module obtains the second position, and returns to execute S24 until the number of analyzed image frames meets the first number allocated to the video.

[0247] Among them, the first number can be determined according to the actual duration of each video and the remaining analysis duration. For example, according to the single-frame analysis duration, the total number of image frames of all videos that can be analyzed within the remaining analysis duration can be calculated. Sort the videos in descending order of duration, and allocate the total number of image frames to each video until the total number is divided up, to obtain the first number allocated to each video. Among them, the first number of image frames can be evenly distributed at various positions in the video.

[0248] When analyzing one image frame of a video, update the remaining analysis duration.

[0249] S30. The policy monitoring module obtains the analysis results of the first number of image frames of all videos.

[0250] After obtaining the analysis results of the first quantity of image frames of all videos, the policy monitoring module can determine the target area of each video according to the analysis results of the image frames in each video. Then, frame-by-frame analysis is performed on the target area of each video to obtain the first highlight segment of each video.

[0251] S31. The policy monitoring module determines the target area of each video based on the analysis results of the first quantity of image frames of all videos.

[0252] In this embodiment, for each video, after obtaining the analysis results of the first quantity of image frames, the policy monitoring module determines the image frame with the highest score according to the scores of each image frame. Among them, the image frame with the highest score may include at least one image frame. If the image frame with the highest score includes one image frame, the image frames within the preset duration 1 before and after this image frame can be directly determined as the target area of video 1. For example, as Figure 9 Figure (a) shows a schematic diagram of a target area. Among them, the image frame with the highest score of video 1 includes image frame 1 (such as 1 in Figure (a) of Figure 9 ), and the segment covered by the preset duration 1 before and after image frame 1 is the target area.

[0253] Alternatively, if the image frame with the highest score includes one image frame, according to the recommended duration of the highlight segment, the segments within 1 / 2 of the recommended duration of the highlight segment before and after the image frame with the highest score can be determined as the candidate highlight segments. The image frames within the preset duration 2 before and after the candidate highlight segments are determined as the target area of video 1. Among them, the duration of the target area can be b times that of the candidate highlight segment. Exemplarily, b can be a number greater than 1 and less than 2.

[0254] For example, as Figure 9 Figure (b) shows another schematic diagram of a target area. Among them, the image frame with the highest score of video 1 includes image frame 1 (such as 1 in Figure (b) of Figure 9 ), the segments within 1 / 2 of the recommended duration of the highlight segment before and after image frame 1 are the candidate highlight segments, and the segment covered by the preset duration 2 before and after the candidate highlight segments is the target area.

[0255] If the image frame with the highest score includes multiple consecutive image frames. According to the recommended duration of the highlight segment, the segment covered by the preset duration 3 before the first image frame of the multiple consecutive first image frames to the preset duration 3 after the last image frame of the multiple consecutive second image frames is determined as the candidate highlight segment.

[0256] For example, as Figure 10 , Figure 10 shows another schematic diagram of a target area. The image frame with the highest score of video 1 includes consecutive image frames 1 (such asFigure 10 1) in the image frame 2 (such as Figure 10 2) in the image frame 3 (such as Figure 10 3) in the above. Then, the segment covered by the preset duration 3 before the image frame 1 to the preset duration 3 after the image frame 3 is the candidate highlight segment of the video 1. The segments covered by the preset duration 4 before and after the candidate highlight segment are the target regions.

[0257] After determining the candidate highlight segments of each video, if the sum of the durations of the candidate highlight segments of all videos is greater than Y times the total recommended duration of the highlight segments, the policy monitoring module needs to adjust the durations of the candidate highlight segments of each video so that the sum of the durations of the candidate highlight segments of all videos is less than or equal to the second multiple of the total recommended duration of the highlight segments.

[0258] Among them, Y times can be a number greater than 1. For example, Y times can be 1.05. Exemplarily, according to the duration (the duration to be adjusted) by which the sum of the durations of the candidate highlight segments of all videos exceeds Y times the total recommended duration of the highlight segments, the duration of the candidate highlight segment is reduced and adjusted according to the actual duration of each video. Exemplarily, the shorter the duration of a video, the more seconds its candidate highlight segment is reduced.

[0259] Suppose two videos are input. The actual duration of video 1 is 10s, and the actual duration of video 2 is 20s. The recommended duration of the highlight segment is 6s, and the total recommended duration of the highlight segments is 10s. The durations of the candidate highlight segments of video 1 and video 2 are both 6s; the durations of the target regions of video 1 and video 2 are both 12s.

[0260] Among them, the sum of the durations of the candidate highlight segments of video 1 and video 2 is 12s, exceeding 1.05 times (10.5s) of the total recommended duration of the highlight segments (10s). The policy monitoring module needs to adjust the durations of the candidate highlight segments of each video. The duration to be adjusted is 12s - 10s * 1.05 = 1.5s. The policy monitoring module distributes the reduction in duration according to the reciprocal of the video duration. The duration of the candidate highlight segment of video 1 (10s) is reduced by 1s; the duration of the candidate highlight segment of video 1 is updated from 6s to 5s, the duration of the candidate highlight segment of video 2 (20s) is reduced by 0.5s, and the duration of the candidate highlight segment of video 2 is updated from 6s to 5.5s. In this way, the sum of the durations of the candidate highlight segments of video 1 and video 2 is 10.5s, which does not exceed 10s * 1.05, and no further adjustment is required.

[0261] Therefore, it is determined that the duration of the candidate highlight segment of video 1 is 5s, and the duration of the target region corresponding to the candidate highlight segment in video 1 is 10s; the duration of the candidate highlight segment of video 2 is 5.5s, and the duration of the target region corresponding to the candidate highlight segment in video 2 is 11s.

[0262] The method provided by S31 can be used to determine the target area of each video. After determining the target area of each video, S31 - S36 can be executed to analyze the image frames of the target area.

[0263] After determining the target area of each video, enter the second analysis stage:

[0264] The policy monitoring module in the media middle - platform framework layer performs frame - by - frame analysis on the target area. Taking video 1 as an example, the process of performing frame - by - frame analysis on the target area of video 1 may include:

[0265] S32, the policy monitoring module sends the file descriptor fd1 of video 1 and the third position of the image frame of the target area of video 1 to the channel interface in the media middle - platform framework layer.

[0266] Among them, the third image frame position can be the position of any one image frame in the target area of video 1.

[0267] S33, the channel interface performs processing operations such as decoding, down - sampling the resolution, and converting the format on video 1 according to the file descriptor of video 1, and stores the processed video 1.

[0268] S34, the channel interface sends the frame data address of video 1 and the third position of the image frame to the highlight segment algorithm interface in the HAL layer through the service interface (such as the one - click video compilation interface) in the FWK layer.

[0269] S35, the highlight segment algorithm interface in the HAL layer obtains the third image frame of video 1 according to the frame data address and the third position of video 1, and then analyzes the image frame at the third position based on a preset highlight segment algorithm to obtain the analysis result of the image frame at the third position.

[0270] S36, the highlight segment algorithm interface in the HAL layer returns the analysis result of the image frame at the third position to the policy monitoring module in the media middle - platform framework layer through the service interface in the FWK layer and the channel interface in the media middle - platform framework layer in sequence.

[0271] The policy monitoring module in the media middle - platform framework layer continues to send the fourth position in the target area of video 1 to the channel interface in the media middle - platform framework layer for analysis until all the second number of image frames in the target area are analyzed.

[0272] Among them, the fourth position can be the position of the next adjacent image frame to the third position in the target area of video 1. The second number of image frames can be evenly distributed at various positions in the target area.

[0273] S37, the policy monitoring module determines the highlight segment of video 1 from the analysis results of the image frames in the target area of video 1.

[0274] In this embodiment, the policy monitoring module obtains the highlight segment of Video 1 based on the recommended duration of the highlight segment of Video 1 and the image frame with the highest score in the target area.

[0275] In some embodiments, the policy monitoring module obtains the image frame with the highest score according to the scores of the image frames. The image frame with the highest score may include at least one image frame. If the image frame with the highest score includes one image frame, the image frames within the first preset duration before and after the image frame with the highest score can be directly determined as the highlight segment of Video 1.

[0276] For example, Figure 11 (a) of gives a schematic diagram of a highlight segment. The image frame with the highest score of Video 1 includes Image Frame 1 (1 in (a) of ). The segment covered by the preset duration 5 before and after Image Frame 1 is the highlight segment of Video 1. The duration of the highlight segment is less than or equal to the recommended duration of the highlight segment. Figure 11 (a) of gives a schematic diagram of a highlight segment. The image frame with the highest score of Video 1 includes Image Frame 1 (1 in (a) of ). The segment covered by the preset duration 5 before and after Image Frame 1 is the highlight segment of Video 1. The duration of the highlight segment is less than or equal to the recommended duration of the highlight segment.

[0277] Alternatively, if the image frame with the highest score includes one image frame, according to the recommended duration of the highlight segment, the segments before and after the image frame with the highest score within 1 / 2 of the recommended duration of the highlight segment can be determined as the highlight segment of Video 1. The duration of the highlight segment is equal to the recommended duration of the highlight segment.

[0278] For example, Figure 11 (b) of gives another schematic diagram of a highlight segment. The image frame with the highest score of Video 1 includes Image Frame 1 (1 in (b) of ). The segments before and after Image Frame 1 within 1 / 2 of the recommended duration of the highlight segment are the highlight segments of Video 1. Figure 11 (b) of gives another schematic diagram of a highlight segment. The image frame with the highest score of Video 1 includes Image Frame 1 (1 in (b) of ). The segments before and after Image Frame 1 within 1 / 2 of the recommended duration of the highlight segment are the highlight segments of Video 1.

[0279] If the image frame with the highest score includes multiple consecutive image frames. According to the recommended duration of the highlight segment, the segment covered by the preset duration 6 before the first image frame of the multiple consecutive image frames with the highest score to the preset duration 6 after the last image frame of the multiple consecutive second image frames is determined as the highlight segment of Video 1.

[0280] For example, Figure 12 , Figure 12 gives another schematic diagram of a highlight segment. The image frame with the highest score of Video 1 includes consecutive Image Frame 1 (1 in ), Image Frame 2 (2 in ), and Image Frame 3 (3 in ). Then, the segment covered by the preset duration 6 before Image Frame 1 to the preset duration 6 after Image Frame 3 is the highlight segment of this Video 1. Figure 12 In, Figure 12 In, Figure 12 In,

[0281] In some embodiments, after analyzing each video, the policy monitoring module updates the remaining analysis duration.

[0282] S38. The policy monitoring module obtains the analysis results of the highlight segments of all videos.

[0283] After the above steps, one highlight segment of each video can be obtained. To avoid missing other highlight segments of the videos in the image material, further, the policy monitoring module can verify the duration of the highlight segments of all videos. When the sum of the durations of all highlight segments does not meet the total recommended duration of the highlight segments, supplementary selection of highlight segments is performed until the output conditions of the highlight segments are met. Among them, meeting the output conditions of the highlight segments includes that the sum of the durations of all highlight segments is equal to or greater than the recommended total duration of the highlight segments, or there are no image frames in the video that can be supplementarily selected.

[0284] Reference Figure 13 , in this embodiment Figure 13 provides a schematic flowchart of the process of supplementarily selecting highlight segments in a video processing method. After the electronic device executes S38, it performs the operation of supplementarily selecting highlight segments:

[0285] S39. The policy monitoring module calculates the first duration of the highlight segments of all videos.

[0286] The first duration refers to the sum of the durations of one highlight segment of each video in the image material obtained after S38. One highlight segment of each video obtained in S38 is determined by the first image frames with a score greater than or equal to the first threshold. That is, the highlight segments in S38 are determined by the first image frames with a "high" scoring result.

[0287] S40. If the first duration does not reach the total recommended duration of the highlight segments, the policy monitoring module supplementarily selects highlight segments from the second image frames of all videos.

[0288] Among them, the second image frames are the image frames with a score greater than or equal to the second threshold and not covered by the already selected highlight segments. That is, the highlight segments supplementarily selected from the second image frames are the segments supplementarily selected from the segments corresponding to the image frames with a "high" and "relatively high" scoring results.

[0289] During the process of supplementing highlight segments from the second image frames, the policy monitoring module can directly screen the second image frames according to the analysis results of each analyzed image frame in the video obtained in S30. During the process of supplementing highlight segments, the policy monitoring module no longer performs the operation of sending the image frames for analysis. It can be understood that if the analysis results obtained in S30 are the analysis results of the key frames of each video, then the supplement of the highlight segments is based on the key frames; if the analysis results obtained in S30 are the analysis results of the ordinary image frames of each video, then the supplement of the highlight segments is based on the ordinary image frames.

[0290] During the process of supplementing highlight segments, to avoid the supplemented segments overlapping with the target area and a highlight segment of the video, resulting in ineffective supplementation. The image frames covered by the highlight segments in each video do not participate in the supplementation operation. That is to say, the policy monitoring module supplements the highlight segments from the second image frames of the non-highlight segments of each video.

[0291] In some embodiments, the policy monitoring module can sort the image frames of the non-highlight segments of all videos in descending order of scores, and process each second image frame in turn from the score result of "high" to the score result of "relatively high".

[0292] When processing each second image frame, obtain the corresponding highlight segment of the second image frame in the video. Among them, the highlight segment corresponding to the second image frame includes the second image frame, and the image frames within the second preset duration before and after the second image frame.

[0293] In some embodiments, the conditions for determining the highlight segment of the second image frame include that the start and end positions of the highlight segment of the second image frame do not exceed the start and end positions of the video itself, and the highlight segment does not contain the segment (or part of the segment) of the target area, the segment (or part of the segment) corresponding to the image frame with a score less than the third threshold, and the segment (or part of the segment) that has been supplemented as a highlight segment.

[0294] Among them, the distance between the start position or the end position of the highlight segment and the position of the second image frame should be greater than or equal to the preset distance. For example, the preset distance can be 500ms, 800ms, 1000ms, etc.

[0295] Among them, the segment corresponding to the image frame with a score less than the third threshold can be understood as the segment from 500ms before the image frame with a score less than the third threshold to 500ms after the image frame. The score of the segment corresponding to the image frame with a score less than the third threshold is 0, and it can also be called the 0-score area or 0-score segment.

[0296] If the duration of the highlight segment is less than or equal to the recommended duration of the highlight segment, determine that the highlight segment is supplemented as a highlight segment, and end the processing of the current second image frame. Sort according to the scores of the second image frames, and continue to process the next second image frame.

[0297] If the duration of the highlight segment is greater than the recommended duration of the highlight segment, the duration of the highlight segment needs to be reduced. The reduction step size can be determined according to the difference between the duration of the highlight segment and the recommended duration of the highlight segment. For example, the first reduction step size can be 1 / 2 of the difference. Exemplarily, if the duration of the highlight segment is 15s and the recommended duration of the highlight segment is 5s, then 5 seconds ((15 - 5) / 2s) are reduced from the midpoint position of the highlight segment to the left and right sides of the second image frame respectively.

[0298] If one side is reduced to the preset distance of the second image frame. For example, the right side position of the highlight segment of the second image frame has been reduced to 500ms from the position of the second image frame. If the duration of the highlight segment is still greater than the recommended duration of the highlight segment, continue to reduce the left side position of the highlight segment, and the right side position will no longer be reduced until the duration of the highlight segment is reduced to less than or equal to the recommended duration of the highlight segment.

[0299] Exemplarily, refer to Figure 14 , Figure 14 Figure 1 shows a schematic diagram of supplementing highlight segments in Video 1. Assume that there is only one video (Video 1) in the image material. The duration of Video 1 is 20s. Video 1 includes a highlight segment 1 with a duration of 4s from 1s to 5s obtained through S37. Video 1 includes Image Frame 1 with a score greater than the second threshold. Among them, the score of Image Frame 1 is 88, and the scoring result is "high". Image Frame 1 is at the 15s position of the video. Taking the recommended duration of the highlight segment as 5s, and the distance between the start position or end position of the effective segment and the position where the image frame is located should be greater than or equal to 500ms as an example for illustration.

[0300] The strategy monitoring module obtains Image Frame 1 for the operation of supplementing highlight segments. Among them, the available segment of Image Frame 1 includes Segment 2 from 5s to 20s of Video 1. The duration of this Segment 2 is 15s, which exceeds the recommended duration of the highlight segment (5s). Therefore, this Segment 2 needs to be reduced.

[0301] Exemplarily, the reduction step size is 5s. Starting from the start and end positions of this segment 2, 5s are reduced on each side of the left and right of the image frame 1. After reducing 5s on the left side of the image frame 1, the start position of the segment 2 is updated from the 5s position of the video 1 to the 10s position; after reducing 5s on the right side of the image frame 1, the end position of the segment 2 overlaps with the position where the image frame 1 is located. According to the requirement that the distance between the end position of the valid segment and the position where the image frame is located should be greater than or equal to 500ms, the end position of the segment 2 is determined to be 500ms to the right of the image frame 1, that is, the 15.5-second position of the video 1. At this time, the duration of the segment 2 is 5.5s (10s - 15.5s), which is still greater than the recommended duration of the highlight segment. Since the right side of the image frame 1 cannot be reduced anymore, therefore, 0.5s is reduced from the left side of the image frame 2, so that the duration of the segment 2 is less than or equal to the recommended duration of the highlight segment.

[0302] Finally, according to the image frame 1 in the video 1, the highlight segment is supplemented and selected, and the obtained highlight segment is segment 2 (10.5s - 15.5s).

[0303] Reference Figure 15 , Figure 15 shows another schematic diagram of supplementing and selecting the highlight segment in the video 1. Combining Figure 14 with the provided example, the video 1 also includes the image frame 2. The score of the image frame 2 is 40, and the scoring result is "low". The image frame 2 is at the 7s position of the video. The segment corresponding to the image frame 2 is segment 3 (6.5s - 7.5s) with 500ms on each side of the left and right of the image frame 2. The available segment of the image frame 1 should not include the segment (segment 3) corresponding to the image frame with a score less than the third threshold. Then the available segment of the image frame 1 is segment 4 of the video 1 from 7.5s to 20s. The duration of this segment 4 is 12.5s, which exceeds the recommended duration of the highlight segment. Therefore, the segment 4 still needs to be reduced.

[0304] Exemplarily, taking the reduction step size of 3 seconds as an example, 3 seconds are reduced respectively from the start and end positions of this segment 4. After reducing 3 seconds on the left side of the image frame 1, the start position of the segment 4 is updated from the 7.5-second position of the video 1 to the 10.5-second position; after reducing 3 seconds on the right side of the image frame 1, the end position of the segment 4 is updated from the end position of the video 1 (20-second position) to the 17-second position. At this time, the duration of the segment 4 is 6.5 seconds (10.5 seconds - 17 seconds), which is still greater than the recommended duration of the highlight segment. Since both sides of the image frame 2 can still be reduced, 1 second is reduced from each side of the image frame 2. The start position of the segment 4 is updated from the 10.5-second position of the video 1 to the 11.5-second position, and the end position of the segment 4 is updated from the 17-second position of the video 1 to the 16-second position. At this time, the duration of the segment 4 is 4.5 seconds (11.5 seconds - 16 seconds), and the duration of the segment 4 is less than the recommended duration of the highlight segment.

[0305] Finally, according to the image frame 1 in Video 1, the highlight segment is selected, and the obtained highlight segment is segment 4 (11.5 seconds - 16 seconds).

[0306] Reference Figure 16 , Figure 16 gives another schematic diagram of selecting highlight segments in Video 1. Combining Figure 14 with the provided example, Video 1 also includes image frame 3. The score of image frame 3 is 75, and the scoring result is "higher". Image frame 3 is at the 17 - second position of the video. The highlight segment corresponding to image frame 1 is segment 2 (10.5 seconds - 15.5 seconds). Then the available segment for image frame 3 is segment 5 of Video 1 from 15.5 seconds to 20 seconds. The duration of this segment 5 is 4.5 seconds, which does not exceed the recommended duration of the highlight segment. Therefore, segment 5 is determined as the highlight segment corresponding to image frame 3. So far, in the case where the image material includes Video 1, and Video 1 includes image frames 1 and 2 with scores greater than the second threshold, the highlight segments of Video 1 are selected, and the obtained highlight segments include segment 2 corresponding to image frame 1 and segment 5 corresponding to image frame 3.

[0307] In some embodiments, if the duration of the highlight segment of the finally obtained image frame 1 is less than the minimum duration of the highlight segment, or the duration of the highlight segment of image frame 1 is less than the preset duration threshold (for example, 500 ms), then image frame 1 is discarded, and the determination of the highlight segment of image frame 1 is no longer performed. In some embodiments, the discarded image frame 1 can be marked to avoid repeating the determination of the highlight segment of image frame 1 during subsequent highlight segment selection, resulting in computational redundancy.

[0308] After each image frame is processed, the selected highlight segment for supplementary selection no longer participates in the subsequent supplementary selection process, and the image frames covered by the highlight segment also do not participate in the subsequent supplementary selection judgment. The newly selected highlight segments in the subsequent selection should not overlap with the previously selected highlight segments.

[0309] When each highlight segment is selected, based on the newly selected highlight segment, the policy monitoring module calculates and updates the sum of the durations (the second duration) of all highlight segments of all videos.

[0310] When the second duration reaches the total recommended duration of the highlight segment, the policy monitoring module no longer performs the highlight segment selection operation and can directly execute S43.

[0311] If the second duration does not reach the total recommended duration of the highlight segment, and all the second image frames with scores greater than the second threshold in all videos of the image material have been processed, the policy monitoring module makes the next judgment.

[0312] Among them, the situation where the second image frame with a score greater than the second threshold has been processed includes that the second image frame has been selected as a highlight segment or the segment corresponding to the second image frame has been discarded because its duration is less than the minimum duration of the highlight segment, or the segment corresponding to the second image frame is covered by an existing highlight segment in the video, and so on.

[0313] The policy monitoring module in the media middle platform framework layer makes a further judgment:

[0314] S41, if the second duration does not reach the total recommended duration of the highlight segment at the first magnification, the policy monitoring module selects a highlight segment from the third image frames of all videos.

[0315] Among them, the first magnification can be a magnification greater than 0.6 and less than 1, such as 0.7, 0.8, 0.9, etc. The third image frame is an image frame with a score greater than or equal to the third threshold.

[0316] Suppose the total recommended duration of the highlight segment is T and the first magnification is 0.8. If the second duration is less than T * 0.8, the policy monitoring module selects a highlight segment from the third image frames with a score greater than the third threshold in all videos.

[0317] After S40, the video of the image material may or may not include a second image frame with a score greater than the second threshold. The policy monitoring module obtains the third image frames with a score greater than the third threshold in the video of the image material, sorts them in descending order of the score, and processes the third image frames with a score greater than the third threshold.

[0318] Similar to the method of selecting a highlight segment from the second image frame, when processing each image frame:

[0319] Obtain the highlight segment corresponding to the third image frame in the video. Among them, the conditions for determining the highlight segment include that the start and end positions of the highlight segment do not exceed the start and end positions of the video itself, and the highlight segment does not include segments of the target area, segments corresponding to image frames with a score less than the third threshold, and segments that have been selected as highlight segments.

[0320] Among them, the distance between the start position or end position of the highlight segment and the position of the image frame should be greater than or equal to a preset distance. For example, the preset distance can be 500ms, 800ms, 1000ms, etc.

[0321] Among them, the segment corresponding to the image frame with a score less than the third threshold can be understood as the segment from 500ms before the image frame with a score less than the third threshold to 500ms after the image frame.

[0322] If the duration of the highlight segment is less than or equal to the recommended duration of the highlight segment, determine that the highlight segment is supplemented as a highlight segment, and end the processing of the current third image frame. Sort according to the scores of the third image frames, and continue to process the next third image frame.

[0323] If the duration of the highlight segment is greater than the recommended duration of the highlight segment, the duration of the highlight segment needs to be reduced. The reduction method can refer to the reduction methods involved in the embodiment of supplementing the highlight segment from the second image frame, which will not be elaborated here. In some embodiments, if the duration of the highlight segment of the finally obtained image frame 1 is less than the minimum duration of the highlight segment, or the duration of the highlight segment of the image frame 1 is less than the preset duration threshold (for example, 500 ms), then discard the image frame 1 and no longer determine the highlight segment of the image frame 1. In some embodiments, the discarded image frame 1 can be marked to avoid repeated determination of the highlight segment of the image frame 1 during subsequent highlight segment supplement selection, resulting in computational redundancy.

[0324] Similarly, after each third image frame is processed, the highlight segment selected as the supplement does not participate in the subsequent supplement selection process, and the third image frames covered by the highlight segment do not participate in the subsequent supplement selection judgment. The newly selected highlight segments in the subsequent supplement selection should not overlap with the previously selected highlight segments.

[0325] When each highlight segment is obtained by supplement selection, the policy monitoring module calculates and updates the sum of the durations of the highlight segments of all videos (the third duration) based on the newly obtained highlight segment by supplement selection.

[0326] When the third duration reaches the total recommended duration of the highlight segment, the policy monitoring module no longer performs the highlight segment supplement selection operation and can directly execute S43.

[0327] If the third duration does not reach the total recommended duration of the highlight segment, and all the third image frames in the image material with scores greater than the third threshold have been processed, the policy monitoring module makes the next judgment.

[0328] Among them, the situation where all the third image frames with scores greater than the third threshold have been processed includes that a third highlight segment or the corresponding segment of the third image frame has been selected by supplement selection, the segment corresponding to the third image frame has been discarded because its duration is less than the minimum duration of the highlight segment, or the segment corresponding to the image frame is covered by the highlight segment of the video, etc.

[0329] The policy monitoring module of the media middle platform framework layer makes a further judgment:

[0330] S42. If the total proposed duration of the highlight segments for the third duration has not reached the total proposed duration of the highlight segments at the second ratio, the strategy monitoring module extends the duration of the highlight segments until there are no highlight segments that can be extended or the sum of the durations of the highlight segments reaches the total proposed duration of the highlight segments at the second ratio.

[0331] Among them, the second ratio is less than the first ratio. Exemplarily, the second ratio can be a value greater than 0 and less than the first ratio, such as 0.4, 0.5, 0.6, etc. Assume that the total proposed duration of the highlight segments is T and the second ratio is 0.6. If the third duration is less than T * 0.6, the strategy monitoring module extends the duration of the highlight segments.

[0332] After S40 and S41, the selection of highlight segments from the second image frame in the video and the selection of highlight segments from the third image frame have been completed, and there are no second image frames and third image frames that can be selected in the video. However, the sum of the durations of all the highlight segments in the video is still very short and does not meet the second ratio of the total proposed duration of the highlight segments. To ensure the duration and quality of the output target video set, the durations of the existing highlight segments can be extended to obtain highlight segments whose durations meet the second ratio of the total proposed duration of the highlight segments.

[0333] It can be understood that when extending the highlight segments, the area where the highlight segments are located and the segment areas corresponding to the image frames with scores less than the third threshold are not used as expandable areas. That is, when extending the existing highlight segments, it cannot overlap with other highlight segment areas and segment areas with low scoring results.

[0334] In some embodiments, if the duration of the highlight segment reaches the maximum duration of the highlight segment; or, there are no available expandable areas on both the left and right sides of the highlight segment; or, the distances from both the left and right sides of the highlight segment to the edge are less than the preset duration threshold (for example, 2s), then these highlight segments are not extended. Refer to Figure 17 , Figure 17 Figure 16 shows a schematic diagram of the expandable area 1 of the highlight segment 1 of video 1. Among them, video 1 also includes the target area 2 of the first highlight segment 2 and the area 3 with a low scoring result. The expandable area of the highlight segment 1 does not overlap or partially overlap with other areas. The boundary of the expandable area 1 is 2s away from the adjacent edge.

[0335] In some embodiments, before extending the duration of the highlight segments, the required extension duration of all videos can be calculated first. Among them, the required extension duration of all videos can be the difference between the total proposed duration T * 0.6 of the highlight segments and the third duration.

[0336] In some embodiments, the policy monitoring module may sort the highlight segments of all videos in the image material according to their scores. Among them, the score of a highlight segment can be calculated as the average of the scores of the image frames covered in the highlight segment. According to the scores of the highlight segments, different weight values are set for the highlight segments with different scores, which are used to allocate the required extended duration. For example, the highlight segments of all videos in the image material are sorted according to their scores. The weight corresponding to the highlight segments in the top 25% of the sorting is 3, the weight corresponding to the highlight segments in the middle 50% of the sorting is 1.5, and the weight corresponding to the highlight segments in the last 25% of the sorting is 1.

[0337] In one example, assume that all videos include 5 highlight segments, the sum of the durations of the 5 highlight segments is 20s, and the total recommended duration T of the highlight segments is 45s. Then, through calculation, the required extended duration is 45s * 60% - 20s = 7s. 7s is the total extended duration required for the 5 videos.

[0338] Assume that highlight segment 1 has reached the maximum duration of the highlight segment, then no extension operation is performed on the duration of this highlight segment 1. The remaining 4 highlight segments are extended.

[0339] Among them, referring to Figure 18 , Figure 18 a schematic diagram of the allocation of the required extended duration is provided. Assume that the score of highlight segment 2 is 90 points, the score of highlight segment 3 is 80 points, the score of highlight segment 4 is 70 points, and the score of highlight segment 5 is 60. Sorted from high to low according to the scores, the top 25% of the 4 highlight segments include highlight segment 2, and the weight assigned to highlight segment 2 is 3; the middle 50% of the 4 highlight segments include highlight segment 3 and highlight segment 4, and the weights assigned to highlight segment 3 and highlight segment 4 are 1.5; the last 25% of the 4 highlight segments include highlight segment 5, and the weight assigned to highlight segment 5 is 1.

[0340] According to the weights assigned to the highlight segments and the required extended duration of 7s, it is calculated that the extended quota duration of highlight segment 2 is 3s, the extended quota duration of highlight segment 3 is 1.5s, the extended quota duration of highlight segment 4 is 1.5s, and the extended quota duration of highlight segment 5 is 1s.

[0341] In one example, referring to Figure 19 , Figure 19Another schematic diagram for allocating the required extended duration is provided. Assume that all videos include 12 highlight segments, and all 12 highlight segments can be extended. Sorted by score from high to low, the top 25% of the 12 highlight segments include 3 highlight segments, and the weight of each of these 3 highlight segments is 3; the middle 50% of the 12 highlight segments include 6 highlight segments, and the weight of each of these 6 highlight segments is 1.5; the last 25% of the 12 highlight segments include 3 highlight segments, and the weight of each of these 3 highlight segments is 1.

[0342] It can be understood that the sorted allocation ratio and the specific values of the weights corresponding to the highlight segments in different ratios can be determined according to the actual situation. The principle of weight allocation is that the higher the score of the highlight segment and the more forward the ranking of the highlight segment, the greater the weight value corresponding to the highlight segment.

[0343] In some embodiments, after calculating the extended quota duration of each highlight segment, the extended quota duration of the highlight segment can be further verified to ensure that the duration of the extended highlight segment does not exceed the maximum duration of the highlight segment in the analysis parameters.

[0344] Specifically, when performing the extension operation on each highlight segment, obtain the fourth duration of the duration of the highlight segment and the already allocated extended quota duration. If the fourth duration is less than the maximum duration of the highlight segment, then take the difference between the fourth duration and the duration of the highlight segment as the extended quota duration of the highlight segment; if the fourth duration is greater than the maximum duration of the highlight segment, then take the difference between the maximum duration of the highlight segment and the duration of the highlight segment as the extended quota duration of the highlight segment.

[0345] When performing the extension operation on each highlight segment, extend 1 / 2 of the extended quota duration to the left and right sides of the highlight segment respectively. If one side extends to a preset distance (such as 1s, 2s) from the edge of the expandable area, then this side will no longer continue to extend; if there is still remaining extended duration, then extend from the other side. If neither side can be extended anymore, or the extended duration on both sides has reached the extended quota duration, then stop the extension operation of this highlight segment. Among them, neither side can be extended anymore can include that the extended boundary distances on both sides reach the preset distance from the edge of the expandable area.

[0346] In an example, refer to Figure 20 , Figure 20 A schematic diagram for extending highlight segment 2 is provided. Assume that the duration of video 1 is 25s, video 1 includes highlight segment 1 (1s - 7s) and highlight segment 2 (13s - 16s), and video 1 also includes a region 0 with a low score result (19s - 20s).

[0347] Assume that the maximum duration of a highlight segment is 6 seconds; the duration of highlight segment 1 has reached the maximum duration of the highlight segment, so no extension operation is performed on highlight segment 1. The duration of highlight segment 2 is 3 seconds, and the allocated extension quota duration for highlight segment 2 is 4s. The fourth duration corresponding to highlight segment 2 is greater than the maximum duration of the highlight segment. Therefore, the difference between the maximum duration of the highlight segment (6s) and the duration of highlight segment 2 (3s) is used as the extension quota duration (3s).

[0348] When performing the extension operation on highlight segment 2, both the left and right sides of highlight segment 2 are extended by 1.5s. The left side can be extended by 1.5s to become 11.5s, and the right side can only be extended by 1s to become 17s. The preset distance (2s) has been reached from the edge of the expandable area (the boundary of area 0). At this point, the extended duration is 2.5s, which has not reached the extension quota duration, and there is still 0.5s remaining.

[0349] Continue to extend highlight segment 2. The right side of highlight segment 2 cannot be extended, and the left side can be extended by 0.5s to become 11s. The actual extended duration reaches the extension quota duration of 3s, and the extension operation of this highlight segment 2 ends.

[0350] According to the order of the scores of the highlight segments from high to low, extend each highlight segment one by one. After each highlight segment is extended, update the remaining expandable highlight segments and the remaining extension quota duration in the image material. Then, according to the updated remaining expandable highlight segments and the remaining extension quota duration, reallocate the extension quota duration and continue to perform a new round of extension processing of the highlight segments.

[0351] Stop the extension operation of the highlight segments until there is no remaining extension quota duration or there are no more expandable highlight segments, and obtain the final results of all the highlight segments in the image material.

[0352] After obtaining the final results of all the highlight segments in the image material, the video processing method provided in Figure 17 can be executed, including:

[0353] S43, the policy monitoring module of the media middle platform framework layer obtains the analysis results of all videos based on the highlight segments of each video after the extension duration.

[0354] Among them, the analysis results of all videos can include the positions of the highlight segments in all videos. The following S45 can be used to report all the picture and video analysis results.

[0355] S44, the policy monitoring module of the media middle platform framework layer reports all the picture and video analysis results to the application function layer through the image highlight segment analysis interface of the media middle platform framework layer.

[0356] S45. The application function layer clips and filters the user-selected materials based on the analysis results of all pictures and videos to obtain all highlight segments.

[0357] S46. The application function layer of the application layer calls the theme summary interface of the media middle platform framework to request and obtain a theme template.

[0358] Among them, this request passes through the theme summary interface and the channel interface of the media middle platform framework and is transmitted to the HAL layer through the FKW layer.

[0359] S47. The HAL layer determines a theme template that matches the scene according to the scenes of the pictures in the highlight segments.

[0360] In some embodiments, the electronic device may be provided with multiple theme templates (style templates). The theme algorithm may recommend a theme template that matches the picture scene based on the picture scenes of the highlight segments, such as people, scenery, food, children, pets, sports, or travel.

[0361] S48. The HAL layer returns the determined theme template to the theme summary interface of the media middle platform framework layer, and the theme summary interface returns the theme template to the application function layer of the application layer.

[0362] Exemplarily, assuming that most of the pictures in the highlight segments are parent-child scenes, then it can be confirmed that the theme template that matches this scene is the parent-child theme category.

[0363] S49. The application function layer distributes the obtained theme and all highlight segments to the basic ability layer.

[0364] The application function layer distributes the theme obtained from S49 and all the highlight segments obtained from S45 to the basic ability layer.

[0365] S50. The basic ability layer generates a target video set according to the theme and all highlight segments.

[0366] That is, the target video set is a video set generated according to all the screened highlight segments and conforming to the recommended theme.

[0367] S51. The basic ability layer sends an indication message for displaying the target video set to the video editing service layer.

[0368] S52. The video editing service layer displays the target video set in the gallery interface.

[0369] In the embodiments of the present application, the media middle platform framework layer is used for decoding video and picture files, converting data formats into a unified format, monitoring the remaining time and adjusting operation strategies, sending data, controlling the operation and termination of algorithms, obtaining results and returning them to the application layer, etc. The FKW layer is used to complete data packaging and provide data and program running services. After receiving the commands sent by the media middle platform framework layer, the HAL layer performs highlight analysis according to the commands and returns the parameter calculation results of the highlight analysis to the media middle platform framework layer. The final results of the algorithms are collected and sorted by the media middle platform framework layer and then sent to the application layer for processing. The application layer can present a clip application interface, video and picture file options, and present the final results of the algorithms.

[0370] After the user activates the one-click video creation function and selects the video and picture files to be edited (for example, up to 30 files are supported), after waiting for a moment, the "one-click video creation" application automatically edits the highlight segments of the video and combines the highlight segments and pictures together according to the algorithm results to generate a clipped short video and preview and play it.

[0371] For the video processing method provided by the embodiments of the present application, the electronic device performs highlight segment analysis on multiple image materials selected by the user. After obtaining the highlight segments of each video in the multiple image materials, if the sum of the durations of the highlight segments is less than the recommended total duration of the highlight segments preset, the electronic device can supplement and select highlight segments from all the videos until the sum of the durations of the highlight segments of all the videos is equal to or greater than the recommended total duration of the highlight segments. Among them, the highlight segments of each video in the multiple image materials are used to splice to obtain a target video set. In this solution, the problem of missing other highlight segments in the video is effectively avoided. At the same time, the duration and quality of the target video set spliced by the highlight segments are ensured.

[0372] It should also be noted that, in the embodiments of the present application, "greater than" can be replaced by "greater than or equal to", "less than or equal to" can be replaced by "less than", or, "greater than or equal to" can be replaced by "greater than", "less than" can be replaced by "less than or equal to".

[0373] Each of the embodiments described herein can be an independent solution or can be combined according to the internal logic, and these solutions all fall within the protection scope of the present application.

[0374] It can be understood that the methods and operations implemented by the electronic device in the above method embodiments can also be implemented by components (such as chips or circuits) available for the electronic device.

[0375] It should be noted that the personal information used in the technical solution of this application is limited to the information for which individual consent has been obtained, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) before the user uses the function, and signing the agreement (authorization) including authorizing the relevant user information. Among them, personal information includes information such as pictures and videos stored by the user.

[0376] In the technical solution disclosed in this application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0377] The method embodiments provided in this application are described above. The device embodiments provided in this application will be described below. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments. Therefore, the content not described in detail can be referred to the above method embodiments. For the sake of brevity, it will not be repeated here.

[0378] The above mainly describes the solution provided in the embodiments of this application from the perspective of method steps. It can be understood that in order to implement the above functions, the electronic device implementing this method includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should be able to realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the protection scope of this application.

[0379] The embodiments of this application can divide the electronic device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other feasible division methods in actual implementation. The following takes the division of each functional module corresponding to each function as an example for illustration.

[0380] This application also provides a chip. The chip is coupled to the memory and is used to read and execute the computer program or instruction stored in the memory to execute the methods in the above embodiments.

[0381] The present application also provides an electronic device, which includes a chip for reading and executing a computer program or instructions stored in a memory, so that the methods in the embodiments are executed.

[0382] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above-mentioned electronic device, the electronic device is enabled to execute each function or step executed by the electronic device in the above method embodiment.

[0383] An embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute each function or step executed by the electronic device in the above method embodiment. For example, the computer may be the above-mentioned electronic device.

[0384] From the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0385] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0386] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0387] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a readable storage medium. Based on such understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0388] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A video processing method, characterized in that, The method includes: Receiving a selection operation of a user for multiple image materials in a picture gallery; wherein, the multiple image materials include videos; Analyzing each video in response to the selection operation to obtain a highlight segment of each video; If the total duration of the highlight segments of all the videos among the multiple image materials is less than the recommended total duration of the highlight segments, supplementarily select highlight segments from the image frames not covered by the highlight segments of all the videos until the output condition of the highlight segments is met; Wherein, the meeting of the output condition of the highlight segments includes that the sum of the durations of all the highlight segments is equal to or greater than the recommended total duration of the highlight segments, or there are no image frames available for supplementary selection in the videos; The highlight segments of each video among the multiple image materials are used for splicing to obtain a target video set; the recommended total duration of the highlight segments is the recommended value of the sum of the durations of the highlight segments of all the videos among the multiple image materials.

2. The method according to claim 1, wherein The analyzing each video to obtain a highlight segment of each video includes: Performing an overview analysis on a preset number of image frames of each video to obtain a score for each image frame; the score is an aesthetic score for the corresponding image frame; Determining the highlight segment of the corresponding video based on the scores of the image frames in each video; Wherein, the highlight segment includes a first image frame, and the image frames within a first preset duration before and after the first image frame; the first image frame is the image frame with the highest score in the overview analysis.

3. The method according to claim 1 or 2, characterized in that, The supplementarily selecting highlight segments from the image frames not covered by the highlight segments of all the videos until the output condition of the highlight segments is met includes: Supplementarily selecting highlight segments from the second image frames of all the videos until the output condition of the highlight segments is met; Wherein, the score of the second image frame is greater than a first threshold, and the second image frame is not covered by the selected highlight segments.

4. The method according to claim 3, wherein The supplementarily selecting highlight segments from the second image frames of all the videos until the output condition of the highlight segments is met includes: Sequentially obtaining the highlight segment corresponding to each second image frame in the order from high to low of the scores of each second image frame; the highlight segment corresponding to the second image frame includes the second image frame, and the image frames within a second preset duration before and after the second image frame; If the sum of the durations of the highlight segments of all the videos after supplementarily selecting from the second image frames is greater than or equal to the recommended total duration of the highlight segments, end the supplementary selection operation of the highlight segments; If the sum of the durations of the highlight segments of all the videos after supplementarily selecting from the second image frames is less than the recommended total duration of the highlight segments, and if the sum of the durations of the highlight segments is greater than or equal to the recommended total duration of the highlight segments at a first multiple, and there are no second image frames available for supplementary selection of highlight segments in the videos, end the supplementary selection operation of the highlight segments; If the sum of the durations of the highlight segments of all the videos after supplementarily selecting from the second image frames is less than the recommended total duration of the highlight segments at the first multiple, supplementarily select highlight segments from the third image frames of all the videos until the output condition of the highlight segments is met; Among them, the third image frame is an image frame with a score greater than or equal to the second threshold and not covered by the highlight segment; the second threshold is less than the first threshold.

5. The method according to claim 4, wherein The method further includes: If the duration of the highlight segment corresponding to the second image frame is greater than the maximum duration of the preset highlight segment, starting from the left and right boundary positions of the highlight segment, reduce a preset distance towards the midpoint position of the highlight segment respectively until the duration of the reduced highlight segment is equal to or less than the maximum duration of the highlight segment; Among them, the maximum duration of the preset highlight segment is the maximum duration allowed for a highlight segment to ensure the complete highlight effect of the video.

6. The method according to claim 4 or 5, characterized in that, The method further includes: If the duration of the highlight segment corresponding to the second image frame is less than the minimum duration of the preset highlight segment, discard the second image frame; Among them, the minimum duration of the preset highlight segment is the minimum duration required for a highlight segment to ensure the complete highlight effect of the video.

7. The method according to claim 4, wherein The step of supplementing highlight segments from the third image frames of all the videos until the output conditions of the highlight segments are met includes: Sequentially obtain the highlight segments corresponding to each third image frame in the order of decreasing scores of the third image frames; the highlight segment corresponding to the third image frame includes the third image frame and the image frames within the third preset duration before and after the third image frame; If the sum of the durations of the highlight segments of all the videos after supplementing from the third image frames is greater than or equal to the recommended total duration of the highlight segments, end the operation of supplementing highlight segments; If the sum of the durations of the highlight segments of all the videos after supplementing from the third image frames is less than the recommended total duration of the highlight segments, and if the sum of the durations is greater than or equal to twice the recommended total duration of the highlight segments and there are no third image frames available for supplementing highlight segments in the video, end the operation of supplementing highlight segments; If the sum of the durations of the highlight segments of all the videos after supplementing from the third image frames is less than twice the recommended total duration of the highlight segments, extend the duration of the highlight segments until the output conditions of the highlight segments are met.

8. The method according to claim 7, wherein The method further includes: If the duration of the highlight segment corresponding to the third image frame is greater than the maximum duration of the preset highlight segment, starting from the left and right boundary positions of the highlight segment, reduce the second preset duration towards the midpoint position of the highlight segment respectively until the duration of the highlight segment after reducing the duration is equal to or less than the maximum duration of the highlight segment; Among them, the maximum duration of the preset highlight segment is the maximum duration allowed for a highlight segment to ensure the complete highlight effect of the video.

9. The method according to claim 7 or 8, characterized in that, The method further includes: If the duration of the highlight segment corresponding to the third image frame is less than the minimum duration of the preset highlight segment, discard the third image frame; Among them, the minimum duration of the preset highlight segment is the minimum duration required for a highlight segment to ensure the complete highlight effect of the video.

10. The method according to claim 7, wherein The step of extending the duration of the highlight segment includes: Select a first target highlight segment from the highlight segments, move the left and right boundary positions of the first target highlight segment to both sides respectively to obtain a second target highlight segment with an extended duration; the second magnification factor is smaller than the first magnification factor; Among them, the first target highlight segment is a highlight segment in the highlight segments whose duration is less than or equal to the maximum duration of the preset highlight segment, and there are expandable regions before and after the target highlight segment, and the distances between the left and right boundary positions of the target highlight segment and the boundary positions of adjacent segments are greater than the preset distance; The expandable region is a connected region in the video other than the highlight segments and the regions with scores less than the third threshold. The connected region is a region with the same score and continuous in time, and there is no overlapping region between the connected regions; The preset maximum duration of the highlight segment is the maximum duration allowed for a highlight segment to ensure the complete highlight effect of the video.

11. The method according to claim 10, wherein The step of moving the left and right boundary positions of the first target highlight segment to both sides respectively to obtain a second target highlight segment with an extended duration includes: Take the difference between the recommended total duration of the highlight segment with the second magnification factor and the sum of the durations of the highlight segments of all videos supplemented and selected from the third image frame as the required extended duration of the highlight segment; According to the required extended duration of the highlight segment, allocate an extended quota duration for each first target highlight segment; the extended quota duration is the duration allowed for the first target highlight segment to expand in the corresponding expandable region; In the expandable region of the target highlight segment, move the left and right boundary positions of the first target highlight segment to both sides by 1 / 2 of the extended quota duration to obtain the second target highlight segment.

12. The method according to claim 11, wherein The step of allocating an extended quota duration for each first target highlight segment according to the required extended duration of the highlight segment includes: Determine the extended quota duration of each target highlight segment according to the required extended duration of the highlight segment and the duration of each first target highlight segment; If the sum of the durations of all highlight segments and the cumulative sum of the extended quota durations are greater than the preset maximum total duration of the highlight segment, update the value of the extended quota duration to the difference between the maximum total duration of the highlight segment and the sum of the durations of all highlight segments; If the sum of the durations of all highlight segments and the cumulative sum of the extended quota durations are less than or equal to the preset maximum total duration of the highlight segment, the value of the extended quota duration remains unchanged; Among them, the preset maximum total duration of the highlight segment is the maximum duration allowed for the highlight segments of all videos.

13. The method according to any one of claims 7-12, characterized in that, The method further includes: After expanding the duration of the highlight segment, if the sum of the durations of all highlight segments is greater than or equal to the recommended total duration of the highlight segment with the second magnification factor, there is no third image frame that can be supplemented and selected in the video, and the highlight segments in the video cannot continue to expand the duration, end the operation of supplementing and selecting the highlight segments.

14. An electronic device, characterized in that, The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory are coupled to the processor; computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1-13.

15. A computer-readable storage medium, characterized in that, including computer instructions, and when the computer instructions run on the electronic device, the electronic device is caused to execute the method according to any one of claims 1-13.

16. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-13 is implemented.

Citation Information

Patent Citations

  • Video automatic generation method and system for resource matching based on materials

    CN111541946A

  • Video editing method, device and equipment and computer readable storage medium

    CN113015005A

  • Video generation method and device, electronic equipment and storage medium

    CN114501058A

  • Automatic video editing method

    CN116634192A

  • Ordering of highlight video segments

    WO2015114196A1