Video processing method, electronic equipment and storage medium

By overviewing the video and determining the target area, the electronic device only analyzes the image frame of the target area, solving the problem of low video processing efficiency and improving the user experience of one-click filming function.

CN120343181AActive Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410039924.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-18
Estimated Expiration
2044-01-10

AI Technical Summary

Technical Problem

When the user selects materials for a long time, the efficiency of electronic devices to analyze videos is low, resulting in the one-click filming function that takes too long and affects the user experience.

Method used

The electronic device performs an overview analysis of the video, determines the target area, and only performs a second number of image frame analysis on the target area, reducing the workload of the overall video frame-by-frame analysis and extracts highlighted fragments.

Benefits of technology

It improves the efficiency of electronic devices to process videos, reduces the time-consuming of one-click filming function, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343181A_ABST
    Figure CN120343181A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method, electronic equipment and a storage medium, and relates to the technical field of video data, and the method comprises the steps that the electronic equipment carries out the overview analysis of a first number of image frames on each video in a plurality of image materials, and obtains a first score of the image frame; determining a target area of the corresponding video based on the first score of the image frame in each video; analyzing a second number of image frames in the target area of each video to obtain a second score of the image frames; and determining a highlight segment of the corresponding video based on the second score of the image frame in the target area in each video. In the scheme, the electronic equipment performs the first quantity and the second quantity of image frame analysis on each video instead of performing frame-by-frame analysis on the whole video, so that the workload of the electronic equipment for performing image frame analysis is reduced, the time consumption of the electronic equipment for realizing a one-key filming function can be reduced, the video processing efficiency of the electronic equipment is improved, and the user experience is improved. And thus, the user experience of the one-key film-forming function is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of video data, and in particular, to a video processing method, an electronic device, and a storage medium. Background Art

[0002] With the development of picture and video processing technologies, users can trigger an electronic device to further process photos and videos in an album. For example, the electronic device can perform splicing processing on multiple image materials (including multiple pictures and / or multiple videos) to obtain a complete spliced video.

[0003] For example, the electronic device or a third-party video processing software in the electronic device may have a one-button video creation function (or a one-button video creation service). The one-button video creation function can automatically analyze and extract the highlight segments of the videos in the image materials selected by the user through an algorithm; then, based on the extracted highlight segments, a clipped video is automatically generated. Among them, the above-mentioned highlight segments are also called wonderful segments, which refer to video segments composed of single-frame images or consecutive multiple-frame images extracted from the above materials and used to record wonderful moments. The wonderful moments can be the moments when wonderful actions corresponding to a person's smiling face, a winning moment, an airplane landing, etc. occur.

[0004] However, when the overall duration of the materials selected by the user is relatively long, it takes a lot of time for the electronic device to analyze the multiple materials selected by the user, and the efficiency of the electronic device in processing videos is relatively low. Summary of the Invention

[0005] Embodiments of the present application provide a video processing method, an electronic device, and a storage medium, which selectively analyze the image frames in each video in the image materials, avoiding the problems of large time consumption and low video processing performance caused by analyzing the complete video, reducing the workload of the electronic device for image frame analysis, thereby reducing the time consumption of the electronic device to implement the one-button video creation function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-button video creation function.

[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions.

[0007] In a first aspect, a video processing method is provided, and the method includes:

[0008] The electronic device receives a selection operation of the user for multiple image materials in the gallery. Among them, the multiple image materials may include videos; the multiple image materials may also include videos and pictures.

[0009] In response to a selection operation, an electronic device performs an overview analysis of the first number of image frames for each video among multiple image materials to obtain a first score. The first score is an aesthetic score for the corresponding image frames. The first number corresponds to the duration of the video.

[0010] Based on the first scores of the image frames in each video, the electronic device determines the target area of the corresponding video. The target area includes the first image frame with the highest score in the overview analysis, and the image frames within the first preset duration before and after the first image frame. The electronic device analyzes the second number of image frames in the target area of each video among the multiple image materials to obtain a second score. The second score is an aesthetic score for the corresponding image frames. The second number is less than or equal to the number of image frames in the target area. Based on the second scores of the image frames in the target area of each video, the electronic device determines the highlight segment of the corresponding video. The highlight segment includes the second image frame with the highest score, and the image frames within the second preset duration before and after the second image frame. The highlight segments of each video among the multiple image materials are used to splice to obtain a target video set.

[0011] In this application, the electronic device first performs an overview analysis on the videos therein, so as to locate the target area that needs further processing. When the electronic device extracts the highlight segment, it only performs the analysis of the second number of image frames for the target area, rather than performing frame-by-frame analysis on the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time-consuming for the electronic device to implement the one-key video generation function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-key video generation function.

[0012] In a possible implementation manner of the first aspect, the method further includes:

[0013] In response to a selection operation, the electronic device obtains the analysis parameters corresponding to the multiple image materials.

[0014] The analysis parameters include the actual duration of the corresponding video, and the recommended value of the upper limit of the total analysis duration, and the recommended value of the upper limit of the total analysis duration represents the recommended value of the maximum duration required for analyzing the multiple image materials.

[0015] The electronic device determines the first number of image frames of each video according to the actual duration of each video among the multiple image materials, the number of videos among the multiple image materials, and the recommended value of the upper limit of the total analysis duration.

[0016] In this application, the electronic device can determine the first number of image frames that can be analyzed for each video within the recommended value of the upper limit of the total analysis duration. Under the limited performance of the electronic device, through the analysis of the first number of image frames, while improving the analysis efficiency of the highlight segment, a relatively reliable analysis effect can also be obtained.

[0017] In a possible implementation of the first aspect, according to the actual duration of each video in multiple image materials, the number of videos in the multiple image materials, and the recommended value of the upper limit of the total analysis duration, determine the first number of image frames of each video, including:

[0018] The electronic device determines the base number and the maximum number of the corresponding video based on the actual duration of each video and the preset corresponding relationship.

[0019] Among them, the base number is the minimum number of image frames required to ensure the analysis effect of the video, and the maximum number is the maximum number of image frames that the duration allows for analyzing the video. The preset corresponding relationship represents the maximum number and the base number of the image frames corresponding to different threshold ranges of the video duration.

[0020] The electronic device determines the total analysis quantity based on the base number and the maximum number of each video, and the recommended value of the upper limit of the total analysis duration; the total analysis quantity is the total number of image frames allowed to be analyzed for all videos in the multiple image materials. According to the actual duration of each video in the multiple image materials and the number of videos in the multiple image materials, distribute the total analysis quantity to each video to obtain the first number of image frames of each video.

[0021] In this application, the first quantity is determined by the maximum number and the base number of the image frames of each video, and the image frames of the first quantity are analyzed, which can not only ensure the analysis effect of the video, but also perform efficient analysis and processing of high-light segments under the limited performance and limited time consumption of the electronic device.

[0022] In a possible implementation of the first aspect, the analysis parameter further includes the single-frame analysis duration; the single-frame analysis duration is the duration required to analyze one image frame.

[0023] The electronic device determines the total analysis quantity based on the base number and the maximum number of each video, and the recommended value of the upper limit of the total analysis duration, including:

[0024] If the recommended value of the upper limit of the total analysis duration is less than the time consumption of analyzing the sum of the image frames of the base number of all videos in the multiple image materials, the total analysis quantity is the ratio of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration.

[0025] If the recommended value of the upper limit of the total analysis duration is greater than the time consumption of analyzing the sum of the image frames of the base number of all videos in the image materials, and the duration of the first multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumption of analyzing the sum of the image frames of the base number of all videos in the multiple image materials, the total analysis quantity is the sum of the base numbers of the image frames of all videos in the multiple image materials.

[0026] If the duration of the first multiple of the recommended upper limit of the total analysis duration is greater than the time taken to analyze the image frames of the sum of the base quantities of all videos in multiple image materials, and the duration of the first multiple of the recommended upper limit of the total analysis duration is less than the time taken to analyze the image frames of the sum of the maximum quantities of all videos in multiple image materials, the total number of analyses is the ratio of the duration of the first multiple of the recommended upper limit of the total analysis duration to the duration of single-frame analysis.

[0027] If the duration of the first multiple of the recommended upper limit of the total analysis duration is greater than the time taken to analyze the image frames of the sum of the maximum quantities of all videos in multiple image materials, the total number of analyses is the sum of the maximum quantities of the image frames of all videos in multiple image materials.

[0028] Among them, the first multiple is greater than 0 and less than 1.

[0029] In this application, the first quantity is determined by the maximum quantity and the base quantity of the image frames of each video, and the image frames of the first quantity are analyzed, which can not only ensure the analysis effect of the video, but also perform efficient analysis and processing of high-light segments under the limited performance and limited time consumption of the electronic device.

[0030] In a possible implementation manner of the first aspect, according to the actual duration of each video in multiple image materials and the number of videos in multiple image materials, the total number of analyses is allocated to each video to obtain the first quantity of the image frames of each video, including:

[0031] The electronic device traverses each video in multiple image materials, updates the first value and the second value of each video until the first value is 0.

[0032] Among them, the initial value of the first value is equal to the total number of analyses, and the initial value of the second value is 0; for each video traversed, the second value of the video is incremented by 1, and the first value is decremented by 1.

[0033] The electronic device takes the second value of each video as the first quantity of the image frames of the video.

[0034] In this application, traversing each video to determine the first quantity of each video can friendly allocate the total number of analyses to each video, so as not to miss the high-light segments of any video.

[0035] In a possible implementation manner of the first aspect, traversing each video in multiple image materials, updating the first value and the second value of each video until the first value is 0, includes:

[0036] Before traversing to the first video in multiple image materials, if the second value of the first video is equal to the maximum quantity of the image frames of the first video, skip the first video and traverse the next video of the first video;

[0037] Among them, skipping the first video means that the second value of the first video is not incremented by 1.

[0038] In this application, if the number of image frames of a video has reached the maximum number, subsequent image frames will not be allocated to this video, but to other videos that have not reached the maximum number, which can improve the effectiveness of video analysis.

[0039] In a possible implementation of the first aspect, the first number of image frames for overview analysis are evenly distributed at various positions of the video.

[0040] In this application, the first number of image frames being evenly distributed at various positions of the video can avoid missing image frames at some positions in the video.

[0041] In a possible implementation of the first aspect, the analysis parameter includes the recommended duration of a highlight segment, and the recommended duration of a highlight segment is the recommended duration of a highlight segment in a video.

[0042] Based on the first score of the image frames in each video, determining the target area of the corresponding video includes:

[0043] For each video, according to the first score of the image frames in the video, obtaining the first image frame. The electronic device determines the candidate highlight segment of the video for the segment with the recommended duration of the highlight segment including the first image frame in the video. The candidate highlight segment and the image frames within the third preset duration before and after the candidate highlight segment are determined as the target area of the video. Among them, the third preset duration is less than the first preset duration.

[0044] In this application, the electronic device can determine a candidate highlight segment that meets the recommended duration of the highlight segment based on the first image frame with the highest first score and the recommended duration of the highlight segment, and thus determine a target area including the candidate highlight segment. Since the target area includes the first image frame, the target area is worthy of further analysis to determine the highlight segment of the video. In this way, the determined highlight segment is relatively accurate.

[0045] In a possible implementation of the first aspect, the analysis parameter includes the total recommended duration of highlight segments, and the total recommended duration of highlight segments is the sum of the recommended durations of the highlight segments of all videos in multiple image materials.

[0046] The method further includes:

[0047] Calculate the sum of the durations of the candidate highlight segments of all videos in multiple image materials; if the sum of the durations of the candidate highlight segments of all videos is greater than the second multiple of the total recommended duration of the highlight segments, adjust the durations of the candidate highlight segments of each video according to the actual duration of each video in the multiple image materials, so that the sum of the durations of the candidate highlight segments of all videos is less than or equal to the second multiple of the total recommended duration of the highlight segments; wherein, the second multiple is greater than 1 and less than 2.

[0048] In this application, when the sum of the durations of the candidate highlight segments is greater than the second multiple of the total recommended duration of the highlight segments, it indicates that the durations of the candidate highlight segments are relatively long, and the durations of the candidate highlight segments can be reduced according to the duration, so as to obtain the durations allowed under limited performance and effective time consumption, thereby effectively analyzing the highlight segments.

[0049] In a possible implementation manner of the first aspect, the analysis parameters include the remaining analysis duration and the single-frame analysis duration; the remaining analysis duration is equal to the duration used for video analysis minus the total duration spent on performing the overview analysis; the single-frame analysis duration is the duration required to analyze one image frame.

[0050] Before analyzing the second number of image frames of the target region of each video in multiple image materials to obtain the second score, the method includes:

[0051] Take the ratio of the remaining analysis duration to the single-frame analysis duration as the analyzable number of the image frames of the target region of all videos in the multiple image materials; allocate the analyzable number to each video according to the duration of the target region of each video in the multiple image materials, and obtain the second number of the image frames of the target region of each video.

[0052] In this application, for the target region, the second number of image frames can also be analyzed instead of analyzing each frame one by one, which can further reduce the time consumption of analyzing the image frames and further improve the analysis efficiency of the highlight segments.

[0053] In a possible implementation manner of the first aspect, allocating the analyzable number to each video according to the duration of the target region of each video in the multiple image materials and obtaining the second number of the image frames of the target region of each video includes:

[0054] Traverse each video in the multiple image materials, update the third value and the fourth value of each video until the third value is 0; wherein, the initial value of the third value is equal to the analyzable number, and the initial value of the fourth value is 0; for each video traversed, the fourth value of the video is incremented by 1 and the third value is decremented by 1; take the fourth value of each video as the second number of the image frames of the target region of the video.

[0055] In this application, each video is traversed to determine the second quantity of each video, and the total analyzable quantity is evenly distributed to each video so as not to miss the highlight segments of any video.

[0056] In a possible implementation of the first aspect, the second quantity of image frames for analysis is evenly distributed at various positions in the target area of the video.

[0057] In this application, the second quantity of image frames being evenly distributed at various positions in the target area of the video can avoid missing the image frames at some positions in the video.

[0058] In a possible implementation of the first aspect, the analysis parameters include the recommended duration of the highlight segment. The recommended duration of the highlight segment is the recommended duration of a highlight segment in a video.

[0059] Based on the second scores of the image frames in the target area of each video, determining the highlight segments of the corresponding video includes:

[0060] For each video, according to the second scores of the image frames in the video, the second image frames are obtained. The second image frames and the image frames within the second preset duration before and after the second image frames are determined as the highlight segments of the video; the duration of the highlight segment is less than or equal to the recommended duration of the highlight segment.

[0061] In this application, the electronic device can determine a highlight segment that meets the recommended duration of the highlight segment based on the second image frame with the highest second score and the recommended duration of the highlight segment. The highlight segment determined based on the second image frame and the recommended duration of the highlight segment is relatively accurate.

[0062] In a possible implementation of the first aspect, the analysis parameters include the total analysis duration and the recommended value of the upper limit of the total analysis duration. The total analysis duration represents the total duration required to analyze multiple image materials, and the recommended value of the upper limit of the total analysis duration represents the recommended value of the maximum duration required to analyze multiple image materials;

[0063] Before obtaining the first scores by performing an overview analysis of the first quantity of image frames for each video in multiple image materials, the method further includes:

[0064] If the total analysis duration is greater than the recommended value of the upper limit of the total analysis duration, randomly select M videos from the multiple image materials; wherein, the duration required to analyze the M videos is less than or equal to the total analysis duration, and M is less than the number of videos in the image materials.

[0065] In this application, when the total analysis duration of multiple image materials is greater than the recommended value of the upper limit of the total analysis duration, some videos can be randomly selected from the multiple image materials for analysis, and efficient analysis and processing of the highlight segments can be achieved under the limited performance and limited time consumption of the electronic device.

[0066] In a possible implementation of the first aspect, an overview analysis of the first number of image frames is performed on each video among multiple image materials to obtain a first score, including:

[0067] If the sum of the actual durations of all videos among the multiple image materials is greater than a preset duration threshold, an overview analysis of the first number of image frames is performed on each video among the multiple image materials to obtain a first score.

[0068] In this application, when the sum of the actual durations of the videos of the multiple image materials is greater than the preset duration threshold, adopting this solution can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumed by the electronic device to implement the one-click video generation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-click video generation function.

[0069] In a second aspect, an electronic device is provided. The electronic device includes a memory, a display screen, and one or more processors; the memory, the display screen are coupled to the processor; computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is made to execute the method described in any one of the above first aspects.

[0070] In a third aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on an electronic device, the electronic device can be made to execute the method described in any one of the above first aspects.

[0071] In a fourth aspect, a computer program product containing instructions is provided. When it runs on an electronic device, the electronic device can be made to execute the method described in any one of the above first aspects.

[0072] In a fifth aspect, an embodiment of the present application provides a chip. The chip includes a processor, and the processor is used to call a computer program in the memory to execute the method in the first aspect.

[0073] It can be understood that the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect can refer to the beneficial effects in the first aspect and any of its possible design manners, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a schematic diagram of an application scenario of a video processing method provided by an embodiment of the present application;

[0075] Figure 2Schematic diagram of an application scenario of another video processing method provided by an embodiment of the present application;

[0076] Figure 3 Schematic diagram of an application scenario of another video processing method provided by an embodiment of the present application;

[0077] Figure 4 Schematic diagram of an application scenario of another video processing method provided by an embodiment of the present application;

[0078] Figure 5 Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;

[0079] Figure 6 Schematic diagram of the software structure of an electronic device provided by an embodiment of the present application;

[0080] Figure 7 Schematic diagram of the method flow for initializing algorithms, obtaining analysis parameters, etc. for each module in the electronic device provided by an embodiment of the present application;

[0081] Figure 8 Schematic diagram of the process of picture analysis in a video processing method provided by an embodiment of the present application;

[0082] Figure 9 Schematic diagram of the technical idea of a video processing method provided by an embodiment of the present application;

[0083] Figure 10 Schematic diagram of the process of video analysis in a video processing method provided by an embodiment of the present application;

[0084] Figure 11 Schematic diagram of extracting image frames provided by an embodiment of the present application;

[0085] Figure 12 Schematic diagram of another way of extracting image frames provided by an embodiment of the present application;

[0086] Figure 13 Schematic diagram of allocating image frames to multiple videos provided by an embodiment of the present application;

[0087] Figure 14 Schematic diagram of a target area provided by an embodiment of the present application;

[0088] Figure 15 Schematic diagram of another target area provided by an embodiment of the present application;

[0089] Figure 16 Schematic diagram of a highlight segment provided by an embodiment of the present application;

[0090] Figure 17 Schematic diagram of another highlight segment provided by an embodiment of the present application;

[0091] Figure 18 This is a schematic flowchart of the post - processing of highlight segments in a video processing method provided by an embodiment of this application. Detailed implementation manners

[0092] In the description of the embodiments of this application, the terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and the appended claims of this application, the singular forms "a", "the", "above", "this" and "such" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one", "one or more" means one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist; for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0093] The reference to "an embodiment" or "some embodiments" etc. described in this specification means that specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of this application. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First", "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.

[0094] In the embodiments of this application, words such as "exemplary" or "for example" are used to mean for example, as an illustration or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0095] First, some nouns or terms involved in this application are explained.

[0096] A highlight segment, also known as an exciting segment, refers to a single-frame image or a video segment composed of consecutive multiple-frame images extracted from a video or a picture for recording exciting moments. The exciting moments can be the moments when a person smiles, wins a championship in a competition, takes off in a sport, an airplane lands, or a goal is scored in a ball game corresponding to the occurrence of exciting actions.

[0097] The one-click video compilation function refers to that an electronic device, in response to a user's selection operation on one or more image materials, automatically analyzes the highlight segments in the image materials through an algorithm and combines the highlight segments into a compiled video set. That is to say, the electronic device can extract multiple highlight segments from one or more image materials and synthesize the multiple highlight segments into a video set. Among them, the image materials selected by the user can be pictures or videos; or the image materials can include pictures and videos.

[0098] Among them, the process of the electronic device selecting highlight segments from a video or an image may include: obtaining aesthetic scoring parameters such as the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate value of each frame of the image, and performing aesthetic scoring on each frame of the image according to these aesthetic scoring parameters. The electronic device can use the single-frame image or consecutive multiple-frame images with the highest score as a highlight segment.

[0099] Currently, the one-click video compilation function supports the editing of materials such as pictures and videos. In the process of implementing the one-click video compilation function, the user may select a relatively large number of pictures or a video with a relatively long duration from the image materials. In this case, it takes a long time for the electronic device to analyze the selected image materials by the user to extract the highlight segments, and the efficiency of the electronic device in processing the video is low. Moreover, the long time-consuming will cause the user to wait too long for the one-click video compilation to output the compiled video, affecting the user experience of the one-click video compilation function.

[0100] In view of the above problems, the embodiments of the present application provide a video processing method. By adopting this solution, in the process of implementing the one-click video compilation function, the electronic device can first perform an overview analysis on the first number of image frames in each video among the multiple image materials selected by the user to obtain the first score of the image frames in each video. Then, the electronic device can determine the target area of the video based on the first image frame with the highest score in a video and the image frames within the first preset duration before and after the first image frame. After that, the electronic device can analyze the second number of images in the target area of the video to obtain the second score of the image frames. Based on the second image frame with the highest second score and the image frames within the second preset duration before and after the second image frame, the highlight segment of the video is determined.

[0101] With this solution, the electronic device first performs an overview analysis on the video therein, so as to locate the target area that needs further processing. When the electronic device extracts the highlight segments, it only analyzes the second number of image frames for the target area, rather than analyzing each frame of the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumed by the electronic device to implement the one-key video compilation function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-key video compilation function.

[0102] The video processing method provided by the embodiments of the present application can be applied to an electronic device with an image processing function. It should be noted that the image materials used for one-key video compilation in the embodiments of the present application can include pictures and videos. The user can use the one-key video compilation function to generate a video set from the highlight segments of multiple pictures, or use the one-key video compilation function to generate a video set from the highlight segments of multiple videos, or use the one-key video compilation function to generate a video set from the highlight segments of multiple pictures and videos.

[0103] The above-mentioned electronic device can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city or wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the electronic device.

[0104] For the sake of easy understanding, the following combines Figures 1-4 , taking the electronic device as a mobile phone as an example, to introduce the application scenario and interface implementation of the electronic device to implement the one-key video compilation function.

[0105] In an application scenario, the user can use the mobile phone to pre-capture multiple image materials, and the mobile phone stores the multiple image materials in the picture library. The mobile phone can also pre-obtain the image materials transmitted from other devices. For exampleFigure 1 As shown in (a) in [description], the icon of the gallery is displayed on the desktop of the mobile phone. When the user wants to use the mobile phone to generate a clipped video set based on multiple image materials, the user can click on the icon of the gallery on the mobile phone desktop. In response to the user's click operation on the icon, the mobile phone displays the gallery interface 101 as shown in Figure 1 (b) in [description]. The gallery interface 101 includes the "One-Click Blockbuster" option. After that, in response to the user's Figure 1 click operation on the "One-Click Blockbuster" option shown in (b) in [description], it can display the gallery interface 102 as shown in Figure 1 (c) in [description]. The gallery interface 102 may include multiple recently taken image materials. The user can select Figure 1 any one or more of the multiple image materials shown in (c) in [description] as candidate image materials for one-click video creation. For example, in response to the user's selection operation on Figure 1 some of the image materials in (c), it can display the gallery interface 103 as shown in Figure 1 (d) in [description]. The gallery interface 103 includes all the image materials in the gallery. The gallery interface 103 may also include a video generation option, such as a "√" check mark option.

[0106] In response to the user's click operation on the Figure 1 "√" check mark option shown in (d) in [description], the mobile phone can analyze the 5 image materials selected by the user in the gallery interface 103, select the highlight segments from each image material, and generate a video set based on the selected highlight segments. During this process, the mobile phone can display Figure 2 the gallery interface 201 shown in [description]. The gallery interface 201 includes the analysis material progress so that the user can intuitively view the analysis progress.

[0107] In one example, after the mobile phone generates the video set, it can display Figure 3 the gallery interface 301 shown in (a) in [description]. The generated video set can be displayed in the gallery interface 301. The mobile phone can automatically play the video set in the gallery interface 301. In addition, as Figure 3 shown in (a) in [description], the gallery interface 301 may also include a video export option 302 for supporting the export of the generated video set. In response to the user's click operation on the video export option 302, the mobile phone can save the video set in the gallery, so that the user can view the video set from the gallery. In response to the user's click operation on the video export option 302, the mobile phone can also display Figure 3 the video export interface 303 shown in (b) in [description].

[0108] In one example, during the process of the user selecting image materials, the mobile phone can prompt the user so that the user can know how many image materials are more appropriate to select. For example, Figure 1 as shown in (d) of

[0109] In one example, during the process of the user selecting image materials, the mobile phone can prompt the user so that the user can know the maximum number of image materials that can be selected. For example, Figure 4 as shown, a prompt message of "Up to 30 image materials can be selected" is displayed in the gallery interface 401, so that the user can know how many image materials can be selected.

[0110] In one example, after generating the video set, the mobile phone can also display other function options on the interface of the video set so that the user can perform operations such as editing, adding special effects, and analyzing the generated video set based on these function options. For example, Figure 3 as shown in 301 of

[0111] Taking the electronic device as a mobile phone as an example below, in combination with Figure 5 the hardware structure of the electronic device will be introduced.

[0112] Figure 5 shows a schematic diagram of the hardware structure of the electronic device 100 provided in the embodiment of the present application. As Figure 5 shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a camera 193, a display screen 194, etc.

[0113] The processor 110 may include one or more processing units. For example, the processor 110 may include a controller, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, the controller may be the nerve center and command center of the mobile phone 100. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions. A memory may also be provided in the processor 110 for storing instructions and data.

[0114] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0115] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0116] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering, such as rendering the Figures 1-4 schematic diagram of the operation interface shown, etc.

[0117] The display screen 194 is used to display the operation interface, screen mirroring images, screen mirroring videos, etc. of the screen mirroring APP. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), and a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0118] In this embodiment, the display screen 194 can be used to display, for example, Figure 1 the gallery interfaces 101, 102, and 103 in Figure 2 ; the display screen 194 can be used to display, for example, Figure 3 the gallery interface 201 in Figure 4 ; the display screen 194 can be used to display, for example, the gallery interface 301 and the video export interface 303 in

[0119] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0120] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0121] The camera 193 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0122] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0123] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0124] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0125] The external memory interface 120 can be used to connect an external memory card to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function.

[0126] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs (APPs) required for at least one function (such as a camera APP, a gallery APP, and third-party video editing software, etc.). The data storage area can store data created during the use of the mobile phone 100 (such as photos or videos taken, mobile phone screenshots, mobile phone screen recording content, images downloaded from other devices, and video sets generated using the one-click video creation function, etc.). In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, and a universal flash storage (UFS), etc.

[0127] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.

[0128] For example, after a video set is generated using the one-click video creation function, the audio module 170 decodes the audio signal of the video set, and then the speaker 170A, also known as the "speaker", converts the audio electrical signal into a sound signal. In this way, the user can hear the background sound synchronized with the video in the highlight segment and the added video background music, etc.

[0129] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0130] The following combines Figure 6 to introduce the software structure of the electronic device.

[0131] Figure 6 It is a schematic diagram of the software structure of the electronic device provided by the embodiments of the present application.

[0132] As Figure 6As shown, the electronic device can adopt a layered architecture, dividing the software into several layers, each layer having a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software layers of the software structure are sequentially divided from top to bottom into: application (APP) layer, media middle platform framework layer, application framework (FWK) layer, and hardware abstract layer (HAL).

[0133] The APP layer, simply referred to as the application layer, may include a series of application program packages, such as cameras, galleries, third-party video editing software, calendars, maps, and navigation. When these application program packages are run, they can access each service module provided by the media middle platform framework layer and the application framework layer through the application programming interface (API), and execute corresponding intelligent services.

[0134] In some embodiments, the camera is used to take photos, videos, slow-motion images, panoramic images, etc. in response to user operations. After these images are taken by the camera, or after the user triggers a mobile phone screenshot, or after the user triggers a mobile phone screen recording, or after the electronic device downloads images from other devices, the electronic device can save these images in the gallery, so that the user can perform video editing operations on the images in the gallery, such as one-click video creation operation.

[0135] In the embodiments of the present application, the gallery is sequentially divided from top to bottom into: business layer, application function layer, and basic function layer.

[0136] Among them, the business layer, also known as the video editing business layer, provides multiple services (which can also be called functions) such as multi-shot video automatic video creation, one-shot multi-gain AI music short film, one-click video creation, and wonderful moments. These services are presented in the form of controls in the user interface (UI) of the gallery. By operating these controls, the user can trigger the gallery to perform corresponding video processing actions. For example, after the user selects image materials (the materials include pictures and / or videos), in response to the user's click operation on the one-click video creation control in the gallery, the gallery can call the underlying module to automatically analyze and extract the highlight segments in the pictures and / or videos through algorithms, and then combine the highlight segments into a clipped video set.

[0137] The application function layer includes an automatic editing framework. Each service in the service layer can call the automatic editing framework to provide automatic editing services for pictures and videos. Exemplarily, the automatic editing framework may include function modules such as segment optimization, storyline organization, layout splicing, and special effect beautification. Segment optimization is used to call the highlight segment analysis interface and policy monitoring interface in the high media middleware framework layer to extract highlight segments from pictures and / or videos. Storyline organization is used to sequentially splice multiple pictures and / or videos in the form of a storyline based on the content of the pictures and / or videos. Layout splicing is used to adjust the interface layout of pictures and / or videos. Special effect beautification is used to adjust the beautification effect of pictures and / or videos, such as adjusting the picture brightness and beautifying the human face, etc.

[0138] The basic function layer is used to perform basic function processing on the clipped picture and / or video segments after the automatic editing framework clips multiple pictures and / or videos. Exemplarily, the basic function layer may include basic function modules such as video splicing, synthesis and saving, video effect rendering, and audio effect processing. Among them, video splicing is used to splice multiple extracted highlight segments (where the highlight segments include pictures and / or videos). Synthesis and saving is used to store the video set obtained after splicing. Video effect rendering is used to add video effects to the video set obtained after splicing, such as adding style filters and themes to the video set. Audio effect processing is used to add sound effects to the video set obtained after splicing, such as adding background music to the video set.

[0139] The media middleware framework layer is a software layer set between the APP layer and the FWK layer. The media middleware framework layer may include an analysis performance query interface, a highlight segment analysis interface, a policy monitoring interface, a pipeline interface, a theme summary interface, and an initialization interface. Among them, the analysis performance query interface is used to calculate the total duration of all videos according to the video analysis speed. The highlight segment analysis interface is used to call the policy monitoring interface to extract highlight segments. The theme summary interface is used to call the underlying algorithm to analyze the picture content of the highlight segments to determine the theme corresponding to the content of the highlight segments. The initialization interface is used to initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, and image super-resolution algorithm, etc. in the HAL layer.

[0140] The policy monitoring interface is used to determine the first quantity of the image frames of each video and the positions of the first quantity of image frames in the video. The channel interface is used to reduce the resolution of the video file according to the file descriptor of each video and the positions of the image frames sent by the policy monitoring interface, and forward the data address of the downscaled video file to the hardware abstraction layer through the application framework layer, and then report the analysis result of the image frames returned by the hardware abstraction layer to the policy monitoring interface. Among them, the analysis result may include the aesthetic score of the image frames.

[0141] The policy monitoring interface is also used to determine the target area of each video according to the first score of the image frame; and determine the second quantity of the image frames in the target area of each video and the positions of the image frames with the second quantity in the target area of the video. Wherein, the target area is the area in the video corresponding to the image frame with the highest score. The channel interface is also used to reduce the resolution of the video file according to the file descriptor of each video and the position of the image frame sent by the policy monitoring interface, and forward the data address of the video file with reduced resolution to the hardware abstraction layer through the application framework layer, and then report the analysis result of the image frame in the target area returned by the hardware abstraction layer to the policy monitoring interface.

[0142] The policy monitoring interface is also used to determine the highlight segment of each video according to the second score of the image frame.

[0143] The theme summary interface is used to call the underlying algorithm to analyze the content of the highlight segment to determine the theme corresponding to the content of the highlight segment. The initialization interface is used to initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, image super-resolution algorithm, etc. in the HAL layer.

[0144] It should be noted that this application is described by taking the gallery providing the one-click video creation function as an example, which does not form a limitation on the embodiments of this application. In actual implementation, a third-party video editing software can adopt the video processing method provided by the embodiments of this application to synthesize multiple pictures and videos selected by the user into a video set with one click.

[0145] The FWK layer, simply referred to as the framework layer, can be used to support the operation of each module in the media middle platform framework layer. For example, the framework layer can include an one-click video creation interface, a parameter management interface, an image data transmission interface, a theme analysis interface, a performance analysis interface, etc.

[0146] The HAL layer is a encapsulation of the Linux kernel driver, providing interfaces upward. It hides the hardware interface details of a specific platform, provides a virtual hardware platform for the operating system, makes it hardware-independent, and can be ported on multiple platforms. For example, the hardware abstraction layer can include a chip analysis speed interface, a highlight segment algorithm, a face detection algorithm, a video acceleration algorithm, and an image super-resolution algorithm. Among them, the highlight segment algorithm is an image processing algorithm provided by the image signal processor. This algorithm can perform aesthetic scoring on each frame of image according to the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate value of each frame of image. The aesthetic scoring can be used as a basis for evaluating whether a frame of image is a highlight segment.

[0147] In this embodiment, the high - light segment algorithm can analyze each image frame in the video to obtain the score of the image frame. For example, if the score of the image frame is greater than the preset score threshold, then the image frame and the image frames within the adjacent preset time period can be used as the high - light segments of the material video. Exemplarily, the preset score threshold can be 90. If the score of the image frame of the video is greater than 90, then the image frame and the image frames within the adjacent preset time period can be used as a high - light segment of the video. Or, if the video includes the scores of N image frames, the image frame with the highest score and the image frames within the preset time period before and after this image frame can be used as the high - light segment of the video.

[0148] In some other feasible embodiments, assuming that the value range of the score of the image frame is between 0 and 100, the scores can also be divided into different levels of score results according to different value ranges of the scores. For example, the first threshold is 80. If the score of the image frame is greater than or equal to the first threshold (the value range of the score is between 80 and 100), it means that the score result of this image frame is "high". The second threshold can be 50. If the score of the image frame is greater than or equal to the second threshold (the value range of the score is between 50 and 79), it means that the score result of this image frame is "relatively high". The third threshold can be 20. If the score of the image frame is greater than or equal to the third threshold (the value range of the score is between 20 and 49), it means that the score result of this image frame is "medium". If the score of the image frame is less than the third threshold (the value range of the score is between 0 and 19), it means that the score result of this image frame is "low". Exemplarily, in one implementation, the image frames with scores greater than or equal to the first threshold / second threshold / third threshold can be used as optional image frames for extracting the high - light segments of the video. In some other implementable ways, the image frame with the highest score can be used as the optional image frame for extracting the high - light segments of the video.

[0149] It should be noted that Figure 6 The layers shown in the software structure and the components included in each layer do not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more layers than shown, such as the system library (FWK LIB) layer and the kernel layer. Each layer may include more or fewer components than shown. In addition, the above - mentioned various functional modules may also be combined into one functional module, and each layer may also be combined into one layer. For example, the high - light segment analysis may include policy monitoring. Another example is that the media middle - platform framework layer may be set in the application framework layer.

[0150] It can be understood that in order for an electronic device to implement the video processing method in the embodiments of the present application, it includes the corresponding hardware and / or software modules for executing various functions. Combining the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments.

[0151] In the video processing method provided by the embodiments of the present application, the policy monitoring module in the media middleware framework layer can determine the first quantity of image frames for overview analysis of each video according to the analysis parameters of the image material. The policy monitoring module can send each image frame in each video to the algorithm module in the HAL layer for analysis, and obtain the first score of each image frame from the algorithm in the HAL layer. Further, based on the first image frame with the highest first score and the image frames within the first preset duration before and after the first image frame, the target area that needs to be further processed for each video is determined. For the target area, an analysis of the second quantity of image frames is performed. The policy monitoring module can send the second quantity of image frames in the target area to the algorithm module in the HAL layer for analysis, and obtain the second score of the image frames from the algorithm in the HAL layer. Based on this, the highlight segment of the video is determined based on the second image frame with the highest second score and the image frames within the second preset duration before and after the second image frame. The electronic device first performs an overview analysis on the videos therein, so as to locate the target area that needs to be further processed. When the electronic device extracts the highlight segment, it only performs an analysis of the second quantity of image frames on the target area, rather than performing a frame-by-frame analysis on the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time-consuming for the electronic device to implement the one-click video compilation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-click video compilation function.

[0152] Next, taking each module in the software structure diagram as shown in Figure 6 as an example of the execution subject of the audio-video processing method, an exemplary description of the video processing method provided by the embodiments of the present application is given.

[0153] Figure 7 A method flow diagram of the operations such as algorithm initialization and obtaining analysis parameters by each module in the electronic device before the policy monitoring module in the electronic device determines the first quantity of image frames for overview analysis of each video is provided. This method can be applied to a one-click video compilation scenario as shown in Figures 1-4 Taking the image material including pictures and videos as an example, as shown in Figure 7 this method may include the following S01 - S16.

[0154] S01, The service layer receives the operation of enabling the one - click video creation function input by the user.

[0155] In this embodiment, the service layer refers to the one - click video creation module of the service layer. That is, the one - click video creation module of the service layer receives the operation of enabling the one - click video creation function input by the user. For example, this operation can specifically be the click operation on the "one - click blockbuster" card as shown in (b) of Figure 1 .

[0156] S02, The service layer loads and displays candidate pictures and candidate videos.

[0157] S03, The service layer receives the operation of the user selecting multiple pictures and videos, and receives the operation of the user inputting to determine the execution of the one - click video creation function.

[0158] For example, the operation of the user selecting multiple pictures and videos can be the click operation on photos and videos as shown in (c) of Figure 1 , and the operation of the user inputting to determine the execution of the one - click video creation function can be the click operation on the checkmark option as shown in (d) of Figure 1 .

[0159] S04, The service layer, through the application function layer, calls the initialization interface of the media middleware framework layer to initialize the relevant algorithms of the HAL layer.

[0160] In this embodiment, the relevant algorithms refer to the algorithms for the functions to be implemented by the service layer. Here, the service function is the one - click video creation function, so the relevant algorithms are the relevant algorithms involved in the one - click video creation function. For example, the relevant algorithms include highlight segment algorithms, face detection algorithms, video acceleration algorithms, and image super - resolution algorithms, etc.

[0161] S05, The initialization interface of the media middleware framework layer sequentially sends the initialization parameters to the algorithm modules of the HAL layer through the channel interface of the media middleware framework layer and the service interface of the FWK layer.

[0162] Among them, in the HAL layer, one algorithm corresponds to one algorithm interface, and the FWK layer is provided with multiple service interfaces. One service interface of the FWK layer corresponds to one algorithm interface of the HAL layer. Each service interface of the FWK layer plays a role in data passthrough between the algorithm interface of the HAL layer and the channel interface of the media middleware framework layer.

[0163] The initialization parameters involved in different algorithms are different. Therefore, the algorithm initialization parameters sent to the algorithm interface through each service interface may be different.

[0164] S06, The algorithm modules of the HAL layer are initialized according to the initialization parameters.

[0165] S07, the algorithm module of the HAL layer returns an initialization success message to the channel interface of the media middle platform framework layer through the service interface of the FWK layer.

[0166] S08, the channel interface of the media middle platform framework layer calls the analysis speed interface of the HAL layer through the service interface of the FWK layer to obtain the chip analysis speed.

[0167] Among them, the service interface of the FWK layer can be a performance analysis interface.

[0168] S09, the analysis speed interface of the HAL layer returns the chip analysis speed to the channel interface of the media middle platform framework layer through the service interface of the FWK layer.

[0169] S10, the channel interface of the media middle platform framework layer returns the chip analysis speed to the initialization interface of the media middle platform framework layer.

[0170] Among them, the chip analysis speed can represent the number of image frames in a picture / video analyzed by the image signal processor per unit time; or, the chip analysis speed can also represent the duration (single-frame analysis duration) of the image signal processor analyzing a picture / analyzing an image frame in a video. Therefore, the single-frame analysis duration of the image signal processor and the processing duration of the image signal processor for a picture can be calculated based on the chip analysis speed. The duration of the image signal processor for a picture and the single-frame analysis duration in a video may be different. Exemplarily, the processing duration of a picture can be 400 ms, and the single-frame analysis duration can be 200 ms.

[0171] It should be understood that since the performances of different image signal processors are different, the chip analysis speeds corresponding to different image signal processors may be different. For an electronic device put on the market, the image signal processor is fixed, so the chip analysis speed corresponding to this image signal processor is also fixed.

[0172] In some embodiments, the channel interface of the media middle platform framework layer can also return other relevant performance parameters of various algorithms to the initialization interface of the media middle platform framework layer.

[0173] S11, the initialization interface of the media middle platform framework layer returns an initialization success message to the service layer through the application function layer.

[0174] Among them, the initialization success message can carry performance parameters of various algorithms, such as the chip analysis speed.

[0175] After obtaining the chip analysis speed, for each video in the image materials selected by the user, the following S12 - S15 can be executed to obtain the analysis parameters for the video and all image materials.

[0176] Taking Video 1 as an example for illustration, Video 1 is one of multiple videos in the image material.

[0177] S12. The service layer sends a query message to the analysis performance query interface of the media middle platform framework layer through the application function layer.

[0178] Among them, the query message includes the file descriptor fd1 of Video 1 selected by the user and the chip analysis speed. Among them, the file descriptor can be used as the unique identifier of the video.

[0179] S13. The analysis performance query interface of the media middle platform framework layer obtains the estimated analysis duration of Video 1 according to the file descriptor fd1 of Video 1.

[0180] Among them, the estimated analysis duration can be the duration required for the image signal processor to analyze a video. Analyzing a video can refer to analyzing a specified number of image frames in a video. For example, the specified number can be 1, or the specified number can also be greater than 1 and less than or equal to the number of image frames included in the video.

[0181] Among them, exemplarily, when the specified number is 1, that is, the estimated analysis duration represents the duration required for the image signal processor to analyze 1 image frame in a video, then the estimated analysis duration can also represent the analysis duration corresponding to the minimum analysis of the image frame. For example, if the single-frame analysis duration of the image signal processor is 200 ms, then the estimated analysis duration of each video (including Video 1) is 200 ms.

[0182] In some embodiments, the specified number can also be the number of image frames included in the video. Then, the estimated analysis duration represents the duration corresponding to the image signal processor analyzing all the image frames in a video. For example, if the single-frame analysis duration of the image signal processor is 200 ms and Video 1 includes 10 image frames, then the estimated analysis duration of Video 1 is 10 * 200 ms = 2000 ms. For example, if Video 2 includes 12 image frames, then the estimated analysis duration of Video 2 is 12 * 200 ms = 2400 ms.

[0183] In some embodiments, the specified number can also be a, where a is greater than 1 and less than the number of image frames included in the video, and a is a natural number. Then, the estimated analysis duration represents the duration corresponding to the image signal processor analyzing a image frames in the video. For example, if the single-frame analysis duration of the image signal processor is 200 ms, Video 1 includes 10 image frames, and a is 5. Then the estimated analysis duration of Video 1 is 5 * 200 ms = 1000 ms.

[0184] S14. The analysis performance query interface of the media middle platform framework layer returns the estimated analysis duration of Video 1 to the service layer through the application function layer.

[0185] After the one-click video compilation module in the service layer obtains the estimated analysis duration of Video 1, it can continue to execute S12 - S15 to obtain the estimated analysis duration of the next video among multiple image materials until the estimated analysis durations of all videos in the multiple image materials are obtained.

[0186] S15. The service layer obtains analysis parameters based on the estimated analysis durations of all image materials.

[0187] Among them, the analysis parameters may include the total analysis duration, the upper limit recommended value of the total analysis duration, the maximum duration of the highlight segment, the minimum duration of the highlight segment, the total recommended duration of the highlight segment, the recommended duration of the highlight segment, whether to force each video to output a highlight segment, whether to enable audio analysis, the selected highlight segments, etc.

[0188] Among them, whether to force each video to output a highlight segment is default to yes, that is, in this embodiment, each video needs to output a highlight segment.

[0189] Among them, the total analysis duration represents the total duration required to complete the analysis of all image materials selected by the user (including all pictures and all videos selected by the user). The total analysis duration includes the sum of the estimated analysis durations of all pictures and the sum of the estimated analysis durations of all videos.

[0190] For the pictures in the image materials, the processing duration of a picture can be directly determined according to the chip speed of the image processor. Therefore, the sum of the estimated analysis durations of all pictures can be directly determined according to the number of pictures in the image materials.

[0191] For the videos in the image materials, the estimated analysis duration of each video can be obtained according to S12 - 15, and the estimated analysis durations of each video are accumulated to obtain the sum of the estimated analysis durations of all videos in the image materials.

[0192] Based on the sum of the estimated analysis durations of all pictures and the sum of the estimated analysis durations of all videos, the total analysis duration of the image materials can be obtained.

[0193] For example, assume that the materials selected by the user include 10 pictures, Video 1, and Video 2. The estimated analysis duration of a picture is 400ms, the estimated analysis duration of Video 1 is 200ms, and the estimated analysis duration of Video 2 is 300ms. Then the above total analysis duration is 10 * 400ms + 200ms + 300ms = 4500ms.

[0194] The recommended value of the upper limit of the total analysis duration represents the recommended value of the maximum duration for completing the analysis of all the above-mentioned pictures and all videos. The recommended value of the upper limit of the total analysis duration can be determined based on the estimated analysis duration of each video and the estimated analysis duration of the pictures. Generally, the recommended value of the upper limit of the total analysis duration is greater than the total analysis duration. For example, calculated by analyzing at least 1 image frame for each video, the total analysis duration is 4400ms. Considering that each video may require analyzing multiple image frames, then the recommended value of the upper limit of the total analysis duration can be much greater than the total analysis duration. For example, the recommended value of the upper limit of the total analysis duration can be preset to 10000ms.

[0195] In some scenarios where the input of some analysis parameters is abnormal, the recommended value of the upper limit of the total analysis duration may also be set to be less than the total analysis duration. In the case where the recommended value of the upper limit of the total analysis duration is less than the total analysis duration, that is, when the recommended value of the upper limit of the total analysis duration is not sufficient to analyze all pictures and all videos (1 image frame of each video), some materials can be selected from the selected image materials for analysis. This part is the method implemented by the policy monitoring module of the media middle platform framework layer, which will be introduced in detail in the following embodiments and will not be elaborated here.

[0196] The maximum duration of a highlight segment represents the maximum allowed duration of a highlight segment in a video. The minimum duration of a highlight segment represents the minimum allowed duration of a highlight segment in a video. The maximum duration of a highlight segment and the minimum duration of a highlight segment can be preset values. Exemplarily, the maximum duration of a highlight segment can be 3000ms, and the minimum duration of a highlight segment can be 1000ms.

[0197] The total recommended duration of highlight segments represents the recommended value of the sum of the recommended durations of the highlight segments of all videos in multiple image materials. The total recommended duration of highlight segments can be determined based on the number of videos, the maximum duration of a highlight segment, and the minimum duration of a highlight segment. Exemplarily, the maximum duration of a highlight segment can be 3000ms, the minimum duration of a highlight segment can be 1000ms, and when the number of videos is 5, the value range of the total recommended duration of highlight segments can be 5000ms - 15000ms. For example, the total recommended duration of highlight segments can be 8000ms.

[0198] The recommended duration of a highlight segment represents the recommended value of the duration of a highlight segment in a video. The highlight segment with this recommended duration value can effectively display the highlight effect. The recommended duration of a highlight segment can be determined based on the maximum duration of a highlight segment and the minimum duration of a highlight segment. Exemplarily, the maximum duration of a highlight segment can be 3000ms, the minimum duration of a highlight segment can be 1000ms, and the value range of the recommended duration of a highlight segment can be 1000ms - 3000ms. For example, the recommended duration of a highlight segment can be 2000ms.

[0199] It should be understood that the maximum duration, minimum duration, total recommended duration, and recommended duration of the highlight segments can all be set according to the actual situation.

[0200] In some embodiments, the analysis parameters may further include the actual duration of each video.

[0201] If the sum of the actual durations of all videos of multiple image materials is greater than a preset duration threshold, the video analysis method provided in S16 - S46 of this solution can be executed to reduce the number of analyzed image frames of the videos in the image materials, reduce the time consumption of the electronic device to implement the one - click video editing function, improve the efficiency of the electronic device in processing videos, and thus improve the user experience of the one - click video editing function.

[0202] S16. The business layer, through the application function layer, sends the file descriptor fd and analysis parameters of all materials to be analyzed to the image highlight segment analysis interface of the media middle - platform framework layer.

[0203] The image highlight segment of the media middle - platform framework layer calls the policy monitoring module of the media middle - platform framework layer to execute the video processing method provided in the following embodiments.

[0204] After receiving the analysis parameters, the policy monitoring module of the media middle - platform framework layer can determine the material analysis strategy according to the upper - limit recommended value of the total analysis duration and the total analysis duration in the analysis parameters. The material analysis strategy refers to analyzing all the image materials selected by the user, or selecting some of the image materials selected by the user for analysis. In the following embodiments, the image materials include all the videos and all the pictures selected by the user.

[0205] After obtaining the total analysis duration of the image materials, the policy monitoring module can determine the number of pictures and videos that can be analyzed according to the total analysis duration of the image materials and the upper - limit recommended value of the total analysis duration. After Figure 7 S16, referring to Figure 8 The method flow for determining the analysis strategy for image materials given below includes:

[0206] S17. The policy monitoring module of the media middle - platform framework layer determines the material analysis strategy according to the total analysis duration and the upper - limit recommended value of the total analysis duration.

[0207] Among them, if the total analysis duration is less than or equal to the upper - limit recommended value of the total analysis duration, the material analysis strategy can be to analyze all pictures and all videos. For example, if the total analysis duration is 4400ms and the upper - limit recommended value of the total analysis duration is 10000ms, the material analysis strategy can be to analyze all the pictures and all the videos in the image materials selected by the user.

[0208] If the total analysis duration is greater than the recommended upper limit of the total analysis duration, the material analysis strategy can be a random selection strategy. For example, if the total analysis duration is 4400 ms and the recommended upper limit of the total analysis duration is 3000 ms, the material analysis strategy can be a random selection strategy. Among them, the random selection strategy means randomly selecting some image materials from the image materials selected by the user for analysis.

[0209] Suppose the analysis value of a single picture is greater than that of one image frame of a video. Then, the random selection strategy can be: for every N pictures selected, M videos are allowed to be selected until the total analysis duration of the selected image materials reaches the recommended value of the analysis duration upper limit or all the pictures in the selected image materials have been taken. Among them, M < N. For example, N can be natural numbers such as 3, 4, 5, etc., and M can be natural numbers less than N such as 1, 2, 3, etc. The specific values of M and N can be determined according to the number of image materials selected by the user. Exemplarily, it can be that for every 4 pictures selected, 1 video is allowed to be selected until the total analysis duration meets 3000 ms; or all the pictures in the selected image materials have been taken.

[0210] For example, the total analysis duration of 10 pictures and 2 videos is 4400 ms, which is greater than the recommended upper limit of the total analysis duration of 3000 ms. According to the random selection strategy of allowing 1 video to be selected for every 4 pictures selected, 4 pictures and 1 video are selected. The total analysis duration of 4 pictures and 1 video is 1800 ms, which does not reach the recommended upper limit of the total analysis duration of 3000 ms. Continue to select image materials. When the 3rd picture is selected in this round, the total analysis duration reaches 3000 ms, and at this time, the selection of image materials stops. Then, the selected image materials are 7 pictures and 1 video. The remaining 3 pictures will not be analyzed.

[0211] In another embodiment, suppose the analysis value of a single picture is less than that of one image frame of a video. Then, the random selection strategy can be that for every P videos selected, Q pictures are allowed to be selected until the total analysis duration of the selected image materials reaches the recommended value of the analysis duration upper limit or all the videos in the selected image materials have been taken. Among them, Q < P. For example, P can be natural numbers such as 3, 4, 5, etc., and Q can be natural numbers less than P such as 1, 2, 3, etc. The specific values of P and Q can be determined according to the number of image materials selected by the user. Exemplarily, it can be that for every 3 videos selected, 1 picture is allowed to be selected until the total analysis duration meets 3000 ms; or all the videos in the selected image materials have been taken.

[0212] For example, the total analysis duration of 10 videos and 3 pictures is 3200 ms, which is greater than the recommended upper limit of the total analysis duration of 3000 ms. According to the random selection strategy of allowing 1 picture to be selected for every 3 videos selected, 3 videos and 1 picture are selected. The total analysis duration of 3 videos and 1 picture is 1000 ms, which does not reach the recommended upper limit of the total analysis duration of 3000 ms. Continue to select image materials. After three rounds of selecting 3 videos and 1 picture, a total of 9 videos and 3 pictures are selected, and the total analysis duration is 3000 ms. At this time, stop selecting image materials. Then, the selected image materials are 9 videos and 3 pictures. The remaining 1 video is not analyzed.

[0213] In some embodiments, the electronic device determines analyzable image materials from the materials selected by the user according to the above random selection strategy (such as a mobile phone). Among them, pictures and videos can be randomly selected from the image materials selected by the user in the order in which the user selects the image materials. Alternatively, random numbers less than the number of materials can also be generated by the Random class in Java to select the corresponding videos or pictures.

[0214] In some embodiments, the image materials only include pictures, and the total analysis duration is the sum of the estimated analysis durations of all pictures. When the total analysis duration is greater than the recommended value of the analysis duration upper limit, the number of analyzable pictures is calculated according to the recommended value of the analysis duration upper limit and the estimated analysis duration of one picture, and the corresponding number of pictures is randomly selected from all pictures for analysis.

[0215] In some embodiments, the image materials only include videos, and the total analysis duration is the sum of the estimated analysis durations of all videos. When the total analysis duration is greater than the recommended value of the analysis duration upper limit, the number of analyzable videos is calculated according to the recommended value of the analysis duration upper limit, and the corresponding number of videos is randomly selected from all videos for analysis. Among them, the duration required for the number of analyzable videos is less than or equal to the total analysis duration.

[0216] After selecting analyzable pictures and / or videos within the recommended value of the analysis duration upper limit from the image materials, the remaining pictures and videos in the image materials selected by the user are not analyzed.

[0217] The image materials include pictures. After the electronic device selects analyzable pictures within the recommended value of the analysis duration upper limit, it executes Figure 8 the S18 - S22 shown to perform highlight segment analysis on the analyzable pictures.

[0218] The following is an example to illustrate the process of performing highlight segment analysis on Picture 1 in combination with S18 - S22, where Picture 1 is one of the analyzable pictures within the recommended value of the analysis duration upper limit.

[0219] S18, The policy monitoring module in the media middle platform framework layer sends an indication message to the channel interface in the media middle platform framework layer. The indication message includes the file descriptor of Picture 1.

[0220] S19, The channel interface in the media middle platform framework layer performs processing operations such as decoding, reducing the resolution, and converting the format on Picture 1 according to the file descriptor of Picture 1, and stores the processed Picture 1.

[0221] S20, The channel interface in the media middle platform framework layer sends the frame data address of Picture 1 to the highlight segment algorithm interface in the HAL layer through the service interface in the FWK layer.

[0222] Among them, the service interface can be the one - click video generation interface.

[0223] S21, The highlight segment algorithm interface in the HAL layer obtains Picture 1 according to the frame data address of Picture 1, and then analyzes Picture 1 based on a preset highlight segment algorithm to obtain an analysis result.

[0224] Exemplarily, the highlight segment algorithm interface can perform aesthetic scoring on Picture 1 according to the image color, image texture features, image quality, edge change rate value, etc. of Picture 1 to obtain an analysis result. Among them, the analysis result can represent the aesthetic score.

[0225] S22, The highlight segment algorithm interface in the HAL layer returns the analysis result of Picture 1 to the policy monitoring module in the media middle platform framework layer through the service interface in the FWK layer and the channel interface in the media middle platform framework layer in sequence.

[0226] After the policy monitoring module in the media middle platform framework layer obtains the analysis result of Picture 1, if there are other pictures, such as Picture 2, then the electronic device can continue to execute the above S18 - S22 to obtain the analysis results of other pictures.

[0227] Among them, the policy monitoring module in the media middle platform framework layer can determine the pictures with scores greater than or equal to the first threshold / second threshold as the highlight segments of multiple image materials according to the analysis results of each picture.

[0228] After obtaining the analysis results of all pictures, if the image materials include videos, the electronic device can use the following S23 - S36 to obtain the analysis results of each video. If the image materials do not include videos, the electronic device outputs a target video set composed of the pictures of the highlight segments.

[0229] Reference Figure 9, which shows a schematic diagram of a technical idea for analyzing high - light segments of a video provided by an embodiment of the present application. During the process of analyzing each video, an electronic device (such as a mobile phone) can extract a first number of image frames from the video for overview analysis to obtain a first score of the image frames. Among them, extracting the first number of image frames can be randomly extracted at any position of the video, or the first number of image frames can be extracted from different segments according to the storyboard points of the video; or the first number of image frames can be extracted at uniform positions of the video. The target area of each video is determined based on the first image frame with the highest first score. Then, a second number of image frames in the target area are analyzed to obtain a second score of the image frames. The high - light segment of the video is determined based on the second image frame with the highest second score. In this solution, instead of analyzing each frame of the entire video, the workload of the electronic device for image - frame analysis can be reduced, thereby reducing the time consumed by the electronic device to implement the one - key video compilation function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one - key video compilation function.

[0230] If the image material includes multiple pictures and videos, since the computational amount of pictures is small and the time consumed is less, the electronic device usually first analyzes the high - light segments of each picture one by one. After completing the analysis of the high - light segments of all pictures, it then analyzes the high - light segments of each video one by one. That is, after the electronic device executes S18 - S22 as shown in Figure 8 , it continues to execute S23 - S37 as shown in Figure 10 .

[0231] If the image material only includes multiple videos, then after the electronic device executes S17, it executes S23 - S37 as shown in Figure 10 , and does not need to execute S19 - S22 as shown in Figure 8 .

[0232] In this embodiment, an example of the case where the image material includes pictures and videos is described. The following videos all refer to videos that can be analyzed within the recommended upper limit of the total analysis duration.

[0233] S23, the policy monitoring module in the media middle - platform framework layer calculates the first number of image frames of each video.

[0234] The first number refers to the number of image frames that can be analyzed in each video within the recommended upper limit of the total analysis duration.

[0235] In some embodiments, the policy monitoring module can determine the first number of image frames of each video according to the actual duration of each video in multiple image materials, the number of videos in multiple image materials, and the recommended upper limit of the total analysis duration.

[0236] For example, according to the recommended value T1 of the total analysis duration upper limit and the single-frame analysis duration t, determine the number of image frames T1 / t that can be analyzed within the recommended value of the total analysis duration upper limit. According to the actual duration of each video and the number of videos, allocate the number of image frames T1 / t to each video to obtain the first quantity of each video.

[0237] Among them, in one example, allocating the number of image frames T1 / t to each video can be to sort the videos in descending order of their actual durations. For the videos ranked in the top 25%, evenly allocate 50% of the number of image frames of T1 / t. For the videos ranked between 25% - 75%, evenly allocate 40% of the number of image frames of T1 / t. For the videos ranked in the bottom 25%, evenly allocate 10% of the number of image frames of T1 / t.

[0238] Among them, in another example, allocating the number of image frames T1 / t to each video can be to allocate the corresponding quantities to each video respectively according to the proportion of the actual duration of the video.

[0239] Or, in another example, the steps for the policy monitoring module to determine the first quantity of the image frames of each video may include:

[0240] S231, the policy monitoring module determines the base quantity and the maximum quantity of the image frames of each video.

[0241] Among them, the base quantity can be understood as the minimum number of image frames that need to be analyzed to ensure the analysis effect of the video when the analysis duration is limited and the analysis performance of the electronic device is limited. The maximum quantity can be understood as the maximum number of image frames that the video can be analyzed for within the allowed duration. Generally, the maximum quantity is greater than the base quantity.

[0242] In some embodiments, the policy monitoring module can determine the base quantity and the maximum quantity of the image frames of each video according to the actual duration of each video and a preset correspondence. Among them, the preset correspondence represents the maximum quantity and the base quantity of the image frames corresponding to different threshold ranges of the video duration.

[0243] The base quantity and the maximum quantity of videos with different durations are different. For example, a preset correspondence between the video duration and the number of image frames can be pre-stored in an electronic device (such as a mobile phone). In the embodiments of the present application, the policy monitoring module can determine the base quantity and the maximum quantity of the image frames of each video according to this preset correspondence.

[0244] Exemplarily, the preset correspondence includes the following (1)-(4):

[0245] (1) The video duration P is less than the first threshold Q1, and the number of image frames of the video is 1.

[0246] (2) The video duration P is equal to the first threshold Q1, and the starting number of image frames of the video is m.

[0247] (3) The video duration P is greater than the first threshold Q1 and less than or equal to the second threshold Q2. The number of video image frames increases based on the starting number. The number of image frames is updated to that is, the integer part of P divided by s plus m. It means that from the start time of the video to Q1, the number of image frames of the video is m, and from the start time of the video to Q2, 1 image frame can correspond to every s seconds.

[0248] (4) The video duration P is greater than the second threshold Q2 and less than the third threshold Q3. The number of video image frames is that is, the integer part of Q2 divided by s, the integer part of (P - Q2) divided by k, plus m. It means that from the start time of the video to Q1, the number of image frames of the video is m; from the start time of the video to Q2, 1 image frame can correspond to every s seconds; from Q2 to Q3, 1 image frame can correspond to every k. Among them, is the floor function symbol.

[0249] Among them, the first threshold < the second threshold < the third threshold.

[0250] It should be noted that if the video duration is very long, more duration thresholds can be set, such as the fourth threshold, the fifth threshold, and so on. When determining the base quantity and the maximum quantity, thresholds such as the first threshold, the second threshold, and the third threshold can be set according to the actual video duration; the interval seconds can be determined according to the analysis performance of the electronic device. The above are just examples for illustration, and no limitations are imposed on the parameter values.

[0251] In some embodiments, the maximum number of image frames is more than the base number, and the frame extraction density for determining the maximum number of image frames is greater than the frame extraction density for determining the base number of image frames. In order to obtain more image frames, generally, the first threshold for determining the maximum number is less than or equal to the first threshold for determining the base number; or, the preset second threshold for determining the maximum number is equal to or greater than the second threshold for determining the base number, or, the third threshold for determining the maximum number is greater than or equal to the third threshold for determining the base number. In some embodiments, the interval seconds for determining the maximum number is less than or equal to the interval seconds for determining the base number.

[0252] The following gives several examples to illustrate the determination process of the base number and the maximum number of image frames of the video.

[0253] For example, the policy monitoring module determines the basic number of image frames of a video. The first threshold can be 3 seconds, s seconds can be 5 seconds, the second threshold can be 30 seconds, k seconds can be 10 seconds, and the third threshold can be 90 seconds.

[0254] Exemplarily, taking the video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example, exemplarily, reference can be made to Figure 11 the schematic diagram of extracting image frames shown.

[0255] The policy monitoring module receives video 1 and determines at least one image frame. At this time, the number of image frames of video 1 is 1 frame. The duration of video 1 (91 seconds) is greater than the first threshold (3 seconds), so the number of image frames of video 1 is updated to 2 frames.

[0256] Starting from the 0th second of the duration of video 1 until the 30th second of the video duration, the number of image frames is increased by one every 5 seconds. As Figure 11 shown, at the 30th second of the video duration, the basic number of image frames of video 1 is 8 frames Starting from the 30th second of the video duration until the 90th second of the video duration, the number of image frames is increased by one every 10 seconds. As Figure 11 shown, at the 90th second of the video duration, the basic number of image frames of video 1 is updated to 14 frames After the 90th second of video 1, the remaining duration of the video is 3 seconds, which does not meet the requirement of extracting 1 image frame every 10 seconds after 90 seconds. Therefore, the number is not increased. So, the basic number of image frames of video 1 with a duration of 93 seconds is 14 frames.

[0257] In some embodiments, after calculating the basic number of image frames of each video, if the total number of the basic number of image frames of all videos does not meet the preset minimum number of frames, it is necessary to adjust the basic number of image frames of all videos. Among them, the adjustment method can be to increase the number of image frames of the video with the basic number of image frames less than the preset value; or, increase the basic number of image frames of the video with a longer duration.

[0258] Exemplarily, in order to ensure obtaining a more accurate video theme, the preset minimum number of frames can be 5 frames.

[0259] In some embodiments, when there is only one video, for example, if the basic number of image frames of Video 1 does not meet 5 frames, directly adjust the basic number of Video 1 to 5 frames. When there are multiple videos, since at least 1 image frame is acquired for each preset video, the preset minimum number of frames should be greater than the number of videos. For example, when the number of videos is 3, the preset minimum number of frames can be 5 frames. Exemplarily, the basic number of image frames of Video 1 is 1 frame, the basic number of image frames of Video 2 is 2 frames, and the basic number of image frames of Video 3 is 1 frame. The sum of the basic numbers of image frames of all videos is 4 frames, which is less than the preset minimum number of frames of 5 frames. The basic number of a video with less than 2 frames of image frames can be adjusted to 2 frames. At this time, the basic number of image frames of Video 1 is 2 frames, the basic number of image frames of Video 2 is 2 frames, and the basic number of image frames of Video 3 is 2 frames. The sum of the basic numbers of image frames of all videos is 6 frames, which is greater than the preset minimum number of frames of 5 frames. If the number of videos is 2 frames, the basic number of image frames of Video 1 is 1 frame, and the basic number of image frames of Video 2 is 1 frame. By adjusting the basic number of a video with less than 2 frames of image frames to 2 frames, the sum of the basic numbers of image frames of all videos is 4 frames, still not meeting the preset minimum number of frames of 5 frames. In this case, the basic number of image frames of the video with the longest duration can be adjusted to 3 frames so that the total number of basic numbers of image frames of all videos meets 5 frames. For example, the duration of Video 1 is 2 seconds, and the duration of Video 2 is 1 second. Then, adjust the basic number of image frames of Video 1 to 3 frames and the basic number of image frames of Video 2 to 2 frames. At this time, the total number of basic numbers of image frames of all videos is 5, meeting the preset minimum number of frames of 5 frames.

[0260] Exemplarily, the policy monitoring module determines the maximum number of image frames of a video. The first threshold can be 2 seconds, s seconds can be 3 seconds, the second threshold can be 30 seconds, k seconds can be 10 seconds, and the third threshold can be 120 seconds.

[0261] Exemplarily, taking the video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example to illustrate, exemplarily, reference can be made to Figure 12 the schematic diagram of image frame extraction shown.

[0262] The policy monitoring module receives Video 1 and extracts at least one image frame. At this time, the maximum number of image frames of Video 1 is 1 frame. The duration of Video 1 (93 seconds) is greater than the first threshold (2 seconds), so the maximum number of image frames of Video 1 is updated to 2 frames.

[0263] Starting from the 0th second of the duration of Video 1 to the 30th second of the video duration, the number of image frames is increased by one every 3 seconds. As Figure 12 shown, at the 30th second of the video duration, the maximum number of image frames of Video 1 is 12 frames Starting from the 30th second of the video duration and within 120 seconds of the video duration, the number of image frames is increased every 10 seconds. As Figure 12 shown, at the 93rd second of the video duration, the maximum number of image frames of Video 1 is updated to 18 frames After the 90th second of Video 1, the remaining duration of the video is 3 seconds, which does not meet the requirement of extracting 1 image frame every 10 seconds after 120 seconds. Therefore, after increasing the number of image frames by 1 at the 90th second, the number of image frames is no longer increased. So, the maximum number of image frames of Video 1 with a duration of 93 seconds is 18 frames.

[0264] S232, the policy monitoring module determines the total number of image frames to be analyzed for all videos according to the recommended value of the analysis duration upper limit.

[0265] In this embodiment, the recommended value of the analysis duration upper limit represents the recommended value of the maximum duration for completing the analysis of all the above pictures and all videos. After consuming the analysis duration for processing the pictures and the videos in the above embodiments, therefore, the total number of image frames to be analyzed for all videos is determined by the first multiple of the recommended value T1 of the analysis duration upper limit.

[0266] The policy monitoring module determines the total number of image frames to be analyzed for all videos according to the first multiple of the recommended value T1 of the analysis duration upper limit. Among them, the first multiple can be 30%. That is, according to 30% of the recommended value of the analysis duration upper limit (T1 * 0.3), the number of analyzable image frames is judged.

[0267] If T1 * 0.3 is less than the time required to analyze the basic number of image frames of all videos, however, T1 is greater than the time required to analyze the basic number of all videos, that is, the duration of T1 * 0.3 is not enough to analyze the basic number of image frames of all videos, then the policy monitoring module can allocate the analysis duration T2 for all videos for analysis. At this time, the total number of image frames to be analyzed for all videos is the sum of the basic numbers of all videos. Among them, the analysis duration T2 is greater than T1 * 0.3, and T2 is less than T1, and T2 is greater than or equal to the time required for the basic number of all videos.

[0268] Or, in the case where the duration of T * 0.3 is not enough to analyze the basic number of image frames of all videos, the policy monitoring module can allocate the entire recommended value T1 of the analysis duration upper limit to the analysis of the basic number of image frames of all videos, and the actual number of image frames processed by the policy monitoring module (total analysis quantity) is the sum of the basic numbers of the image frames of all videos.

[0269] If the recommended value T1 of the analysis duration upper limit is less than the time required to analyze the basic number of image frames of all videos, assuming that the single-frame analysis duration is t, at this time, the actual number of image frames processed by the policy monitoring module (total analysis quantity) is T1 / t (frames).

[0270] If T1 * 0.3 is greater than the time taken to analyze the maximum number of image frames of all videos, that is, the duration of T1 * 0.3 is sufficient to analyze the maximum number of image frames of all videos. Then, the policy monitoring module allocates T1 * 0.3 for analyzing the maximum number of image frames of all videos, and the number of image frames actually processed by the policy monitoring module (total analysis quantity) is the sum of the maximum numbers of image frames of all videos.

[0271] If T1 * 0.3 is greater than the time taken to analyze the basic quantity of image frames of all videos, and T1 * 0.3 is less than the time taken to analyze the maximum number of all videos. Assuming the analysis duration of a single frame is t, the number of image frames actually processed by the policy monitoring module (total analysis quantity) is T1 * 0.3 / t (frames).

[0272] Exemplarily, assume the input video durations are 15s, 40s, 100s, 150s respectively, the basic quantities of image frames corresponding to each video are 5 frames, 9 frames, 14 frames, 14 frames respectively, with a total of 42 frames. The maximum quantities of image frames corresponding to each video are 7 frames, 13 frames, 19 frames, 21 frames respectively, with a total of 60 frames.

[0273] Assume the analysis duration of a single frame is 200ms. Then, the time taken to analyze the basic quantity of image frames of all videos is the time taken for 42 frames, which is 8400ms; the time taken to analyze the maximum quantity of image frames of all videos is the time taken for 60 frames, which is 12000ms.

[0274] If T1 * 0.3 is less than 8400ms, that is, T1 * 0.3 is less than the time taken to analyze the basic quantity of all videos, the policy monitoring module allocates 8400ms for processing 42 image frames of all videos according to the time taken to process 42 frames. Or, the policy monitoring module allocates the analysis duration upper limit recommended value T1 for analyzing the image frames of all videos. At this time, the number of image frames actually processed by the policy monitoring module (total analysis quantity) is 42 frames. If T1 is less than 8400ms, the number of image frames actually processed by the policy monitoring module (total analysis quantity) is T1 / 200 (frames).

[0275] If T1 * 0.3 is greater than 12000ms, that is, T1 * 0.3 is greater than the time taken to analyze the maximum number of all videos, the policy monitoring module allocates T1 * 0.3 for analyzing the image frames of all videos, and the number of image frames actually processed (total analysis quantity) is 60 frames.

[0276] If T1 * 0.3 is greater than 8400ms, and T1 * 0.3 is less than 12000ms, within T1 * 0.3, the number of image frames actually processed by the policy monitoring module (total analysis quantity) is T1 * 0.3 / 200ms (frames).

[0277] In this embodiment, the upper limit of the analysis duration is only a recommended value for planning and is not strongly verified. In this embodiment, the estimated analysis duration of video analysis is allowed to exceed the recommended value of the upper limit of the analysis duration.

[0278] S233. The policy monitoring module in the media middle platform framework layer determines the first quantity of the image frames of each video according to the total analysis quantity of all videos.

[0279] In some embodiments, after obtaining the total analysis quantity of the image frames of all videos, the policy monitoring module can sequentially allocate the corresponding quantity of image frames to be analyzed (the first quantity) for each video.

[0280] Exemplarily, the videos can be sorted in descending order of their actual durations, and each video is traversed in turn. Each time a video is traversed, the first value and the second value of the video are updated.

[0281] Among them, the initial value of the first value is the total analysis quantity; the initial value of the second value of the video is 0.

[0282] During the process of traversing the videos, each time a video is traversed, the second value of the video is incremented by 1, and the first value is decremented by 1 until the first value is 0, that is, until the total analysis quantity is allocated. When the total analysis quantity is allocated, the second value of each video is the corresponding first quantity.

[0283] It should be noted that when allocating the first quantity of image frames for each video, the first quantity of image frames allocated to each video should be less than or equal to the maximum quantity of the image frames of the video, and the first quantity of image frames allocated to each video should be greater than or equal to the basic quantity of the image frames of the video. If the second value of a video is equal to the maximum quantity before traversing to this video, this video and all subsequent traversal processes will skip this video and traverse the next video. The videos skipped without traversing do not increase the second value. In this way, the first quantity corresponding to the actual duration of each video can be obtained.

[0284] Exemplarily, refer to Figure 13 , Figure 13A schematic diagram for allocating image frames to multiple videos is provided. Assume the input videos are Video 1 (15s), Video 2 (20s), Video 3 (30s), and Video 4 (35s), and the maximum number of image frames corresponding to each video is 7 frames, 8 frames, 12 frames, and 12 frames respectively. Assume the total number of frames to be analyzed for all videos is 38 frames. The policy monitoring module traverses these 4 videos and allocates the 38 frames to these 4 videos in sequence. Starting from the first round, 1 image frame is allocated to each video in each round. That is, when traversing each video, the corresponding second value is incremented by 1. After the 7th round, the second value of the image frames of Video 1 has reached the maximum number of its corresponding image frames (7 frames), and no more image frames will be allocated to it subsequently. That is, this video and subsequent traversal processes will skip this video. In the 8th round, only Video 2, Video 3, and Video 4 are traversed. After the 8th round, the second value of the image frames of Video 2 has reached the maximum number of its corresponding image frames (8 frames), and no more image frames will be allocated to it subsequently. In the 9th round, only Video 3 and Video 4 are traversed. Until the end of the 11th round, the number of allocated frames for each video is 7 frames, 8 frames, 11 frames, and 11 frames. Among them, Video 1 and Video 2 have reached the maximum number of image frames, and Video 3 and Video 4 can still be allocated. At this time, 37 frames out of the total number of frames to be analyzed have been allocated, and there is 1 remaining image frame, which is not enough to be allocated to all videos (Video 3 and Video 4). The policy monitoring module allocates the remaining 1 image frame to Video 4 with a longer duration in the 12th round according to the duration of the videos. After the 38 frames are allocated, the first quantity of the image frames of Video 1 is 7 frames, the first quantity of the image frames of Video 2 is 8 frames, the first quantity of the image frames of Video 3 is 11 frames, and the first quantity of the image frames of Video 4 is 12 frames.

[0285] After obtaining the first quantity of the image frames of each video, the policy monitoring module can perform highlight segment analysis on the image frames of each video. Among them, the policy monitoring module performing highlight segment analysis on each video can include two analysis stages. Among them, the first analysis stage includes performing an overview analysis on the first quantity of image frames, and the second analysis stage includes further analysis based on the target area.

[0286] Among them, the first analysis stage refers to performing an overview analysis on the first quantity of image frames for each video. Specifically, for each video, the policy monitoring module issues the position of one image frame to the channel interface of the media middle platform framework layer each time, enabling it to analyze the image frame at that position and return the analysis result; until the number of image frames issued for analysis reaches the first quantity of the video. After the first analysis stage, the policy monitoring module can obtain the analysis results of the first quantity of image frames of each video. This analysis result includes the first score of the image frame. Thus, the target area of each video can be determined based on the first image frame with the highest score.

[0287] The second analysis stage refers to the analysis of a second quantity of image frames for the target area. The policy monitoring module sequentially sends the position of an image frame of the target area to the channel interface of the media middle platform framework layer to obtain the analysis result of the image frame at that position, where the analysis result includes the second score of the image frame. Thus, the highlight segment of the video is determined based on the second image frame with the highest second score.

[0288] Taking Video 1 as an example, the process of analyzing the highlight segment of the video is illustrated below in combination with S24 - S37. Among them, S24 - S29 are the first analysis stage, and S30 - S37 are the second analysis stage.

[0289] S24, the policy monitoring module sends the file descriptor fd1 of Video 1 and the position of the first image frame of Video 1 to the channel interface of the media middle platform framework layer.

[0290] Among them, the positions of the first quantity of image frames in the video can be evenly distributed. Then the position of the first image frame can be the position of the first image frame starting from the start time of the video.

[0291] In some embodiments, the positions of the first quantity of image frames in the video can also be distributed according to the scene division points of the video and the similar picture areas of the video.

[0292] The position of the first image frame can be the position of any image frame in Video 1. For example, the position of the first image frame can be the position of the first image frame of Video 1; or, the position of the first image frame can also be other specified positions in Video 1.

[0293] S25, the channel interface of the media middle platform framework layer performs processing operations such as decoding, reducing the resolution, and converting the format of Video 1 according to the file descriptor of Video 1, and stores the processed Video 1.

[0294] S26, the channel interface of the media middle platform framework layer sends the frame data address of Video 1 and the corresponding position of the first image frame in Video 1 to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one - click video compilation interface).

[0295] S27, the highlight segment algorithm interface of the HAL layer obtains the first image frame of Video 1 according to the frame data address of Video 1 and the corresponding position of the first image frame in Video 1, and then analyzes the image frame at the position of the first image frame based on the preset highlight segment algorithm to obtain the analysis result of the image frame at the position of the first image frame.

[0296] S28, the highlight segment algorithm interface of the HAL layer sequentially returns the analysis result of the image frame at the position of the first image frame to the policy monitoring module through the service interface of the FWK layer and the channel interface of the media middle platform framework layer.

[0297] Among them, the analysis result can represent the scoring value and the scoring result of the image frame.

[0298] S29. The policy monitoring module obtains the position of the second image frame and returns to execute S24 until the number of analyzed image frames meets the first quantity of the video.

[0299] In this way, the policy monitoring module can obtain the analysis results of the first quantity of image frames of all videos. Among them, the analysis result includes the first scoring of the image frame. Based on the first scoring of the image frame, the solution of the second analysis stage is executed.

[0300] S30. The policy monitoring module determines the target area of each video based on the analysis results of the image frames of all videos.

[0301] In this embodiment, for each video, after the policy monitoring module obtains the analysis results of the first quantity of image frames, according to the first scoring of each image frame, it determines the first image frame with the highest scoring. Among them, the first image frame can include at least one image frame. If the first image frame includes one image frame, the image frames within the first preset duration before and after the first image frame can be directly determined as the target area of Video 1. For example, Figure 14 (a) of Figure 14 gives a schematic diagram of a target area. Among them, the first image frame of Video 1 includes Image Frame 1 (1 in (a) of

[0302] Or, if the first image frame includes one image frame, according to the highlight segment recommended duration, the segments within 1 / 2 of the highlight segment recommended duration before and after the first image frame can be determined as the candidate highlight segments. The image frames within the third preset duration before and after the candidate highlight segments are determined as the target area of Video 1. Among them, the duration of the target area can be b times that of the candidate highlight segment. Exemplarily, b can be a number greater than 1 and less than 2.

[0303] For example, Figure 14 (b) of Figure 14 gives another schematic diagram of a target area. Among them, the first image frame of Video 1 includes Image Frame 1 (1 in (b) of

[0304] If the first image frame includes a plurality of consecutive image frames, according to the recommended duration of the highlight segment, the segment covered by the fourth preset duration before the first image frame of the plurality of consecutive first image frames to the fourth preset duration after the last image frame of the plurality of consecutive second image frames is determined as the candidate highlight segment.

[0305] For example, as Figure 15 , Figure 15 shows another schematic diagram of the target area. The first image frame of Video 1 includes consecutive Image Frame 1 (such as 1 in Figure 15 ), Image Frame 2 (such as 2 in Figure 15 ), and Image Frame 3 (such as 3 in Figure 15 ). Then, the segment covered by the fourth preset duration before Image Frame 1 to the fourth preset duration after Image Frame 3 is the candidate highlight segment of Video 1. The segments covered by the fourth preset duration before and after the candidate highlight segment are the target areas.

[0306] After determining the candidate highlight segment of each video, if the sum of the durations of the candidate highlight segments of all videos is greater than the second multiple of the total recommended duration of the highlight segment, the policy monitoring module needs to adjust the durations of the candidate highlight segments of each video so that the sum of the durations of the candidate highlight segments of all videos is less than or equal to the second multiple of the total recommended duration of the highlight segment.

[0307] Among them, the second multiple can be a number greater than 1. For example, the second multiple can be 1.05. Exemplarily, according to the duration by which the sum of the durations of the candidate highlight segments of all videos exceeds the second multiple of the total recommended duration of the highlight segment (the duration to be adjusted), the durations of the candidate highlight segments are reduced and adjusted according to the actual durations of each video. Exemplarily, the shorter the duration of a video, the more seconds its candidate highlight segment is reduced.

[0308] Suppose two videos are input. The actual duration of Video 1 is 10s, and the actual duration of Video 2 is 20s. The recommended duration of the highlight segment is 6s, and the total recommended duration of the highlight segment is 10s. The durations of the candidate highlight segments of Video 1 and Video 2 are both 6s; the durations of the target areas of Video 1 and Video 2 are both 12s.

[0309] Among them, the sum of the durations of the candidate highlight segments of Video 1 and Video 2 is 12s, which exceeds 1.05 times (10.5s) of the total recommended duration of the highlight segments (10s). The strategy monitoring module needs to adjust the durations of the candidate highlight segments of each video. The duration to be adjusted is 12s - 10s * 1.05 = 1.5s. The strategy monitoring module distributes the weights according to the reciprocals of the video durations. The duration of the candidate highlight segment of Video 1 (10s) is reduced by 1s; the duration of the candidate highlight segment of Video 1 is updated from 6s to 5s. The duration of the candidate highlight segment of Video 2 (20s) is reduced by 0.5s, and the duration of the candidate highlight segment of Video 2 is updated from 6s to 5.5s. In this way, the sum of the durations of the candidate highlight segments of Video 1 and Video 2 is 10.5s, which does not exceed 10s * 1.05, and no further adjustment is required.

[0310] Therefore, it is determined that the duration of the candidate highlight segment of Video 1 is 5s, and the duration of the target area corresponding to the candidate highlight segment in Video 1 is 10s; the duration of the candidate highlight segment of Video 2 is 5.5s, and the duration of the target area corresponding to the candidate highlight segment in Video 2 is 11s.

[0311] The method provided by S30 can be used to determine the target area of each video. After determining the target area of each video, S31 - S36 can be executed to analyze the second number of image frames of the target area.

[0312] S31. The strategy monitoring module determines the second number of image frames of the target area of each video.

[0313] Among them, the second number refers to the number of image frames that can be analyzed in the target area of each video during the remaining analysis duration. Among them, the remaining analysis duration is equal to the duration used for video analysis minus the total duration spent on performing the overview analysis.

[0314] In some embodiments, the analyzable number W of the image frames of the target areas of all videos can be determined according to the ratio of the remaining analysis duration to the single-frame analysis duration. Further, the analyzable number is allocated to each video according to the duration of the target area of each video to obtain the second number of image frames of the target area of each video.

[0315] In a feasible method, the method for determining the second number of the target area can refer to the method for determining the first number provided by S233.

[0316] In some other embodiments, the second number of the image frames of each target area can also be determined according to the ratio of the number of image frames included in each target area to the number of image frames of all target areas. For example, according to Determine the second number of each target area, where i iis the ratio of the number of image frames included in the i-th target region to the number of image frames of all target regions, and I is the number of image frames of all target regions. is the floor symbol.

[0317] Exemplarily, assume there are 3 videos, and all of each video transmits 10 image frames per second (10 FPS). The durations of the target regions of the 3 videos are 5s, 10s, and 15s respectively. Assume the single-frame analysis duration is 200 ms and the remaining analysis duration is 25s.

[0318] During the remaining analysis duration, the analyzable quantity W is: 25s / 200ms = 125 (frames). The numbers of image frames included in the target regions of the 3 videos are 50, 100, and 150 respectively; according to the ratio of 1:2:3, 125 frames are allocated to the 3 videos. When allocating 20 image frames to Video 1, 40 image frames to Video 2, and 60 image frames to Video 3, there are still 5 image frames left. Image frames can be preferentially allocated to the videos with longer durations according to the actual durations of the videos. For example, allocate 3 image frames to Video 3 and 2 image frames to Video 2.

[0319] Alternatively, in some embodiments, the remaining image frames can also be allocated according to the scores of the target regions. Among them, the score of the target region is the average value of the first scores of the first image frames. For example, there are still 5 image frames left. According to the score situations of the target regions, the scores corresponding to the target regions of the 3 videos are 80, 90, and 70 respectively. Preferentially increase the second quantity of the image frames of the video with a higher score. Therefore, 3 image frames can be allocated to Video 2, and 1 image frame can be allocated to each of Video 1 and Video 3. Or, 2 image frames can be allocated to each of Video 2 and Video 1, and 1 image frame can be allocated to Video 3. Until 125 image frames are allocated, the second quantity of the image frames of the target region of each video is obtained.

[0320] After the policy monitoring module in the media middle platform framework layer determines the second quantity of the image frames of the target region of each video, taking Video 1 as an example, the process of performing frame-by-frame analysis on the target region of Video 1 may include:

[0321] S32, the policy monitoring module sends the file descriptor fd1 of Video 1 and the position of the third image frame of the target region of Video 1 to the channel interface of the media middle platform framework layer.

[0322] Among them, the position of the third image frame can be the position of any image frame in the target region of Video 1. The second quantity of image frames is evenly distributed at each position of the target region.

[0323] S33. The channel interface of the media middleware framework layer processes operations such as decoding, downscaling the resolution, and converting the format of Video 1 based on the file descriptor of Video 1, and stores the processed Video 1.

[0324] S34. The channel interface of the media middleware framework layer sends the frame data address and the position of the third image frame of Video 1 to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one-click video editing interface).

[0325] S35. The highlight segment algorithm interface of the HAL layer obtains the third image frame of Video 1 according to the frame data address and the position of the third image frame of Video 1, and then analyzes the image frame at the position of the third image frame based on a preset highlight segment algorithm to obtain the analysis result of the image frame at the position of the third image frame.

[0326] S36. The highlight segment algorithm interface of the HAL layer returns the analysis result of the image frame at the position of the third image frame to the policy monitoring module of the media middleware framework layer through the service interface of the FWK layer and the channel interface of the media middleware framework layer in sequence.

[0327] Among them, the analysis result includes the second score of the image frame.

[0328] The policy monitoring module of the media middleware framework layer continues to send the position of the fourth image frame in the target area of Video 1 to the channel interface of the media middleware framework layer for analysis until all the second number of image frames in the target area are analyzed.

[0329] Among them, the position of the fourth image frame can be the position of any one image frame in the target area of Video 1.

[0330] S37. The policy monitoring module of the media middleware framework layer determines the highlight segment of Video 1 based on the analysis result of the image frames in the target area of Video 1.

[0331] In this embodiment, the policy monitoring module obtains the second image frame with the highest second score according to the second score of the image frame. Among them, the second image frame can include at least one image frame. If the second image frame includes one image frame, the image frames within the second preset duration before and after the second image frame can be directly determined as the highlight segment of Video 1.

[0332] For example, Figure 16 (a) of Figure 16 gives a schematic diagram of a highlight segment. Among them, the second image frame of Video 1 includes Image Frame 1 (such as

[0333] Alternatively, if the second image frame includes one image frame, according to the recommended duration of the highlight segment, the segments before and after the second image frame with a duration of 1 / 2 of the recommended duration of the highlight segment can be determined as the highlight segments of Video 1. The duration of the highlight segment is equal to the recommended duration of the highlight segment.

[0334] For example, as Figure 16 (b) of gives another schematic diagram of the highlight segment. Among them, the second image frame of Video 1 includes Image Frame 1 (such as 1 in (b) of Figure 16 ), and the segments before and after Image Frame 1 with a duration of 1 / 2 of the recommended duration of the highlight segment are the highlight segments of Video 1.

[0335] If the second image frame includes multiple consecutive image frames. According to the recommended duration of the highlight segment, the segment covered by the fifth preset duration before the first image frame of the multiple consecutive second image frames to the fifth preset duration after the last image frame of the multiple consecutive second image frames is determined as the highlight segment of Video 1.

[0336] For example, as Figure 17 , Figure 17 gives another schematic diagram of the highlight segment. The first image frame of Video 1 includes consecutive Image Frame 1 (such as 1 in Figure 17 ), Image Frame 2 (such as 2 in Figure 17 ), and Image Frame 3 (such as 3 in Figure 17 ). Then, the segment covered by the fifth preset duration before Image Frame 1 to the fifth preset duration after Image Frame 3 is the highlight segment of this Video 1.

[0337] If there are other videos, such as Video 2, then the electronic device can continue to execute the above S32 - S37 to obtain the analysis results of the highlight segments of other videos.

[0338] After obtaining the analysis results of the highlight segments of all videos, in some embodiments, referring to Figure 18 , gives a schematic flowchart of the post - processing of the highlight segment in a video processing method. After the electronic device executes S37, it can report all the picture and video analysis results by using the following S38.

[0339] S38, the policy monitoring module of the media middle - platform framework layer reports all the picture and video analysis results to the application function layer through the image highlight segment analysis interface of the media middle - platform framework layer.

[0340] Among them, the analysis results can include the positions of all highlight segments. For example, the start time and end time of the highlight segments of each video.

[0341] S39. The application function layer clips and filters the user-selected materials based on the analysis results of all pictures and videos to obtain all highlight segments.

[0342] S40. The application function layer of the application layer calls the theme summary interface of the media middle platform framework to request and obtain a theme template.

[0343] Among them, this request passes through the theme summary interface and the channel interface of the media middle platform framework and is transmitted to the HAL layer through the FKW layer.

[0344] S41. The HAL layer determines a theme template that matches the scene according to the scenes of the pictures in the highlight segments.

[0345] In some embodiments, the electronic device may be provided with multiple theme templates (style templates). The theme algorithm can recommend a theme template that matches the picture scene based on the picture scenes of the highlight segments, such as people, scenery, food, children, pets, sports, or travel.

[0346] S42. The HAL layer returns the determined theme template to the theme summary interface of the media middle platform framework layer, and the theme summary interface returns the theme template to the application function layer of the application layer.

[0347] Exemplarily, assuming that most of the pictures in the highlight segments are parent-child scenes, then it can be confirmed that the theme template that matches this scene is of the parent-child theme category.

[0348] S43. The application function layer distributes the obtained theme and all highlight segments to the basic capability layer.

[0349] The application function layer distributes the theme obtained from S42 and all the highlight segments obtained from S39 to the basic capability layer.

[0350] S44. The basic capability layer generates a target video set based on the theme and all highlight segments.

[0351] That is, the target video set is a video set generated according to all the filtered highlight segments and conforming to the recommended theme.

[0352] S45. The basic capability layer sends an indication message for displaying the target video set to the video editing service layer.

[0353] S46. The video editing service layer displays the video set in the gallery interface.

[0354] In the embodiments of the present application, the media middle platform framework layer is used for decoding video and picture files, converting the data format into a unified format, monitoring the remaining time and adjusting the operation strategy, sending data, controlling the running and termination of algorithms, obtaining results and returning them to the application layer, etc. The FKW layer is used to complete data packaging and provide data and program running services. After receiving the commands sent by the media middle platform framework layer, the HAL layer performs highlight analysis according to the commands and returns the parameter calculation results of the highlight analysis to the media middle platform framework layer. The final results of the algorithms are collected and sorted by the media middle platform framework layer and then sent to the application layer for processing. The application layer can present the editing application interface, video and picture file options, and present the final results of the algorithms.

[0355] After the user activates the one - click video compilation function, select the video and picture files to be edited (for example, up to 30 files are supported). After waiting for a moment, the "one - click video compilation" application automatically edits the highlight segments of the video and combines the highlight segments and pictures together according to the algorithm results to generate a compiled video set and preview and play it.

[0356] For the video processing method provided by the embodiments of the present application, the electronic device can first perform overview analysis on the first number of image frames in each video among the multiple image materials selected by the user to obtain the first score of the image frames in each video. Then, the electronic device can determine the target area of the video based on the first image frame with the highest score in a video, and the image frames within the first preset duration before and after the first image frame. After that, the electronic device can analyze the second number of images in the target area of the video to obtain the second score of the image frames. Based on the second image frame with the highest second score, and the image frames within the second preset duration before and after the second image frame, the highlight segment of the video is determined. By adopting this solution, the electronic device first performs overview analysis on the videos therein, so that the target area that needs further processing can be located. When the electronic device extracts the highlight segment, it only performs analysis on the second number of image frames for the target area, rather than performing frame - by - frame analysis on the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumption of the electronic device to implement the one - click video compilation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one - click video compilation function.

[0357] It should also be noted that, in the embodiments of the present application, "greater than" can be replaced by "greater than or equal to", "less than or equal to" can be replaced by "less than", or, "greater than or equal to" can be replaced by "greater than", "less than" can be replaced by "less than or equal to".

[0358] Each of the embodiments described herein can be an independent solution or can be combined according to the internal logic, and these solutions all fall within the protection scope of the present application.

[0359] It can be understood that the methods and operations implemented by the electronic device in the above method embodiments can also be implemented by components (such as chips or circuits) available for the electronic device.

[0360] It should be noted that the personal information used in the technical solution of this application is limited to the information for which individual consent has been obtained, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) before the user uses this function, and signing the agreement (authorization) that authorizes the relevant user information. Among them, personal information includes information such as pictures and videos stored by the user.

[0361] In the technical solution disclosed in this application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0362] The above describes the method embodiments provided by this application. The following will describe the device embodiments provided by this application. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments. Therefore, for the content not described in detail, reference can be made to the above method embodiments. For the sake of brevity, it will not be repeated here.

[0363] The above mainly describes the solution provided by the embodiments of this application from the perspective of method steps. It can be understood that in order to implement the above functions, the electronic device implementing this method includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should be able to realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the protection scope of this application.

[0364] The embodiments of this application can divide the electronic device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other feasible division methods in actual implementation. The following takes the example of dividing each functional module corresponding to each function for illustration.

[0365] The present application also provides a chip, which is coupled to a memory and is used to read and execute a computer program or instructions stored in the memory to execute the methods in the above embodiments.

[0366] The present application also provides an electronic device, which includes a chip used to read and execute a computer program or instructions stored in a memory, so that the methods in the embodiments are executed.

[0367] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above electronic device, the electronic device is caused to execute each function or step that the electronic device executes in the above method embodiments.

[0368] An embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute each function or step that the electronic device executes in the above method embodiments. For example, the computer may be the above electronic device.

[0369] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0370] In the several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the module or unit is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0371] The unit described as a separate component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0372] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a readable storage medium. Based on such understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0373] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A video processing method, characterized in that, The method includes: Receiving a selection operation by a user for multiple image materials in a picture gallery; wherein, the multiple image materials include videos; In response to the selection operation, performing an overview analysis on each video among the multiple image materials for a first number of image frames to obtain a first score; wherein, the first score is an aesthetic score for the corresponding image frames; the first number corresponds to the duration of the video; Based on the first scores of the image frames in each video, determining a target area for the corresponding video; wherein, the target area includes a first image frame, and the image frames within a first preset duration before and after the first image frame; the first image frame is the image frame with the highest first score; Performing an analysis on the target area of each video among the multiple image materials for a second number of image frames to obtain a second score; wherein, the second score is an aesthetic score for the corresponding image frames; the second number is less than or equal to the number of image frames in the target area; Based on the second scores of the image frames in the target area of each video, determining a highlight segment for the corresponding video; wherein, the highlight segment includes a second image frame, and the image frames within a second preset duration before and after the second image frame; the second image frame is the image frame with the highest second score in the target area; the highlight segments of each video among the multiple image materials are used for splicing to obtain a target video set.

2. The method according to claim 1, wherein The method further includes: In response to the selection operation, obtaining analysis parameters corresponding to the multiple image materials; wherein, the analysis parameters include the actual duration of the corresponding video, and a recommended upper limit value of the total analysis duration, and the recommended upper limit value of the total analysis duration represents a recommended value of the maximum duration required for analyzing the multiple image materials; Determining the first number of image frames for each video according to the actual duration of each video among the multiple image materials, the number of videos among the multiple image materials, and the recommended upper limit value of the total analysis duration.

3. The method according to claim 2, wherein The determining the first number of image frames for each video according to the actual duration of each video among the multiple image materials, the number of videos among the multiple image materials, and the recommended upper limit value of the total analysis duration includes: Based on the actual duration of each video and a preset corresponding relationship, determining a basic number and a maximum number for the corresponding video; wherein, the basic number is the minimum number of image frames required to ensure the analysis effect of the video, and the maximum number is the maximum number of image frames that can be analyzed for the video within the duration; the preset corresponding relationship represents the maximum number and the basic number of image frames corresponding to different threshold ranges of the video duration; Based on the basic number and the maximum number of each video, and the recommended upper limit value of the total analysis duration, determining a total analysis number; the total analysis number is the total number of image frames allowed to be analyzed for all videos among the multiple image materials; Allocating the total analysis number to each video according to the actual duration of each video among the multiple image materials, and the number of videos among the multiple image materials, to obtain the first number of image frames for each video.

4. The method according to claim 3, wherein The analysis parameter further includes the single-frame analysis duration; the single-frame analysis duration is the duration required to analyze an image frame. Determining the total analysis quantity based on the base quantity and the maximum quantity of each video, and the recommended value of the upper limit of the total analysis duration includes: If the recommended value of the upper limit of the total analysis duration is less than the time consumed for analyzing the sum of the base quantity of image frames of all videos in the plurality of image materials, the total analysis quantity is the ratio of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration. If the recommended value of the upper limit of the total analysis duration is greater than the time consumed for analyzing the sum of the base quantity of image frames of all videos in the image materials, and the duration of the first multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumed for analyzing the sum of the base quantity of image frames of all videos in the plurality of image materials, the total analysis quantity is the sum of the base quantity of image frames of all videos in the plurality of image materials. If the duration of the first multiple of the recommended value of the upper limit of the total analysis duration is greater than the time consumed for analyzing the sum of the base quantity of image frames of all videos in the plurality of image materials, and the duration of the first multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumed for analyzing the sum of the maximum quantity of image frames of all videos in the plurality of image materials, the total analysis quantity is the ratio of the duration of the first multiple of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration. If the duration of the first multiple of the recommended value of the upper limit of the total analysis duration is greater than the time consumed for analyzing the sum of the maximum quantity of image frames of all videos in the plurality of image materials, the total analysis quantity is the sum of the maximum quantity of image frames of all videos in the plurality of image materials. Wherein, the first multiple is greater than 0 and less than 1.

5. The method according to claim 3 or 4, characterized in that, Allocating the total analysis quantity to each video according to the actual duration of each video in the plurality of image materials and the number of videos in the plurality of image materials to obtain the first quantity of image frames of each video includes: Traversing each video in the plurality of image materials, updating the first value and the second value of each video until the first value is 0; wherein, the initial value of the first value is equal to the total analysis quantity, and the initial value of the second value is 0; for each video traversed, the second value of the video is incremented by 1, and the first value is decremented by 1. Taking the second value of each video as the first quantity of image frames of the video.

6. The method according to claim 5, wherein The traversing each video in the plurality of image materials, updating the first value and the second value of each video until the first value is 0 includes: Before traversing to the first video in the plurality of image materials, if the second value of the first video is equal to the maximum quantity of image frames of the first video, skip the first video and traverse the next video of the first video. Wherein, skipping the first video means that the second value of the first video is not incremented by 1.

7. The method according to any one of claims 1-6, characterized in that The first quantity of image frames for overview analysis is evenly distributed at various positions of the video.

8. The method according to any one of claims 2-7, characterized in that, The analysis parameter includes the recommended duration of a highlight segment, and the recommended duration of a highlight segment is the recommended duration of a highlight segment in a video. Determining a target region corresponding to a video based on the first score of the image frames in each video includes: For each of the videos, obtaining the first image frame according to the first score of the image frames in the video; Determining a candidate highlight segment of the video for a segment of the video that includes the first image frame and has a duration of a highlight segment suggestion duration; Determining the target region of the video as the candidate highlight segment and the image frames within a third preset duration before and after the candidate highlight segment; the third preset duration is less than the first preset duration.

9. The method according to claim 8, characterized in that The analysis parameter includes a total highlight segment suggestion duration, and the total highlight segment suggestion duration is the sum of the suggestion durations of the highlight segments of all videos in the plurality of image materials; The method further includes: Calculating the sum of the durations of the candidate highlight segments of all videos in the plurality of image materials; If the sum of the durations of the candidate highlight segments of all videos is greater than a second multiple of the total highlight segment suggestion duration, adjusting the durations of the candidate highlight segments of each video according to the actual duration of each video in the plurality of image materials, such that the sum of the durations of the candidate highlight segments of all videos is less than or equal to the second multiple of the total highlight segment suggestion duration; wherein the second multiple is greater than 1 and less than 2.

10. The method according to any one of claims 2-9, characterized in that, The analysis parameters include a remaining analysis duration and a per-frame analysis duration; the remaining analysis duration is equal to the duration for video analysis minus the total duration spent on performing the overview analysis; the per-frame analysis duration is the duration required to analyze one image frame; Before analyzing a second number of image frames in the target region of each video in the plurality of image materials to obtain a second score, the method includes: Taking the ratio of the remaining analysis duration to the per-frame analysis duration as the analyzable number of image frames in the target regions of all videos in the plurality of image materials; Allocating the analyzable number to each video according to the duration of the target region of each video in the plurality of image materials to obtain the second number of image frames in the target region of each video.

11. The method according to claim 10, wherein The allocating the analyzable number to each video according to the duration of the target region of each video in the plurality of image materials to obtain the second number of image frames in the target region of each video includes: Traversing each video in the plurality of image materials and updating a third value and a fourth value for each video until the third value is 0; wherein the initial value of the third value is equal to the analyzable number, and the initial value of the fourth value is 0; for each video traversed, the fourth value of the video is incremented by 1 and the third value is decremented by 1; Taking the fourth value of each video as the second number of image frames in the target region of the video.

12. The method according to any one of claims 1-11, characterized in that, The second number of image frames for analysis is evenly distributed at various positions in the target region of the video.

13. The method according to any one of claims 2-12, characterized in that, The analysis parameter includes a highlight segment suggestion duration, and the highlight segment suggestion duration is the suggestion duration of a highlight segment in a video; Determining a highlight segment corresponding to a video based on the second score of the image frames in the target region of each video includes: For each of the videos, obtain the second image frames according to the second scoring of the image frames in the video. Determine the highlight segments of the video as the second image frames and the image frames within the second preset duration before and after the second image frames; the duration of the highlight segments is less than or equal to the recommended duration of the highlight segments.

14. The method according to any one of claims 2-13, characterized in that, The analysis parameters include the total analysis duration and the recommended upper limit value of the total analysis duration. The total analysis duration represents the total duration required to analyze the multiple image materials, and the recommended upper limit value of the total analysis duration represents the recommended value of the maximum duration required to analyze the multiple image materials. Before obtaining the first score by performing an overview analysis of the first number of image frames for each video in the multiple image materials, the method further includes: If the total analysis duration is greater than the recommended upper limit value of the total analysis duration, randomly select M videos from the multiple image materials; where the duration required to analyze the M videos is less than or equal to the total analysis duration, and M is less than the number of videos in the image materials.

15. The method according to any one of claims 1 to 14, characterized in that, The obtaining of the first score by performing an overview analysis of the first number of image frames for each video in the multiple image materials includes: If the sum of the actual durations of all the videos in the multiple image materials is greater than the preset duration threshold, perform an overview analysis of the first number of image frames for each video in the multiple image materials to obtain the first score.

16. An electronic device, characterized in that, The electronic device includes a display screen, a memory, and one or more processors; the display screen and the memory are coupled to the processor; the memory stores computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1-15.

17. A computer-readable storage medium, characterized in that, It includes computer instructions that, when running on an electronic device, cause the electronic device to execute the method according to any one of claims 1-15.

18. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-15 is implemented.

Citation Information

Patent Citations

  • Video editing method and device, electronic equipment and computer readable storage medium

    CN110996169A

  • Target image generation method, target image generation device, medium and electronic equipment

    CN111464833A

  • Video segment extraction method, video segment extraction device and storage medium

    CN112069952A

  • Video processing method and device, electronic equipment and readable storage medium

    CN113286194A

  • Video cover picture selection method and device, equipment and storage medium

    CN116309236A