Video processing method, electronic device, and storage medium

By performing overview analysis on image materials and extracting highlight fragments from target areas, the problem of low efficiency in analyzing long materials on electronic devices is solved, improving the processing efficiency and user experience of the one-click image creation function.

CN120343181BActive Publication Date: 2026-04-10HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-01-10
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

When the user selects a long video clip, the electronic device takes a long time to analyze multiple image clips, resulting in low video processing efficiency and affecting the user experience of the one-click video creation function.

Method used

Electronic devices identify target areas by performing overview analysis on each video in the image material, and then perform a second number of image frame analyses only on the target areas to extract highlight segments, thus reducing the workload of image frame analysis.

Benefits of technology

It improves the efficiency of electronic devices in processing video, reduces the time required to achieve one-click video creation, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343181B_ABST
    Figure CN120343181B_ABST
Patent Text Reader

Abstract

The application discloses a video processing method, an electronic device and a storage medium, and relates to the technical field of video data, and comprises the following steps: an electronic device performs overview analysis on a first number of image frames of each video in a plurality of image materials, and obtains a first score of the image frames. Based on the first score of the image frames in each video, a target region of the corresponding video is determined; and the target region of each video is analyzed for a second number of image frames to obtain a second score of the image frames. Based on the second score of the image frames in the target region of each video, a highlight segment of the corresponding video is determined. In the scheme, the electronic device performs image frame analysis on a first number and a second number of image frames of each video, instead of frame-by-frame analysis on the entire video, thereby reducing the workload of the electronic device in image frame analysis, reducing the time consumption of the electronic device in realizing a one-key film-making function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-key film-making function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of video data, and in particular, to a video processing method, an electronic device, and a storage medium. BACKGROUND

[0002] With the development of picture and video processing technology, a user can trigger an electronic device to further process photos and videos in a photo album. For example, the electronic device can perform stitching processing on multiple image materials (including multiple pictures and / or multiple videos) to obtain a complete stitched video.

[0003] For example, the electronic device or third-party video processing software in the electronic device can have a one-click video generation function (or one-click video generation service). The one-click video generation function can automatically analyze and extract highlight clips of videos in multiple image materials selected by a user through an algorithm, and then automatically generate a video clip based on the extracted highlight clips. The highlight clips, also referred to as highlight moments, are single-frame images or video clips composed of continuous multiple frames extracted from the above-mentioned materials to record highlight moments. The highlight moments can be moments when a character smiles, wins a championship, or an airplane lands.

[0004] However, in the case where the overall duration of the materials selected by the user is long, the electronic device will consume a large amount of time to analyze the multiple materials selected by the user, and the efficiency of the electronic device in processing videos is low. SUMMARY

[0005] Embodiments of the present application provide a video processing method, an electronic device, and a storage medium, which selectively analyze image frames in each video in image materials, avoid the problem of long time consumption and low video processing performance caused by analyzing complete videos, reduce the workload of the electronic device in analyzing image frames, thereby reducing the time consumption of the electronic device in implementing the one-click video generation function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-click video generation function.

[0006] To achieve the above object, embodiments of the present application adopt the following technical solutions.

[0007] In a first aspect, a video processing method is provided, and the method comprises:

[0008] The electronic device receives a selection operation of a user on multiple image materials in a gallery. The multiple image materials can include videos, and the multiple image materials can also include videos and pictures.

[0009] The electronic device performs overview analysis on each video in the plurality of image materials in a first number of image frames in response to the selection operation, to obtain a first score. The first score is an aesthetic score of the corresponding image frame. The first number corresponds to a time length of the video.

[0010] The electronic device determines a target region of each video based on the first score of the image frame in the video, wherein the target region includes a first image frame with the highest overview analysis score, and image frames within a first preset time length before and after the first image frame. The electronic device performs analysis on the target region of each video in the plurality of image materials in a second number of image frames to obtain a second score. The second score is an aesthetic score of the corresponding image frame. The second number is less than or equal to the number of image frames in the target region. The electronic device determines a highlight segment of each video based on the second score of the image frame in the target region of the video. The highlight segment includes a second image frame with the highest second score, and image frames within a second preset time length before and after the second image frame. The highlight segments of the videos in the plurality of image materials are used to splice a target video set.

[0011] In this application, the electronic device first performs overview analysis on the videos therein, which can locate the target region that needs further processing. When the electronic device extracts the highlight segment, it only performs analysis on the target region in a second number of image frames, instead of frame-by-frame analysis on the entire video, which can reduce the workload of the electronic device in image frame analysis, thereby reducing the time consumption of the electronic device in implementing the one-key film function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-key film function.

[0012] In a possible implementation manner of the first aspect, the method further includes:

[0013] The electronic device obtains analysis parameters corresponding to the plurality of image materials in response to the selection operation.

[0014] The analysis parameters include an actual time length of the corresponding video, and an analysis total time length upper limit suggestion value, which represents a suggestion value of the maximum time length required for analyzing the plurality of image materials.

[0015] The electronic device determines the first number of image frames of each video according to the actual time length of each video in the plurality of image materials, the number of videos in the plurality of image materials, and the analysis total time length upper limit suggestion value.

[0016] In this application, the electronic device can determine the first number of image frames of each video within the analysis total time length upper limit suggestion value. Under the limited performance of the electronic device, by analyzing the first number of image frames, the analysis efficiency of the highlight segment can be improved, and a relatively reliable analysis effect can also be obtained.

[0017] In a possible implementation manner of the first aspect, the first quantity of image frames of each video is determined according to the actual time length of each video in the plurality of image materials, the number of videos in the plurality of image materials, and the upper limit recommended value of the total analysis time length, and includes the following steps.

[0018] The electronic device determines the basic quantity and the maximum quantity of the corresponding video based on the actual time length of each video and a preset corresponding relationship.

[0019] The basic quantity is the minimum quantity of image frames required to ensure the analysis effect of the video, and the maximum quantity is the maximum quantity of image frames allowed to analyze the video in the time length. The preset corresponding relationship represents the maximum quantity and the basic quantity of image frames corresponding to different threshold ranges of video time length.

[0020] The electronic device determines the total analysis quantity based on the basic quantity and the maximum quantity of each video and the upper limit recommended value of the total analysis time length. The total analysis quantity is the total quantity of image frames allowed to analyze all videos in the plurality of image materials. The total analysis quantity is distributed to each video according to the actual time length of each video in the plurality of image materials and the number of videos in the plurality of image materials, to obtain the first quantity of image frames of each video.

[0021] In this application, the first quantity is determined by the maximum quantity and the basic quantity of image frames of each video, and the first quantity of image frames is analyzed, which can not only ensure the analysis effect of the video, but also efficiently analyze and process the highlight segment under the limited performance and limited time consumption of the electronic device.

[0022] In a possible implementation manner of the first aspect, the analysis parameter further includes a single-frame analysis time length; the single-frame analysis time length is the time length required for analyzing one image frame.

[0023] The electronic device determines the total analysis quantity based on the basic quantity and the maximum quantity of each video and the upper limit recommended value of the total analysis time length, including:

[0024] If the upper limit recommended value of the total analysis time length is less than the time consumption of analyzing the sum of the basic quantity image frames of all videos in the plurality of image materials, the total analysis quantity is the ratio of the upper limit recommended value of the total analysis time length to the single-frame analysis time length.

[0025] If the upper limit recommended value of the total analysis time length is greater than the time consumption of analyzing the sum of the basic quantity image frames of all videos in the image material, and the time length of the first multiple of the upper limit recommended value of the total analysis time length is less than the time consumption of analyzing the sum of the basic quantity image frames of all videos in the plurality of image materials, the total analysis quantity is the sum of the basic quantity image frames of all videos in the plurality of image materials.

[0026] If the time length of the first multiple of the total length upper limit suggestion value is greater than the time length of the sum of the image frames of the base number of all videos in the plurality of image materials, and the time length of the first multiple of the total length upper limit suggestion value is less than the time length of the sum of the image frames of the maximum number of all videos in the plurality of image materials, the total number is the ratio of the time length of the first multiple of the total length upper limit suggestion value to the single-frame analysis time length.

[0027] If the time length of the first multiple of the total length upper limit suggestion value is greater than the time length of the sum of the image frames of the base number of all videos in the plurality of image materials, and the time length of the first multiple of the total length upper limit suggestion value is less than the time length of the sum of the image frames of the maximum number of all videos in the plurality of image materials, the total number is the ratio of the time length of the first multiple of the total length upper limit suggestion value to the single-frame analysis time length.

[0028] The first multiple is greater than 0 and less than 1.

[0029] In the present application, the first number is determined by the maximum number and the base number of the image frames of each video, and the image frames of the first number are analyzed, which can not only ensure the analysis effect of the video, but also can efficiently analyze and process the highlight fragments under the limited performance and limited time consumption of the electronic device.

[0030] In a possible implementation manner of the first aspect, the total number of analyses is allocated to each video according to the actual time length of each video in the plurality of image materials and the number of videos in the plurality of image materials, and the first number of image frames of each video is obtained, including:

[0031] The electronic device traverses each video in the plurality of image materials, and updates the first value and the second value of each video until the first value is 0.

[0032] The initial value of the first value is equal to the total number of analyses, and the initial value of the second value is 0; each time a video is traversed, the second value of the video is increased by 1, and the first value is decreased by 1.

[0033] The electronic device takes the second value of each video as the first number of image frames of the video.

[0034] In the present application, each video is traversed to determine the first number of each video, and the total number of analyses can be friendly allocated to each video, so that the highlight fragments of any video are not missed.

[0035] In a possible implementation manner of the first aspect, the first value and the second value of each video in the plurality of image materials are updated until the first value is 0, including:

[0036] Before traversing the first video in the plurality of image materials, if the second value of the first video is equal to the maximum number of image frames of the first video, the first video is skipped, and the next video of the first video is traversed.

[0037] wherein, skipping the first video means that the second value of the first video is not added 1.

[0038] In the present application, if the number of image frames of a video reaches the maximum number, the video will not be allocated image frames in the future, and image frames will be allocated to other videos that have not reached the maximum number, which can improve the effectiveness of video analysis.

[0039] In a possible implementation of the first aspect, the first number of image frames for overview analysis are evenly distributed in various positions of the video.

[0040] In the present application, the first number of image frames are evenly distributed in various positions of the video, which can avoid missing image frames in some positions of the video.

[0041] In a possible implementation of the first aspect, the analysis parameter includes a highlight segment suggested duration, and the highlight segment suggested duration is a suggested duration of a highlight segment in one video.

[0042] Based on the first scores of the image frames in each video, a target region of the corresponding video is determined, including:

[0043] For each video, a first image frame is obtained according to the first scores of the image frames in the video. The electronic device determines a candidate highlight segment of the video, which includes a segment of the video with a highlight segment suggested duration of the first image frame. The candidate highlight segment and image frames within a third preset duration before and after the candidate highlight segment are determined as a target region of the video. The third preset duration is less than the first preset duration.

[0044] In the present application, the electronic device can determine a candidate highlight segment that meets the highlight segment suggested duration according to the first image frame with the highest first score and the highlight segment suggested duration, and determine a target region that includes the candidate highlight segment. The target region is worth further analysis because it includes the first image frame, so as to determine the highlight segment of the video. In this way, the determined highlight segment is more accurate.

[0045] In a possible implementation of the first aspect, the analysis parameter includes a total highlight segment suggested duration, and the total highlight segment suggested duration is a sum of suggested durations of highlight segments of all videos in a plurality of image materials.

[0046] The method further includes:

[0047] The sum of the lengths of the candidate highlight segments of all videos in the plurality of image materials is calculated; if the sum of the lengths of the candidate highlight segments of all videos is greater than a second multiple of the total recommended length of highlight segments, the length of the candidate highlight segment of each video is adjusted according to the actual length of each video in the plurality of image materials, so that the sum of the lengths of the candidate highlight segments of all videos is less than or equal to the second multiple of the total recommended length of highlight segments; wherein the second multiple is greater than 1 and less than 2.

[0048] In the present application, when the sum of the lengths of the candidate highlight segments is greater than the second multiple of the total recommended length of highlight segments, it indicates that the length of the candidate highlight segment is relatively long, and the length of the candidate highlight segment can be reduced to obtain a length allowed under limited performance and effective time consumption, so that effective highlight segment analysis can be realized.

[0049] In a possible implementation of the first aspect, the analysis parameters include a remaining analysis length and a single-frame analysis length; the remaining analysis length is equal to the length used for video analysis minus the total length spent on performing overview analysis; and the single-frame analysis length is the length required for analyzing one image frame.

[0050] Before the analysis of the target region of each video in the plurality of image materials is performed for the second number of image frames to obtain the second score, the method comprises:

[0051] The ratio of the remaining analysis length to the single-frame analysis length is taken as the analyzable number of image frames of the target region of all videos in the plurality of image materials; and the analyzable number is allocated to each video according to the length of the target region of each video in the plurality of image materials to obtain the second number of image frames of the target region of each video.

[0052] In the present application, the second number of image frame analysis can also be performed on the target region, rather than frame-by-frame analysis, which can further reduce the time consumption of analyzing image frames and further improve the analysis efficiency of highlight segments.

[0053] In a possible implementation of the first aspect, the allocation of the analyzable number to each video according to the length of the target region of each video in the plurality of image materials to obtain the second number of image frames of the target region of each video comprises:

[0054] Each video in the plurality of image materials is traversed, and a third value and a fourth value of each video are updated until the third value is 0; wherein the initial value of the third value is equal to the analyzable number, and the initial value of the fourth value is 0; each time a video is traversed, the fourth value of the video is incremented by 1, and the third value is decremented by 1; and the fourth value of each video is taken as the second number of image frames of the target region of the video.

[0055] In the present application, each video is traversed to determine the second number of each video, and the total number of analyzable is evenly distributed to each video, so that no high light segment of any video is missed.

[0056] In a possible implementation of the first aspect, the second number of image frames analyzed are evenly distributed at various positions of the target region of the video.

[0057] In the present application, the second number of image frames are evenly distributed at various positions of the target region of the video, which can avoid missing image frames at some positions in the video.

[0058] In a possible implementation of the first aspect, the analysis parameter includes a high light segment suggestion duration, and the high light segment suggestion duration is a suggested duration of a high light segment in a video.

[0059] Based on the second scores of the image frames in the target region in each video, a high light segment of the corresponding video is determined, including:

[0060] For each video, a second image frame is obtained according to the second scores of the image frames in the video. The second image frame and the image frames within a second preset time duration before and after the second image frame are determined as a high light segment of the video; and the duration of the high light segment is less than or equal to the high light segment suggestion duration.

[0061] In the present application, the electronic device can determine a high light segment conforming to the high light segment suggestion duration according to the second image frame with the highest second score and the high light segment suggestion duration, and the high light segment determined based on the second image frame and the high light segment suggestion duration is more accurate.

[0062] In a possible implementation of the first aspect, the analysis parameter includes an analysis total duration and an analysis total duration upper limit suggestion value, the analysis total duration represents a total duration required for analyzing the plurality of image materials, and the analysis total duration upper limit suggestion value represents a suggested value of a maximum duration required for analyzing the plurality of image materials.

[0063] Before the overview analysis of the first number of image frames of each video in the plurality of image materials is performed to obtain the first scores, the method further includes:

[0064] If the analysis total duration is greater than the analysis total duration upper limit suggestion value, M videos are randomly selected from the plurality of image materials; wherein a duration required for analyzing the M videos is less than or equal to the analysis total duration, and M is less than the number of videos in the image materials.

[0065] In the present application, when the analysis total duration of the plurality of image materials is greater than the analysis total duration upper limit suggestion value, a part of the videos in the plurality of image materials can be randomly selected for analysis, and under the limited performance and limited time consumption of the electronic device, efficient analysis and processing of the high light segment can be realized.

[0066] In a possible implementation manner of the first aspect, the overview analysis of each video in the plurality of image materials is performed on the first number of image frames to obtain the first score, including:

[0067] If the sum of the actual time lengths of all videos in the plurality of image materials is greater than the preset time length threshold, the overview analysis of each video in the plurality of image materials is performed on the first number of image frames to obtain the first score.

[0068] In the present application, when the sum of the actual time lengths of the videos of the plurality of image materials is greater than the preset time length threshold, the present solution can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumption of the electronic device for implementing the one-key film making function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-key film making function.

[0069] In a second aspect, an electronic device is provided, which includes a memory, a display screen and one or more processors; the memory, the display screen and the processor are coupled; the memory stores computer program code, and the computer program code includes computer instructions, which, when executed by the processor, causes the electronic device to perform the method in any one of the first aspect.

[0070] In a third aspect, a computer readable storage medium is provided, which stores instructions, and when the instructions are run on an electronic device, the electronic device can perform the method in any one of the first aspect.

[0071] In a fourth aspect, a computer program product is provided, which includes instructions, and when the instructions are run on an electronic device, the electronic device can perform the method in any one of the first aspect.

[0072] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor, and the processor is configured to call a computer program in a memory to execute the method in the first aspect.

[0073] It can be understood that the electronic device of the second aspect, the computer readable storage medium of the third aspect, the computer program product of the fourth aspect, and the chip of the fifth aspect can achieve the beneficial effects of the first aspect and any possible design manner thereof, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS

[0074] Figure 1 An application scenario schematic diagram of a video processing method provided by an embodiment of the present application is shown in the following figure;

[0075] Figure 2Another application scenario diagram of the video processing method provided by the embodiment of the present application is shown in FIG. 6;

[0076] Figure 3 Another application scenario diagram of the video processing method provided by the embodiment of the present application is shown in FIG. 6;

[0077] Figure 4 Another application scenario diagram of the video processing method provided by the embodiment of the present application is shown in FIG. 6;

[0078] Figure 5 A hardware structure diagram of an electronic device provided by the embodiment of the present application is shown in FIG. 7;

[0079] Figure 6 A software structure diagram of an electronic device provided by the embodiment of the present application is shown in FIG. 8;

[0080] Figure 7 A method flow diagram of operations such as algorithm initialization and analysis parameter acquisition of each module in an electronic device provided by the embodiment of the present application is shown in FIG. 9;

[0081] Figure 8 A flow diagram of picture analysis in a video processing method provided by the embodiment of the present application is shown in FIG. 10;

[0082] Figure 9 A technical idea diagram of a video processing method provided by the embodiment of the present application is shown in FIG. 11;

[0083] Figure 10 A flow diagram of video analysis in a video processing method provided by the embodiment of the present application is shown in FIG. 12;

[0084] Figure 11 A diagram of extracting image frames provided by the embodiment of the present application is shown in FIG. 13;

[0085] Figure 12 A diagram of extracting image frames provided by the embodiment of the present application is shown in FIG. 13;

[0086] Figure 13 A diagram of assigning image frames to multiple videos provided by the embodiment of the present application is shown in FIG. 14;

[0087] Figure 14 A diagram of a target region provided by the embodiment of the present application is shown in FIG. 15;

[0088] Figure 15 A diagram of a target region provided by the embodiment of the present application is shown in FIG. 15;

[0089] Figure 16 A diagram of a highlight segment provided by the embodiment of the present application is shown in FIG. 16;

[0090] Figure 17 A diagram of a highlight segment provided by the embodiment of the present application is shown in FIG. 16;

[0091] Figure 18 A flowchart of a post-processing process for a highlight segment in a video processing method is provided in an embodiment of the present application. DETAILED DESCRIPTION

[0092] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments and are not intended to be limiting of the present application. As used in the specification and the appended claims of the application, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in the following embodiments, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term “and / or” used in the context of the following embodiments refers to a conjunctive relationship and / or a disjunctive relationship unless otherwise stated.

[0093] Reference in the specification to “one embodiment” or “some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases “in one embodiment” or “in some embodiments” in various places in the specification are not necessarily all referring to the same embodiment, although they can. The terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” and the like are used synonymously to denote a certain feature, structure, or characteristic that is included in, or that is in the nature of, the application. The term “connected” is used synonymously with “coupled” and / or “in communication,” unless otherwise explicitly stated. The terms “first,” “second,” and the like, do not denote any order, quantity, combination, or importance, but rather are used to nomenclature the different elements of the application.

[0094] In the embodiments of the present application, the word “exemplary” or “for example” is used to mean serving as an example, instance, or illustration. Any embodiment or design described in the embodiments of the present application as “exemplary” or “for example” should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the word “exemplary” or “for example” is used to present concepts in a concrete manner.

[0095] First, some terms or phrases involved in the present application are explained.

[0096] A highlight clip, also referred to as a highlight moment, refers to a single frame image or a video clip composed of a plurality of continuous frames of images extracted from a video or an image, and used to record a highlight moment. The highlight moment can be a moment when a character smiles, a race wins, a sports jumps, a plane lands, or a ball scores.

[0097] One-click video generation refers to that, in response to a selection operation of a user on one or more image materials, an electronic device automatically analyzes highlight clips in the image materials by using an algorithm, and combines the highlight clips into a video clip. That is, the electronic device can extract a plurality of highlight clips from one or more image materials, and combine the plurality of highlight clips into a video clip. The image materials selected by the user can be images or videos, or the image materials can include images and videos.

[0098] In this case, the electronic device can obtain aesthetic score parameters of each frame of image, such as image color, image texture feature, image quality, frame interpolation with previous and next frames, and edge change rate, and score each frame of image according to the aesthetic score parameters. The electronic device can select a single frame of image or a plurality of continuous frames of images with the highest score as a highlight clip.

[0099] At present, the one-click video generation function supports editing of image materials such as images and videos. In the process of implementing the one-click video generation function, the user can select a large number of images or a video with a long time length from the image materials. In this case, the electronic device takes a long time to analyze the image materials selected by the user to extract highlight clips, and the efficiency of processing the video by the electronic device is low. In addition, the long time consumption causes the user to wait for a long time for the output of the one-click video generation, which affects the user experience of the one-click video generation function.

[0100] In view of the above problems, the embodiments of the present application provide a video processing method. In the process of implementing the one-click video generation function, the electronic device can first perform overview analysis on a first number of image frames in each video in a plurality of image materials selected by the user, to obtain a first score of the image frames in each video. Then, the electronic device can determine a target region of the video based on a first image frame with the highest score in the video and image frames within a first preset time length before and after the first image frame. Subsequently, the electronic device can analyze a second number of images in the target region of the video to obtain a second score of the image frames. Based on a second image frame with the highest second score and image frames within a second preset time length before and after the second image frame, the electronic device can determine a highlight clip of the video.

[0101] By adopting the scheme, the electronic device first performs overview analysis on the video, so that the target region needing further processing can be located. When the electronic device extracts the highlight segment, the second number of image frames are analyzed only for the target region instead of frame-by-frame analysis on the entire video, so that the workload of the electronic device in image frame analysis can be reduced, the time consumption of the electronic device in implementing the one-key-to-clip function can be reduced, the efficiency of the electronic device in processing the video can be improved, and the user experience of the one-key-to-clip function can be improved.

[0102] The video processing method provided in the embodiments of the present application can be applied to an electronic device with an image processing function. It should be noted that the image material for one-key-to-clip in the embodiments of the present application can include pictures and videos. The user can use the one-key-to-clip function to generate highlight segments of multiple pictures into a video set, use the one-key-to-clip function to generate highlight segments of multiple videos into a video set, or use the one-key-to-clip function to generate highlight segments of multiple pictures and videos into a video set.

[0103] The electronic device can also be referred to as a terminal, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. The electronic device can be a mobile phone, a smart television, a wearable device, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, or a wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technology and specific device form of the electronic device.

[0104] For ease of understanding, the following describes the application scenarios and interface implementation of the electronic device in implementing the one-key-to-clip function with reference to the accompanying drawings. Figures 1-4 Taking the electronic device as a mobile phone, the application scenarios and interface implementation of the electronic device in implementing the one-key-to-clip function are introduced.

[0105] In one application scenario, the user can use the mobile phone to pre-shoot multiple image materials, and the mobile phone stores the multiple image materials in the gallery. The mobile phone can also pre-acquire image materials transmitted by other devices. For example, the user can use the mobile phone to scan a plurality of pictures, and the mobile phone can acquire the scanned pictures.Figure 1 As shown in (a) of FIG. 10, an icon of the gallery is displayed in the desktop of the mobile phone. When the user wants to generate a video set based on multiple image materials using the mobile phone, the user can click the icon of the gallery in the desktop of the mobile phone. In response to the clicking operation of the user on the icon, the mobile phone displays the gallery interface 101 as shown in (b) of FIG. 10. The gallery interface 101 includes a "one-click big picture" option. Then, in response to the clicking operation of the user on the "one-click big picture" option as shown in (b) of FIG. 10, the mobile phone can display the gallery interface 102 as shown in (c) of FIG. 10. The gallery interface 102 can include multiple recently taken image materials. The user can select any one or more of the multiple image materials as shown in (c) of FIG. 10 as candidate image materials for one-click picture generation. For example, in response to the selection operation of the user on part of the image materials in (c), the mobile phone can display the gallery interface 103 as shown in (d) of FIG. 10. The gallery interface 103 includes all image materials in the gallery. The gallery interface 103 can also include a video generation option, such as a "√" check option. Figure 1 Figure 1 Figure 1 Figure 1 Figure 1 Figure 1

[0106] In response to the clicking operation of the user on the "√" check option as shown in (d) of FIG. 10, the mobile phone can analyze the 5 image materials selected by the user in the gallery interface 103, select highlight segments from the image materials, and generate a video set according to the selected highlight segments. In this process, the mobile phone can display the gallery interface 201 as shown in (a) of FIG. 20. The gallery interface 201 includes an analysis material progress to allow the user to intuitively view the analysis progress. Figure 1 Figure 2

[0107] In one example, after the mobile phone generates the video set, the mobile phone can display the gallery interface 301 as shown in (a) of FIG. 30. The gallery interface 301 can display the generated video set. The mobile phone can automatically play the video set in the gallery interface 301. In addition, as shown in (a) of FIG. 30, the gallery interface 301 can also include a video export option 302 for supporting export of the generated video set. In response to the clicking operation of the user on the video export option 302, the mobile phone can save the video set in the gallery, so that the user can view the video set from the gallery. In response to the clicking operation of the user on the video export option 302, the mobile phone can also display the video export interface 303 as shown in (b) of FIG. 30. Figure 3 Figure 3 Figure 3

[0108] ​​​​​​​​​​​In one example, as the user selects image materials, the phone can provide prompts to help the user determine the appropriate number of image materials to choose. For example... Figure 1 As shown in (d), the image library interface 103 displays the prompt message "6 or more image materials will produce better results", so that users can know at least how many image materials to select to generate a better video set.

[0109] In one example, as the user selects image assets, the phone can provide prompts to let the user know the maximum number of image assets they can select. For example... Figure 4 As shown, the gallery interface 401 displays the message "A maximum of 30 image materials can be selected", so that users can know how many image materials they can select.

[0110] In one example, after generating the video set, the phone can also display an interface showing other function options, allowing users to edit, add effects, analyze, and perform other operations on the generated video set based on these options. Figure 3 As shown in 301, other function options may include, but are not limited to, templates, music, clips, sharing, etc.

[0111] The following example uses a mobile phone as an electronic device, combined with... Figure 5 The hardware structure of electronic devices will be introduced.

[0112] Figure 5 A schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of this application is shown. Figure 5 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a camera 193, a display screen 194, etc.

[0113] The processor 110 can include one or more processing units, for example: the processor 110 can include a controller, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, the controller can be the nerve center and command center of the mobile phone 100. The controller can generate operation control signals according to instruction operation codes and timing signals, complete the control of fetching instructions and executing instructions. The processor 110 can also be provided with a memory for storing instructions and data.

[0114] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. Avoiding repeated access, reducing the waiting time of the processor 110, thus improving the efficiency of the system.

[0115] The wireless communication function of the electronic device 100 can be realized by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0116] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering, such as rendering the operation interface diagram shown in Figures 1-4 , etc.

[0117] The display screen 194 is configured to display the operation interface of the screen projection APP, the screen projection image, the screen projection video, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a quantum dot light-emitting diode (QLED), or the like. In some embodiments, the mobile phone 100 can include one or N display screens 194, where N is a positive integer greater than 1.

[0118] In this embodiment, the display screen 194 can be configured to display the gallery interface 101, the gallery interface 102, and the gallery interface 103 in FIG. 1. Figure 1 In this embodiment, the display screen 194 can be configured to display the gallery interface 201 in FIG. 2. Figure 2 In this embodiment, the display screen 194 can be configured to display the gallery interface 301 and the video export interface 303 in FIG. 3. Figure 3 In this embodiment, the display screen 194 can be configured to display the gallery interface 401 in FIG. 4. Figure 4

[0119] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor.

[0120] The ISP is configured to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electric signal, and the camera photosensitive element transmits the electric signal to the ISP for processing, and the image visible to the naked eye is converted. The ISP can also optimize the algorithm of the noise, brightness, and skin color of the image. The ISP can also optimize the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be arranged in the camera 193.

[0121] ​The camera 193 is used to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or the like format image signal. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than one.

[0122] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0123] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0124] The NPU is a neural-network (NN) computing processor, which is inspired by the structure of biological neural networks, such as the transmission mode between human brain neurons, and can quickly process input information and continuously self-learn. Through the NPU, the electronic device 100 can realize intelligent cognition applications, such as image recognition, face recognition, speech recognition, text understanding, etc.

[0125] The external memory interface 120 can be used to connect an external memory card to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions.

[0126] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 performs various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (APP) required by a function (for example, a camera APP, a gallery APP, and a third-party video editing software, etc.), and the like. The data storage area can store data created during use of the electronic device 100 (for example, a photo or a video taken, a mobile phone screenshot, a mobile phone screen recording content, an image downloaded from another device, and a video set generated by using a one-key photo function, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.

[0127] The electronic device 100 can realize an audio function through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, an application processor, and the like.

[0128] For example, after a video set is generated by using a one-key photo function, the audio module 170 decodes an audio signal of the video set, and then the speaker 170A, also called a “loudspeaker”, converts an audio electrical signal into a sound signal. In this way, the user can hear background sound synchronized with the video in the highlight segment and added video music, and the like.

[0129] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0130] The following will be described in combination with Figure 6 The software structure of the electronic device is introduced.

[0131] Figure 6 The software structure of the electronic device provided in the embodiments of the present application is illustrated.

[0132] As Figure 6As shown, the electronic device can employ a layered architecture, which divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the software layers of the software structure are divided from top to bottom into an application (APP) layer, a media middleware framework layer, an application framework (FWK) layer, and a hardware abstraction layer (HAL).

[0133] The APP layer, referred to as the application layer for short, can include a series of application packages, such as a camera, a gallery, third-party video editing software, a calendar, a map, and navigation, etc. When these application packages are run, each service module provided by the media middleware framework layer and the application framework layer can be accessed through an application programming interface (API), and corresponding intelligentized businesses can be performed.

[0134] In some embodiments, the camera is used to take photos, videos, slow-motion images, panoramic images, and the like in response to user operations. After these images are taken by the camera, or after the user triggers a phone screenshot, or after the user triggers a phone screen recording, or after the electronic device downloads images from other devices, the electronic device can save these images in the gallery, so that the user can perform video editing operations on the images in the gallery, such as one-key video generation operations.

[0135] The embodiments of the present application divide the gallery from top to bottom into a business layer, an application function layer, and a basic function layer.

[0136] The business layer, also referred to as a video editing business layer, provides a plurality of businesses (which can also be referred to as functions), such as multi-camera video automatic video generation, one-way multi-AI music short video, one-key video generation, and highlight moments. These businesses are presented in the form of controls in the user interface (UI) of the gallery. The user can trigger the gallery to perform corresponding video processing actions by operating these controls. For example, after the user selects image materials (materials including pictures and / or videos), in response to the user's click operation on the one-key video generation control in the gallery, the gallery can call underlying modules to automatically analyze and extract highlight segments in the pictures and / or videos through algorithms, and then combine the highlight segments into a well-edited video set.

[0137] The application function layer includes an automatic clipping framework. Each service in the service layer can call the automatic clipping framework to provide automatic clipping services for pictures and videos. Illustratively, the automatic clipping framework can include functional modules such as highlight clip analysis, story line organization, layout splicing, and special effect beautification. The highlight clip analysis is used to call the highlight clip analysis interface and the policy monitoring interface in the high media platform framework layer to extract highlight clips from pictures and / or videos. The story line organization is used to sequentially splice multiple pictures and / or videos in the form of a story line based on the content of the pictures and / or videos. The layout splicing is used to adjust the interface layout of the pictures and / or videos. The special effect beautification is used to adjust the beautification effect of the pictures and / or videos, such as adjusting the brightness of the picture and beautifying the face of a person.

[0138] The basic function layer is used to perform basic function processing on the clipped pictures and / or video clips after the automatic clipping framework clips the multiple pictures and / or videos. Illustratively, the basic function layer can include basic function modules such as video splicing, synthesis saving, video effect rendering, and audio effect processing. The video splicing is used to splice the extracted multiple highlight clips (wherein the highlight clips include pictures and / or videos). The synthesis saving is used to store the video set obtained after splicing. The video effect rendering is used to add video effects to the video set obtained after splicing, such as adding style filters and themes to the video set. The audio effect processing is used to add sound effects to the video set obtained after splicing, such as adding background music to the video set.

[0139] The media platform framework layer is a software layer arranged between the APP layer and the FWK layer. The media platform framework layer can include an analysis performance query interface, a highlight clip analysis interface, a policy monitoring interface, a pipeline interface, a theme summary interface, and an initialization interface. The analysis performance query interface is used to calculate the total duration of all videos according to the video analysis speed. The highlight clip analysis interface is used to call the policy monitoring interface to extract highlight clips. The theme summary interface is used to call the underlying algorithm to analyze the picture content of the highlight clips to determine the theme corresponding to the content of the highlight clips. The initialization interface is used to initialize the highlight clip algorithm, the face detection algorithm, the video acceleration algorithm, and the image super-resolution algorithm of the HAL layer.

[0140] The policy monitoring interface is used to determine the first number of image frames of each video and the position of the first number of image frames in the video. The pipeline interface is used to downsample the video file according to the file descriptor of each video and the position of the image frame issued by the policy monitoring interface, and forward the data address of the downsampled video file to the hardware abstraction layer through the application framework layer, and then report the analysis result of the image frame returned by the hardware abstraction layer to the policy monitoring interface. The analysis result can include an aesthetic score of the image frame.

[0141] The policy monitoring interface is further configured to determine a target region of each video according to the first scores of the image frames, and determine a second number of image frames of the target region of each video and positions of the second number of image frames in the target region of the video. The target region is a region corresponding to the image frame with the highest score in the video. The channel interface is further configured to downsample the video file according to the file descriptor of each video and the positions of the image frames issued by the policy monitoring interface, and forward a data address of the downsampled video file to the hardware abstraction layer through the application framework layer, and then report an analysis result of the image frames in the target region returned by the hardware abstraction layer to the policy monitoring interface.

[0142] The policy monitoring interface is further configured to determine a highlight segment of each video according to the second scores of the image frames.

[0143] The theme summarization interface is configured to call a bottom algorithm to analyze a picture content of the highlight segment to determine a theme corresponding to the content of the highlight segment. The initialization interface is configured to initialize the highlight segment algorithm, the face detection algorithm, the video acceleration algorithm and the image super-resolution algorithm of the HAL layer.

[0144] It should be noted that the present application is described by taking the gallery providing one-key video generation function as an example, and the present application is not limited to the embodiments. In actual implementation, the third-party video editing software can use the video processing method provided by the embodiments to generate a video set from a plurality of pictures and videos selected by a user in one key.

[0145] The FWK layer, referred to as the framework layer, can be used to support running of each module in the media middle station framework layer. For example, the framework layer can include a one-key video generation interface, a parameter management interface, an image data transmission interface, a theme analysis interface and a performance analysis interface.

[0146] The HAL layer is a packaging of a Linux kernel driver, which provides an interface upward, and hides hardware interface details of a specific platform, so as to provide a virtual hardware platform for an operating system, and make it hardware-independent, and portable on multiple platforms. For example, the hardware abstraction layer can include a chip analysis speed interface, a highlight segment algorithm, a face detection algorithm, a video acceleration algorithm and an image super-resolution algorithm. The highlight segment algorithm is an image processing algorithm provided by an image signal processor. The algorithm can perform aesthetic scoring on each image according to image color, image texture features, image quality, frame interpolation with previous and subsequent frames and edge change rate values of each image. The aesthetic scoring can be used as a basis for evaluating whether a frame of image is a highlight segment.

[0147] In this embodiment, the highlight segment algorithm can analyze each image frame in the video to obtain a score of the image frame. For example, if the score of the image frame is greater than a preset score threshold, the image frame and image frames within a preset time length adjacent to the image frame can be used as a highlight segment of the material video. For example, the preset score threshold can be 90, and if the score of the image frame of the video is greater than 90, the image frame and image frames within a preset time length adjacent to the image frame can be used as a highlight segment of the video. Alternatively, the video includes scores of N image frames, and the image frame with the highest score and image frames within a preset time length before and after the image frame are used as a highlight segment of the video.

[0148] In some other possible embodiments, assuming that the score of the image frame is in the range of 0-100, the score can be divided into different rating results according to different value ranges of the score. For example, the first threshold is 80, and the score of the image frame is greater than or equal to the first threshold (the value range of the score is 80-100), which indicates that the score result of the image frame is "high". The second threshold can be 50, and the score of the image frame is greater than or equal to the second threshold (the value range of the score is 50-79), which indicates that the score result of the image frame is "relatively high". The third threshold can be 20, and the score of the image frame is greater than or equal to the third threshold (the value range of the score is 20-49), which indicates that the score result of the image frame is "medium". The score of the image frame is less than the third threshold (the value range of the score is 0-19), which indicates that the score result of the image frame is "low". For example, in an implementation manner, the image frame with a score greater than or equal to the first threshold / second threshold / third threshold can be used as a selectable image frame for extraction of a highlight segment of the video. In some other implementable manners, the image frame with the highest score can be used as a selectable image frame for extraction of a highlight segment of the video.

[0149] It should be noted that, Figure 6 It should be noted that,

[0150] It can be understood that, in order to implement the video processing method in the embodiments of the present application, the electronic device contains the hardware and / or software modules corresponding to the execution of each function. The algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in hardware or a combination of hardware and computer software. Whether a certain function is executed in hardware or computer software driven hardware depends on the specific application of the technical solution and the design constraints. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments.

[0151] The video processing method provided by the embodiments of the present application can determine the first number of image frames for overview analysis of each video according to the analysis parameters of the image material by the strategy monitoring module of the media middleware framework layer. The strategy monitoring module can send each image frame in each video to the algorithm module of the HAL layer for analysis, and obtain the first score of each image frame from the algorithm of the HAL layer. Further, based on the first image frame with the highest first score and the image frames within the first preset time length before and after the first image frame, the target region of each video that needs further processing is determined. For the target region, a second number of image frames are analyzed, and the strategy monitoring module can send the second number of image frames in the target region to the algorithm module of the HAL layer for analysis, and obtain the second score of the image frames from the algorithm of the HAL layer. Based on the second image frame with the highest second score and the image frames within the second preset time length before and after the second image frame, the highlight segment of the video is determined. The electronic device first performs overview analysis on the video, so that the target region that needs further processing can be located. When the electronic device extracts the highlight segment, only the second number of image frames in the target region are analyzed, instead of frame-by-frame analysis of the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumption of the electronic device to realize the one-key filming function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one-key filming function.

[0152] The following takes the execution subject of the audio and video processing method as an example, and each module in the software structure diagram as shown in Figure 6 The video processing method provided by the embodiments of the present application is exemplarily described.

[0153] Figure 7 A method flow diagram of each module in the electronic device performing algorithm initialization, obtaining analysis parameters, etc. before the strategy monitoring module of the electronic device determines the first number of image frames for overview analysis of each video is provided. The method can be applied in the one-key filming scene as shown in Figures 1-4 As shown in Figure 7 As shown in

[0154] S01, the service layer receives an operation of enabling the one-key photo function input by a user.

[0155] In this embodiment, the service layer refers to a one-key photo module of the service layer. That is, the operation of enabling the one-key photo function input by the user is received by the one-key photo module of the service layer. For example, the operation can be a click operation on the "one-key photo" card as shown in (b) of FIG. 1. Figure 1

[0156] S02, the service layer loads and displays candidate pictures and candidate videos.

[0157] S03, the service layer receives an operation of selecting multiple pictures and videos by a user, and receives an operation of determining to execute the one-key photo function input by the user.

[0158] For example, the operation of selecting multiple pictures and videos by the user can be a click operation on the photos and videos as shown in (c) of FIG. 1, and the operation of determining to execute the one-key photo function input by the user can be a click operation on the OK option as shown in (d) of FIG. 1. Figure 1 Figure 1

[0159] S04, the service layer initializes related algorithms of the HAL layer by calling an initialization interface of the media middleware framework layer through the application function layer.

[0160] In this embodiment, the related algorithms refer to algorithms of the function to be implemented by the service layer. Here, the service function is the one-key photo function, and the related algorithms refer to related algorithms involved in the one-key photo function. For example, the related algorithms include a highlight segment algorithm, a face detection algorithm, a video acceleration algorithm, and an image super-resolution algorithm, etc.

[0161] S05, the initialization interface of the media middleware framework layer successively issues initialization parameters to the algorithm modules of the HAL layer through a channel interface of the media middleware framework layer and a service interface of the FWK layer.

[0162] In the HAL layer, one algorithm corresponds to one algorithm interface, and the FWK layer is provided with multiple service interfaces. One service interface of the FWK layer corresponds to one algorithm interface of the HAL layer. Each service interface of the FWK layer plays a role of data transparent transmission between the algorithm interface of the HAL layer and the channel interface of the media middleware framework layer.

[0163] Different algorithms involve different initialization parameters, and therefore the algorithm initialization parameters issued to the algorithm interfaces through the service interfaces can be different.

[0164] S06, the algorithm modules of the HAL layer are initialized according to the initialization parameters.

[0165] ​​​S07, the algorithm module of the HAL layer returns an initialization success message to the channel interface of the media middleware framework layer through the service interface of the FWK layer.

[0166] S08, the channel interface of the media middleware framework layer calls the analysis speed interface of the HAL layer through the service interface of the FWK layer to obtain the chip analysis speed.

[0167] Among them, the service interface of the FWK layer can be a performance analysis interface.

[0168] S09, the HAL layer's analysis speed interface returns the chip analysis speed to the media middleware framework layer's channel interface through the FWK layer's service interface.

[0169] S10, the channel interface of the media middleware framework layer returns the chip analysis speed to the initialization interface of the media middleware framework layer.

[0170] The chip analysis speed can characterize the number of image frames analyzed by the image signal processor per unit time; alternatively, it can characterize the duration of analysis of a single image frame (single-frame analysis time). Therefore, the single-frame analysis time and the processing time for a single image can be calculated based on the chip analysis speed. The processing time for a single image and the single-frame analysis time for a video may differ. For example, the processing time for a single image could be 400ms, and the single-frame analysis time could be 200ms.

[0171] It should be understood that because different image signal processors have different performance characteristics, the chip analysis speed corresponding to different image signal processors may vary. For a commercially available electronic device, the image signal processor is fixed, and therefore the chip analysis speed corresponding to that image signal processor is also fixed.

[0172] In some embodiments, the channel interface of the media middleware framework layer can also return other relevant performance parameters of various algorithms to the initialization interface of the media middleware framework layer.

[0173] S11, the initialization interface of the media middleware framework layer returns an initialization success message to the business layer through the application function layer.

[0174] The initialization success message can carry performance parameters of various algorithms, such as chip analysis speed.

[0175] After obtaining the chip analysis speed, for each video in the image materials selected by the user, the following steps S12-S15 can be executed to obtain the analysis parameters for the video and all image materials.

[0176] Taking video 1 as an example, video 1 is one of the plurality of videos in the image material.

[0177] S12, the service layer sends a query message to the analysis performance query interface of the media middleware framework layer through the application function layer.

[0178] The query message includes the file descriptor fd1 of the video 1 selected by the user and the chip analysis speed. The file descriptor can be used as a unique identifier of the video.

[0179] S13, the analysis performance query interface of the media middleware framework layer obtains the estimated analysis duration of video 1 according to the file descriptor fd1 of video 1.

[0180] The estimated analysis duration can be the duration required by the image signal processor to analyze one video. Analyzing one video can refer to analyzing a specified number of image frames in one video. For example, the specified number can be 1, or the specified number can be greater than 1 and less than or equal to the number of image frames included in the video.

[0181] For example, when the specified number is 1, that is, the estimated analysis duration represents the duration required by the image signal processor to analyze 1 image frame in one video, the estimated analysis duration can also represent the analysis duration corresponding to the minimum limit of the image frame. For example, the single-frame analysis duration of the image signal processor is 200ms, and the estimated analysis duration of each video (including video 1) is 200ms.

[0182] In some embodiments, the specified number can also be the number of image frames included in the video, and the estimated analysis duration represents the duration corresponding to the analysis of all image frames in one video by the image signal processor. For example, the single-frame analysis duration of the image signal processor is 200ms, and video 1 includes 10 image frames, so the estimated analysis duration of video 1 is 10*200ms=2000ms. For example, video 2 includes 12 image frames, so the estimated analysis duration of video 2 is 12*200ms=2400ms.

[0183] In some embodiments, the specified number can also be a, a is greater than 1 and less than the number of image frames included in the video, and a is a natural number. Then, the estimated analysis duration represents the duration corresponding to the analysis of a image frames in the video by the image signal processor. For example, the single-frame analysis duration of the image signal processor is 200ms, video 1 includes 10 image frames, and a is 5. Then the estimated analysis duration of video 1 is 5*200ms=1000ms.

[0184] S14, the analysis performance query interface of the media middleware framework layer returns the estimated analysis duration of video 1 to the service layer through the application function layer.

[0185] After the one-key patch module of the service layer obtains the estimated analysis duration of video 1, it can continue to return to execute S12-S15 to obtain the estimated analysis duration of the next video in the plurality of image materials, until the estimated analysis duration of all videos in the plurality of image materials is obtained.

[0186] S15, the service layer obtains the analysis parameters according to the estimated analysis duration of all image materials.

[0187] The analysis parameters can include an analysis total duration, an upper limit recommended value of the analysis total duration, a maximum duration of a highlight segment, a minimum duration of a highlight segment, a total recommended duration of a highlight segment, a recommended duration of a highlight segment, whether to force each video to output a highlight segment, whether to enable audio analysis, a selected highlight segment, and the like.

[0188] The whether to force each video to output a highlight segment is by default yes, that is, in this embodiment, a highlight segment needs to be output for each video.

[0189] The analysis total duration represents the total duration required to complete the analysis of all image materials (including all pictures and all videos selected by the user) selected by the user, and the analysis total duration includes the sum of the estimated analysis durations of all pictures and the sum of the estimated analysis durations of all videos.

[0190] For pictures in the image materials, the processing duration of a picture can be directly determined according to the chip speed of the image processor, and therefore, the sum of the estimated analysis durations of all pictures can be directly determined according to the number of pictures in the image materials.

[0191] For videos in the image materials, the estimated analysis duration of each video can be obtained according to S12-15, and the sum of the estimated analysis durations of all videos in the image materials can be obtained by accumulating the estimated analysis duration of each video.

[0192] Based on the sum of the estimated analysis durations of all pictures and the sum of the estimated analysis durations of all videos, the analysis total duration of the image materials can be obtained.

[0193] For example, assuming that the materials selected by the user include 10 pictures, video 1, and video 2, the estimated analysis duration of a picture is 400 ms, the estimated analysis duration of video 1 is 200 ms, and the estimated analysis duration of video 2 is 300 ms, the above analysis total duration is 10*400 ms+200 ms+300 ms=4500 ms.

[0194] The analysis total time upper limit suggestion value represents a suggested value of the maximum time length for completing all the picture analysis and all the video analysis. The analysis total time upper limit suggestion value can be determined according to the estimated analysis time length of each video and the estimated analysis time length of the picture. Generally, the analysis total time upper limit suggestion value is greater than the analysis total time. For example, the analysis total time is 4400 ms according to the minimum analysis of 1 image frame of each video, and considering that each video can need to analyze multiple image frames, the analysis total time upper limit suggestion value can be much greater than the analysis total time, for example, the analysis total time upper limit suggestion value can be preset to 10000 ms.

[0195] In some scenarios where the analysis parameter input is abnormal, the analysis total time upper limit suggestion value can also be set to be less than the analysis total time. In the case where the analysis total time upper limit suggestion value is less than the analysis total time, that is, in the case where the analysis total time upper limit suggestion value is not enough to analyze all the pictures and all the 1 image frame of each video, part of the selected image material can be selected for analysis. The method is implemented by the strategy monitoring module of the media center layer, which will be described in detail in the following embodiments, and will not be described here.

[0196] The maximum time length of the highlight segment represents the maximum allowed time length of a highlight segment in a video. The minimum time length of the highlight segment represents the minimum allowed time length of a highlight segment in a video. The maximum time length of the highlight segment and the minimum time length of the highlight segment can be preset values. For example, the maximum time length of the highlight segment can be 3000 ms, and the minimum time length of the highlight segment can be 1000 ms.

[0197] The total highlight segment suggestion time length represents a suggested value of the sum of the suggested time lengths of all the highlight segments of multiple image materials of videos. The total highlight segment suggestion time length can be determined according to the number of videos, the maximum time length of the highlight segment, and the minimum time length of the highlight segment. For example, the maximum time length of the highlight segment can be 3000 ms, the minimum time length of the highlight segment can be 1000 ms, and when the number of videos is 5, the total highlight segment suggestion time length can be in the range of 5000 ms-15000 ms, for example, the total highlight segment suggestion time length can be 8000 ms.

[0198] The highlight segment suggestion time length represents a suggested value of the time length of a highlight segment of a video, and the highlight segment with the suggested value of the time length can effectively show the highlight effect. The highlight segment suggestion time length can be determined according to the maximum time length of the highlight segment and the minimum time length of the highlight segment. For example, the maximum time length of the highlight segment can be 3000 ms, the minimum time length of the highlight segment can be 1000 ms, and the highlight segment suggestion time length can be in the range of 1000 ms-3000 ms, for example, the highlight segment suggestion time length can be 2000 ms.

[0199] It needs to be understood that the maximum duration of the highlight segment, the minimum duration of the highlight segment, the total recommended duration of the highlight segment, and the recommended duration of the highlight segment can be set according to actual conditions.

[0200] In some embodiments, the analysis parameters can further include the actual duration of each video.

[0201] If the sum of the actual durations of all videos of the plurality of image materials is greater than the sum and is greater than the preset duration threshold, the video analysis method provided in S16-S46 of the present solution can be performed to reduce the number of image frames of the videos in the image materials for analysis, to reduce the time consumption of the electronic device in implementing the one-key film-making function, to improve the efficiency of the electronic device in processing videos, and to further improve the user experience of the one-key film-making function.

[0202] S16, the service layer issues, through the application function layer, the file descriptors fd and analysis parameters of all to-be-analyzed materials to the image highlight segment analysis interface of the media middle station framework layer.

[0203] The image highlight segment of the media middle station framework layer calls the policy monitoring module of the media middle station framework layer to perform the video processing method provided in the following embodiments.

[0204] After receiving the analysis parameters, the policy monitoring module of the media middle station framework layer can determine the material analysis strategy according to the upper limit recommended value of the total analysis duration and the total analysis duration in the analysis parameters. The material analysis strategy refers to analyzing all image materials selected by the user, or selecting part of the image materials selected by the user for analysis. In the following embodiments, the image materials include all videos and all pictures selected by the user.

[0205] After obtaining the total analysis duration of the image materials, the policy monitoring module can determine the number of analyzable pictures and the number of analyzable videos according to the total analysis duration of the image materials and the upper limit recommended value of the total analysis duration. Figure 7 After S16, refer to the method flow for determining the analysis strategy of the image materials given in Figure 8

[0206] S17, the policy monitoring module of the media middle station framework layer determines the material analysis strategy according to the total analysis duration and the upper limit recommended value of the total analysis duration.

[0207] If the total analysis duration is less than or equal to the upper limit recommended value of the total analysis duration, the material analysis strategy can be to analyze all pictures and all videos in the image materials selected by the user. For example, if the total analysis duration is 4400 ms and the upper limit recommended value of the total analysis duration is 10000 ms, the material analysis strategy can be to analyze all pictures and all videos in the image materials selected by the user.

[0208] ​If the total analysis time is greater than the upper limit of the total analysis time, the material analysis strategy can be a random extraction strategy. For example, if the total analysis time is 4400 ms and the upper limit of the total analysis time is 3000 ms, the material analysis strategy can be a random extraction strategy. The random extraction strategy refers to randomly selecting part of the image materials from the selected image materials by the user for analysis.

[0209] Suppose the analysis value of one picture is greater than the analysis value of one image frame of a video. Then, the random extraction strategy can be: every N pictures selected, M videos are allowed to be selected, until the total analysis time of the selected image materials reaches the upper limit of the analysis time or the pictures in the selected image materials have been taken. Wherein, M < N, for example, N can be 3, 4, 5, etc. Natural number, M can be 1, 2, 3, etc. Natural number less than N. The specific values of M and N can be determined according to the number of image materials selected by the user. For example, every 4 pictures are allowed to select 1 video, until the total analysis time reaches 3000 ms; or the pictures in the selected image materials have been taken.

[0210] For example, the total analysis time of 10 pictures and 2 videos is 4400 ms, which is greater than the upper limit of the total analysis time 3000 ms. According to the random extraction strategy of every 4 pictures selected, 1 video is allowed to be selected, 4 pictures and 1 video are selected, and the total analysis time of 4 pictures and 1 video is 1800 ms, which does not reach the upper limit of the total analysis time 3000 ms. Continue to select image materials, when the third picture in this round is selected, the total analysis time reaches 3000 ms, at which point the selection of image materials is stopped. Then, the selected image materials are 7 pictures and 1 video. The remaining 3 pictures are not analyzed.

[0211] In another embodiment, suppose the analysis value of one picture is less than the analysis value of one image frame of a video. Then, the random extraction strategy can be: every P video selected, Q picture is allowed to be selected, until the total analysis time of the selected image materials reaches the upper limit of the analysis time or the videos in the selected image materials have been taken. Wherein, Q < P. For example, P can be 3, 4, 5, etc. Natural number, Q can be 1, 2, 3, etc. Natural number less than P. The specific values of P and Q can be determined according to the number of image materials selected by the user. For example, every 3 videos are allowed to select 1 picture, until the total analysis time reaches 3000 ms; or the videos in the selected image materials have been taken.

[0212] For example, the total analysis time of 10 videos and 3 pictures is 3200ms, which is greater than the upper limit of the total analysis time of 3000ms. According to the random selection strategy of selecting 3 videos and 1 picture, 3 videos and 1 picture are selected, and the total analysis time of 3 videos and 1 picture is 1000ms, which does not reach the upper limit of the total analysis time of 3000ms. Continue to select image materials, and after three rounds of selection of 3 videos and 1 picture, a total of 9 videos and 3 pictures are selected, and the total analysis time is 3000ms. At this time, the selection of image materials is stopped. Therefore, the selected image materials are 9 videos and 3 pictures. The remaining 1 video is not analyzed.

[0213] In some embodiments, the electronic device determines the analyzable image materials from the user-selected materials according to the above random selection strategy (such as a mobile phone). Among them, the pictures and videos can be randomly selected from the user-selected image materials according to the order of the user-selected image materials. Or, a random number less than the number of materials can also be generated through the Java Random class, so as to select the video or picture corresponding to the random number.

[0214] In some embodiments, the image materials only include pictures, and the total analysis time is the sum of the estimated analysis time of all pictures. When the total analysis time is greater than the upper limit of the recommended analysis time, the number of analyzable pictures is calculated according to the upper limit of the recommended analysis time and the estimated analysis time of one picture, and a corresponding number of pictures are randomly selected from all pictures for analysis.

[0215] In some embodiments, the image materials only include videos, and the total analysis time is the sum of the estimated analysis time of all videos. When the total analysis time is greater than the upper limit of the recommended analysis time, the number of analyzable videos is calculated according to the upper limit of the recommended analysis time, and a corresponding number of videos are randomly selected from all videos for analysis. The time required for the number of analyzable videos is less than or equal to the total analysis time.

[0216] After selecting the pictures and / or videos analyzable within the upper limit of the recommended analysis time from the image materials, the remaining pictures and videos in the user-selected image materials are not analyzed.

[0217] The image materials include pictures, and the electronic device performs Figure 8 S18-S22 shown in the figure to analyze the highlight segments of the analyzable pictures.

[0218] The following will illustrate the process of analyzing the highlight segments of picture 1 in S18-S22, wherein picture 1 is one of the pictures analyzable within the upper limit of the recommended analysis time.

[0219] S18, the policy monitoring module of the media middleware framework layer sends an indication message to the channel interface of the media middleware framework layer, the indication message including a file descriptor of picture 1.

[0220] S19, the channel interface of the media middleware framework layer decodes, down-samples and converts the format of picture 1 according to the file descriptor of picture 1, and stores the processed picture 1.

[0221] S20, the channel interface of the media middleware framework layer sends the frame data address of picture 1 to the highlight clip algorithm interface of the HAL layer through the service interface of the FWK layer.

[0222] The service interface can be a one-key clip interface.

[0223] S21, the highlight clip algorithm interface of the HAL layer obtains picture 1 according to the frame data address of picture 1, and analyzes picture 1 based on a preset highlight clip algorithm to obtain an analysis result.

[0224] Illustratively, the highlight clip algorithm interface can perform aesthetic scoring on picture 1 according to image color, image texture features, image quality and edge change rate value of picture 1 to obtain the analysis result. The analysis result can represent the aesthetic score.

[0225] S22, the highlight clip algorithm interface of the HAL layer returns the analysis result of picture 1 to the policy monitoring module of the media middleware framework layer through the service interface of the FWK layer and the channel interface of the media middleware framework layer in sequence.

[0226] After the policy monitoring module of the media middleware framework layer obtains the analysis result of picture 1, if there are other pictures, such as picture 2, the electronic device can continue to perform S18-S22 to obtain the analysis result of the other pictures.

[0227] The policy monitoring module of the media middleware framework layer can determine the pictures with a score greater than or equal to the first threshold / second threshold as the highlight clips of the multiple image materials according to the analysis results of the pictures.

[0228] After obtaining the analysis results of all pictures, if the image materials include videos, the electronic device can obtain the analysis results of the videos by performing S23-S36. If the image materials do not include videos, the electronic device outputs a target video set composed of the pictures of the highlight clips.

[0229] Reference Figure 9Fig. 1 shows a schematic diagram of a technical idea of a highlight segment analysis method for a video according to an embodiment of the present application. In the process of analyzing each video, the electronic device (such as a mobile phone) can extract a first number of image frames in the video for overview analysis to obtain a first score of the image frames. The first number of image frames can be randomly extracted at any position of the video, or the first number of image frames can be extracted according to the shot points of the video, or the first number of image frames can be extracted at uniform positions of the video. The target region of each video is determined based on the first image frame with the highest first score. Then, a second number of image frames in the target region are analyzed to obtain a second score of the image frames. The highlight segment of the video is determined based on the second image frame with the highest second score. In this solution, instead of frame-by-frame analysis of the entire video, the workload of the electronic device for image frame analysis can be reduced, thereby reducing the time consumption of the electronic device for implementing the one-key film making function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-key film making function.

[0230] If the image material includes multiple pictures and videos, since the calculation amount of the pictures is small and the time consumption is low, the electronic device usually analyzes the highlight segments of the pictures one by one, and after the highlight segment analysis of all the pictures is completed, the highlight segments of the videos are analyzed one by one. That is, after the electronic device executes S18-S22 as shown in Fig. 1, the electronic device continues to execute S23-S37 as shown in Fig. 1. Figure 8 Figure 10 If the image material includes multiple pictures and videos, since the calculation amount of the pictures is small and the time consumption is low, the electronic device usually analyzes the highlight segments of the pictures one by one, and after the highlight segment analysis of all the pictures is completed, the highlight segments of the videos are analyzed one by one. That is, after the electronic device executes S18-S22 as shown in Fig. 1, the electronic device continues to execute S23-S37 as shown in Fig. 1.

[0231] If the image material includes multiple pictures and videos, since the calculation amount of the pictures is small and the time consumption is low, the electronic device usually analyzes the highlight segments of the pictures one by one, and after the highlight segment analysis of all the pictures is completed, the highlight segments of the videos are analyzed one by one. That is, after the electronic device executes S18-S22 as shown in Fig. 1, the electronic device continues to execute S23-S37 as shown in Fig. 1. Figure 10 Figure 8 If the image material includes multiple pictures and videos, since the calculation amount of the pictures is small and the time consumption is low, the electronic device usually analyzes the highlight segments of the pictures one by one, and after the highlight segment analysis of all the pictures is completed, the highlight segments of the videos are analyzed one by one. That is, after the electronic device executes S18-S22 as shown in Fig. 1, the electronic device continues to execute S23-S37 as shown in Fig. 1.

[0232] In this embodiment, the case where the image material includes pictures and videos is exemplified. The following videos refer to videos that can be analyzed within the upper limit of the total analysis time.

[0233] S23, the strategy monitoring module of the media center framework layer calculates the first number of image frames of each video.

[0234] The first number refers to the number of image frames in each video that can be analyzed within the upper limit of the total analysis time.

[0235] In some embodiments, the strategy monitoring module can determine the first number of image frames of each video according to the actual time length of each video in the multiple image materials, the number of videos in the multiple image materials, and the upper limit of the total analysis time.

[0236] ​​For example, the number of analyzable image frames T1 / t is determined according to the analysis total time upper limit suggestion value T1 and the single frame analysis time t. According to the actual time length of each video and the number of videos, the number of analyzable image frames T1 / t is allocated to each video to obtain the first number of image frames of each video.

[0237] In an example, the number of analyzable image frames T1 / t is allocated to each video, which can be sorted according to the actual time length of the video from long to short, and for the video ranked in the top 25%, the number of image frames T1 / t*50% is allocated on average. For the video ranked in 25%-75%, the number of image frames T1 / t*40% is allocated on average. For the video ranked in the last 25%, the number of image frames T1 / t*10% is allocated on average.

[0238] In another example, the number of analyzable image frames T1 / t is allocated to each video, which can be allocated to each video according to the proportion of the actual time length of the video.

[0239] Alternatively, in another example, the step of determining the first number of image frames of each video by the strategy monitoring module can include:

[0240] S231, the strategy monitoring module determines the basic number and the maximum number of image frames of each video.

[0241] The basic number can be understood as the minimum number of image frames required to ensure the analysis effect of the video when the analysis time is limited and the analysis performance of the electronic device is limited. The maximum number can be understood as the maximum number of image frames allowed to analyze the video. Generally, the maximum number is greater than the basic number.

[0242] In some embodiments, the strategy monitoring module can determine the basic number and the maximum number of image frames of each video according to the actual time length of each video and a preset correspondence relationship. The preset correspondence relationship represents the maximum number and the basic number of image frames corresponding to different threshold ranges of video time length.

[0243] The basic number and the maximum number of videos of different time lengths are different. For example, the electronic device (such as a mobile phone) can pre-store a preset correspondence relationship between video time length and image frame number. In the embodiment of the application, the strategy monitoring module can determine the basic number and the maximum number of image frames of each video according to the preset correspondence relationship.

[0244] For example, the preset correspondence relationship includes the following (1)-(4):

[0245] (1) The video duration P is less than the first threshold Q1, and the number of image frames of the video is 1.

[0246] (2) The video duration P is equal to the first threshold Q1, and the starting number of image frames of the video is m.

[0247] (3) The video duration P is greater than the first threshold Q1 and less than or equal to the second threshold Q2, and the number of image frames of the video is increased on the basis of the starting number. The number of image frames is updated as that is, P divided by s and rounded down plus m. It is indicated that from the starting time of the video to Q1, the number of image frames of the video is m, and from the starting time of the video to Q2, 1 image frame can be corresponded every s seconds.

[0248] (4) The video duration P is greater than the second threshold Q2 and less than the third threshold Q3, and the number of image frames of the video is that is, Q2 divided by s and rounded down, P-Q2 divided by k and rounded down, plus m. It is indicated that from the starting time of the video to Q1, the number of image frames of the video is m; from the starting time of the video to Q2, 1 image frame can be corresponded every s seconds; and from Q2 to Q3, 1 image frame can be corresponded every k. Wherein, is a down rounding symbol.

[0249] The first threshold is less than the second threshold, and the second threshold is less than the third threshold.

[0250] It should be noted that if the video duration is very long, more thresholds of time duration can be set, such as a fourth threshold, a fifth threshold, and the like. In determining the basic number and the maximum number, the thresholds such as the first threshold, the second threshold and the third threshold can be set according to the actual video duration; and the interval seconds can be determined according to the analysis performance of the electronic device. The above are only given as an example, and the parameter values are not limited.

[0251] In some embodiments, the maximum number of image frames is more than the basic number, and the frame extraction density for determining the maximum number of image frames is greater than the frame extraction density for determining the basic number of image frames. In order to obtain more image frames, generally, the first threshold for determining the maximum number is less than or equal to the first threshold for determining the basic number; or the preset second threshold for determining the maximum number is equal to or greater than the second threshold for determining the basic number, or the third threshold for determining the maximum number is greater than or equal to the third threshold for determining the basic number. In some embodiments, the interval seconds for determining the maximum number is less than or equal to the interval seconds for determining the basic number.

[0252] Several examples are given below to illustrate the determination process of the basic number and the maximum number of image frames of the video.

[0253] For example, the policy monitoring module determines the base number of image frames of the video, the first threshold value can be 3 seconds, s seconds can be 5 seconds, the second threshold value can be 30 seconds, k seconds can be 10 seconds, and the third threshold value can be 90 seconds.

[0254] For example, the policy monitoring module calculates the video 1 with a duration of 93 seconds, and for example, the image frame extraction diagram shown in FIG. 1 can be referred to. Figure 11

[0255] The policy monitoring module receives the video 1 and determines at least one image frame, at this time, the number of image frames of the video 1 is 1 frame. The duration of the video 1 (91 seconds) is greater than the first threshold value (3 seconds), so the number of image frames of the video 1 is updated to 2 frames.

[0256] From the 0th second of the duration of the video 1, within 30 seconds of the duration of the video, an image frame is added every 5 seconds, as shown in FIG. 1. Figure 11 As shown in FIG. 1, at the 30th second of the duration of the video, the base number of image frames of the video 1 is 8 frames From the 30th second of the duration of the video, within 90 seconds of the duration of the video, an image frame is added every 10 seconds, as shown in FIG. 1. Figure 11 As shown in FIG. 1, at the 90th second of the duration of the video, the base number of image frames of the video 1 is updated to 14 frames After the 90th second of the video 1, the remaining duration of the video is 3 seconds, which does not meet the requirement of extracting an image frame every 10 seconds after 90 seconds, so the number is not increased. Therefore, the base number of image frames of the video 1 with a duration of 93 seconds is 14 frames.

[0257] In some embodiments, after calculating the base number of image frames of each video, if the total number of the base number of image frames of all videos does not meet the preset minimum frame number, the base number of image frames of all videos needs to be adjusted. The adjustment method can be to increase the number of image frames of the video whose base number of image frames is less than the preset value; or to increase the base number of image frames of the video with a longer duration.

[0258] For example, in order to ensure that a more accurate video theme is obtained, the preset minimum frame number can be 5 frames.

[0259] ​In some embodiments, if there is only one video, such as video 1, the base number of image frames of video 1 is directly adjusted to 5 frames if the base number of image frames of video 1 does not meet 5 frames. If there are multiple videos, since each video is preset to obtain at least one image frame, the preset minimum frame number should be greater than the number of videos. For example, if the number of videos is 3, the preset minimum frame number can be 5 frames. For example, the base number of image frames of video 1 is 1 frame, the base number of image frames of video 2 is 2 frames, and the base number of image frames of video 3 is 1 frame. The sum of the base number of image frames of all videos is 4 frames, which is less than the preset minimum frame number 5 frames. The base number of image frames of the video whose base number of image frames is less than 2 frames can be adjusted to 2 frames. At this time, the base number of image frames of video 1 is 2 frames, the base number of image frames of video 2 is 2 frames, and the base number of image frames of video 3 is 2 frames. The sum of the base number of image frames of all videos is 6 frames, which is greater than the preset minimum frame number 5 frames. If the number of videos is 2, the base number of image frames of video 1 is 1 frame, and the base number of image frames of video 2 is 1 frame. By adjusting the base number of image frames of the video whose base number of image frames is less than 2 frames to 2 frames, the sum of the base number of image frames of all videos is 4 frames, which still does not meet the preset minimum frame number 5 frames. In this case, the base number of image frames of the video with the longest duration can be adjusted to 3 frames, so that the total number of the base number of image frames of all videos meets 5 frames. For example, the duration of video 1 is 2 seconds, and the duration of video 2 is 1 second. Then, the base number of image frames of video 1 is adjusted to 3 frames, and the base number of image frames of video 2 is adjusted to 2 frames. At this time, the total number of the base number of image frames of all videos is 5, which meets the preset minimum frame number 5 frames.

[0260] For example, the policy monitoring module determines the maximum number of image frames of the video, the first threshold value can be 2 seconds, s seconds can be 3 seconds, the second threshold value can be 30 seconds, k seconds can be 10 seconds, and the third threshold value can be 120 seconds.

[0261] For example, taking video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example, for example, reference can be made to the image frame extraction diagram shown in Figure 12

[0262] The policy monitoring module receives video 1 and extracts at least one image frame. At this time, the maximum number of image frames of video 1 is 1 frame. The duration of video 1 (93 seconds) is greater than the first threshold value (2 seconds), so the maximum number of image frames of video 1 is updated to 2 frames.

[0263] From the 0th second of the duration of video 1, one image frame is added every 3 seconds within 30 seconds of the duration of the video, as shown in Figure 12 At 30 seconds of the duration of the video, the maximum number of image frames of video 1 is 12 frames ​Starting from the 30th second of the video duration, and continuing until the 120th second, the number of image frames increases by one every 10 seconds, such as... Figure 12 As shown, at 93 seconds into the video, the maximum number of image frames in video 1 is updated to 18 frames. After the 90th second of video 1, the remaining video duration is 3 seconds, which does not meet the requirement of extracting one image frame every 10 seconds after 120 seconds. Therefore, after adding one image frame at the 90th second, the number of image frames is not increased further. So, the maximum number of image frames in video 1 with a duration of 93 seconds is 18 frames.

[0264] S232, the strategy monitoring module determines the total number of image frames to be analyzed for all videos based on the recommended upper limit of the analysis duration.

[0265] In this embodiment, the upper limit of the analysis time is a suggested maximum time for completing the analysis of all images and videos described above. Considering the time required for processing the images and videos as described in the above embodiments, the total number of image frames to be analyzed for all videos is determined by multiplying the upper limit of the analysis time by the first factor T1.

[0266] The strategy monitoring module determines the total number of image frames to be analyzed across all videos based on the first multiplier of the recommended upper limit of analysis duration, T1. This first multiplier can be 30%. In other words, the number of image frames that can be analyzed is determined based on 30% (T1*0.3) of the recommended upper limit of analysis duration.

[0267] If T1*0.3 is less than the time required to analyze the base number of image frames for all videos, but T1 is greater than the time required to analyze the base number of image frames for all videos (meaning T1*0.3 is insufficient to analyze the base number of image frames for all videos), then the strategy monitoring module can allocate an analysis time T2 to all videos for analysis. In this case, the total number of image frames analyzed for all videos is the sum of the base number for all videos. Specifically, the analysis time T2 is greater than T1*0.3, less than T1, and greater than or equal to the time required to analyze the base number of image frames for all videos.

[0268] Alternatively, if the duration of T*0.3 is insufficient to analyze the basic number of image frames in all videos, the policy monitoring module can allocate the entire recommended upper limit value of analysis duration T1 to the analysis of the basic number of image frames in all videos. The actual number of image frames processed by the policy monitoring module (total number of analyses) is the sum of the basic number of image frames in all videos.

[0269] If the upper limit of the analysis time T1 is less than the time taken to analyze the basic number of image frames of all videos, assuming the analysis time per frame is t, then the actual number of image frames processed by the policy monitoring module (total number of analyses) is T1 / t (frames).

[0270] If T1*0.3 is greater than the time consumption of analyzing the maximum number of image frames of all videos, that is, the time length of T1*0.3 is enough to analyze the maximum number of image frames of all videos, then the strategy monitoring module allocates T1*0.3 for analyzing the maximum number of image frames of all videos, and the number of image frames actually processed by the strategy monitoring module (total analysis number) is the sum of the maximum number of image frames of all videos.

[0271] If T1*0.3 is greater than the time consumption of analyzing the basic number of image frames of all videos, and T1*0.3 is less than the time consumption of analyzing the maximum number of image frames of all videos, assuming that the single-frame analysis time length is t, the number of image frames actually processed by the strategy monitoring module (total analysis number) is T1*0.3 / t (frames).

[0272] Exemplarily, assuming that the input video time lengths are 15s, 40s, 100s, and 150s respectively, and the basic number of image frames corresponding to each video is 5 frames, 9 frames, 14 frames, and 14 frames respectively, the total is 42 frames. The maximum number of image frames corresponding to each video is 7 frames, 13 frames, 19 frames, and 21 frames respectively, and the total is 60 frames.

[0273] Assuming that the single-frame analysis time length is 200ms, then the time consumption of analyzing the basic number of image frames of all videos is the time consumption of 42 frames, which is 8400ms; the time consumption of analyzing the maximum number of image frames of all videos is the time consumption of 60 frames, which is 12000ms.

[0274] If T1*0.3 is less than 8400ms, that is, T1*0.3 is less than the time consumption of analyzing the basic number of image frames of all videos, the strategy monitoring module allocates 8400ms for processing 42 image frames of all videos according to the time consumption of processing 42 frames. Alternatively, the strategy monitoring module allocates the analysis time length upper limit suggestion value T1 for analyzing the image frames of all videos, at this time, the number of image frames actually processed by the strategy monitoring module (total analysis number) is 42 frames. If T1 is less than 8400ms, the number of image frames actually processed by the strategy monitoring module (total analysis number) is T1 / 200 (frames).

[0275] If T1*0.3 is greater than 12000ms, that is, T1*0.3 is greater than the time consumption of analyzing the maximum number of image frames of all videos, the strategy monitoring module allocates T1*0.3 for analyzing the image frames of all videos, and the number of image frames actually processed (total analysis number) is 60 frames.

[0276] If T1*0.3 is greater than 8400ms and T1*0.3 is less than 12000ms, the number of image frames actually processed by the strategy monitoring module (total analysis number) is T1*0.3 / 200ms (frames) within T1*0.3.

[0277] In this embodiment, the upper limit of the analysis duration is only a recommended value for planning, but not for strict checking. The estimated analysis duration of the video analysis in this embodiment allows the recommended value of the upper limit of the analysis duration to be exceeded.

[0278] S233, the policy monitoring module of the media center framework layer determines the first number of image frames of each video according to the total analysis number of all videos.

[0279] In some embodiments, after obtaining the total analysis number of image frames of all videos, the policy monitoring module can assign a corresponding number of image frames to be analyzed (the first number) to each video in turn.

[0280] For example, the videos can be sorted according to their actual durations from long to short, and each video can be traversed in turn. After traversing each video in turn, the first value and the second value of the video are updated.

[0281] The initial value of the first value is the total analysis number, and the initial value of the second value of the video is 0.

[0282] During the traversal of the videos, the second value of the video is incremented by 1 and the first value is decremented by 1 after traversing each video, until the first value is 0, that is, until the total analysis number is assigned. When the total analysis number is assigned, the second value of each video is the corresponding first number.

[0283] It should be noted that when assigning the first number of image frames to each video, the first number of image frames assigned to each video should be less than or equal to the maximum number of image frames of the video, and the first number of image frames assigned to each video should be greater than or equal to the basic number of image frames of the video. If the second value of a video is equal to the maximum number before the video is traversed, the video is skipped in this and subsequent traversal processes, and the next video of the video is traversed. The video that is not traversed is not incremented by the second value. In this way, the first number corresponding to the actual duration of the video can be obtained for each video.

[0284] For example, referring to Figure 13 , Figure 13A method for allocating image frames to multiple videos is provided. Assume that the input videos are video 1 (15s), video 2 (20s), video 3 (30s), and video 4 (35s), and the maximum number of image frames corresponding to each video is 7, 8, 12, and 12 respectively. Assume that the total number of analysis is 38 frames. The strategy monitoring module iterates through the four videos and allocates the 38 frames to the four videos in turn. Starting from the first round, one image frame is allocated to each video in each round, i.e., when each video is iterated, the second value corresponding to the video is increased by 1. After the seventh round, the second value of the image frames of video 1 has reached the maximum number of image frames (7 frames) corresponding to the video, and the video is skipped in the subsequent iteration. In the eighth round, only video 2, video 3, and video 4 are iterated. After the eighth round, the second value of the image frames of video 2 has reached the maximum number of image frames (8 frames) corresponding to the video, and the video is skipped in the subsequent iteration. In the ninth round, only video 3 and video 4 are iterated. Until the eleventh round ends, the number of allocated frames of each video is 7, 8, 11, and 11 respectively, among which video 1 and video 2 have reached the maximum number of image frames, and video 3 and video 4 can continue to be allocated. At this time, 37 frames of the total number of analysis have been allocated, leaving 1 image frame, which is not enough to be allocated to all videos (video 3 and video 4). In the twelfth round, the strategy monitoring module allocates the remaining 1 image frame to video 4 according to the time length of the videos. After the allocation of 38 frames is completed, the first number of image frames of video 1 is 7, the first number of image frames of video 2 is 8, the first number of image frames of video 3 is 11, and the first number of image frames of video 4 is 12.

[0285] After obtaining the first number of image frames of each video, the strategy monitoring module can perform highlight segment analysis on the image frames of each video. The highlight segment analysis of each video by the strategy monitoring module can include two analysis stages. The first analysis stage includes overview analysis of the first number of image frames, and the second analysis stage includes further analysis based on a target region.

[0286] The first analysis stage refers to overview analysis of the first number of image frames for each video. Specifically, for each video, the strategy monitoring module issues a position of an image frame to the channel interface of the media center framework layer each time, so that the media center framework layer analyzes the image frame at the position and returns an analysis result. This process is repeated until the number of image frames analyzed reaches the first number of the video. After the first analysis stage, the strategy monitoring module can obtain the analysis result of the first number of image frames of each video. The analysis result includes a first score of the image frame. Thus, the target region of each video can be determined based on the first image frame with the highest score.

[0287] The second analysis stage refers to the analysis of a second number of image frames of the target region. The policy monitoring module sequentially issues a position of one image frame of the target region to the channel interface of the media middle station framework layer to obtain an analysis result of the image frame at the position, wherein the analysis result comprises a second score of the image frame. Thus, the highlight segment of the video is determined based on a second image frame with the highest second score.

[0288] Taking the video 1 as an example, the process of the highlight segment analysis of the video is exemplarily illustrated below in combination with S24-S37. Among them, S24-S29 is the first analysis stage, and S30-S37 is the second analysis stage.

[0289] S24, the policy monitoring module issues a file descriptor fd1 of the video 1 and a first image frame position of the video 1 to the channel interface of the media middle station framework layer.

[0290] The positions of the first number of image frames in the video can be evenly distributed. Thus, the first image frame position can be the position of the first image frame from the start time of the video.

[0291] In some embodiments, the positions of the first number of image frames in the video can also be distributed according to the split points of the video and the similar picture regions of the video.

[0292] The first image frame position can be the position of any one image frame in the video 1. For example, the first image frame position can be the position of the first image frame of the video 1; or the first image frame position can also be other specified positions in the video 1.

[0293] S25, the channel interface of the media middle station framework layer decodes, reduces the resolution, and converts the format of the video 1 according to the file descriptor of the video 1, and stores the processed video 1.

[0294] S26, the channel interface of the media middle station framework layer sends the frame data address of the video 1 and the corresponding first image frame position in the video 1 to the highlight segment algorithm interface of the HAL layer through the service interface (such as the one-key segment interface) of the FWK layer.

[0295] S27, the highlight segment algorithm interface of the HAL layer obtains the first image frame of the video 1 according to the frame data address of the video 1 and the corresponding first image frame position in the video 1, and analyzes the image frame at the first image frame position based on the preset highlight segment algorithm to obtain the analysis result of the image frame at the first image frame position.

[0296] S28, the highlight segment algorithm interface of the HAL layer returns the analysis result of the image frame at the first image frame position to the policy monitoring module through the service interface of the FWK layer and the channel interface of the media middle station framework layer in sequence.

[0297] The analysis result can represent a score of the image frame and the score result.

[0298] S29, the policy monitoring module acquires the position of the second image frame, and returns to perform S24 until the number of analyzed image frames meets the first number of videos.

[0299] In this way, the policy monitoring module can obtain the analysis result of the first number of image frames of all videos. The analysis result includes a first score of the image frame. Based on the first score of the image frame, a scheme of the second analysis stage is performed.

[0300] S30, the policy monitoring module determines the target region of each video based on the analysis result of the image frame of all videos.

[0301] In the embodiment, for each video, after obtaining the analysis result of the first number of image frames, the policy monitoring module determines the first image frame with the highest score according to the first score of each image frame. The first image frame can include at least one image frame. If the first image frame includes one image frame, the image frames within the first preset time length before and after the first image frame can be directly determined as the target region of video 1. For example, as shown in (a) of FIG. 1, the first image frame of video 1 includes image frame 1 (e.g., 1 in (a) of FIG. 1), and the segments within the first preset time length before and after image frame 1 are the target region. Figure 14 Figure 14

[0302] Alternatively, if the first image frame includes one image frame, candidate high-light segments can be determined according to the half high-light segment suggestion time length before and after the first image frame. The image frames within the third preset time length before and after the candidate high-light segments are determined as the target region of video 1. The time length of the target region can be b times of the candidate high-light segments. For example, b can be a number greater than 1 and less than 2.

[0303] For example, as shown in (b) of FIG. 1, the first image frame of video 1 includes image frame 1 (e.g., 1 in (b) of FIG. 1), and the segments within the half high-light segment suggestion time length before and after image frame 1 are the candidate high-light segments, and the segments within the third preset time length before and after the candidate high-light segments are the target region. Figure 14 Figure 14

[0304] ​​​​If the first image frame comprises a plurality of continuous image frames. According to the highlight segment proposal duration, the segment covered by the fourth preset duration before the first image frame of the plurality of continuous first image frames to the fourth preset duration after the last image frame of the plurality of continuous second image frames is determined as the candidate highlight segment.

[0305] For example, as Figure 15 , Figure 15 Another schematic diagram of the target region is given. The first image frame of video 1 comprises image frame 1 (such as 1 in Figure 15 ), image frame 2 (such as 2 in Figure 15 ), and image frame 3 (such as 3 in Figure 15 ). Then, the segment covered by the fourth preset duration before image frame 1 to the fourth preset duration after image frame 3 is the candidate highlight segment of the video 1. The segments covered by the fourth preset duration before and after the candidate highlight segment are the target regions.

[0306] After determining the candidate highlight segment of each video, if the sum of the lengths of the candidate highlight segments of all videos is greater than the second multiple of the total highlight segment proposal duration, the strategy monitoring module needs to adjust the lengths of the candidate highlight segments of the videos, so that the sum of the lengths of the candidate highlight segments of all videos is less than or equal to the second multiple of the total highlight segment proposal duration.

[0307] The second multiple can be a number greater than 1, for example, the second multiple can be 1.05. Illustratively, according to the length of the candidate highlight segment of each video that exceeds the second multiple of the total highlight segment proposal duration (the length to be adjusted), the length of the candidate highlight segment is reduced according to the actual length of each video. Illustratively, the shorter the length of the video, the more seconds the candidate highlight segment is reduced.

[0308] Suppose two videos are input, the actual length of video 1 is 10s, and the actual length of video 2 is 20s. The highlight segment proposal duration is 6s, and the total highlight segment proposal duration is 10s. The length of the candidate highlight segment of video 1 and video 2 is 6s; the length of the target region of video 1 and video 2 is 12s.

[0309] The sum of the time length of the candidate highlight segment of video 1 and video 2 is 12s, which exceeds 1.05 times (10.5s) of the total recommended time length of the highlight segment (10s). The strategy monitoring module needs to adjust the time length of the candidate highlight segment of each video. The time length to be adjusted is 12s-10s*1.05=1.5s. The strategy monitoring module reduces the time length of the candidate highlight segment of video 1 (10s) by 1s according to the inverse weighted distribution of the video time length; the time length of the candidate highlight segment of video 1 is updated from 6s to 5s, and the time length of the candidate highlight segment of video 2 (20s) is reduced by 0.5s, and the time length of the candidate highlight segment of video 2 is updated from 6s to 5.5s. In this way, the sum of the time length of the candidate highlight segment of video 1 and video 2 is 10.5s, which does not exceed 10s*1.05, and there is no need to continue adjusting.

[0310] Therefore, the time length of the candidate highlight segment of video 1 is determined to be 5s, and the time length of the target region corresponding to the candidate highlight segment in video 1 is 10s; the time length of the candidate highlight segment of video 2 is 5.5s, and the time length of the target region corresponding to the candidate highlight segment in video 2 is 11s.

[0311] The target region of each video can be determined by the method provided in S30. After determining the target region of each video, S31-S36 can be executed to analyze the second number of image frames of the target region.

[0312] In S31, the strategy monitoring module determines the second number of image frames of the target region of each video.

[0313] The second number refers to the number of image frames that can be analyzed in the target region of each video within the remaining analysis time length. The remaining analysis time length is equal to the time length for video analysis minus the total time length spent on performing the overview analysis.

[0314] In some embodiments, the analyzable number W of image frames of the target region of all videos can be determined according to the ratio of the remaining analysis time length to the single-frame analysis time length. Further, the analyzable number is distributed to each video according to the time length of the target region of the video to obtain the second number of image frames of the target region of the video.

[0315] In a feasible method, the method for determining the second number of the target region can refer to the method for determining the first number provided in S233.

[0316] In some other embodiments, the second number of image frames of each target region can also be determined according to the ratio of the number of image frames contained in each target region to the number of image frames of all target regions. For example, the second number of image frames of each target region can be determined according to the ratio of the number of image frames contained in each target region to the number of image frames of all target regions. The second number of each target region is determined, wherein i iThe ratio of the number of image frames contained in the i-th target region to the total number of image frames of all target regions, and I is the total number of image frames of all target regions. The floor symbol.

[0317] For example, assuming there are 3 videos, each video has a total of 10 image frames per second (10 FPS), and the target region duration of the 3 videos is 5s, 10s, and 15s respectively, and the single-frame analysis duration is 200ms, and the remaining analysis duration is 25s.

[0318] In the remaining analysis duration, the number of images W that can be analyzed is: 25s / 200ms = 125 (frames). The number of image frames contained in the target region of the 3 videos is 50, 100, and 150 respectively; 125 frames are allocated to the 3 videos according to the ratio of 1:2:3. After allocating 20 image frames to video 1, 40 image frames to video 2, and 60 image frames to video 3, there are still 5 image frames left. According to the actual duration of the video, the image frames are preferentially allocated to the video with longer duration, for example, 3 image frames are allocated to video 3 and 2 image frames are allocated to video 2.

[0319] Alternatively, in some embodiments, the remaining image frames can also be allocated according to the score of the target region. The score of the target region is the average of the first score of the first image frame. For example, there are 5 image frames left, and according to the score of the target region, the scores corresponding to the target regions of the 3 videos are 80, 90, and 70 respectively. The second number of image frames of the video with a high score is preferentially increased, so 3 image frames are allocated to video 2, and 1 image frame is allocated to video 1 and video 3 respectively. Alternatively, 2 image frames are allocated to video 2 and video 1, and 1 image frame is allocated to video 3. Until 125 image frames are allocated, the second number of image frames of the target region of each video is obtained.

[0320] After the policy monitoring module of the media middleware framework layer determines the second number of image frames of the target region of each video, the process of frame-by-frame analysis of the target region of video 1 is illustrated by taking video 1 as an example, which can include:

[0321] S32, the policy monitoring module issues the file descriptor fd1 of video 1 and the third image frame position of the target region of video 1 to the channel interface of the media middleware framework layer.

[0322] The third image frame position can be the position of any image frame in the target region of video 1. The second number of image frames is evenly distributed in each position of the target region.

[0323] ​​S33, the channel interface of the media middle platform framework layer performs decoding, resolution reduction and format conversion on video 1 according to the file descriptor of video 1, and stores the processed video 1.

[0324] S34, the channel interface of the media middle platform framework layer sends the frame data address of video 1 and the position of the third image frame to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one-click video creation interface).

[0325] S35, the HAL layer's highlight fragment algorithm interface obtains the third image frame of video 1 based on the frame data address and the position of the third image frame. Then, it analyzes the image frame at the position of the third image frame based on the preset highlight fragment algorithm, and obtains the analysis results of the image frame at the position of the third image frame.

[0326] The S36,HAL layer's highlight fragment algorithm interface sequentially returns the analysis results of the image frame at the third image frame position to the media platform framework layer's strategy monitoring module through the FWK layer's service interface and the media platform framework layer's channel interface.

[0327] The analysis results include a second score for the image frames.

[0328] The strategy monitoring module of the media middle platform framework layer continues to send the fourth image frame position of the target area of ​​video 1 to the channel interface of the media middle platform framework layer for analysis, until the analysis of the second number of image frames in the target area is completed.

[0329] The position of the fourth image frame can be any image frame in the target area of ​​video 1.

[0330] S37, the strategy monitoring module of the media middle platform framework layer determines the highlight segment of video 1 based on the analysis results of the image frames of the target area of ​​video 1.

[0331] In this embodiment, the strategy monitoring module obtains the second image frame with the highest second score based on the second score of the image frame. The second image frame may include at least one image frame. If the second image frame includes only one image frame, the image frames before and after the second image frame within a second preset time period can be directly identified as the highlight segments of video 1.

[0332] For example, such as Figure 16 (a) provides a schematic diagram of a highlight segment. The second image frame of video 1 includes image frame 1 (e.g., ...). Figure 16 In (a) 1), the segment covered by the second preset duration before and after image frame 1 is the highlight segment of video 1. The duration of the highlight segment is less than or equal to the suggested duration of the highlight segment.

[0333] Alternatively, if the second image frame includes one image frame, according to the highlight segment suggestion duration, the segment before and after the second image frame by 1 / 2 of the highlight segment suggestion duration can be determined as the highlight segment of the video 1. Wherein, the duration of the highlight segment is equal to the highlight segment suggestion duration.

[0334] For example, as shown in (b) of FIG. 6, the first image frame of the video 1 includes one image frame 1 (e.g., 1 in (b) of FIG. 6), and the segment before and after the image frame 1 by 1 / 2 of the highlight segment suggestion duration is the highlight segment of the video 1. Figure 16 Figure 16 For example, as shown in (b) of FIG. 6, the first image frame of the video 1 includes one image frame 1 (e.g., 1 in (b) of FIG. 6), and the segment before and after the image frame 1 by 1 / 2 of the highlight segment suggestion duration is the highlight segment of the video 1.

[0335] For example, as shown in (b) of FIG. 6, the first image frame of the video 1 includes one image frame 1 (e.g., 1 in (b) of FIG. 6), and the segment before and after the image frame 1 by 1 / 2 of the highlight segment suggestion duration is the highlight segment of the video 1.

[0336] For example, as shown in (b) of FIG. 6, the first image frame of the video 1 includes one image frame 1 (e.g., 1 in (b) of FIG. 6), and the segment before and after the image frame 1 by 1 / 2 of the highlight segment suggestion duration is the highlight segment of the video 1. Figure 17 Figure 17 For example, as shown in (b) of FIG. 6, the first image frame of the video 1 includes one image frame 1 (e.g., 1 in (b) of FIG. 6), and the segment before and after the image frame 1 by 1 / 2 of the highlight segment suggestion duration is the highlight segment of the video 1. Figure 17 Figure 17 For example, as shown in (b) of FIG. 6, the first image frame of the video 1 includes one image frame 1 (e.g., 1 in (b) of FIG. 6), and the segment before and after the image frame 1 by 1 / 2 of the highlight segment suggestion duration is the highlight segment of the video 1. Figure 17 If there are other videos, such as video 2, the electronic device can continue to perform the above S32-S37 to obtain the analysis results of the highlight segments of the other videos.

[0337] After obtaining the analysis results of the highlight segments of all videos, in some embodiments, referring to (b) of FIG. 6, a flowchart of post-processing of the highlight segments in a video processing method is given. After performing S37, the electronic device can report all picture and video analysis results by using the following S38.

[0338] Figure 18 S38, the strategy monitoring module of the media middleware framework layer reports all picture and video analysis results to the application function layer through the image highlight segment analysis interface of the media middleware framework layer.

[0339] S38, the strategy monitoring module of the media middleware framework layer reports all picture and video analysis results to the application function layer through the image highlight segment analysis interface of the media middleware framework layer.

[0340] Wherein, the analysis results can include the positions of all highlight segments. For example, the start time and the end time of the highlight segment of each video.

[0341] ​​​​S39, the application function layer performs editing and screening on the selected material according to the analysis results of all pictures and videos, and obtains all highlight clips.

[0342] S40, the application function layer of the application layer calls the theme summary interface of the media platform framework to request a theme template.

[0343] The request is transmitted to the HAL layer through the FKW layer via the theme summary interface and the channel interface of the media platform framework.

[0344] S41, the HAL layer determines the theme template that meets the scene of the picture in the highlight clip.

[0345] In some embodiments, the electronic device can be provided with multiple theme templates (style templates). The theme algorithm can recommend a theme template that meets the scene of the picture based on the scene of the picture in the highlight clip, such as a person, a landscape, a food, a child, a pet, a sport, or a travel.

[0346] S42, the HAL layer returns the determined theme template to the theme summary interface of the media platform framework layer, and the theme summary interface returns the theme template to the application function layer of the application layer.

[0347] For example, assuming that the pictures in the highlight clip are mostly parent-child scenes, it can be confirmed that the theme template that meets the scene is a parent-child theme.

[0348] S43, the application function layer distributes the obtained theme and all highlight clips to the basic capability layer.

[0349] The application function layer distributes the theme obtained from S42 and all highlight clips obtained from S39 to the basic capability layer.

[0350] S44, the basic capability layer generates a target video set according to the theme and all highlight clips.

[0351] That is, the target video set is a video set that meets the recommended theme and is generated according to all screened highlight clips.

[0352] S45, the basic capability layer sends an instruction message to display the target video set to the video editing service layer.

[0353] S46, the video editing service layer displays the video set in the gallery interface.

[0354] In the embodiments of the present application, the media center framework layer is used to decode video and picture files, convert data format to a unified format, monitor remaining time, adjust operation strategy, send data, control algorithm running and termination, obtain results and return to the application layer, and the like. The FKW layer is used to complete data packaging and provide data and program running services. After accepting the command sent by the media center framework layer, the HAL layer performs highlight analysis according to the command, and returns the parameter calculation result of the highlight analysis to the media center framework layer. The final result of the algorithm is collected and sorted by the media center framework layer, and then sent to the application layer for processing. The application layer can present a clip application interface, video and picture file options, and the final result of the presentation algorithm.

[0355] After the user starts the one-key video generation function, selects the video and picture files to be clipped (for example, a maximum of 30 files), and waits for a moment, the one-key video generation application automatically clips the highlight segment of the video, combines the highlight segment and the pictures according to the algorithm result, generates a clipped video set, and plays a preview.

[0356] The video processing method provided in the embodiments of the present application can first perform overview analysis on a first quantity of image frames in each video in a plurality of image materials selected by a user, to obtain a first score of the image frames in each video. Then, the electronic device can determine a target region of the video based on a first image frame with the highest score in the video and image frames within a first preset time period before and after the first image frame. After that, the electronic device can perform analysis on a second quantity of images in the target region of the video, to obtain a second score of the image frames. Based on a second image frame with the highest second score and image frames within a second preset time period before and after the second image frame, the highlight segment of the video is determined. With the present solution, the electronic device first performs overview analysis on the videos, so that the target region that needs to be further processed can be located. When the electronic device extracts the highlight segment, only the second quantity of image frames in the target region are analyzed, instead of frame-by-frame analysis on the entire video, so that the workload of the electronic device for image frame analysis can be reduced, thereby the time consumption of the electronic device for implementing the one-key video generation function can be reduced, the efficiency of the electronic device for processing videos can be improved, and the user experience of the one-key video generation function can be improved.

[0357] It should also be noted that in the embodiments of the present application, "greater than" can be replaced by "greater than or equal to", "less than or equal to" can be replaced by "less than", or "greater than or equal to" can be replaced by "greater than", and "less than" can be replaced by "less than or equal to".

[0358] The various embodiments described herein can be independent solutions or combined according to inherent logic, and all fall within the protection scope of the present application.

[0359] It can be understood that the methods and operations realized by the electronic device in each of the above method embodiments can also be realized by components (such as chips or circuits) that can be used in the electronic device.

[0360] It should be noted that the personal information used in the technical solutions of the present application is limited to information that has obtained individual consent of the person, including but not limited to informing and reminding the user to read the relevant user agreement (notification) before the user uses the function, and signing the agreement (authorization) including authorization of relevant user information. Among them, the personal information includes the information such as pictures and videos stored by the user.

[0361] In the technical solutions disclosed in the present application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0362] The method embodiments provided by the present application are described above, and the device embodiments provided by the present application will be described below. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments, and therefore, the content not described in detail can be referred to the method embodiments described above, and for brevity, will not be described here.

[0363] The above mainly describes the solutions provided by the embodiments of the present application from the perspective of method steps. It can be understood that in order to realize the above functions, the electronic device implementing the method contains the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solutions. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the protection scope of the present application.

[0364] The embodiments of the present application can divide the functional modules of the electronic device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated in one processing module. The integrated module can be realized in the form of hardware or software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical functional division. When actually implemented, there can be other feasible division manners. The following will be described taking the division of each functional module corresponding to each function as an example.

[0365] The application further provides a chip coupled with the memory, the chip being configured to read and execute the computer program or instructions stored in the memory to perform the method in each of the embodiments.

[0366] The application further provides an electronic device comprising a chip configured to read and execute the computer program or instructions stored in the memory so that the method in each of the embodiments is performed.

[0367] The application further provides a computer readable storage medium comprising computer instructions, when the computer instructions are run on the electronic device, the electronic device is caused to perform each function or step performed by the electronic device in the method embodiments.

[0368] The application further provides a computer program product, when the computer program product is run on a computer, the computer is caused to perform each function or step performed by the electronic device in the method embodiments. For example, the computer can be the electronic device.

[0369] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0370] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0371] The units described as separate components can or can not be physically separated, and the components shown as units can be one physical unit or multiple physical units, that is, can be located in one place, or can be distributed to multiple different places. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.

[0372] In addition, each of the function units in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software function unit. When the integrated unit is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present application essentially, or the part that contributes to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, including a number of instructions to enable an apparatus (which can be a single chip, a chip, etc.) or a processor to perform all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store program codes.

[0373] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of video processing, the method comprising: The method comprises: receiving a selection operation of a user on a plurality of image materials in a gallery; wherein the plurality of image materials comprises videos; in response to the selection operation, performing overview analysis on each video in the plurality of image materials for a first number of image frames to obtain a first score; wherein the first score is an aesthetic score of a corresponding image frame; the first number corresponds to a time length of a video, and the first number of each video is less than a number of image frames of the corresponding video; based on the first score of the image frames in each video, determining a target region of the corresponding video; wherein the target region comprises a first image frame and image frames within a first preset time length before and after the first image frame; the first image frame is the image frame with the highest first score; performing analysis on the target region of each video in the plurality of image materials for a second number of image frames to obtain a second score; wherein the second score is an aesthetic score of a corresponding image frame; the second number is less than or equal to a number of image frames of the target region; based on the second score of the image frames in the target region of each video, determining a highlight segment of the corresponding video; wherein the highlight segment comprises a second image frame and image frames within a second preset time length before and after the second image frame; the second image frame is the image frame with the highest second score in the target region; the highlight segments of the videos in the plurality of image materials are used to splice to obtain a target video set.

2. The method of claim 1, wherein, The method further comprises: in response to the selection operation, obtaining analysis parameters corresponding to the plurality of image materials; wherein the analysis parameters comprise an actual time length of the corresponding video and an upper limit suggestion value of an analysis total time length, the upper limit suggestion value of the analysis total time length representing a suggestion value of a maximum time length required for analyzing the plurality of image materials; determining the first number of image frames of each video in the plurality of image materials according to the actual time length of each video, the number of videos in the plurality of image materials, and the upper limit suggestion value of the analysis total time length.

3. The method of claim 2, wherein, The determination of the first number of image frames of each video according to the actual time length of each video, the number of videos in the plurality of image materials, and the upper limit suggestion value of the analysis total time length comprises: based on the actual time length of each video and a preset corresponding relationship, determining a basic number and a maximum number of the corresponding video; wherein the basic number is a minimum number of image frames required for ensuring the analysis effect of the video, and the maximum number is a maximum number of image frames allowed to be analyzed for the video length; the preset corresponding relationship represents the maximum number and the basic number of image frames corresponding to different threshold ranges of video time lengths; based on the basic number and the maximum number of each video and the upper limit suggestion value of the analysis total time length, determining an analysis total number; the analysis total number is a total number of image frames allowed to be analyzed for all videos in the plurality of image materials; distributing the analysis total number to each video in the plurality of image materials according to the actual time length of each video and the number of videos in the plurality of image materials to obtain the first number of image frames of each video.

4. The method of claim 3, wherein, The analysis parameter further comprises a single-frame analysis duration; the single-frame analysis duration is a duration required for analyzing one image frame; The analysis total quantity is determined based on the basis quantity and the maximum quantity of each video and the upper limit of the analysis total duration suggestion value, comprising: If the upper limit of the analysis total duration suggestion value is less than a time consumption of analyzing image frames of the basis quantity of all videos in the multiple image materials, the analysis total quantity is a ratio of the upper limit of the analysis total duration suggestion value and the single-frame analysis duration; If the upper limit of the analysis total duration suggestion value is greater than a time consumption of analyzing image frames of the basis quantity sum of all videos in the image materials, and a first multiple of the upper limit of the analysis total duration suggestion value is less than a time consumption of analyzing image frames of the basis quantity sum of all videos in the multiple image materials, the analysis total quantity is a sum of image frames of the basis quantity of all videos in the multiple image materials; If the first multiple of the upper limit of the analysis total duration suggestion value is greater than a time consumption of analyzing image frames of the basis quantity sum of all videos in the multiple image materials, and the first multiple of the upper limit of the analysis total duration suggestion value is less than a time consumption of analyzing image frames of the maximum quantity sum of all videos in the multiple image materials, the analysis total quantity is a ratio of the first multiple of the upper limit of the analysis total duration suggestion value and the single-frame analysis duration; If the first multiple of the upper limit of the analysis total duration suggestion value is greater than a time consumption of analyzing image frames of the maximum quantity sum of all videos in the multiple image materials, the analysis total quantity is a maximum quantity sum of image frames of all videos in the multiple image materials; Wherein, the first multiple is greater than 0 and less than 1.

5. The method according to claim 3 or 4, characterized in that, The first quantity of image frames of each video is obtained by distributing the analysis total quantity to each video according to the actual duration of each video in the multiple image materials and the video quantity in the multiple image materials, comprising: Each video is traversed, and a first value and a second value of each video are updated until the first value is 0; wherein, an initial value of the first value is equal to the analysis total quantity, and an initial value of the second value is 0; each time a video is traversed, the second value of the video is added by 1, and the first value is reduced by 1; The second value of each video is taken as the first quantity of image frames of the video.

6. The method of claim 5, wherein, Each video is traversed, and a first value and a second value of each video are updated until the first value is 0, comprising: Before traversing a first video in the multiple image materials, if the second value of the first video is equal to the maximum quantity of image frames of the first video, the first video is skipped, and a next video of the first video is traversed; Wherein, skipping the first video means that the second value of the first video is not added by 1.

7. The method according to any one of claims 1-4, characterized in that, The first quantity of image frames for overview analysis is uniformly distributed in each position of the video.

8. The method according to any one of claims 2-4, characterized in that, The analysis parameter corresponding to the multiple image materials comprises a highlight segment suggestion duration, and the highlight segment suggestion duration is a suggestion duration of one highlight segment in one video. The target region of each video is determined based on the first score of the image frames in the video, including: For each video, the first image frame is obtained according to the first score of the image frames in the video; A segment in the video containing a highlight segment proposal duration of the first image frame is determined as a candidate highlight segment of the video; The candidate highlight segment and the image frames within a third preset duration before and after the candidate highlight segment are determined as the target region of the video; the third preset duration is less than the first preset duration.

9. The method of claim 8, wherein, The analysis parameters include a total highlight segment proposal duration, which is the sum of the proposal durations of the highlight segments of all videos in the plurality of image materials; The method further includes: The sum of the durations of the candidate highlight segments of all videos in the plurality of image materials is calculated; If the sum of the durations of the candidate highlight segments of all videos is greater than a second multiple of the total highlight segment proposal duration, the durations of the candidate highlight segments of each video are adjusted according to the actual durations of the videos in the plurality of image materials, so that the sum of the durations of the candidate highlight segments of all videos is less than or equal to the second multiple of the total highlight segment proposal duration; the second multiple is greater than 1 and less than 2.

10. The method of any one of claims 2-4, wherein, The analysis parameters corresponding to the plurality of image materials include a remaining analysis duration and a single-frame analysis duration; the remaining analysis duration is equal to the duration for video analysis minus the total duration spent on performing the overview analysis; the single-frame analysis duration is the duration required for analyzing one image frame; Before the analysis of the target region of each video in the plurality of image materials is performed for a second number of image frames to obtain a second score, the method includes: The ratio of the remaining analysis duration to the single-frame analysis duration is taken as the analyzable number of image frames of the target region of all videos in the plurality of image materials; The analyzable number is distributed to each video according to the duration of the target region of each video, to obtain a second number of image frames of the target region of each video.

11. The method of claim 10, wherein, The distribution of the analyzable number to each video according to the duration of the target region of each video to obtain a second number of image frames of the target region of each video includes: Each video in the plurality of image materials is traversed, and a third value and a fourth value of each video are updated until the third value is 0; the initial value of the third value is equal to the analyzable number, and the initial value of the fourth value is 0; each time a video is traversed, the fourth value of the video is incremented by 1, and the third value is decremented by 1; The fourth value of each video is taken as the second number of image frames of the target region of the video.

12. The method of any one of claims 1-4, wherein, The second number of image frames for analysis is uniformly distributed at various positions of the target region of the video.

13. The method of any one of claims 2-4, wherein, The analysis parameters corresponding to the plurality of image materials include a highlight segment proposal duration, which is the proposal duration of a highlight segment in a video; The highlight segment of each video is determined based on the second score of the image frames in the target region of the video, including: For each video, a second image frame is obtained according to a second score of image frames in the video; The second image frame and image frames within the second preset time length before and after the second image frame are determined as a highlight segment of the video; and a time length of the highlight segment is less than or equal to the highlight segment suggestion time length.

14. The method of any one of claims 2-4, wherein, The analysis parameters corresponding to the plurality of image materials include an analysis total time length and an analysis total time length upper limit suggestion value; the analysis total time length represents a total time length required for analyzing the plurality of image materials; and the analysis total time length upper limit suggestion value represents a suggestion value of a maximum time length required for analyzing the plurality of image materials; Before the overview analysis of each video in the plurality of image materials is performed on the first number of image frames to obtain the first score, the method further includes: If the analysis total time length is greater than the analysis total time length upper limit suggestion value, M videos are randomly selected from the plurality of image materials; wherein a time length required for analyzing the M videos is less than or equal to the analysis total time length, and M is less than a number of videos in the image materials.

15. The method of any one of claims 1-4, wherein, The overview analysis of each video in the plurality of image materials is performed on the first number of image frames to obtain the first score, including: If a sum of actual time lengths of all videos in the plurality of image materials is greater than a preset time length threshold, the overview analysis of each video in the plurality of image materials is performed on the first number of image frames to obtain the first score.

16. An electronic device, comprising: The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processor are coupled; the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the method of any one of claims 1-15.

17. A computer-readable storage medium, characterized in that, The computer program code includes computer instructions, and when the computer instructions are executed on the electronic device, the electronic device executes the method of any one of claims 1-15.

18. A computer program product comprising computer programs / instructions, characterized in that, The computer program / code is executed by the processor to implement the method of any one of claims 1-15.

Citation Information

Patent Citations

  • Target image generation method, target image generation device, medium and electronic equipment

    CN111464833A

  • Image capture device and method for image score-based video quality enhancement

    US20200036889A1