Video processing method, electronic equipment and storage medium

By performing keyframe overview analysis of the video and frame-by-frame analysis of the target area, the problem of low video processing efficiency under long materials is solved, and the processing speed and user experience of the one-click film function is improved.

CN120343182AActive Publication Date: 2025-07-18HONOR DEVICE CO LTD

Patent Information

Application Number
CN202410040101.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-18
Estimated Expiration
2044-01-10

AI Technical Summary

Technical Problem

When the material selected by the user is long, the efficiency of electronic devices to analyze videos is low, resulting in too long time-consuming and affecting the user experience.

Method used

By performing partial keyframe overview analysis on each video in the image material, the target area is determined, and only the target area is analyzed frame by frame, highlighted fragments are extracted, and the overall video analysis workload is reduced.

Benefits of technology

It improves the efficiency of electronic devices to process videos, reduces the time-consuming of one-click filming function, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343182A_ABST
    Figure CN120343182A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method, electronic equipment and a storage medium, and relates to the technical field of video data, and the method comprises the steps that the electronic equipment carries out the overview analysis of a first number of key frames on each video in a plurality of image materials, and obtains a first score of the image frame; determining a target area of the corresponding video based on the first score of the key frame in each video; analyzing the target area of each video frame by frame to obtain a second score of each frame; and determining a highlight segment of the corresponding video based on the second score of each frame in the target area in each video. In the scheme, the electronic equipment carries out overview analysis and frame-by-frame analysis on each video instead of carrying out frame-by-frame analysis on the whole video, so that the workload of the electronic equipment for carrying out image frame analysis is reduced, the time consumption of the electronic equipment for realizing a one-key filming function can be reduced, the video processing efficiency of the electronic equipment is improved, and the user experience is improved. And thus, the user experience of the one-key film-forming function is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of video data, and in particular, to a video processing method, an electronic device, and a storage medium. Background Art

[0002] With the development of image and video processing technologies, users can trigger an electronic device to further process photos and videos in an album. For example, the electronic device can splice multiple image materials (including multiple pictures and / or multiple videos) to obtain a complete spliced video.

[0003] For example, the electronic device or a third-party video processing software in the electronic device may have a function of generating a video with one click (or referred to as a service of generating a video with one click). The function of generating a video with one click can automatically analyze and extract highlight segments of videos in the selected multiple image materials by an algorithm; then, automatically generate a clipped video based on the extracted highlight segments. Among them, the above-mentioned highlight segments are also called wonderful segments, which refer to video segments composed of single-frame images or consecutive multiple-frame images extracted from the above materials and used to record wonderful moments. The wonderful moments can be the moments when wonderful actions such as a person's smiling face, a moment of winning a championship, an airplane landing, etc. occur.

[0004] However, when the overall duration of the selected materials by the user is relatively long, it takes a lot of time for the electronic device to analyze the selected multiple materials, and the efficiency of the electronic device in processing videos is low. Summary of the Invention

[0005] Embodiments of the present application provide a video processing method, an electronic device, and a storage medium, which avoid the problems of large time consumption and low video processing performance caused by analyzing a complete video by performing partial key-frame analysis on each video in the image materials.

[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions.

[0007] In a first aspect, a video processing method is provided, and the method includes:

[0008] The electronic device receives a selection operation of the user for multiple image materials in the gallery. Among them, the multiple image materials include videos; or, the multiple image materials may further include videos and pictures.

[0009] The electronic device, in response to the selection operation, performs an overview analysis of the first number of I-frames for each video in the multiple image materials to obtain a first score. The first score is an aesthetic score for the corresponding image frame, the first number corresponds to the duration of the video; the I-frame is an image frame including complete image information.

[0010] The electronic device determines the target region of the corresponding video based on the first score of the I-frame in each video; the target region includes the first I-frame, and the I-frames within the first preset duration before and after the first I-frame; the first I-frame is the I-frame with the highest first score. Frame-by-frame analysis is performed on the target region of each video in multiple image materials to obtain the second score of each image frame in the target region. Among them, the frame-by-frame analysis includes I-frame-by-frame analysis or frame-by-frame analysis of each image frame, and the second score is the aesthetic score of the corresponding image frame.

[0011] The electronic device determines the highlight segment of the corresponding video based on the second score corresponding to each video; wherein, the highlight segment includes the second image frame, and the frames within the second preset duration before and after the second image frame.

[0012] Among them, the second image frame is the image frame with the highest second score in the target region; the highlight segments of each video in multiple image materials are used to splice to obtain the target video set.

[0013] In this application, the electronic device first performs an overview analysis on the videos therein, so as to locate the target region that needs to be processed frame by frame. When the electronic device extracts the highlight segment, it only performs frame-by-frame analysis on the target region, rather than on the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time-consuming of the electronic device to implement the one-click video creation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-click video creation function.

[0014] In another possible implementation manner of the first aspect, performing an overview analysis on the first number of I-frames of each video in multiple image materials to obtain the first score includes:

[0015] Obtain the first position of the analyzed I-frame in the video;

[0016] Based on the first position of the analyzed I-frame and the first score of the analyzed I-frame, determine the second position of the I-frame to be analyzed in the video, and obtain the first score of the I-frame at the second position until the number of analyzed I-frames reaches the first number.

[0017] In this application, the electronic device can determine the first number of image frames that can be analyzed for each video within the recommended value of the total analysis duration limit. Under the limited performance of the electronic device, through the analysis of the first number of image frames, while improving the analysis efficiency of the highlight segment, a relatively reliable analysis effect can also be obtained.

[0018] In another possible implementation manner of the first aspect, the first position includes the first preset position and the second preset position of the video.

[0019] Determining a second position of an I-frame to be analyzed in a video based on a first position of an analyzed I-frame and a first score of the analyzed I-frame includes:

[0020] Regarding the area between the midpoint position between a first preset position and a second preset position and the start position of the video as a first area; the first area includes the I-frames at the first position, and the scores of the I-frames in the first area are used as the first scores of the I-frames at the first preset position.

[0021] Regarding the area between the midpoint position and the end position of the video as a second area; the second area includes the I-frames at the second preset position, and the scores of the second area are used as the first scores of the I-frames at the second preset position.

[0022] Obtaining the product of the score of each area and the duration of the area, and regarding the midpoint position of the area with the highest product as the second position.

[0023] In this application, based on the score of the analyzed I-frame and the first position of the I-frame, a region with a relatively high score can be determined from the video, so as to determine the second position of the I-frame to be analyzed from this region with a relatively high score, improving the quality of the I-frame to be analyzed and indirectly improving the effectiveness of the analysis result of the I-frame.

[0024] In another possible implementation manner of the first aspect, the first preset position is the 1 / 3 position of the video duration, and the second preset position is the 2 / 3 position of the video duration.

[0025] In this application, regarding the 1 / 3 position of the video duration as the first preset position and the 2 / 3 position of the video duration as the second preset position can directly and effectively determine the quality of the area from the start position to the midpoint position of the video and the quality of the area from the midpoint position to the end position of the video, so as to further determine the second position of the I-frame to be analyzed. This avoids the problem of low effectiveness of the analysis result caused by randomly selecting the first position.

[0026] In another possible implementation manner of the first aspect, the first position is a non-preset position. Then, the first position is the position of the I-frame that is not the first one sent for analysis in the video, and the first position is the position of an I-frame during the I-frame analysis process of the video.

[0027] Determining a second position of an I-frame to be analyzed in a video based on a first position of an analyzed I-frame and a first score of the analyzed I-frame includes:

[0028] Determining a third area including the I-frame at the first position according to the adjacent boundaries before and after the I-frame at the first position.

[0029] Among them, the adjacent boundaries include the position of the analyzed I-frame adjacent to the I-frame at the first position, the starting position of the video, the ending position of the video, and one of the boundaries of the connected region. The connected region is a region with the same score and continuous in time, and there is no region coverage between connected regions. The third region is one of the connected regions of the video.

[0030] Obtain multiple connected regions of the video and calculate the scores of each connected region; the score of the connected region is the average of the first score of the included I-frames and the score of the connected region. Obtain the product of each region score and the region duration, and use the midpoint position of the connected region with the highest product as the second position.

[0031] In this application, during the I-frame overview analysis of the video, the connected regions in the video can be dynamically determined according to the scores of the analyzed I-frames and the first positions of the I-frames. Determine the second position of the I-frame to be analyzed from the connected region with the largest product of the region score and the region duration. The score of the I-frame at this second position is not too low, which improves the quality of the I-frames to be analyzed from the side and further improves the effectiveness of the analysis results of the I-frames.

[0032] In another possible implementation manner of the first aspect, determining the third region corresponding to the I-frame at the first position according to the adjacent boundaries before and after the I-frame at the first position includes:

[0033] Determine the first boundary of the third region according to the first score of the I-frame at the first position and the corresponding first weight, and the score and the corresponding second weight of the adjacent boundary before the I-frame at the first position; determine the second boundary of the third region according to the first score and the first weight of the I-frame at the first position, and the score and the corresponding third weight of the adjacent boundary after the I-frame at the first position. Among them, the score and the weight have a corresponding relationship.

[0034] In this application, determining the boundaries of the third region of the I-frame according to the I-frame score and weight, and the scores and weights of the adjacent boundaries of the I-frame makes the obtained third region better represent the scoring situation of the I-frame. During the subsequent process of determining the second position of the I-frame to be analyzed based on the connected region, the scores of each region are more accurate, and the I-frames to be analyzed determined are more effective. The target region determined based on the I-frame is more accurate, and thus the highlight segment determined by frame-by-frame analysis based on the target region has a better effect.

[0035] In another possible implementation manner of the first aspect, determining the target region of the corresponding video based on the first score of the I-frame in each video includes:

[0036] For each video, obtain a first I-frame based on the first score of the I-frame in the video. Determine a candidate highlight segment of the video for a segment that includes the first I-frame in the video and has a duration equal to the recommended duration of the highlight segment. Determine a target region of the video for the candidate highlight segment and the I-frames within a third preset duration before and after the candidate highlight segment; the third preset duration is less than the first preset duration.

[0037] In this application, the electronic device can determine a candidate highlight segment that meets the recommended duration of the highlight segment based on the first image frame with the highest first score and the recommended duration of the highlight segment, and thus determine a target region that includes the candidate highlight segment. Since the target region includes the first image frame, the target region is worthy of further analysis to determine the highlight segment of the video. In this way, the determined highlight segment is relatively accurate.

[0038] In another possible implementation of the first aspect, the method further includes:

[0039] In response to a selection operation, obtain analysis parameters corresponding to a plurality of image materials; wherein, the analysis parameters include the actual duration of the corresponding video and a recommended value of the upper limit of the total analysis duration, and the recommended value of the upper limit of the total analysis duration represents a recommended value of the maximum duration required to analyze the plurality of image materials.

[0040] Determine a first quantity of I-frames for each video according to the actual duration of each video in the plurality of image materials, the number of videos in the plurality of image materials, and the recommended value of the upper limit of the total analysis duration.

[0041] In this application, the electronic device can determine a first quantity of image frames that can be analyzed for each video within the recommended value of the upper limit of the total analysis duration. Under the limited performance of the electronic device, by analyzing the first quantity of image frames, while improving the analysis efficiency of the highlight segment, a relatively reliable analysis effect can also be obtained.

[0042] In another possible implementation of the first aspect, determining a first quantity of I-frames for each video according to the actual duration of each video in the plurality of image materials, the number of videos in the plurality of image materials, and the recommended value of the upper limit of the total analysis duration includes:

[0043] Based on the actual duration of each video and a preset correspondence, determine a base quantity and a maximum quantity for the corresponding video; wherein, the base quantity is the minimum quantity of I-frames that need to be analyzed to ensure the analysis effect of the video, and the maximum quantity is the maximum quantity of I-frames that the duration allows for analyzing the video; the preset correspondence represents the maximum quantity and the base quantity of I-frames corresponding to different threshold ranges of the video duration.

[0044] Based on the base quantity and the maximum quantity of each video, as well as the recommended value of the upper limit of the total analysis duration, determine the total analysis quantity; the total analysis quantity is the total quantity of I-frames allowed to be analyzed for all videos in multiple image materials.

[0045] According to the actual duration of each video in multiple image materials and the number of videos in multiple image materials, distribute the total analysis quantity to each video to obtain the first quantity of I-frames for each video.

[0046] In this application, the first quantity is determined by the maximum quantity and the base quantity of the image frames of each video, and the image frames of the first quantity are analyzed, which can not only ensure the analysis effect of the video, but also efficiently analyze and process the highlight segments under the limited performance and limited time consumption of the electronic device.

[0047] In another possible implementation manner of the first aspect, the analysis parameter further includes the single-frame analysis duration; the single-frame analysis duration is the duration required to analyze one image frame.

[0048] Based on the base quantity and the maximum quantity of each video, as well as the recommended value of the upper limit of the total analysis duration, determining the total analysis quantity includes:

[0049] If the recommended value of the upper limit of the total analysis duration is less than the time consumption for analyzing the sum of the base quantities of I-frames of all videos in multiple image materials, the total analysis quantity is the ratio of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration; if the recommended value of the upper limit of the total analysis duration is greater than the time consumption for analyzing the sum of the base quantities of I-frames of all videos in the image materials, and the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumption for analyzing the sum of the base quantities of I-frames of all videos in multiple image materials, the total analysis quantity is the sum of the base quantities of I-frames of all videos in multiple image materials.

[0050] If the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is greater than the time consumption for analyzing the sum of the base quantities of I-frames of all videos in multiple image materials, and the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumption for analyzing the sum of the maximum quantities of I-frames of all videos in multiple image materials, the total analysis quantity is the ratio of the duration of the second multiple of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration; if the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is greater than the time consumption for analyzing the sum of the maximum quantities of I-frames of all videos in multiple image materials, the total analysis quantity is the sum of the maximum quantities of I-frames of all videos in multiple image materials.

[0051] Wherein, the second multiple is greater than 0 and less than 1.

[0052] In this application, the first quantity is determined based on the maximum quantity and the base quantity of the image frames of each video, and the image frames of the first quantity are analyzed, which can not only ensure the analysis effect of the video, but also perform efficient analysis and processing of high-light segments under the limited performance and limited time consumption of the electronic device.

[0053] In another possible implementation manner of the first aspect, according to the actual duration of each video in multiple image materials and the number of videos in the multiple image materials, the total analysis quantity is allocated to each video to obtain the first quantity of I frames of each video, including:

[0054] Traverse each video in the multiple image materials, and update the first value and the second value of each video until the first value is 0; wherein, the initial value of the first value is equal to the total analysis quantity, and the initial value of the second value is 0.

[0055] Wherein, for each video traversed, the second value of the video is incremented by 1, and the first value is decremented by 1.

[0056] Use the second value of each video as the first quantity of I frames of the video.

[0057] In this application, traversing each video to determine the first quantity of each video can evenly allocate the total analysis quantity to each video, so as not to miss the high-light segments of any video.

[0058] In another possible implementation manner of the first aspect, traversing each video in the multiple image materials and updating the first value and the second value of each video until the first value is 0 includes:

[0059] Before traversing to the first video in the multiple image materials, if the second value of the first video is equal to the maximum quantity of I frames of the first video, skip the first video and traverse the next video of the first video. Wherein, skipping the first video means that the second value of the first video is not incremented by 1.

[0060] In this application, if the number of image frames of a video has reached the maximum quantity, no more image frames will be allocated to this video subsequently, and the image frames will be allocated to other videos that have not reached the maximum quantity, which can improve the effectiveness of video analysis.

[0061] In another possible implementation manner of the first aspect, perform frame-by-frame analysis on the target area of each video in the multiple image materials to obtain the second score of each image frame in the target area, including:

[0062] Based on the preset single-frame analysis duration and the number of all I-frames in the target regions of all videos in the image material, obtain the sum of the first durations for analyzing the I-frames. Here, the single-frame analysis duration is the duration required to analyze one image frame. Based on the single-frame analysis duration and the number of all image frames in the target regions of all videos in the image material, obtain the sum of the second durations for analyzing the image frames.

[0063] If the remaining analysis duration is less than the sum of the first durations, perform frame-by-frame analysis on the I-frames in the target regions of each video in multiple image materials to obtain the second score of the I-frames in the target regions; the remaining analysis duration is equal to the duration for video analysis minus the total duration spent on performing the overview analysis.

[0064] If the remaining analysis duration is greater than or equal to the sum of the first durations, or if the remaining analysis duration is greater than the sum of the second durations, perform frame-by-frame analysis on the image frames in the target regions of each video in multiple image materials to obtain the second score of the image frames in the target regions.

[0065] In this application, based on the analysis time of all I-frames in the target region, the analysis time of all image frames in the target region, and the remaining analysis duration, determine the frame-by-frame analysis strategy for the target region, which can analyze a limited number of image frames in the target region within the remaining analysis duration, thereby improving the analysis efficiency of highlight segments while achieving a relatively reliable analysis effect under the limited performance of the electronic device.

[0066] In another possible implementation manner of the first aspect, the method further includes:

[0067] Arrange the target regions in descending order of their scores, and allocate the remaining analysis duration to each target region according to the time taken to analyze all image frames in the target region until the remaining analysis duration is fully allocated. If there are target regions that have not been allocated analysis duration, the target regions that have not been allocated analysis duration will not perform frame-by-frame analysis.

[0068] In this application, if there are target regions that have not been allocated analysis duration, that is, the remaining analysis duration is not sufficient to analyze all target regions, then the target regions that have not been allocated analysis duration will not perform frame-by-frame analysis. Among them, the target regions that have not been allocated analysis duration are determined based on the score ranking. The target regions with lower scores have lower value, and the quality of their corresponding highlight segments is also poor, so the analysis of them is abandoned, which can further reduce the duration consumption of the target regions. Or, allocate more duration to the target regions with higher scores to obtain highlight segments with higher value, which can ensure the effectiveness of the highlight segments.

[0069] In another possible implementation manner of the first aspect, based on the second score corresponding to each video, determine the highlight segment of the corresponding video, including:

[0070] For each video, obtain a second image frame according to the second score of the image frames in the target area;

[0071] Determine the second image frame and the image frames within a second preset time period before and after the second image frame as the highlight segment of the video; the duration of the highlight segment is less than or equal to the recommended duration of the highlight segment.

[0072] In this application, the electronic device can determine a highlight segment that meets the recommended duration of the highlight segment based on the second image frame with the highest second score and the recommended duration of the highlight segment. The highlight segment determined based on the second image frame and the recommended duration of the highlight segment is relatively accurate.

[0073] In another possible implementation manner of the first aspect, before obtaining the first score by performing an overview analysis of the first number of I-frames on each video in multiple image materials, the method further includes:

[0074] If the total analysis duration of multiple image materials is greater than the recommended value of the upper limit of the preset total analysis duration, randomly select M videos from the multiple image materials. Wherein, the total analysis duration represents the total duration required to analyze multiple image materials, and the recommended value of the upper limit of the preset total analysis duration represents the recommended value of the maximum duration required to analyze multiple image materials; the duration required to analyze M videos is less than or equal to the total analysis duration, and M is less than the number of videos in the image materials.

[0075] In this application, when the total analysis duration of multiple image materials is greater than the recommended value of the upper limit of the total analysis duration, some videos can be randomly selected from the multiple image materials for analysis. With the limited performance and limited time consumption of the electronic device, efficient analysis and processing of the highlight segment can be achieved.

[0076] In another possible implementation manner of the first aspect, performing an overview analysis of the first number of I-frames on each video in multiple image materials to obtain the first score includes:

[0077] If the sum of the actual durations of all videos in multiple image materials is greater than the preset duration threshold, perform an overview analysis of the first number of I-frames on each video in multiple image materials to obtain the first score.

[0078] In this application, when the sum of the actual durations of the videos in multiple image materials is greater than the preset duration threshold, adopting this solution can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumption of the electronic device to implement the one-click video compilation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-click video compilation function.

[0079] In a second aspect, an electronic device is provided. The electronic device includes a memory, a display screen, and one or more processors; the memory, the display screen are coupled to the processor; computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method according to any one of the above-mentioned first aspects.

[0080] In a third aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on an electronic device, the electronic device can be caused to execute the method according to any one of the above-mentioned first aspects.

[0081] In a fourth aspect, a computer program product including instructions is provided. When it runs on an electronic device, the electronic device can be caused to execute the method according to any one of the above-mentioned first aspects.

[0082] In a fifth aspect, an embodiment of the present application provides a chip. The chip includes a processor, and the processor is used to call a computer program in a memory to execute the method according to the first aspect.

[0083] It can be understood that for the beneficial effects that can be achieved by the electronic device described in the above-mentioned second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect, reference may be made to the beneficial effects in the first aspect and any one of its possible design manners, which will not be elaborated here. Description of the Drawings

[0084] Figure 1 It is a schematic diagram of an application scenario of a video processing method provided by an embodiment of the present application;

[0085] Figure 2 It is a schematic diagram of another application scenario of a video processing method provided by an embodiment of the present application;

[0086] Figure 3 It is a schematic diagram of another application scenario of a video processing method provided by an embodiment of the present application;

[0087] Figure 4 It is a schematic diagram of another application scenario of a video processing method provided by an embodiment of the present application;

[0088] Figure 5 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;

[0089] Figure 6 It is a schematic diagram of the software structure of an electronic device provided by an embodiment of the present application;

[0090] Figure 7Schematic diagram of the method for initializing algorithms, obtaining analysis parameters, etc. for each module in the electronic device provided by the embodiments of the present application;

[0091] Figure 8 Schematic diagram of the process of picture analysis in a video processing method provided by the embodiments of the present application;

[0092] Figure 9 Schematic diagram of the technical idea of a video processing method provided by the embodiments of the present application;

[0093] Figure 10 Schematic diagram of the process of video analysis in a video processing method provided by the embodiments of the present application;

[0094] Figure 11 Schematic diagram of extracting I-frames provided by the embodiments of the present application;

[0095] Figure 12 Schematic diagram of another method for extracting I-frames provided by the embodiments of the present application;

[0096] Figure 13 Schematic diagram of allocating I-frames for multiple videos provided by the embodiments of the present application;

[0097] Figure 14 Schematic diagram of the first preset position and the second preset position in a 15-second video provided by the embodiments of the present application;

[0098] Figure 15 Schematic diagram of a first position provided by the embodiments of the present application;

[0099] Figure 16 Schematic diagram of area 3 provided by the embodiments of the present application;

[0100] Figure 17 Schematic diagram of the connected area of a video provided by the embodiments of the present application;

[0101] Figure 18 Schematic diagram of a target area provided by the embodiments of the present application;

[0102] Figure 19 Schematic diagram of another target area provided by the embodiments of the present application;

[0103] Figure 20 Schematic diagram of a high-light segment provided by the embodiments of the present application;

[0104] Figure 21 Schematic diagram of another high-light segment provided by the embodiments of the present application;

[0105] Figure 22 Schematic diagram of the post-processing process for high-light segments in a video processing method provided by the embodiments of the present application. Detailed implementation manners

[0106] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "the", "above-mentioned", "this" and "this one" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship.

[0107] The reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure or characteristic described in combination with this embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.

[0108] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0109] First, some nouns or terms involved in the present application are explained.

[0110] Key frames, also known as intra-coded picture frames (I-frames). The characteristic of I-frames is that they carry a complete image information, and only this frame needs to be decoded during the decoding process to extract a complete picture. I-frames are usually the first frames of each group of pictures (GOP), and after being moderately compressed, they serve as reference points for random access. I-frames (key frames) are a type of image frame that includes complete image information.

[0111] Highlight segments, also known as exciting segments, refer to single-frame images or video segments composed of consecutive multiple frames extracted from videos or pictures for recording exciting moments. The exciting moments can be the moments when a person smiles, wins a championship in a competition, takes off in sports, lands an airplane, or scores a goal in a ball game corresponding to the exciting actions.

[0112] The one-key video compilation function means that in response to a user's selection operation on one or more image materials, the electronic device automatically analyzes the highlight segments in the image materials through an algorithm and combines the highlight segments into a compiled video set. That is to say, the electronic device can extract multiple highlight segments from one or more image materials and synthesize them into a video set. Among them, the image materials selected by the user can be pictures or videos; or the image materials can include pictures and videos.

[0113] Among them, the process of the electronic device selecting highlight segments from videos or images can include: obtaining aesthetic scoring parameters such as the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate values of each frame of the image, and performing aesthetic scoring on each frame of the image according to these aesthetic scoring parameters. The electronic device can use the single-frame image or consecutive multiple frames with the highest score as a highlight segment.

[0114] Currently, the one-key video compilation function supports the editing of materials such as pictures and videos. During the implementation of the one-key video compilation function, users may select a relatively large number of pictures or videos with a relatively long duration from the image materials. In this case, when the electronic device analyzes the selected image materials to extract highlight segments, it takes a long time, and the efficiency of the electronic device in processing videos is low. Moreover, the long time-consuming will cause users to wait too long for the one-key video compilation to output the compiled video, affecting the user experience of the one-key video compilation function.

[0115] In view of the above problems, an embodiment of the present application provides a video processing method. With this solution, in the process of implementing the one-click video generation function, the electronic device can first perform an overview analysis on the first number of key frames in each video among multiple image materials to obtain the first score of the key frames of each video. Then, the electronic device can determine the target area of the video based on the first key frame with the highest first score and the key frames within the first preset duration before and after the first key frame. After that, the electronic device can perform a frame-by-frame analysis on the target area in the video to obtain the second score of each image frame. Based on the second image frame with the highest second score and the image frames within the second preset duration before and after the second image frame, it is determined as the highlight segment of the video.

[0116] With this solution, the electronic device first performs an overview analysis on the videos therein, so that the target area that needs to be processed frame by frame can be located. When the electronic device extracts the highlight segment, it only performs a frame-by-frame analysis on the target area instead of the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time consumed by the electronic device to implement the one-click video generation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-click video generation function.

[0117] The video processing method provided by the embodiment of the present application can be applied to an electronic device with image processing functions. It should be noted that the image materials for one-click video generation in the embodiment of the present application can include pictures and videos. The user can use the one-click video generation function to generate a video set from the highlight segments of multiple pictures, or use the one-click video generation function to generate a video set from the highlight segments of multiple videos, or use the one-click video generation function to generate a video set from the highlight segments of multiple pictures and videos.

[0118] The above-mentioned electronic device can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city or wireless terminal in smart home, etc. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the electronic device.

[0119] Taking the electronic device as a mobile phone as an example below, in combination with Figures 1 - 4 Taking the electronic device as a mobile phone as an example, this paper introduces the application scenario and interface implementation for the electronic device to achieve the function of generating a ready-made video in one click.

[0120] In an application scenario, the user can use the mobile phone to pre-capture multiple image materials in advance, and the mobile phone stores the multiple image materials in the picture library. The mobile phone can also pre-obtain image materials transmitted from other devices. As Figure 1 shown in (a) in Figure 1 , the icon of the picture library is displayed on the desktop of the mobile phone. When the user wants to use the mobile phone to generate a clipped video set based on multiple image materials, the user can click the icon of the picture library on the mobile phone desktop. In response to the user's click operation on the icon, the mobile phone displays the picture library interface 101 as shown in (b) in Figure 1 . The picture library interface 101 includes an option of "One-click Blockbuster". After that, in response to the user's click operation on the "One-click Blockbuster" option shown in (b) in Figure 1 , the mobile phone can display the picture library interface 102 as shown in (c) in Figure 1 . The picture library interface 102 may include multiple recently captured image materials. The user can select any one or more of the multiple image materials shown in (c) in Figure 1 as candidate image materials for generating a ready-made video in one click. For example, in response to the user's selection operation on some of the image materials in (c) in Figure 1The gallery interface 103 shown in (d) thereof. The gallery interface 103 includes all image materials in the gallery. The gallery interface 103 may also include video generation options, such as the "√" check option.

[0121] The mobile phone responds to the user's Figure 1 click operation on the "√" check option shown in (d) thereof, and can analyze 5 image materials selected by the user in the gallery interface 103, select highlight segments from each image material, and generate a video set based on the selected highlight segments. During this process, the mobile phone can display Figure 2 the gallery interface 201 shown in. The gallery interface 201 includes the analysis material progress, so that the user can intuitively view the analysis progress.

[0122] In one example, after the mobile phone generates a video set, it can display Figure 3 the gallery interface 301 shown in (a) thereof. The generated video set can be displayed in the gallery interface 301. The mobile phone can automatically play the video set in the gallery interface 301. In addition, as Figure 3 shown in (a) thereof, the gallery interface 301 may also include a video export option 302 for supporting the export of the generated video set. In response to the user's click operation on the video export option 302, the mobile phone can save the video set in the gallery, so that the user can view the video set from the gallery. In response to the user's click operation on the video export option 302, the mobile phone can also display Figure 3 the video export interface 303 shown in (b) thereof.

[0123] In one example, during the process of the user selecting image materials, the mobile phone can prompt the user, so that the user can know how many image materials are more appropriate to select. For example, as Figure 1 shown in (d) thereof, the prompt message "Better effect with more than 6 image materials" is displayed in the gallery interface 103, so that the user can know at least how many image materials to select to generate a video set with better effect.

[0124] In one example, during the process of the user selecting image materials, the mobile phone can prompt the user, so that the user can know the maximum number of image materials that can be selected. For example, as Figure 4 shown, the prompt message "Up to 30 image materials can be selected" is displayed in the gallery interface 401, so that the user can know how many image materials can be selected.

[0125] In one example, after generating the video set, the mobile phone can also display other function options on the interface of the video set, so that the user can perform operations such as editing, adding special effects, and analyzing on the generated video set based on these function options. For example, as Figure 3As shown by 301 in [the figure], other function options may include, but are not limited to, options such as templates, music, clips, sharing, etc.

[0126] Taking the electronic device as a mobile phone as an example below, in combination with Figure 5 the hardware structure of the electronic device will be introduced.

[0127] Figure 5 FIG. shows a schematic diagram of the hardware structure of the electronic device 100 provided by an embodiment of the present application. As Figure 5 shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a camera 193, a display screen 194, etc.

[0128] The processor 110 may include one or more processing units. For example, the processor 110 may include a controller, an application processor (AP), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, the controller may be the nerve center and command center of the mobile phone 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions. A memory may also be provided in the processor 110 for storing instructions and data.

[0129] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use this instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0130] The wireless communication function of the electronic device 100 can be implemented by antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor, baseband processor, etc.

[0131] The electronic device 100 implements the display function through the GPU, display screen 194, application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and is used for graphics rendering, such as rendering Figures 1 - 4 the schematic diagram of the operation interface shown, etc.

[0132] The display screen 194 is used to display the operation interface, projection screen image, projection screen video, etc. of the projection screen APP. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), and a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 100 may include 1 or N display screens 194, and N is a positive integer greater than 1.

[0133] In this embodiment, the display screen 194 can be used to display Figure 1 the gallery interface 101, gallery interface 102, gallery interface 103 in; the display screen 194 can be used to display Figure 2 the gallery interface 201 in; the display screen 194 can be used to display Figure 3 the gallery interface 301 and video export interface 303 in; the display screen 194 can be used to display Figure 4 the gallery interface 401 in, etc.

[0134] The electronic device 100 can implement the shooting function through the ISP, camera 193, video codec, GPU, display screen 194, application processor, etc.

[0135] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.

[0136] The camera 193 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0137] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0138] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0139] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0140] The external memory interface 120 can be used to connect to an external memory card to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function.

[0141] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs (APPs) required for at least one function (such as a camera APP, a gallery APP, and third-party video editing software, etc.). The data storage area can store data created during the use of the mobile phone 100 (such as photos or videos taken, mobile phone screenshots, mobile phone screen recording content, images downloaded from other devices, and video sets generated using the one-click video generation function, etc.). In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, and a universal flash storage (UFS), etc.

[0142] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.

[0143] For example, after a video set is generated using the one-click video generation function, the audio module 170 decodes the audio signal of the video set, and then the speaker 170A, also known as the "speaker", converts the audio electrical signal into a sound signal. In this way, the user can hear the background sound synchronized with the video in the highlight segment and the added video background music, etc.

[0144] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0145] Next, in combination with Figure 6 the software structure of the electronic device will be introduced.

[0146] Figure 6 is a schematic diagram of the software structure of the electronic device provided by the embodiments of the present application.

[0147] As Figure 6As shown, the electronic device can adopt a layered architecture, dividing the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software layers of the software structure are sequentially divided from top to bottom into: application (APP) layer, media middle platform framework layer, application framework (FWK) layer, and hardware abstract layer (HAL).

[0148] The APP layer, simply referred to as the application layer, can include a series of application packages, such as cameras, galleries, third-party video editing software, calendars, maps, and navigation. When these application packages are run, they can access various service modules provided by the media middle platform framework layer and the application framework layer through the application programming interface (API), and execute corresponding intelligent services.

[0149] In some embodiments, the camera is used to capture photos, videos, slow-motion images, panoramic images, etc. in response to user operations. After these images are captured by the camera, or after the user triggers a mobile phone screenshot, or after the user triggers a mobile phone screen recording, or after the electronic device downloads images from other devices, the electronic device can save these images in the gallery, so that the user can perform video editing operations on the images in the gallery, such as one-click video creation operations.

[0150] In the embodiments of the present application, the gallery is sequentially divided from top to bottom into: business layer, application function layer, and basic function layer.

[0151] Among them, the business layer, also known as the video editing business layer, provides multiple services (which can also be called functions) such as multi-shot video automatic video creation, one-shot multi-gain AI music short film, one-click video creation, and wonderful moments. These services are presented in the form of controls in the user interface (UI) of the gallery. By operating these controls, the user can trigger the gallery to perform corresponding video processing actions. For example, after the user selects image materials (the materials include pictures and / or videos), in response to the user's click operation on the one-click video creation control in the gallery, the gallery can call the underlying module to automatically analyze and extract the highlight segments from the pictures and / or videos through algorithms, and then combine the highlight segments into a clipped video set.

[0152] The application function layer includes an automatic editing framework. Each service in the service layer can call the automatic editing framework to provide automatic editing services for pictures and videos. Exemplarily, the automatic editing framework may include function modules such as segment optimization, storyline organization, layout splicing, and special effect beautification. Segment optimization is used to call the light segment analysis interface and policy monitoring interface in the high media middleware framework layer to extract high-light segments from pictures and / or videos. Storyline organization is used to sequentially splice multiple pictures and / or videos in the form of a storyline based on the content of the pictures and / or videos. Layout splicing is used to adjust the interface layout of pictures and / or videos. Special effect beautification is used to adjust the beautification effect of pictures and / or videos, such as adjusting the picture brightness and beautifying the human face, etc.

[0153] The basic function layer is used to perform basic function processing on the clipped picture and / or video segments after the automatic editing framework clips multiple pictures and / or videos. Exemplarily, the basic function layer may include basic function modules such as video splicing, synthesis and storage, video effect rendering, and audio effect processing. Among them, video splicing is used to splice multiple extracted high-light segments (where the high-light segments include pictures and / or videos). Synthesis and storage is used to store the video set obtained after splicing. Video effect rendering is used to add video effects to the video set obtained after splicing, such as adding style filters and themes to the video set. Audio effect processing is used to add sound effects to the video set obtained after splicing, such as adding background music to the video set.

[0154] The media middleware framework layer is a software layer set between the application layer and the application framework. The media middleware framework layer may include an analysis performance query interface, a high-light segment analysis interface, a policy monitoring interface, a pipeline interface, a theme summary interface, and an initialization interface. Among them, the analysis performance query interface is used to calculate the total duration of all videos according to the video analysis speed. The high-light segment analysis interface is used to call the policy monitoring interface to extract high-light segments. The theme summary interface is used to call the underlying algorithm to analyze the picture content of the high-light segments to determine the theme corresponding to the content of the high-light segments. The initialization interface is used to initialize the high-light segment algorithm, face detection algorithm, video acceleration algorithm, and image super-resolution algorithm in the HAL layer, etc.

[0155] The policy monitoring interface is used to calculate the first quantity of key frames (I-frames) of each video and the position of the key frames in the video (I-frame position). The channel interface is used to reduce the resolution of the video file according to the file descriptor and I-frame position of each video sent by the policy monitoring interface, and forward the data address of the downscaled video file to the hardware abstraction layer through the application framework layer, and then report the analysis result of the I-frame returned by the hardware abstraction layer to the policy monitoring interface. Among them, the analysis result may include the aesthetic score of the image frame.

[0156] The policy monitoring interface is also used to determine the second position of the I-frame based on the analysis result of the I-frame at the first position. When the number of analyzed I-frames in the video meets the first quantity of the video, the target area of the video is determined according to the analysis result of each I-frame. Among them, the target area is the area corresponding to the I-frame with the highest score in the video. The channel interface is also used to reduce the resolution of the video file according to the file descriptor of each video and the frame (I-frame or image frame) in the target area issued by the policy monitoring interface, and forward the data address of the video file with reduced resolution to the hardware abstraction layer through the application framework layer, and then report the analysis result of the frame (I-frame or image frame) in the target area returned by the hardware abstraction layer to the policy monitoring interface. In this embodiment, the image frame represents a common image frame different from the key frame.

[0157] The policy monitoring interface is also used to determine the highlight segment of each video according to the analysis result of the frame (I-frame or image frame) in the target area.

[0158] The theme summary interface is used to call the underlying algorithm to analyze the picture content of the highlight segment to determine the theme corresponding to the content of the highlight segment. The initialization interface is used to initialize the highlight segment algorithm, face detection algorithm, video acceleration algorithm, image super-resolution algorithm, etc. in the HAL layer.

[0159] It should be noted that this application is described by taking the gallery providing the one-click video creation function as an example, which does not limit the embodiments of this application. In actual implementation, third-party video editing software can adopt the video processing method provided by the embodiments of this application to synthesize multiple pictures and videos selected by the user into a video set with one click.

[0160] The FWK layer, simply referred to as the framework layer, can be used to support the operation of each module in the media middle platform framework layer. For example, the framework layer can include a one-click video creation interface, a parameter management interface, an image data transmission interface, a theme analysis interface, a performance analysis interface, etc.

[0161] The HAL layer is an encapsulation of the Linux kernel driver, providing an interface upward. It hides the hardware interface details of a specific platform, provides a virtual hardware platform for the operating system, makes it hardware-independent, and can be transplanted on multiple platforms. For example, the hardware abstraction layer can include a chip analysis speed interface, a highlight segment algorithm, a face detection algorithm, a video acceleration algorithm, and an image super-resolution algorithm. Among them, the highlight segment algorithm is an image processing algorithm provided by the image signal processor. This algorithm can perform aesthetic scoring on each frame of image according to the image color, image texture features, image quality, frame interpolation with the previous and next frames, and edge change rate value of each frame of image. The aesthetic scoring can be used as a basis for evaluating whether a frame of image is a highlight segment.

[0162] In this embodiment, the high-light segment algorithm can score each frame (I-frame or ordinary image frame) in the video. For example, if the score of a frame is greater than a preset score threshold, then the frame and the frames within a preset duration adjacent thereto can be used as the high-light segments of the source video. Exemplarily, the preset score threshold can be 90. If the score of a frame in the video is greater than 90, then the frame and the frames within a preset duration adjacent thereto can be used as a high-light segment of the video. Alternatively, for a video including the scores of N frames, the frame with the highest score and the frames within a preset duration before and after this frame can be used as the high-light segments of the video.

[0163] In some other feasible embodiments, assuming that the value range of the score of a frame is 0-100, the scores can also be divided into different levels of score results according to different value ranges of the scores. For example, the first threshold is 80. If the score of a frame is greater than or equal to the first threshold (the value range of the score is 80-100), it indicates that the score result of this frame is "high". The second threshold can be 50. If the score of a frame is greater than or equal to the second threshold (the value range of the score is 50-79), it indicates that the score result of this frame is "relatively high". The third threshold can be 20. If the score of a frame is greater than or equal to the third threshold (the value range of the score is 20-49), it indicates that the score result of this image frame is "medium". If the score of a frame is less than the third threshold (the value range of the score is 0-19), it indicates that the score result of this frame is "low". Exemplarily, in one implementation manner, the frames with scores greater than or equal to the first threshold / second threshold / third threshold can be used as candidate frames for extracting the high-light segments of the video. In some other implementable manners, the frame with the highest score can be used as a candidate frame for extracting the high-light segments of the video.

[0164] It should be noted that Figure 6 The layers shown in the software structure and the components included in each layer do not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more layers than those shown in the figure, such as the system library (FWK LIB) layer and the kernel layer. Each layer may include more or fewer components than those shown in the figure. In addition, the above-mentioned various functional modules may also be combined into one functional module, and each layer may also be combined into one layer. For example, the high-light segment analysis may include policy monitoring. Another example is that the media middle platform framework layer may be set in the application framework layer.

[0165] It can be understood that in order for an electronic device to implement the video processing method in the embodiments of the present application, it includes the corresponding hardware and / or software modules for performing various functions. Combining the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments.

[0166] In the video processing method provided by the embodiments of the present application, the policy monitoring module in the media middle platform framework layer can determine the first number of key frames for overview analysis of each video according to the analysis parameters of the image material, and dynamically determine the positions of the key frames of each video. The policy monitoring module can send the positions of each key frame in each video to the algorithm module in the HAL layer for analysis, obtain the first scores of each key frame from the algorithm in the HAL layer, and further determine the key frame with the highest first score and its corresponding area as the target area. Perform frame-by-frame analysis on the target area to obtain the second scores of each frame in the target area. The policy monitoring module can send each image frame in the target area to the algorithm module in the HAL layer for analysis, obtain the scoring results of each frame in the target area from the algorithm in the HAL layer, and thus determine the image frame with the highest second score and the area corresponding to the image frame as the highlight segment. Through the combined analysis of overview analysis and frame-by-frame analysis, rather than performing frame-by-frame analysis on the entire video, the workload of the electronic device for image frame analysis can be reduced, thereby reducing the time-consuming for the electronic device to implement the one-click video creation function, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the one-click video creation function.

[0167] Next, taking each module in the software structure diagram as shown Figure 6 as an example, the video processing method provided by the embodiments of the present application will be described exemplarily.

[0168] Figure 7 A method flow schematic diagram is provided for the various modules in the electronic device to perform operations such as algorithm initialization and obtaining analysis parameters before the policy monitoring module in the electronic device determines the first number of key frames for overview analysis of each video. This method can be applied to a one-click video creation scenario as shown Figures 1 - 4 as. As shown Figure 7 by taking the image material including pictures and videos as an example, this method can include the following S01-S16.

[0169] S01, the business layer receives the operation of the user enabling the one-click video creation function.

[0170] In this embodiment, the service layer refers to the one-click video creation module in the service layer. That is, the one-click video creation module in the service layer receives the operation of enabling the one-click video creation function input by the user. For example, this operation can specifically be the click operation on the "One-click Blockbuster" card as shown in (b) of Figure 1 .

[0171] S02. The service layer loads and displays candidate pictures and candidate videos.

[0172] S03. The service layer receives the operation of the user selecting multiple pictures and videos, and receives the operation of the user inputting to determine to execute the one-click video creation function.

[0173] For example, the operation of the user selecting multiple pictures and videos can be the click operation on the photos and videos as shown in (c) of Figure 1 , and the operation of the user inputting to determine to execute the one-click video creation function can be the click operation on the check mark option as shown in (d) of Figure 1 .

[0174] S04. The service layer calls the initialization interface of the media middle platform framework layer through the application function layer to initialize the relevant algorithms of the HAL layer.

[0175] In this embodiment, the relevant algorithms refer to the algorithms for the functions to be implemented by the service layer. Here, the service function is the one-click video creation function, so the relevant algorithms are the relevant algorithms involved in the one-click video creation function. For example, the relevant algorithms include highlight segment algorithms, face detection algorithms, video acceleration algorithms, and image super-resolution algorithms, etc.

[0176] S05. The initialization interface of the media middle platform framework layer sequentially sends the initialization parameters to the algorithm module of the HAL layer through the channel interface of the media middle platform framework layer and the service interface of the FWK layer.

[0177] Among them, in the HAL layer, one algorithm corresponds to one algorithm interface. The FWK layer is provided with multiple service interfaces, and one service interface of the FWK layer corresponds to one algorithm interface of the HAL layer. Each service interface of the FWK layer plays a role in data pass-through between the algorithm interface of the HAL layer and the channel interface of the media middle platform framework layer.

[0178] The initialization parameters involved in different algorithms are different. Therefore, the algorithm initialization parameters sent to the algorithm interface through each service interface may be different.

[0179] S06. The algorithm module of the HAL layer initializes according to the initialization parameters.

[0180] S07. The algorithm module of the HAL layer returns an initialization success message to the channel interface of the media middle platform framework layer through the service interface of the FWK layer.

[0181] S08, The channel interface of the media middle platform framework layer calls the analysis speed interface of the HAL layer through the service interface of the FWK layer to obtain the chip analysis speed.

[0182] Among them, the service interface of the FWK layer can be a performance analysis interface.

[0183] S09, The analysis speed interface of the HAL layer returns the chip analysis speed to the channel interface of the media middle platform framework layer through the service interface of the FWK layer.

[0184] S10, The channel interface of the media middle platform framework layer returns the chip analysis speed to the initialization interface of the media middle platform framework layer.

[0185] Among them, the chip analysis speed can represent the number of image frames in a picture / video analyzed by the image signal processor per unit time; or, the chip analysis speed can also represent the single-frame analysis duration of the image signal processor. Among them, the single-frame analysis duration refers to the duration of analyzing an image frame in a video. The chip analysis speed can also represent the duration of the image signal processor analyzing a picture.

[0186] Based on this, the single-frame analysis duration of the image signal processor and the processing duration of the image signal processor for a picture can be calculated according to the chip analysis speed. The single-frame analysis duration of the image signal processor and the processing duration of the image signal processor for a picture may be different. Exemplarily, the processing duration of a picture can be 400 ms, and the single-frame analysis duration can be 200 ms.

[0187] It should be understood that since the performances of different image signal processors are different, the chip analysis speeds corresponding to different image signal processors may be different. For an electronic device put on the market, the image signal processor is fixed, so the chip analysis speed corresponding to this image signal processor is also fixed.

[0188] In some embodiments, the channel interface of the media middle platform framework layer can also return other relevant performance parameters of various algorithms to the initialization interface of the media middle platform framework layer.

[0189] S11, The initialization interface of the media middle platform framework layer sends an initialization success message to the service layer through the application function layer.

[0190] Among them, the initialization success message can carry performance parameters of various algorithms, such as the chip analysis speed.

[0191] After obtaining the chip analysis speed, for each video in the image materials selected by the user, the following S12 - S15 can be executed to obtain the analysis parameters for the video and all image materials.

[0192] Taking Video 1 as an example for illustration, Video 1 is one of multiple videos in the image materials.

[0193] S12. The service layer sends a query message to the analysis performance query interface of the media middleware framework layer through the application function layer.

[0194] Among them, the query message includes the file descriptor fd1 of Video 1 selected by the user and the chip analysis speed. Among them, the file descriptor can be used as the unique identifier of the video.

[0195] S13. The analysis performance query interface of the media middleware framework layer obtains the estimated analysis duration of Video 1 according to the file descriptor fd1 of Video 1.

[0196] Among them, the estimated analysis duration can be the duration required for the image signal processor to analyze a video. Analyzing a video can refer to analyzing a specified number of key frames in a video. For example, the specified number can be 1, or the specified number can also be greater than 1 and less than or equal to the number of key frames included in the video.

[0197] Among them, exemplarily, when the specified number is 1, that is, the estimated analysis duration represents the duration required for the image signal processor to analyze 1 key frame in a video, then the estimated analysis duration can also represent the analysis duration corresponding to the minimum analysis of key frames. For example, if the single-frame analysis duration of the image signal processor is 200 ms, then the estimated analysis duration of each video (including Video 1) is 200 ms.

[0198] In some embodiments, the specified number can also be the number of key frames included in the video. Then, the estimated analysis duration can be the duration corresponding to analyzing all key frames of the video. For example, if the single-frame analysis duration of the image signal processor is 200 ms and Video 1 includes 10 key frames, then the estimated analysis duration of Video 1 is 10 * 200 ms = 2000 ms. For example, if Video 2 includes 12 key frames, then the estimated analysis duration of Video 2 is 12 * 200 ms = 2400 ms.

[0199] In some embodiments, the specified number can also be a, where a is greater than 1 and less than the number of image frames included in the video, and a is a natural number. Then, the estimated analysis duration can be the duration corresponding to analyzing a key frames in the video. For example, if the single-frame analysis duration of the image signal processor is 200 ms, Video 1 includes 10 key frames, and N is 5. Then the estimated analysis duration of Video 1 is 5 * 200 ms = 1000 ms.

[0200] S14. The analysis performance query interface of the media middleware framework layer returns the estimated analysis duration of Video 1 to the service layer through the application function layer.

[0201] After the one - click video composition module at the service layer obtains the estimated analysis duration of Video 1, it can return to continue executing S12 - S15 to obtain the estimated analysis duration of the next video among multiple image materials until the estimated analysis durations of all videos among multiple image materials are obtained.

[0202] S15. The service layer obtains analysis parameters according to the estimated analysis durations of all image materials.

[0203] Among them, the analysis parameters may include the total analysis duration, the upper limit recommended value of the total analysis duration, the maximum duration of the highlight segment, the minimum duration of the highlight segment, the total recommended duration of the highlight segment, the recommended duration of the highlight segment, whether to force each video to output a highlight segment, whether to enable audio analysis, the selected highlight segments, etc.

[0204] Among them, whether to force each video to output a highlight segment is defaulted to yes, that is, in this embodiment, each video needs to output a highlight segment.

[0205] Among them, the total analysis duration represents the total duration required to complete the analysis of all image materials selected by the user (including all pictures and all videos selected by the user). The total analysis duration includes the sum of the estimated analysis durations of all pictures and the sum of the estimated analysis durations of all videos.

[0206] For the pictures in the image materials, the processing duration of a picture can be directly determined according to the chip speed of the image processor. Therefore, the sum of the estimated analysis durations of all pictures can be directly determined according to the number of pictures in the image materials.

[0207] For the videos in the image materials, the estimated analysis duration of each video can be obtained according to S12 - 15, and the estimated analysis durations of each video are accumulated to obtain the sum of the estimated analysis durations of all videos in the image materials.

[0208] Based on the sum of the estimated analysis durations of all pictures and the sum of the estimated analysis durations of all videos, the total analysis duration of the image materials can be obtained.

[0209] For example, assume that the materials selected by the user include 10 pictures, Video 1, and Video 2. The estimated analysis duration of a picture is 400ms, the estimated analysis duration of Video 1 is 200ms, and the estimated analysis duration of Video 2 is 300ms. Then the above - mentioned total analysis duration is 10 * 400ms + 200ms + 300ms = 4500ms.

[0210] The recommended value of the upper limit of the total analysis duration represents the recommended value of the maximum duration for completing the analysis of all the above-mentioned pictures and all videos. The recommended value of the upper limit of the total analysis duration can be determined based on the estimated analysis duration of each video and the estimated analysis duration of the pictures. Generally, the recommended value of the upper limit of the total analysis duration is greater than the total analysis duration. For example, calculated by analyzing at least 1 image frame for each video, the total analysis duration is 4400ms. Considering that each video may require analyzing multiple image frames, then the recommended value of the upper limit of the total analysis duration can be much greater than the total analysis duration. For example, the recommended value of the upper limit of the total analysis duration can be preset to 10000ms.

[0211] In some scenarios where the input of some analysis parameters is abnormal, the recommended value of the upper limit of the total analysis duration may also be set to be less than the total analysis duration. In the case where the recommended value of the upper limit of the total analysis duration is less than the total analysis duration, that is, when the recommended value of the upper limit of the total analysis duration is not sufficient to analyze all pictures and all videos (1 image frame of each video), some materials can be selected from the selected image materials for analysis. This part is the method implemented by the policy monitoring module of the media middle platform framework layer and will be introduced in detail in the following embodiments and will not be elaborated here.

[0212] The maximum duration of a highlight segment represents the maximum allowed duration of a highlight segment in a video. The minimum duration of a highlight segment represents the minimum required duration of a highlight segment in a video. The maximum duration of a highlight segment and the minimum duration of a highlight segment can be preset values. Exemplarily, the maximum duration of a highlight segment can be 3000ms, and the minimum duration of a highlight segment can be 1000ms.

[0213] The total recommended duration of highlight segments represents the recommended value of the sum of the recommended durations of the highlight segments of all videos in multiple image materials. The total recommended duration of highlight segments can be determined based on the number of videos, the maximum duration of a highlight segment, and the minimum duration of a highlight segment. Exemplarily, the maximum duration of a highlight segment can be 3000ms, the minimum duration of a highlight segment can be 1000ms, and when the number of videos is 5, the value range of the total recommended duration of highlight segments can be 5000ms - 15000ms. For example, the total recommended duration of highlight segments can be 8000ms.

[0214] The recommended duration of a highlight segment represents the recommended value of the duration of a highlight segment in a video. The highlight segment with this recommended duration value can effectively display the highlight effect. The recommended duration of a highlight segment can be determined based on the maximum duration of a highlight segment and the minimum duration of a highlight segment. Exemplarily, the maximum duration of a highlight segment can be 3000ms, the minimum duration of a highlight segment can be 1000ms, and the value range of the recommended duration of a highlight segment can be 1000ms - 3000ms. For example, the recommended duration of a highlight segment can be 2000ms.

[0215] It should be understood that the maximum duration, minimum duration, total recommended duration, and recommended duration of the highlight segments can all be set according to the actual situation.

[0216] In some embodiments, the analysis parameters may further include the actual duration of each video.

[0217] If the sum of the actual durations of all the videos of multiple image materials is greater than a preset duration threshold, the video analysis method provided in S16 - S46 of this solution can be executed to reduce the number of frames of the videos in the image materials to be analyzed, reduce the time consumption of the electronic device to implement the one - click video creation function, improve the efficiency of the electronic device in processing videos, and thereby improve the user experience of the one - click video creation function.

[0218] S16. The business layer, through the application function layer, sends the file descriptor fd and analysis parameters of all the materials to be analyzed to the image highlight segment analysis interface of the media middle - platform framework layer.

[0219] After receiving the analysis parameters, the policy monitoring module of the media middle - platform framework layer can determine the material analysis policy according to the upper - limit recommended value of the total analysis duration and the total analysis duration in the analysis parameters. The material analysis policy refers to analyzing all the image materials selected by the user, or selecting some of the image materials selected by the user for analysis. In the following embodiments, the image materials include all the videos and all the pictures selected by the user.

[0220] After obtaining the total analysis duration of the image materials, the policy monitoring module can determine the number of pictures and videos that can be analyzed according to the total analysis duration of the image materials and the upper - limit recommended value of the total analysis duration in the analysis parameters. After Figure 7 S16, referring to Figure 8 The method flow for determining the analysis policy for image materials given below includes:

[0221] Specifically, in one implementation, after S16, execute:

[0222] S17. The policy monitoring module of the media middle - platform framework layer determines the analysis policy according to the total analysis duration and the upper - limit recommended value of the total analysis duration in the analysis parameters.

[0223] Among them, if the total analysis duration of the image materials is less than or equal to the upper - limit recommended value of the total analysis duration, the material analysis policy can be to analyze all the pictures and all the videos. For example, if the total analysis duration is 4400ms and the upper - limit recommended value of the total analysis duration is 10000ms, the material analysis policy can be to analyze all the pictures and all the videos in the image materials selected by the user.

[0224] If the total analysis duration of the image materials is greater than the recommended upper limit of the total analysis duration, the material analysis strategy can be a random selection strategy. For example, if the total analysis duration is 4400 ms and the recommended upper limit of the total analysis duration is 3000 ms, the material analysis strategy can be a random selection strategy. Among them, the random selection strategy means randomly selecting some of the image materials selected by the user for analysis.

[0225] Suppose the analysis value of a single picture is greater than that of one I-frame of a video. Then, the random selection strategy can be: for every N pictures selected, M videos are allowed to be selected until the total analysis duration of the selected image materials reaches the recommended value of the analysis duration upper limit or all the pictures in the selected image materials have been selected. Among them, M < N. For example, N can be natural numbers such as 3, 4, 5, etc., and M can be natural numbers less than N such as 1, 2, 3, etc. The specific values of M and N can be determined according to the number of image materials selected by the user. Exemplarily, it can be that for every 4 pictures selected, 1 video is allowed to be selected until the total analysis duration meets 3000 ms; or all the pictures in the selected image materials have been selected.

[0226] For example, the total analysis duration of 10 pictures and 2 videos is 4400 ms, which is greater than the recommended upper limit of the total analysis duration of 3000 ms. According to the random selection strategy of allowing 1 video to be selected for every 4 pictures selected, 4 pictures and 1 video are selected. The total analysis duration of 4 pictures and 1 video is 1800 ms, which does not reach the recommended upper limit of the total analysis duration of 3000 ms. Continue to select image materials. When the 3rd picture is selected in this round, the total analysis duration reaches 3000 ms, and at this time, the selection of image materials stops. Then, the selected image materials are 7 pictures and 1 video. The remaining 3 pictures will not be analyzed.

[0227] In another embodiment, suppose the analysis value of a single picture is less than that of one I-frame of a video. Then, the random selection strategy can be that for every P videos selected, Q pictures are allowed to be selected until the total analysis duration of the selected image materials reaches the recommended value of the analysis duration upper limit or all the videos in the selected image materials have been selected. Among them, Q < P. For example, P can be natural numbers such as 3, 4, 5, etc., and Q can be natural numbers less than P such as 1, 2, 3, etc. The specific values of P and Q can be determined according to the number of image materials selected by the user. Exemplarily, it can be that for every 3 videos selected, 1 picture is allowed to be selected until the total analysis duration meets 3000 ms; or all the videos in the selected image materials have been selected.

[0228] For example, the total analysis duration of 10 videos and 3 pictures is 3200 ms, which is greater than the recommended upper limit of the total analysis duration of 3000 ms. According to the random selection strategy that allows selecting 1 picture for every 3 selected videos, 3 videos and 1 picture are selected. The total analysis duration of 3 videos and 1 picture is 1000 ms, which does not reach the recommended upper limit of the total analysis duration of 3000 ms. Continue to select image materials. After three rounds of selecting 3 videos and 1 picture, a total of 9 videos and 3 pictures are selected, and the total analysis duration is 3000 ms. At this time, stop selecting image materials. Then, the selected image materials are 9 videos and 3 pictures. The remaining 1 video is not analyzed.

[0229] In some embodiments, the electronic device determines the analyzable image materials from the materials selected by the user according to the above random selection strategy (such as a mobile phone). Among them, pictures and videos can be randomly selected from the image materials selected by the user in the order in which the user selects the image materials. Or, it can also be randomly selected. For example, by generating a random number less than the number of materials through the Random class in Java, so as to select the video or picture corresponding to the random number.

[0230] In some embodiments, the image materials only include pictures, and the total analysis duration is the sum of the estimated analysis durations of all pictures. When the total analysis duration is greater than the recommended value of the analysis duration upper limit, the number of analyzable pictures is calculated according to the recommended value of the analysis duration upper limit and the estimated analysis duration of analyzing one picture, and the corresponding number of pictures is randomly selected from all pictures for analysis.

[0231] In some embodiments, the image materials only include videos, and the total analysis duration is the sum of the estimated analysis durations of all videos. When the total analysis duration is greater than the recommended value of the analysis duration upper limit, the number of analyzable videos is calculated according to the recommended value of the analysis duration upper limit, and the corresponding number of videos is randomly selected from all videos for analysis. Among them, the duration required for the number of analyzable videos is less than or equal to the total analysis duration.

[0232] After selecting the pictures and / or videos that can be analyzed within the recommended value of the analysis duration upper limit from the image materials, the remaining pictures and videos in the image materials selected by the user are not analyzed.

[0233] The image materials include pictures. After the electronic device selects the pictures that can be analyzed within the recommended value of the analysis duration upper limit, it executes Figure 8 the S18 - S22 shown to perform highlight segment analysis on the analyzable pictures.

[0234] The following is an example of the process of performing highlight segment analysis on Picture 1 in combination with S18 - S22, where Picture 1 is one of the pictures that can be analyzed within the recommended value of the analysis duration upper limit.

[0235] S18. The policy monitoring module in the media middle platform framework layer sends an indication message to the channel interface in the media middle platform framework layer. The indication message includes the file descriptor of Picture 1.

[0236] S19. The channel interface in the media middle platform framework layer performs processing operations such as decoding, resolution reduction, and format conversion on Picture 1 according to the file descriptor of Picture 1, and stores the processed Picture 1.

[0237] S20. The channel interface in the media middle platform framework layer sends the frame data address of Picture 1 to the highlight segment algorithm interface in the HAL layer through the service interface in the FWK layer.

[0238] Among them, the service interface can be a one-click video generation interface.

[0239] S21. The highlight segment algorithm interface in the HAL layer obtains Picture 1 according to the frame data address of Picture 1, and then analyzes Picture 1 based on a preset highlight segment algorithm to obtain an analysis result.

[0240] Exemplarily, the highlight segment algorithm interface can perform aesthetic scoring on Picture 1 according to the image color, image texture features, image quality, edge change rate value, etc. of Picture 1 to obtain an analysis result. Among them, the analysis result can represent the aesthetic score.

[0241] In some embodiments, the analysis result may include a scoring value, and the analysis result may also include a scoring value and a scoring result.

[0242] S22. The highlight segment algorithm interface in the HAL layer returns the analysis result of Picture 1 to the policy monitoring module in the media middle platform framework layer through the service interface in the FWK layer and the channel interface in the media middle platform framework layer in sequence.

[0243] After the policy monitoring module in the media middle platform framework layer obtains the analysis result of Picture 1, if there are other pictures, such as Picture 2, the electronic device can continue to execute the above S18 - S22 to obtain the analysis results of other pictures.

[0244] Among them, the policy monitoring module in the media middle platform framework layer can determine the pictures with scores greater than or equal to the first threshold / second threshold as the highlight segments of multiple image materials according to the analysis results of each picture.

[0245] After obtaining the analysis results of all pictures, if the image materials include videos, the electronic device can use the following S23 - S36 to obtain the analysis results of each video. If the image materials do not include videos, the electronic device outputs a target video set composed of the pictures of the highlight segments.

[0246] Reference Figure 9, which shows a schematic diagram of a technical idea for analyzing high - light segments of a video provided by an embodiment of the present application. During the process of analyzing each video, the electronic device can first obtain a first number of I - frames (key frames) in the video for analysis to obtain a first score for the I - frames. Among them, obtaining the first number of I - frames can be randomly obtained from the video. For example, the I - frames of different segments can be randomly obtained according to the storyboard points of the video. Or, the positions of the first number of I - frames can also be dynamically determined based on the first score. The target area of each video is determined based on the first I - frame with the highest first score. Then, frame - by - frame (I - frame by I - frame or image - frame by image - frame) analysis is performed on the target area to obtain a second score for each image frame in the target area. The high - light segment of the video is obtained based on the second image frame with the highest second score. In this solution, instead of performing frame - by - frame analysis on the entire video, the workload of frame analysis by the electronic device can be reduced, thereby reducing the time consumption of the electronic device to implement the one - key video compilation function, improving the efficiency of the electronic device in processing videos, and further improving the user experience of the one - key video compilation function.

[0247] If the image material includes multiple pictures and videos, since the calculation amount of pictures is small and the time consumption is less, the electronic device usually first analyzes the high - light segments of pictures one by one. After completing the analysis of the high - light segments of all pictures, it then analyzes the high - light segments of videos one by one. That is, after the electronic device finishes executing S18 - S22 as shown in Figure 8 , it continues to execute S23 - S37 as shown in Figure 10 .

[0248] If the image material only includes multiple videos, then after the electronic device executes S17, it executes S23 - S37 as shown in Figure 10 , and does not need to execute S19 - S22 as shown in Figure 8 .

[0249] In this embodiment, an example of the case where the image material includes pictures and videos is described. The following videos all refer to the videos that can be analyzed within the recommended upper limit of the total analysis duration.

[0250] S23, the policy monitoring module in the media middle - platform framework layer calculates the first number of I - frames of each video.

[0251] The first number refers to the number of I - frames that can be analyzed in each video within the recommended upper limit of the total analysis duration.

[0252] In some embodiments, the policy monitoring module can determine the first number of I - frames of each video according to the actual duration of each video in multiple image materials, the number of videos in multiple image materials, and the recommended upper limit of the total analysis duration.

[0253] For example, according to the recommended value T1 of the total analysis duration upper limit and the single-frame analysis duration t, determine the number of I-frames T1 / t that can be analyzed within the recommended value of the total analysis duration upper limit. According to the actual duration of each video and the number of videos, allocate the number of I-frames T1 / t to each video to obtain the first quantity of each video.

[0254] Among them, in one example, allocating the number of I-frames T1 / t to each video can be to sort the videos in descending order according to their actual durations. For the videos ranked in the top 25%, evenly allocate 50% of the number of I-frames of T1 / t. For the videos ranked between 25% and 75%, evenly allocate 40% of the number of I-frames of T1 / t. For the videos ranked in the bottom 25%, evenly allocate 10% of the number of I-frames of T1 / t.

[0255] Among them, in another example, allocating the number of I-frames T1 / t to each video can be to allocate corresponding quantities to each video respectively according to the proportion of the actual duration of the video.

[0256] Or, in another example, the steps for the policy monitoring module to determine the first quantity of the image frames of each video may include:

[0257] S231, the policy monitoring module calculates the base quantity and the maximum quantity of the I-frames of each video.

[0258] Among them, the base quantity can be understood as the minimum number of I-frames that need to be analyzed to ensure the analysis effect of the video when the analysis duration is limited and the analysis performance of the electronic device is limited. The maximum quantity can be understood as the maximum number of I-frames that the duration allows to analyze the video. Generally, the maximum quantity is greater than the base quantity.

[0259] In some embodiments, the policy monitoring module can determine the base quantity and the maximum quantity of the I-frames of each video according to the actual duration of each video and a preset corresponding relationship. Among them, the preset corresponding relationship represents the maximum quantity and the base quantity of the I-frames corresponding to different threshold ranges of the video duration.

[0260] The base quantity and the maximum quantity of videos with different durations are different. For example, a preset corresponding relationship between the video duration and the number of I-frames can be pre-stored in an electronic device (such as a mobile phone). In the embodiments of the present application, the policy monitoring module can determine the base quantity and the maximum quantity of the I-frames of each video according to this preset corresponding relationship.

[0261] Exemplarily, the preset corresponding relationship includes the following (1)-(4):

[0262] (1) If the video duration P is less than the first threshold Q1, the number of I-frames of the video is 1.

[0263] (2) The video duration P is equal to the first threshold Q1, and the starting number of I-frames of the video is m.

[0264] (3) The video duration P is greater than the first threshold Q1 and less than or equal to the second threshold Q2, and the number of I-frames of the video increases based on the starting number. The number of I-frames is updated to That is, the integer part of P divided by s plus m. It means that from the start time of the video to Q1, the number of I-frames of the video is m, and from the start time of the video to Q2, 1 I-frame can be corresponding to every s seconds.

[0265] (4) The video duration P is greater than the second threshold Q2 and less than the third threshold Q3, and the number of I-frames of the video is That is, the integer part of Q2 divided by s, the integer part of (P - Q2) divided by k, plus m. It means that from the start time of the video to Q1, the number of I-frames of the video is m; from the start time of the video to Q2, 1 I-frame can be corresponding to every s seconds; from Q2 to Q3, 1 I-frame can be corresponding to every k. Among them, is the floor symbol.

[0266] Among them, the first threshold < the second threshold < the third threshold.

[0267] It should be noted that if the video duration is very long, more duration thresholds can be set, such as the fourth threshold, the fifth threshold, and so on. When determining the base number and the maximum number, thresholds such as the first threshold, the second threshold, and the third threshold can be set according to the actual video duration; the interval seconds can be determined according to the analysis performance of the electronic device. The above are just examples for illustration, and the parameter values are not limited.

[0268] In some embodiments, the maximum number of I-frames is more than the base number, and the frame extraction density for determining the maximum number of I-frames is greater than the frame extraction density for determining the base number of I-frames. In order to obtain more I-frame numbers, generally, the first threshold for determining the maximum number is less than or equal to the first threshold for determining the base number; or, the preset second threshold for determining the maximum number is equal to or greater than the second threshold for determining the base number, or, the third threshold for determining the maximum number is greater than or equal to the third threshold for determining the base number. In some embodiments, the interval seconds for determining the maximum number is less than or equal to the interval seconds for determining the base number.

[0269] The density of the maximum number of I-frames in the video is greater and the number is more; when the analysis duration is sufficient, more I-frames of the video can be analyzed, and the maximum number of I-frames can be more. The above are just examples for illustration, and the parameter values are not limited.

[0270] The following gives several examples to illustrate the determination process of the base number and the maximum number of I-frames of the video.

[0271] For example, the policy monitoring module determines the basic number of I-frames of a video. The first threshold can be 3 seconds, s seconds can be 5 seconds, the second threshold can be 30 seconds, k seconds can be 10 seconds, and the third threshold can be 90 seconds.

[0272] Exemplarily, taking video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example, exemplarily, reference can be made to Figure 11 the schematic diagram of I-frame extraction shown.

[0273] The policy monitoring module receives video 1 and determines at least one I-frame. At this time, the number of I-frames of video 1 is 1 frame. The duration of video 1 (91 seconds) is greater than the first threshold (3 seconds), so the number of I-frames of video 1 is updated to 2 frames.

[0274] Starting from the 0th second of the duration of video 1 to the 30th second of the video duration, the number of I-frames is increased by one every 5 seconds. As Figure 11 shown, at the 30th second of the video duration, the basic number of I-frames of video 1 is 8 frames Starting from the 30th second of the video duration to the 90th second of the video duration, the number of I-frames is increased by one every 10 seconds. As Figure 11 shown, at the 90th second of the video duration, the basic number of video 1 is updated to 14 frames. After the 90th second of video 1, the remaining duration of the video is 3 seconds, which does not meet the requirement of extracting 1 I-frame every 10 seconds after 90 seconds. Therefore, the number is not increased. So, the basic number of video 1 with a duration of 93 seconds is 14 frames

[0275] In some embodiments, after calculating the basic number of each video, if the total number of the basic numbers of all videos does not meet the preset minimum number of frames, it is necessary to adjust the basic numbers of all videos. Among them, the adjustment method can be to increase the number of frames of the video whose basic number is less than the preset value; or, increase the basic number of the video with a longer duration.

[0276] Exemplarily, in order to ensure obtaining a more accurate video theme, the preset minimum number of frames can be 5 frames.

[0277] In some embodiments, when there is only one video, for example, when the basic number of Video 1 is less than 5 frames, the basic number of Video 1 is directly adjusted to 5 frames. When there are multiple videos, since at least one I-frame is acquired for each preset video, the preset minimum number of frames should be greater than the number of videos. For example, when the number of videos is 3, the preset minimum number of frames can be 5 frames. Exemplarily, the basic number of Video 1 is 1 frame, the basic number of Video 2 is 2 frames, and the basic number of Video 3 is 1 frame. The total of the basic numbers of all videos is 4 frames, which is less than the preset minimum number of frames, 5 frames. The basic number of the video with less than 2 frames can be adjusted to 2 frames. At this time, the basic number of Video 1 is 2 frames, the basic number of Video 2 is 2 frames, and the basic number of Video 3 is 2 frames. The total of the basic numbers of all videos is 6 frames, which is greater than the preset minimum number of frames, 5 frames. If the number of videos is 2 frames, the basic number of Video 1 is 1 frame, and the basic number of Video 2 is 1 frame. By adjusting the basic number of the video with less than 2 frames to 2 frames, the total of the basic numbers of all videos is still 4 frames, which does not meet the preset minimum number of frames, 5 frames. In this case, the basic number of the video with the longest duration can be adjusted to 3 frames so that the total of the basic numbers of all videos meets 5 frames. For example, the duration of Video 1 is 2 seconds, and the duration of Video 2 is 1 second. Then, the basic number of Video 1 is adjusted to 3 frames, and the basic number of Video 2 is adjusted to 2 frames. At this time, the total of the basic numbers of all videos is 5, meeting the preset minimum number of frames, 5 frames.

[0278] Exemplarily, the policy monitoring module determines the maximum number of I-frames of a video. The first threshold can be 2 seconds, s seconds can be 3 seconds, the second threshold can be 30 seconds, k seconds can be 10 seconds, and the third threshold can be 120 seconds.

[0279] Exemplarily, taking the video 1 with a duration of 93 seconds calculated by the policy monitoring module as an example, it can be referred to Figure 12 the schematic diagram of extracting I-frames as shown.

[0280] The policy monitoring module receives Video 1 and extracts at least one I-frame. At this time, the maximum number of Video 1 is 1 frame. The duration of Video 1 (93 seconds) is greater than the first threshold (2 seconds). Thus, the maximum number of Video 1 is updated to 2 frames.

[0281] Starting from the 0th second of the duration of Video 1 to the 30th second of the video duration, the number of I-frames is increased by one every 3 seconds. As Figure 12 shown, at the 30th second of the video duration, the maximum number of Video 1 is 12 frames Starting from the 30th second of the video duration to the 120th second of the video duration, the number of I-frames is increased by one every 10 seconds. As Figure 12 shown, at the 93rd second of the video duration, the maximum number of Video 1 is updated to 18 frames After the 90th second of Video 1, the remaining duration of the video is 3 seconds, which does not meet the requirement of extracting 1 I-frame every 10 seconds after 120 seconds. Therefore, after increasing the number of I-frames by 1 at the 90th second, the number of I-frames is no longer increased. So, the maximum number of frames of Video 1 with a duration of 93 seconds is 18 frames.

[0282] S232. The policy monitoring module determines the total number of I-frames to be analyzed for all videos according to the recommended value of the analysis duration upper limit.

[0283] Among them, the recommended value of the analysis duration upper limit represents the recommended value of the maximum duration for completing the analysis of all the above pictures and all videos. After the processing of pictures and videos in the above embodiments requires the consumption of analysis duration, therefore, the total number of I-frames to be analyzed for all videos is determined by the first magnification of the recommended value T1 of the analysis duration upper limit.

[0284] The policy monitoring module determines the total number of I-frames to be analyzed for all videos according to the first magnification of the recommended value T1 of the analysis duration upper limit. Among them, the first magnification can be 30%. That is, according to 30% of the recommended value of the analysis duration upper limit (T1 * 0.3), the number of analyzable I-frames is judged.

[0285] If T1 * 0.3 is less than the time required to analyze the basic number of I-frames of all videos, that is, the duration of T1 * 0.3 is not enough to analyze the basic number of I-frames of all videos, then the policy monitoring module can allocate the analysis duration T2 for all videos for analysis. At this time, the total number of analyzed frames for all videos is the sum of the basic numbers of all videos. Among them, the analysis duration T2 is greater than T * 0.3, less than T1, and greater than or equal to the time required for the basic number of all videos.

[0286] Or, in the case where the duration of T * 0.3 is not enough to analyze the basic number of I-frames of all videos, the policy monitoring module can allocate the entire recommended value T1 of the analysis duration upper limit to the analysis of the basic number of I-frames of all videos, and the actual number of I-frames processed by the policy monitoring module (total number of analyzed frames) is the sum of the basic numbers of I-frames of all videos.

[0287] If the recommended value T1 of the analysis duration upper limit is less than the time required to analyze the basic number of I-frames of all videos, assuming that the analysis duration of a single frame is t, at this time, the actual number of I-frames processed by the policy monitoring module (total number of analyzed frames) is T1 / t (frames).

[0288] If T1 * 0.3 is greater than the time required to analyze the maximum number of I-frames of all videos, that is, the duration of T1 * 0.3 is enough to analyze the maximum number of I-frames of all videos. Then, the policy monitoring module only allocates T1 * 0.3 for the analysis of the maximum number of I-frames of all videos, and the actual number of I-frames processed by the policy monitoring module (total number of analyzed frames) is the sum of the maximum numbers of all videos.

[0289] If T1 * 0.3 is greater than the time taken to analyze the base number of I-frames of all videos and, at the same time, T1 * 0.3 is less than the time taken to analyze the base number of I-frames of all videos, assuming the analysis time per frame is t, the number of I-frames (total analysis quantity) actually processed by the policy monitoring module is T1 * 0.3 / t (frames).

[0290] Exemplarily, assume the input video durations are 15s, 40s, 100s, and 150s respectively, and the corresponding base number of I-frames for each video are 5 frames, 9 frames, 14 frames, and 14 frames respectively, with a total of 42 frames. The corresponding maximum number of I-frames for each video are 7 frames, 13 frames, 19 frames, and 21 frames respectively, with a total of 60 frames.

[0291] Assume the analysis time per frame is 200 ms. Then, the time taken to analyze the base number of I-frames of all videos, which is the time taken for 42 frames, is 8400 ms; the time taken to analyze the maximum number of I-frames of all videos, which is the time taken for 60 frames, is 12000 ms.

[0292] If T1 * 0.3 is less than 8400 ms, that is, T1 * 0.3 is less than the time taken to analyze the base number of I-frames of all videos, the policy monitoring module allocates 8400 ms for processing the 42 I-frames of all videos according to the time taken to process 42 frames. Alternatively, the policy monitoring module allocates the entire analysis duration upper limit recommended value T1 for I-frame analysis of all videos. At this time, the number of I-frames (total analysis quantity) actually processed by the policy monitoring module is 42 frames. If T1 is less than 8400 ms, the number of I-frames (total analysis quantity) actually processed by the policy monitoring module is T1 / 200 (frames).

[0293] If T1 * 0.3 is greater than 12000 ms, that is, T1 * 0.3 is greater than the time taken to analyze the maximum number of I-frames of all videos, the policy monitoring module allocates T1 * 0.3 for I-frame analysis of all videos, and the number of I-frames (total analysis quantity) actually processed is 60 frames.

[0294] If T1 * 0.3 is greater than 8400 ms and, at the same time, T1 * 0.3 is less than 12000 ms, within T1 * 0.3, the number of I-frames (total analysis quantity) actually processed by the policy monitoring module is T1 * 0.3 / 200 ms (frames).

[0295] In this embodiment, the analysis duration upper limit is only a recommended value for planning and is not strongly verified. In this embodiment, the estimated analysis duration of video analysis is allowed to exceed the recommended value of the analysis duration upper limit.

[0296] S233. The policy monitoring module in the media middle platform framework layer determines the first quantity of I-frames for each video according to the total analysis quantity of all videos.

[0297] After the policy monitoring module obtains the total analysis quantity of the image frames of all videos, it can sequentially allocate the corresponding number of I-frames to be analyzed (the first quantity) for each video.

[0298] Exemplarily, the videos can be sorted in descending order of their actual durations, and each video is traversed in turn. Each time a video is traversed, the first value and the second value of the video are updated.

[0299] Among them, the initial value of the first value is the total analysis quantity; the initial value of the second value of the video is 0.

[0300] During the process of traversing the videos, each time a video is traversed, the second value of the video is incremented by 1, and the first value is decremented by 1 until the first value is 0, that is, until the total analysis quantity is allocated completely. When the total analysis quantity is allocated completely, the second value of each video is the corresponding first quantity.

[0301] It should be noted that when allocating the first quantity of I-frames for each video, the first quantity of I-frames allocated to each video should be less than or equal to the maximum quantity of I-frames of the video, and the first quantity of I-frames allocated to each video should be greater than or equal to the base quantity of I-frames of the video. If, before traversing a certain video, the second value of the video is equal to the maximum quantity, then this video and all subsequent traversal processes will skip this video and traverse the next video of this video. The videos skipped without traversing do not increase the second value. In this way, the first quantity corresponding to the actual duration of each video can be obtained.

[0302] Exemplarily, refer to Figure 13 , Figure 13A schematic diagram for allocating I-frames to multiple videos is provided. Assume the input videos are Video 1 (15 s), Video 2 (20 s), Video 3 (30 s), and Video 4 (35 s), and the maximum number of I-frames corresponding to each video is 7 frames, 8 frames, 12 frames, and 12 frames respectively. Assume the total number of frames to be analyzed for all videos is 38 frames. The policy monitoring module traverses these 4 videos and allocates the 38 frames to these 4 videos in sequence. Starting from the first round, 1 frame is allocated to each video in each round. That is, when traversing each video, the corresponding second value is incremented by 1. After the 7th round, the second value of the I-frames of Video 1 has reached the maximum number of its corresponding I-frames (7 frames), and no more I-frames will be allocated to it subsequently. That is, this video and all subsequent traversal processes will skip this video. In the 8th round, only Video 2, Video 3, and Video 4 are traversed. After the 8th round, the second value of Video 2 has reached its corresponding maximum number (8 frames), and no more I-frames will be allocated to it subsequently. In the 9th round, only Video 3 and Video 4 are traversed. Until the end of the 11th round, the number of allocated frames for each video is 7 frames, 8 frames, 11 frames, and 11 frames. Among them, Video 1 and Video 2 have reached the maximum number, and Video 3 and Video 4 can still be allocated. At this time, 37 frames of the total number of frames to be analyzed have been allocated, and there is 1 frame remaining, which is not enough to be allocated to all videos (Video 3 and Video 4). The policy monitoring module allocates the remaining 1 frame to Video 4 with a longer duration in the 12th round according to the duration of the videos. After the 38 frames are allocated, the first quantity of Video 1 is 7 frames, the first quantity of Video 2 is 8 frames, the first quantity of Video 3 is 11 frames, and the first quantity of Video 4 is 12 frames.

[0303] After obtaining the first quantity of each video, the policy monitoring module can perform highlight segment analysis on the I-frames of each video. Among them, the policy monitoring module performing highlight segment analysis on each video can include two analysis stages. Among them, the first analysis stage includes performing an overview analysis based on the I-frames of the first quantity, and the second analysis stage includes performing frame-by-frame analysis based on the target area.

[0304] Among them, the first analysis stage refers to the overview analysis of the first number of I-frames for each video. Specifically, for each video, the policy monitoring module sends the position of one I-frame to the channel interface of the media middleware framework layer each time, so that it analyzes the I-frame at that position and returns the analysis result; until the number of I-frames sent for analysis reaches the first number of the video. After the first analysis stage, the policy monitoring module can obtain the analysis results of the first number of I-frames of each video. The analysis result includes the first score of the I-frame. Thus, the target area of each video can be determined based on the first I-frame with the highest score. Optionally, exemplarily, when the policy monitoring module receives the first score of each I-frame, it can determine the score of the corresponding area according to the first score of the I-frame. Different weights are set for areas with different scores. In areas with lower scores, the positions of I-frames are selected for analysis in a sparse frame extraction manner; or, in areas with lower scores, no I-frames are extracted and sent for analysis. In areas with higher scores, the positions of I-frames are selected for analysis in a dense frame extraction manner. In this way, the area with the highest score is finally determined as the target area.

[0305] The second analysis stage refers to the frame-by-frame analysis of the target area. Specifically, for all videos, the frame-by-frame analysis strategy for the target area is determined according to the duration of the target area. For example, the frame-by-frame analysis strategy can be to perform frame-by-frame analysis of I-frames for the target area of each video; or, the frame-by-frame analysis strategy can be to perform frame-by-frame analysis of ordinary image frames for the target area of each video. The policy monitoring module sequentially sends each frame of the target area to the channel interface of the media middleware framework layer to obtain the second score of each frame. Thus, based on the second scores of each frame of the target area of each video, the second image frame (I-frame or image frame) with the highest second score and the corresponding area are determined as the highlight segment of the video.

[0306] Taking Video 1 as an example, the process of analyzing the highlight segment of the video is illustrated below in combination with S24 - S37. Among them, S24 - S29 are the first analysis stage, and S30 - S37 are the second analysis stage.

[0307] In S24, the policy monitoring module of the media middleware framework layer sends the file descriptor fd1 of Video 1 and the corresponding first position in Video 1 to the channel interface of the media middleware framework layer.

[0308] Among them, the first position refers to the position of the first I-frame in Video 1. The first position may include one or more preset positions. For example, the first position may include the position at the 1 / 3 time point of the entire video duration. The first position also includes a first preset position and a second preset position, where one of the preset thresholds is the position at the 1 / 3 time point of the entire video duration, and the second preset position is the position at the 2 / 3 time point. The preset position is generally the position of the I-frame for the first time the video is sent for analysis. After receiving the first score of the I-frame sent for the first time, the I-frame at the second position is dynamically determined for analysis.

[0309] For example, with reference to Figure 14 as shown, Figure 14 Figure 7 shows a schematic diagram of the first preset position (1 / 3 time point of the video duration) and the second preset position (2 / 3 time point of the video duration) in a 15-second video.

[0310] S25. The channel interface of the media middle platform framework layer performs processing operations such as decoding, reducing the resolution, and converting the format of Video 1 according to the file descriptor of Video 1, and stores the processed Video 1.

[0311] S26. The channel interface of the media middle platform framework layer sends the frame data address of Video 1 and the corresponding first position in Video 1 to the highlight segment algorithm interface of the HAL layer through the service interface (such as the one-click video compilation interface) of the FWK layer.

[0312] S27. The highlight segment algorithm interface of the HAL layer obtains the I-frame at the first position of Video 1 according to the frame data address of Video 1 and the corresponding first position in Video 1, and then analyzes the I-frame at the first position based on the preset highlight segment algorithm to obtain the analysis result of the I-frame at the first position.

[0313] In this embodiment, the highlight segment algorithm interface of the HAL layer obtains the first position used to indicate the time point. In fact, there may not necessarily be an I-frame at the first position. In this case, the highlight segment algorithm interface of the HAL layer can obtain the I-frame closest to the first position at a time point near the first position as the I-frame at the first position for analysis, and obtain the analysis result of the I-frame at the first position.

[0314] S28. The highlight segment algorithm interface of the HAL layer returns the analysis result of the I-frame at the first position to the policy monitoring module of the media middle platform framework layer through the service interface of the FWK layer and the channel interface of the media middle platform framework layer in sequence.

[0315] Among them, the analysis result of the I-frame at the first position can represent the second score of the I-frame.

[0316] In S29, the policy monitoring module in the media middle platform framework layer determines a second position based on the analysis result of the I-frame at the first position, and returns to execute S24 until the number of analyzed I-frames meets the first quantity of the video.

[0317] In this way, the policy monitoring module can obtain the analysis results of the first quantity of I-frames of all videos. Among them, the analysis result includes the first score of the I-frame. Based on the first score of the I-frame, the solution of the second analysis stage is executed.

[0318] In some embodiments, for the above S29, the policy monitoring module determines the second position according to the analysis result of the I-frame at the first position, and further includes:

[0319] If the first position is a preset position, and when the preset position includes the first preset position, the first preset position divides the video into region 1 and region 2, and takes the midpoint position of region 1 or region 2 as the second position for sending the I-frame at the second position for analysis.

[0320] If the first position is a preset position, and the preset position includes the first preset position and the second preset position. The region formed by the starting position of the video and the midpoint position between the first preset position and the second preset position is used as the first region. The first region includes the first preset position, and the score of the first region is the score of the I-frame at the first preset position.

[0321] The region formed by the midpoint position between the first preset position and the second preset position and the end position of the video is used as the second region. The second region includes the second preset position, and the score of the second region is the score of the I-frame at the second preset position.

[0322] Calculate the product of the score of the first region and the region duration, and the product of the score of the second region and the region duration. Take the midpoint position of the region with the higher product as the second position for sending the I-frame at the second position for analysis.

[0323] Exemplarily, taking the first preset position as the 1 / 3 time point of the video duration (corresponding to I-frame 1) and the second preset position as the 2 / 3 time point of the video duration (corresponding to I-frame 2) as an example for illustration. After obtaining the first scores of the I-frames at the 1 / 3 time point and 2 / 3 time point of the video, the second position of the video is determined.

[0324] Exemplarily, referring to Figure 15 , Figure 15 Another schematic diagram of the first position is given.

[0325] After the policy monitoring module obtains the analysis results of I-frame 1 and I-frame 2, it can divide the region by taking the midpoint of the positions where I-frame 1 and I-frame 2 are located according to the positions of these I-frame 1 and I-frame 2. Referring to Figure 15, taking the mid - position between the 1 / 3 time - point and the 2 / 3 time - point as the dividing point, the area from the starting position of the video to the mid - point position is regarded as the area 1 corresponding to I - frame 1, and the area from the mid - point position of the video to the end position of the video is regarded as the area 2 corresponding to I - frame 2. For example, if the first score of I - frame 1 is 70, then the score of area 1 is 70; if the first score of I - frame 2 is 90, then the score of area 2 is 90.

[0326] In this embodiment, areas with the same score and continuous in time are regarded as connected areas, and there is no area overlap between connected areas. Combining Figure 15 with the given example, the connected areas of the video include area 1 and area 2. Among them, the product of the score of area 1 and the area duration is 70*(7.5 - 0)=525; the product of the score of area 2 and the area duration is 90*(15 - 7.5)=675. The policy monitoring module determines that the area with the largest product is area 2, and selects a position of an I - frame in area 2, for example, the mid - point position of area 2, as the second position for sending for analysis.

[0327] After obtaining the first score of the I - frame, if the first score is less than the third threshold, then the scores of the areas formed within the preset duration before and after the time - point of this I - frame are 0. Among them, the preset duration can be 500 ms. That is, the scores of the areas within 1 second before and after the mid - point position of this I - frame are 0. When determining the position of the next I - frame subsequently, it is not determined from the areas with a score of 0.

[0328] If the first score of the I - frame is greater than the third threshold, that is, the scoring result of the I - frame is "medium", "relatively high", "high". The policy monitoring module then obtains the adjacent boundaries before and after the position of this I - frame. Among them, the adjacent boundaries include one or two of the positions of the adjacent I - frames that have been analyzed, the boundaries of the areas with a score of 0, the starting position of the video, the end position of the video, etc.

[0329] If the adjacent boundary is the position of an adjacent I - frame that has been analyzed, the score of this boundary is the score of this I - frame; if the adjacent boundary is the boundary of an area with a score of 0, the score of this boundary is 0; if the adjacent boundary is the starting position or the end position of the video, the score of this boundary is the score of the original area where this position is located. The scores of each connected area can include the average value of the first scores of the I - frames it contains. The connected area does not include the 0 - score area.

[0330] In this embodiment, different weights can be set for scores in different value ranges. For example, for scores greater than or equal to the first threshold, the corresponding weight is 1; for scores greater than or equal to the second threshold and less than the first threshold, the corresponding weight is 2; for scores greater than the third threshold, the corresponding weight is 3. Among them, weight 1 is greater than weight 2, and weight 2 is greater than weight 3. For example, weight 1 can be 3, weight 2 can be 2, and weight 1 can be 1.

[0331] The policy monitoring module can determine the first boundary and the second boundary of the third region corresponding to the I-frame according to the scores and corresponding weights of the adjacent boundaries of the I-frame, the first score and the corresponding weight of the I-frame.

[0332] Exemplarily, the adjacent boundary of I-frame 3 includes the position (P1, such as the position at 10s of the video) of the adjacent analyzed I-frame 2 before I-frame 3, and the end position of the video after I-frame 3. Among them, the first score of I-frame 2 is 90, the first score is greater than the first threshold, and the corresponding weight is 1 (w1, such as 3).

[0333] The boundary position determined after the time point corresponding to I-frame 3 is the position of the video end time (P2, such as at the time point 15s). There is no analyzed I-frame at the time point 15s, but the score of the original region 2 where the time point 15s is located is 90, which is greater than the first threshold, and the corresponding weight is 1 (w2, such as 3).

[0334] If the first position is Figure 15 the position (P3, such as the position at 11.75s of the video) of I-frame 3 shown, the first score of I-frame 3 is 60, which is greater than the second threshold and less than the first threshold, and the corresponding weight is 2 (w3, such as 2).

[0335] The policy monitoring module then determines the boundary position of the third region of I-frame 3 according to the scores and corresponding weights of the adjacent boundaries of I-frame 3, and the first score and the corresponding weight of I-frame 3, including:

[0336] The calculation method of the first boundary P11 of the third region (region 3) corresponding to I-frame 3 can be obtained by the following formula:

[0337] P11 = (P3 - P1) / (w1 + w3) * w1 + P1.

[0338] For example, referring to Figure 16 , Figure 16 a schematic diagram of region 3 is provided. Among them, P1 is 10s, w1 is 3, P3 is 11.75s, and w3 is 2. The position of the first boundary P11 is 10.75s.

[0339] Alternatively, calculate the first boundary P11 through P11 = P3 - (P3 - P1) / (w3 + w1) * w3.

[0340] It can be understood that when w1 is greater than w3, P11 can be calculated by the above formula. If the score of w1 is less than w3, the calculation method of P11 is:

[0341] P11 = (P3 - P1) / (w3 + w1) * w3 + P1; or, P11 = P3 - (P3 - P1) / (w3 + w1) * w1.

[0342] The calculation method of the second boundary P22 of the third region (region 3) corresponding to I-frame 3 can be obtained by the following formula:

[0343] P22 = (P2 - P3) / (w3 + w2) * w3 + P3.

[0344] For example, referring to Figure 16 , P2 is 15s, w2 is 3, P3 is 11.75s, and w3 is 2. The position of the second boundary P22 is 12.75s.

[0345] Alternatively, the second boundary P22 is calculated by P22 = P2 - (P2 - P3) / (w3 + w2) * w2.

[0346] It can be understood that when w3 is less than w2, P11 can be calculated by the above formula. If the score of w3 is greater than w2, the calculation method of P11 is:

[0347] P22 = (P2 - P3) / (w3 + w1) * w2 + P3; or, P22 = P2 - (P2 - P3) / (w3 + w1) * w3.

[0348] In this way, the first boundary of region 3 corresponding to I-frame 3 can be determined to be at the position of 10.75s, and the second boundary is at the position of 12.75s. That is, referring to Figure 16 , the region of 10.75s - 12.75s of the video is region 3 corresponding to I-frame 3.

[0349] After determining region 3 of I-frame 3, the second position is determined according to the scores of all connected regions in the video.

[0350] After determining region 3 corresponding to I-frame 3 based on the first score of I-frame 3, the existing connected regions in the video are determined.

[0351] For example, referring to Figure 17 , Figure 17 shows a schematic diagram of the connected regions of a video. Figure 17 In, the original region 2 where I-frame 3 is located is divided into two connected regions, region 4 and region 5, by region 3. The existing connected regions in the video include region 1, region 4, region 3, and region 5.

[0352] Among them, region 1 only includes an I-frame 1 (at the 5s position, the first score is 70) to be analyzed, and since region 1 is the original region determined in the first round, the score of region 1 is the first score of this I-frame 1; the score of region 4 is 90; there is no analyzed I-frame in region 5 for the time being, so the score of region 5 is still the score of the original region 2, which is 90.

[0353] The score of Region 3 can be calculated by the following formula:

[0354] Score of Region 3 = (First score of I-frame 3 + Score of the original region where Region 3 is located) / 2.

[0355] Among them, the original region where Region 3 (10.75s - 12.75s) is located is Region 2, the score of Region 2 is 90, and the score of I-frame 3 is 60. Therefore, the score of Region 3 is 75.

[0356] Obtain the product of the score and the region duration of each connected region, and select the second position from the connected regions with the largest product for analysis. For example, the midpoint position of the connected region with the largest product can be used as the second position.

[0357] Figure 17 Among them, the product corresponding to Region 1 is 70 * (7.5 - 0) = 525, the product corresponding to Region 3 is 75 * (12.75 - 10.75) = 150, the product corresponding to Region 4 is 90 * (10.75 - 7.5) = 292.5, and the product corresponding to Region 5 is 90 * (15 - 12.75) = 202.5.

[0358] Among them, the connected region with the largest product is Region 1, and the strategy monitoring module can determine the midpoint position of Region 1 as the second position for analysis, referring to Figure 17 the position of 3.75s.

[0359] The strategy monitoring module sends the I-frame at the second position for analysis. Based on the first score of I-frame 4 at the second position, it determines Region 4 corresponding to I-frame 4, thereby updating the connected regions of the video. Determine the third position from the connected regions with the largest product, and repeat this process until the number of I-frames sent for analysis reaches the first number.

[0360] After the number of Video 1 sent by the strategy monitoring module in the media middle platform framework layer meets the first number of Video 1, if there are other videos, such as Video 2, then the electronic device can continue to execute the above S23 - S29 to obtain the analysis results of the first number of I-frames of other videos. In this way, by repeatedly executing S23 - S29 for each video, the strategy monitoring module can obtain the analysis results of the first number of I-frames in each video among multiple image materials.

[0361] After obtaining the analysis results of the first number of I-frames in each video among the image materials, the policy monitoring module can determine the target area of each video according to the analysis results of the first number of I-frames in each video, and perform frame-by-frame analysis on the target area of each video to obtain the first highlight segment of each video. Among them, before the policy monitoring module of the media middle platform framework layer performs frame-by-frame analysis on the target area of each video, it can also determine the frame-by-frame analysis policy of the target area based on the duration of the target area of each video.

[0362] Still taking Video 1 as an example, the second analysis stage of a video processing method in this embodiment will be described in combination with Figure 10 shown in S30-S37.

[0363] S30, the policy monitoring module of the media middle platform framework layer determines the target area of each video based on the analysis results of the I-frames of all videos.

[0364] Among them, the analysis results include the first scores of the first number of I-frames of each video.

[0365] In this embodiment, for each video, after the policy monitoring module obtains the first scores of the first number of I-frames, it can determine the first I-frame with the highest first score. Based on this first I-frame and the recommended highlight segment duration, the candidate highlight segment of the video is determined, and the area formed within the third preset duration before and after the candidate highlight segment is determined as the target area. Exemplarily, the duration of the target area can be 2 times the duration of the candidate highlight segment.

[0366] In some embodiments, the first I-frame may include at least one I-frame. If the first I-frame includes one I-frame, the I-frames within the first preset duration before and after the first I-frame can be directly determined as the target area of Video 1. For example, as Figure 18 (a) of gives a schematic diagram of a target area. Among them, the first I-frame of Video 1 includes I-frame 1 (such as 1 in (a) of Figure 18 ), and the segment covered by the first preset duration before and after I-frame 1 is the target area.

[0367] Or, if the first I-frame includes one I-frame, according to the recommended highlight segment duration, the segments of 1 / 2 of the recommended highlight segment duration before and after the first I-frame can be determined as the candidate highlight segment. The I-frames within the third preset duration before and after the candidate highlight segment are determined as the target area of Video 1. Among them, the duration of the target area can be b times the duration of the candidate highlight segment. Exemplarily, b can be a number greater than 1 and less than or equal to 2.

[0368] For example, as Figure 18(b) of it gives another schematic diagram of the target area. Among them, the first I-frame of Video 1 includes I-frame 1 (such as 1 in Figure 18 in (b)), and the segments with a duration of 1 / 2 of the recommended highlight segment duration before and after I-frame 1 are candidate highlight segments. The segments covered by the third preset duration before and after the candidate highlight segments are the target area.

[0369] If the first I-frame includes multiple consecutive I-frames. According to the recommended highlight segment duration, the segment covered by the fourth preset duration before the first I-frame among the multiple consecutive first I-frames to the fourth preset duration after the last I-frame among the multiple consecutive second I-frames is determined as the candidate highlight segment.

[0370] For example, such as Figure 19 , Figure 19 is another schematic diagram of the target area. The first I-frame of Video 1 includes consecutive I-frame 1 (such as 1 in Figure 19 ), I-frame 2 (such as 2 in Figure 19 ), and I-frame 3 (such as 3 in Figure 19 ). Then, the segment covered by the fourth preset duration before I-frame 1 to the fourth preset duration after I-frame 3 is the candidate highlight segment of this Video 1. The segments covered by the fourth preset duration before and after the candidate highlight segment are the target area.

[0371] After determining the candidate highlight segments of each video, if the sum of the durations of the candidate highlight segments of all videos is greater than the second multiple of the total recommended highlight segment duration, the policy monitoring module needs to adjust the durations of the candidate highlight segments of each video so that the sum of the durations of the candidate highlight segments of all videos is less than or equal to the second multiple of the total recommended highlight segment duration.

[0372] Among them, the second multiple can be a number greater than 1. For example, the second multiple can be 1.05. Exemplarily, according to the duration by which the sum of the durations of the candidate highlight segments of all videos exceeds the second multiple of the total recommended highlight segment duration (the duration to be adjusted), the durations of the candidate highlight segments are reduced and adjusted according to the actual durations of each video. Exemplarily, the shorter the duration of a video, the more seconds its candidate highlight segment is reduced.

[0373] Suppose two videos are input. The actual duration of Video 1 is 10s, and the actual duration of Video 2 is 20s. The recommended highlight segment duration is 6s, and the total recommended highlight segment duration is 10s. The durations of the candidate highlight segments of Video 1 and Video 2 are both 6s; the durations of the target areas of Video 1 and Video 2 are both 12s.

[0374] Among them, the sum of the durations of the candidate highlight segments of Video 1 and Video 2 is 12s, exceeding 1.05 times (10.5s) of the total recommended duration of the highlight segments (10s). The strategy monitoring module needs to adjust the durations of the candidate highlight segments of each video. The duration to be adjusted is 12s - 10s * 1.05 = 1.5s. The strategy monitoring module allocates weights according to the reciprocal of the video duration. The duration of the candidate highlight segment of Video 1 (10s) is reduced by 1s; the duration of the candidate highlight segment of Video 1 is updated from 6s to 5s, the duration of the candidate highlight segment of Video 2 (20s) is reduced by 0.5s, and the duration of the candidate highlight segment of Video 2 is updated from 6s to 5.5s. In this way, the sum of the durations of the candidate highlight segments of Video 1 and Video 2 is 10.5s, not exceeding 10s * 1.05, and no further adjustment is required.

[0375] Therefore, it is determined that the duration of the candidate highlight segment of Video 1 is 5s, and the duration of the target area corresponding to the candidate highlight segment in Video 1 is 10s; the duration of the candidate highlight segment of Video 2 is 5.5s, and the duration of the target area corresponding to the candidate highlight segment in Video 2 is 11s.

[0376] The method provided by S30 can be used to determine the target area of each video. After determining the target area of each video, S31 - S36 can be executed for frame-by-frame analysis of the target area.

[0377] S31, the strategy monitoring module of the media middle platform framework layer determines the frame-by-frame analysis strategy for the target area of each video.

[0378] Among them, the frame-by-frame analysis strategy for the target area includes analyzing all image frames of the target area (high-precision analysis), or analyzing all I-frames of the target area (low-precision analysis).

[0379] If the remaining analysis duration satisfies the time consumption for analyzing the image frames in the target areas of all videos, then analyze all the image frames in the target areas of each video; if the remaining analysis duration does not satisfy the time consumption for analyzing the image frames in the target areas of all videos, and the remaining analysis duration satisfies the time consumption for analyzing the I-frames in the target areas of all videos, then analyze all the image frames in the target areas of each video; if the remaining analysis duration does not satisfy the time consumption for analyzing the I-frames in the target areas of all videos, then analyze the I-frames in the target areas of each video.

[0380] Among them, the time consumption for analyzing the I-frames in the target areas of all videos can be determined according to the number of I-frames in the target areas of all videos and the single-frame analysis time consumption. The time consumption for analyzing all the image frames in the target areas of each video can be determined according to the number of image frames in the target areas of all videos and the single-frame analysis time consumption. The remaining analysis duration refers to the duration for video analysis minus the total duration spent on the overview analysis of the key frames.

[0381] Exemplarily, assume that the entire input video transmits 10 image frames per second (10 FPS) and 1 I-frame per second; the durations of the target regions of each video are 5s, 10s, and 15s respectively. Assume that it takes 200 ms to process 1 frame and the remaining analysis duration is 25s.

[0382] The durations taken to analyze all the image frames in the target regions of the 3 videos are respectively: 5s * 10 FPS * 200 ms = 10s, 10s * 10 FPS * 200 ms = 20s, 15s * 10 FPS * 200 ms = 30s, with a total of 60s.

[0383] The durations taken to analyze all the I-frames in the target regions of the 3 videos are respectively: 5s * 1 FPS * 200 ms = 1s, 10s * 1 FPS * 200 ms = 2s, 15s * 1 FPS * 200 ms = 3s, with a total of 6s.

[0384] The remaining analysis duration of 25s is not sufficient for the time taken to analyze all the image frames in the target regions of the 3 videos, but it is sufficient for the time taken to analyze all the I-frames in the target regions of the 3 videos. Then, the policy monitoring module analyzes all the image frames in the target regions of the 3 videos.

[0385] In some embodiments, after determining the per-frame analysis policy, the policy monitoring module of the media middle platform framework layer can also allocate the durations in descending order according to the scores of the target regions of each video. For example, the score of target region 1 of video 1 is 80, the score of target region 2 of video 2 is 90, and the score of target region 3 of video 3 is 90. The per-frame analysis policy is to analyze each image frame, and the remaining analysis duration is 25s. According to the principle of preferentially allocating durations to target regions with higher scores. The score of target region 2 is the highest, and the analysis duration of the image frames in target region 2 is 20s, so 20s of duration is allocated to video 2. At this time, the remaining analysis duration is 5s. The score of target region 1 is 80, and the analysis duration of the image frames in target region 1 is 10s. The remaining 5s of analysis duration is all allocated to video 1. The analysis duration is exhausted. Video 3 does not receive the allocated duration, so the per-frame analysis of target region 3 is not performed.

[0386] After the policy monitoring module of the media middle platform framework layer determines the per-frame analysis policy for the target region of each video, for example, the per-frame analysis policy indicates analyzing all the I-frames in the target region of each video. Taking video 1 as an example, the process of performing per-I-frame analysis on the target region of video 1 may include:

[0387] S32. The policy monitoring module of the media middle platform framework layer sends the file descriptor fd1 of video 1 and the positions of the I-frames in the target region of video 1 to the channel interface of the media middle platform framework layer.

[0388] Among them, if the frame-by-frame analysis strategy indicates to analyze the I-frames in the target regions of all videos, the position of the I-frame can be the position of an I-frame in the target region of Video 1.

[0389] It can be understood that if the frame-by-frame analysis strategy indicates to analyze the image frames in the target regions of all videos, the position of the I-frame can be the position of an image frame in the target region of Video 1.

[0390] S33. The channel interface of the media middle platform framework layer performs processing operations such as decoding, downscaling the resolution, and converting the format on Video 1 according to the file descriptor of Video 1, and stores the processed Video 1.

[0391] S34. The channel interface of the media middle platform framework layer sends the frame data address of Video 1 and the position of the I-frame to the highlight segment algorithm interface of the HAL layer through the service interface of the FWK layer (such as the one-click video creation interface).

[0392] S35. The highlight segment algorithm interface of the HAL layer obtains the I-frame of Video 1 according to the frame data address of Video 1 and the position of the I-frame, and then analyzes the I-frame based on a preset highlight segment algorithm to obtain the analysis result of the I-frame.

[0393] S36. The highlight segment algorithm interface of the HAL layer returns the analysis result of the I-frame to the policy monitoring module of the media middle platform framework layer through the service interface of the FWK layer and the channel interface of the media middle platform framework layer in sequence.

[0394] The policy monitoring module of the media middle platform framework layer continues to send the position of the next I-frame in the target region of Video 1 to the channel interface of the media middle platform framework layer for analysis until all the I-frames in the target region are analyzed.

[0395] S37. The policy monitoring module of the media middle platform framework layer determines the highlight segments of Video 1 from the analysis results of the I-frames in the target region of Video 1.

[0396] Among them, the analysis result includes the second score of the I-frame in the target region.

[0397] In this embodiment, the policy monitoring module obtains the second I-frame with the highest second score according to the second score of the I-frame. Among them, the second I-frame can include at least one I-frame. If the second I-frame includes one I-frame, the I-frames within the second preset duration before and after the second I-frame can be directly determined as the highlight segments of Video 1.

[0398] For example, Figure 20 (a) of gives a schematic diagram of a highlight segment. Among them, the second I-frame of Video 1 includes I-frame 1 (such as Figure 20In 1) of (a) below, the segments covered by the second preset duration before and after I-frame 1 are the highlight segments of Video 1. Among them, the duration of the highlight segment is less than or equal to the recommended duration of the highlight segment.

[0399] Alternatively, if the second I-frame includes one I-frame, according to the recommended duration of the highlight segment, the segments of 1 / 2 of the recommended duration of the highlight segment before and after the second I-frame can be determined as the highlight segments of Video 1. Among them, the duration of the highlight segment is equal to the recommended duration of the highlight segment.

[0400] For example, as Figure 20 (b) below gives another schematic diagram of the highlight segment. Among them, the second I-frame of Video 1 includes I-frame 1 (as Figure 20 in 1) of (b) below), and the segments of 1 / 2 of the recommended duration of the highlight segment before and after I-frame 1 are the highlight segments of Video 1.

[0401] If the second I-frame includes multiple consecutive I-frames. According to the recommended duration of the highlight segment, the segment covered by the fifth preset duration before the first I-frame of the multiple consecutive second I-frames to the fifth preset duration after the last I-frame of the multiple consecutive second I-frames is determined as the highlight segment of Video 1.

[0402] For example, as Figure 21 , Figure 21 gives another schematic diagram of the highlight segment. The first I-frame of Video 1 includes consecutive I-frame 1 (as Figure 21 in 1 below), I-frame 2 (as Figure 21 in 2 below), I-frame 3 (as Figure 21 in 3 below). Then, the segment covered by the fifth preset duration before I-frame 1 to the fifth preset duration after I-frame 3 is the highlight segment of this Video 1.

[0403] If there are other videos, such as Video 2, then the electronic device can continue to execute the above S32 - S37 to obtain the analysis results of the highlight segments of other videos.

[0404] After obtaining the analysis results of the highlight segments of all videos, in some embodiments, referring to Figure 22 , gives a schematic flowchart of the post-processing of the highlight segments in a video processing method. After the electronic device executes S37, it can use the following S38 to report all the image and video analysis results.

[0405] S38, the policy monitoring module of the media middle platform framework layer reports all the image and video analysis results to the application function layer through the image highlight segment analysis interface of the media middle platform framework layer.

[0406] Among them, the analysis results can include the positions of all the highlight segments. For example, the start time and end time of the highlight segments of each video.

[0407] S39. The application function layer clips and filters the user-selected materials according to the analysis results of all pictures and videos to obtain all highlight segments.

[0408] S40. The application function layer of the application layer calls the theme summary interface of the media middle platform framework to request and obtain the theme template.

[0409] Among them, this request passes through the theme summary interface and the channel interface of the media middle platform framework and is transmitted to the HAL layer through the FKW layer.

[0410] S41. The HAL layer determines the theme template that conforms to the scene according to the scene of the picture in the highlight segment.

[0411] In some embodiments, the electronic device may be provided with multiple theme templates (style templates). The theme algorithm can recommend the theme template that conforms to the picture scene based on the picture scene of the highlight segment, such as people, scenery, food, children, pets, sports or travel.

[0412] S42. The HAL layer returns the determined theme template to the theme summary interface of the media middle platform framework layer, and the theme summary interface returns the theme template to the application function layer of the application layer.

[0413] Exemplarily, assuming that most of the pictures in the highlight segment are parent-child scenes, then it can be confirmed that the theme template that conforms to this scene is the parent-child theme category.

[0414] S43. The application function layer distributes the obtained theme and all highlight segments to the basic capability layer.

[0415] The application function layer distributes the theme obtained from S42 and all the highlight segments obtained from S39 to the basic capability layer.

[0416] S44. The basic capability layer generates a target video set according to the theme and all highlight segments.

[0417] That is, the target video set is a video set generated according to all the filtered highlight segments and conforming to the recommended theme.

[0418] S45. The basic capability layer sends an indication message for displaying the target video set to the video editing service layer.

[0419] S46. The video editing service layer displays the target video set in the gallery interface.

[0420] In the embodiments of the present application, the media middle platform framework layer is used for decoding video and picture files, converting the data format into a unified format, monitoring the remaining time and adjusting the operation strategy, sending data, controlling the operation and termination of the algorithm, obtaining the results and returning them to the application layer, etc. The FKW layer is used to complete data packaging and provide data and program operation services. After receiving the commands sent by the media middle platform framework layer, the HAL layer performs highlight analysis according to the commands and returns the parameter calculation results of the highlight analysis to the media middle platform framework layer. The final results of the algorithm are collected and sorted by the media middle platform framework layer and then sent to the application layer for processing. The application layer can present the editing application interface, video and picture file options, and present the final results of the algorithm.

[0421] After the user starts the function of generating a video with one click, select the video and picture files to be edited (for example, up to 30 files are supported). After waiting for a while, the "Generate a Video with One Click" application automatically edits the highlight segments of the video, combines the highlight segments and pictures together according to the algorithm results, generates the edited short video, and previews and plays it.

[0422] For the video processing method provided in the embodiments of the present application, the electronic device can first perform an overview analysis on the first number of key frames in each video among multiple image materials to obtain the first score of the key frames of each video. Then, the electronic device can determine the target area of the video based on the first key frame with the highest first score and the key frames within the first preset duration before and after the first key frame. After that, the electronic device can perform frame-by-frame analysis on the target area in the video to obtain the second score of each frame. Based on the second image frame with the highest second score and the frames within the second preset duration before and after the second image frame, they are determined as the highlight segments of the video. With this solution, the electronic device first performs an overview analysis on the videos therein, so that the target area that needs to be processed frame by frame can be located. When the electronic device extracts the highlight segments, it only performs frame-by-frame analysis on the target area instead of the entire video, which can reduce the workload of the electronic device for image frame analysis, thereby reducing the time-consuming of the electronic device to implement the function of generating a video with one click, improving the efficiency of the electronic device for processing videos, and further improving the user experience of the function of generating a video with one click.

[0423] It should also be noted that, in the embodiments of the present application, "greater than" can be replaced by "greater than or equal to", "less than or equal to" can be replaced by "less than", or, "greater than or equal to" can be replaced by "greater than", and "less than" can be replaced by "less than or equal to".

[0424] Each of the embodiments described herein can be an independent solution or can be combined according to the internal logic, and these solutions all fall within the protection scope of the present application.

[0425] It can be understood that the methods and operations implemented by the electronic device in the above method embodiments can also be implemented by components (such as chips or circuits) available for the electronic device.

[0426] It should be noted that the personal information used in the technical solution of this application is limited to the information obtained with the individual consent of the individual, including but not limited to, before the user uses this function, notifying and reminding the user to read the relevant user agreement (notification), and signing the agreement (authorization) including authorizing the relevant user information. Among them, personal information includes information such as pictures and videos stored by the user.

[0427] In the technical solution disclosed in this application, the processing of the user's personal information such as collection, storage, use, processing, transmission, provision, and disclosure complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0428] The above describes the method embodiments provided in this application. The following will describe the device embodiments provided in this application. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments. Therefore, the content not described in detail can be referred to the above method embodiments. For the sake of brevity, it will not be repeated here.

[0429] The above mainly describes the solution provided in the embodiments of this application from the perspective of method steps. It can be understood that, in order to implement the above functions, the electronic device implementing this method includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should be able to realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the protection scope of this application.

[0430] The embodiments of this application can divide the electronic device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other feasible division methods in actual implementation. The following takes the division of each functional module corresponding to each function as an example for illustration.

[0431] The present application also provides a chip, which is coupled to a memory and is configured to read and execute a computer program or instructions stored in the memory to execute the methods in the above embodiments.

[0432] The present application also provides an electronic device, which includes a chip configured to read and execute a computer program or instructions stored in a memory, so that the methods in the embodiments are executed.

[0433] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above electronic device, the electronic device is caused to execute each function or step that the electronic device executes in the above method embodiments.

[0434] An embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute each function or step that the electronic device executes in the above method embodiments. For example, the computer may be the above electronic device.

[0435] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0436] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0437] The units described as separate components may or may not be physically separated. The components displayed as units may be a physical unit or multiple physical units, that is, they may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0438] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0439] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A video processing method, characterized in that, The method includes: Receiving a selection operation of a user for a plurality of image materials in a picture gallery; wherein, the plurality of image materials includes videos; In response to the selection operation, performing an overview analysis of a first number of I-frames for each video among the plurality of image materials to obtain a first score; the first score is an aesthetic score for the corresponding image frame, and the first number corresponds to the duration of the video; the I-frame is an image frame including complete image information; Based on the first scores of the I-frames in each video, determining a target region for the corresponding video; the target region includes a first I-frame, and I-frames within a first preset duration before and after the first I-frame; the first I-frame is the I-frame with the highest first score; Performing frame-by-frame analysis on the target region of each video among the plurality of image materials to obtain a second score for each image frame in the target region; the frame-by-frame analysis includes I-frame by I-frame analysis or each image frame by frame-by-frame analysis, and the second score is an aesthetic score for the corresponding image frame; Based on the second scores corresponding to each video, determining a highlight segment for the corresponding video; wherein, the highlight segment includes a second image frame, and frames within a second preset duration before and after the second image frame; the second image frame is the image frame with the highest second score in the target region; the highlight segments of each video among the plurality of image materials are used to splice to obtain a target video set.

2. The method according to claim 1, characterized in that, The performing an overview analysis of a first number of I-frames for each video among the plurality of image materials to obtain a first score includes: Obtaining a first position of the analyzed I-frame in the video; Based on the first position of the analyzed I-frame and the first score of the analyzed I-frame, determining a second position of the I-frame to be analyzed for the video, and obtaining the first score of the I-frame at the second position, until the number of analyzed I-frames reaches the first number.

3. The method according to claim 2, wherein The first position includes a first preset position and a second preset position of the video; The determining a second position of the I-frame to be analyzed for the video based on the first position of the analyzed I-frame and the first score of the analyzed I-frame includes: Taking the region between the midpoint position between the first preset position and the second preset position and the start position of the video as a first region; the first region includes the I-frame at the first position, and the score of the I-frame in the first region is used as the first score of the I-frame at the first preset position; Taking the region between the midpoint position and the end position of the video as a second region; the second region includes the I-frame at the second preset position, and the score of the second region is used as the first score of the I-frame at the second preset position; Obtaining the product of the score of each region and the duration of the region, and taking the midpoint position of the region with the highest product as the second position.

4. The method according to claim 3, wherein The first preset position is the 1 / 3 position of the video duration, and the second preset position is the 2 / 3 position of the video duration.

5. The method according to claim 2, wherein The first position is a non-preset position; The determining a second position of the I-frame to be analyzed for the video based on the first position of the analyzed I-frame and the first score of the analyzed I-frame includes: Determine a third region including the I-frame at the first position according to the adjacent boundaries before and after the I-frame at the first position; the adjacent boundaries include the position of the I-frame adjacent to and already analyzed with respect to the I-frame at the first position, the start position of the video, the end position of the video, and the boundary of the connected region; wherein, the connected region is a region with the same score and continuous in time, and there is no region coverage between the connected regions, and the third region is one of the connected regions of the video; Obtain multiple connected regions of the video and calculate the scores of each connected region; the score of the connected region is the average of the first score of the included I-frame and the score of the connected region; Obtain the product of each region score and the region duration, and use the midpoint position of the connected region with the highest product as the second position.

6. The method according to claim 5, wherein The determining the third region corresponding to the I-frame at the first position according to the adjacent boundaries before and after the I-frame at the first position includes: Determine the first boundary of the third region according to the first score of the I-frame at the first position and the corresponding first weight, and the score and the corresponding second weight of the adjacent boundary before the I-frame at the first position; Determine the second boundary of the third region according to the first score of the I-frame at the first position and the first weight, and the score and the corresponding third weight of the adjacent boundary after the I-frame at the first position; Wherein, there is a corresponding relationship between the score and the weight.

7. The method according to any one of claims 1-6, characterized in that, The determining the target region of the corresponding video based on the first scores of the I-frames in each video includes: For each video, obtain the first I-frame according to the first scores of the I-frames in the video; Determine the candidate highlight segment of the video as the segment including the first I-frame in the video and having a duration of the preset highlight segment recommended duration; Determine the candidate highlight segment and the I-frames within the third preset duration before and after the candidate highlight segment as the target region of the video; the third preset duration is less than the first preset duration.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: In response to the selection operation, obtain the analysis parameters corresponding to the multiple image materials; wherein, the analysis parameters include the actual duration of the corresponding video and the analysis total duration upper limit recommended value, and the analysis total duration upper limit recommended value represents the recommended value of the maximum duration required to analyze the multiple image materials; Determine the first number of I-frames of each video according to the actual duration of each video in the multiple image materials, the number of videos in the multiple image materials, and the analysis total duration upper limit recommended value.

9. The method according to claim 8, wherein The determining the first number of I-frames of each video according to the actual duration of each video in the multiple image materials, the number of videos in the multiple image materials, and the analysis total duration upper limit recommended value includes: Based on the actual duration of each video and the preset corresponding relationship, determine the base quantity and the maximum quantity of the corresponding video; wherein, the base quantity is the minimum number of I-frames required to ensure the analysis effect of the video, and the maximum quantity is the maximum number of I-frames that the duration allows for analyzing the video; the preset corresponding relationship represents the maximum quantity and the base quantity of the I-frames corresponding when the video duration is in different threshold ranges; Based on the base quantity and the maximum quantity of each video, and the recommended value of the upper limit of the total analysis duration, determine the total analysis quantity; the total analysis quantity is the total number of I-frames allowed to be analyzed for all videos in the multiple image materials; According to the actual duration of each video in the multiple image materials and the number of videos in the multiple image materials, allocate the total analysis quantity to each video to obtain the first quantity of I-frames for each video.

10. The method according to claim 9, characterized in that, The analysis parameter further includes the single-frame analysis duration; the single-frame analysis duration is the duration required to analyze one image frame; The determining the total analysis quantity based on the base quantity and the maximum quantity of each video, and the recommended value of the upper limit of the total analysis duration includes: If the recommended value of the upper limit of the total analysis duration is less than the time consumed for analyzing the sum of the base quantities of the I-frames of all videos in the multiple image materials, the total analysis quantity is the ratio of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration; If the recommended value of the upper limit of the total analysis duration is greater than the time consumed for analyzing the sum of the base quantities of the I-frames of all videos in the image materials, and the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumed for analyzing the sum of the base quantities of the I-frames of all videos in the multiple image materials, the total analysis quantity is the sum of the base quantities of the I-frames of all videos in the multiple image materials; If the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is greater than the time consumed for analyzing the sum of the base quantities of the I-frames of all videos in the multiple image materials, and the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is less than the time consumed for analyzing the sum of the maximum quantities of the I-frames of all videos in the multiple image materials, the total analysis quantity is the ratio of the duration of the second multiple of the recommended value of the upper limit of the total analysis duration to the single-frame analysis duration; If the duration of the second multiple of the recommended value of the upper limit of the total analysis duration is greater than the time consumed for analyzing the sum of the maximum quantities of the I-frames of all videos in the multiple image materials, the total analysis quantity is the sum of the maximum quantities of the I-frames of all videos in the multiple image materials; Wherein, the second multiple is greater than 0 and less than 1.

11. The method according to claim 9 or 10, characterized in that, The allocating the total analysis quantity to each video according to the actual duration of each video in the multiple image materials and the number of videos in the multiple image materials to obtain the first quantity of I-frames for each video includes: Traverse each video among the multiple image materials, and update the first value and the second value of each video until the first value is 0; wherein, the initial value of the first value is equal to the total number of analyses, and the initial value of the second value is 0; for each traversed video, the second value of the video is incremented by 1, and the first value is decremented by 1; Use the second value of each video as the first quantity of I-frames of the video.

12. The method according to claim 11, wherein The traversing each video among the multiple image materials and updating the first value and the second value of each video until the first value is 0 includes: Before traversing to the first video among the multiple image materials, if the second value of the first video is equal to the maximum quantity of I-frames of the first video, skip the first video and traverse the next video of the first video; Wherein, skipping the first video means that the second value of the first video is not incremented by 1.

13. The method according to any one of claims 1-12, characterized in that, The per-frame analysis of the target region of each video among the multiple image materials to obtain the second score of each image frame in the target region includes: Based on the preset single-frame analysis duration and the quantity of all I-frames in the target regions of all videos in the image material, obtain the sum of the first durations for analyzing I-frames; the single-frame analysis duration is the duration required to analyze one image frame; Based on the single-frame analysis duration and the quantity of all image frames in the target regions of all videos in the image material, obtain the sum of the second durations for analyzing image frames; If the remaining analysis duration is less than the sum of the first durations, perform per-frame analysis on the I-frames of the target region of each video among the multiple image materials to obtain the second score of the I-frames in the target region; the remaining analysis duration is equal to the duration for video analysis minus the total duration spent on performing the overview analysis; If the remaining analysis duration is greater than or equal to the sum of the first durations, or the remaining analysis duration is greater than the sum of the second durations, perform per-frame analysis on the image frames of the target region of each video among the multiple image materials to obtain the second score of the image frames in the target region.

14. The method according to claim 13, wherein The method further includes: Arrange the target regions in descending order of the scores of the target regions, and allocate the remaining analysis duration to each of the target regions according to the time consumed for analyzing all image frames in the target region until the remaining analysis duration is completely allocated; If there are target regions that have not been allocated analysis duration, the target regions that have not been allocated analysis duration will not perform per-frame analysis.

15. The method according to any one of claims 1-14, characterized in that The determining the highlight segment of the corresponding video based on the second score corresponding to each video includes: For each video, obtain the second image frame according to the second score of the image frames of the target region; Determine the second image frame and the image frames within the second preset duration before and after the second image frame as the highlight segment of the video; the duration of the highlight segment is less than or equal to the highlight segment recommended duration.

16. The method according to any one of claims 1 to 15, characterized in that, Before performing the overview analysis of the first quantity of I-frames of each video among the multiple image materials to obtain the first score, the method further includes: If the total analysis duration of the multiple image materials is greater than the recommended value of the preset upper limit of the total analysis duration, randomly select M videos from the multiple image materials; wherein, the total analysis duration represents the total duration required to analyze the multiple image materials, and the recommended value of the preset upper limit of the total analysis duration represents the recommended value of the maximum duration required to analyze the multiple image materials; the duration required to analyze the M videos is less than or equal to the total analysis duration, and M is less than the number of videos in the image materials.

17. The method according to any one of claims 1 to 16, characterized in that, The overview analysis of the first number of I-frames for each video in the multiple image materials to obtain the first score includes: If the sum of the actual durations of all the videos in the multiple image materials is greater than the preset duration threshold, perform the overview analysis of the first number of I-frames for each video in the multiple image materials to obtain the first score.

18. An electronic device, characterized in that, The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory are coupled to the processor; the memory stores computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-17.

19. A computer-readable storage medium, characterized in that, It includes computer instructions. When the computer instructions run on the electronic device, the electronic device executes the method according to any one of claims 1-17.

20. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-17 is implemented.

Citation Information

Patent Citations

  • Video editing method and device, equipment and storage medium

    CN113709560A

  • Video processing method and electronic equipment

    CN115567660A

  • Mobile terminal short video highlight moment editing method based on key behavior recognition

    CN116095363A

  • Method and system for selecting highlight segments

    US20230230378A1

Cited By

  • Material video duplicate checking method and device, electronic equipment and computer readable medium

    CN119089003A